Unifying explainable agency via a ladder of intentions

Explainable Agency (XAg) is a subfield of Explainable AI (XAI) dedicated to explaining agents that perceive, reflect, reason, and act on an environment. However, the process of decision that brings about action is very heterogeneous across different architectures, and thus results in similarly heter...

ver descrição completa

Detalhes bibliográficos
Autores: Giménez Ábalos, Víctor, Tormos Llorente, Adrián, Vázquez Salceda, Javier|||0000-0003-1732-9446, Edström, Filip, Brännström, Mattias, Lindqvist, John, Cortés García, Claudio Ulises|||0000-0003-0192-3096, Álvarez Napagao, Sergio|||0000-0001-9946-9703
Tipo de documento: artigo
Data de publicação:2026
País:España
Recursos:Universitat Politècnica de Catalunya (UPC)
Repositório:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglês
OAI Identifier:oai:upcommons.upc.edu:2117/459254
Acesso em linha:https://hdl.handle.net/2117/459254
https://dx.doi.org/10.1007/s10796-026-10710-w
Access Level:Acceso aberto
Palavra-chave:XAI
Intentions
Agent explainability
Knowledge representation
Knowledge transfer
Cognitive architecture
Telic explanations
Explainable agency
RL
BDI
Agentic AI
Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Agents intel·ligents
Descrição
Resumo:Explainable Agency (XAg) is a subfield of Explainable AI (XAI) dedicated to explaining agents that perceive, reflect, reason, and act on an environment. However, the process of decision that brings about action is very heterogeneous across different architectures, and thus results in similarly heterogeneous ways of producing explanations. This becomes an issue in the absence of standards for evaluating explanations, both whether they are understandable by humans, and also whether they are trustworthy: the paradigm of agent explainability is riddled with different definitions on what are good metrics for evaluating explanations of an agent of a particular architecture or agent paradigm. Furthermore, when those metrics require of user studies, much work needs to be redone when applied to a different agent architecture, cohort of explainees, or context. Within this work, we propose a unifying way of examining agent behaviour, stratified into different abstraction levels (the rungs of a ladder), so as to be able to compare different agent architecture components. This architecture serves as an engineering pattern, decoupling the problem of providing trustworthy abstractions for giving explanations from the problem of formulating those explanations in an understandable manner. This paper is a first step in bringing about a world where architectures provide – and evaluate their trustworthiness on – abstractions that have been tested and evaluated by user studies, as well as enabling explanation providers to innovate on methods for formulating explanations from abstractions without needing expert knowledge on every agent architecture.