Full text
INTERNATIONAL DOCTORAL SCHOOL OF THE USC Ilia Stepin PhD Thesis Argumentative Conversational Agents for Explainable Artificial Intelligence Santiago de Compostela, 2023 Doctoral Programme in Information Technology Research
DOCTORAL THESIS ARGUMENTATIVE CONVERSATIONAL AGENTS FOR EXPLAINABLE ARTIFICIAL INTELLIGENCE Ilia Stepin INTERNATIONAL PHD SCHOOL OF THE UNIVERSITY OF SANTIAGO DE COMPOSTELA DOCTORAL PROGRAMME IN INFORMATION TECHNOLOGY RESEARCH SANTIAGO DE COMPOSTELA 2023
Declaración del autor de la tesis D. Ilia Stepin Título de la tesis: Argumentative Conversational Agents for Explainable Artificial Intelligence Presento mi tesis, siguiendo el procedimiento adecuado al Reglamento, y declaro que: 1. La tesis abarca los resultados de la elaboración de mi trabajo. 2. De ser el caso, en la tesis se hace referencia a las colaboraciones que tuvo este trabajo. 3. Confirmo que la tesis no incurre en ningún tipo de plagio de otros autores ni de trabajos presentados por mí para la obtención de otros títulos. 4. La tesis es la versión definitiva presentada para su defensa y coincide la versión impresa con la presentada en formato electrónico. Y me comprometo a presentar el Compromiso Documental de Supervisión en caso de que el original no esté en la Escuela. En Santiago de Compostela, 30 de Junio de 2023 Fdo. Ilia Stepin
Autorización del Director/Tutor de la Tesis Argumentative Conversational Agents for Explainable Artificial Intelligence D. José María Alonso Moral, Profesor Titular de la Universidad de Santiago de Compostela D. Alejandro Catalá, Profesor Ayudante Doctor de la Universidad de Santiago de Compostela INFORMAN: Que la presente tesis, se corresponde con el trabajo realizado por D. Ilia Stepin, bajo nuestra dirección/tutorización, y autorizamos su presentación, considerando que reúne los requisitos exigidos en el Reglamento de Estudios de Doctorado de la USC, y que como directores/tutores de esta no incurre en las causas de abstención establecidas en la Ley 40/2015. De acuerdo con lo indicado en el Reglamento de Estudios de Doctorado, declaramos también que la presente tesis doctoral es idónea para ser defendida en base a la modalidad de COMPENDIO DE PUBLICACIONES, en las que la participación del doctorando/a fue decisiva para su elaboración y las publicaciones se ajustan al Plan de Investigación. La utilización de estos artículos en esta memoria, está en conocimiento de que ninguno de los trabajos aquí reunidos podrá ser presentado en ninguna otra tesis doctoral. En Santiago de Compostela, 30 de Junio de 2023 Fdo. José María Alonso Moral Director/a tesis Fdo. Alejandro Catalá Director/a tesis
ACKNOWLEDGEMENTS First and foremost, I am extremely grateful to my supervisors, Dr. José María Alonso Moral and Dr. Alejandro Catalá for their guidance and support throughout my entire doctoral study. I would like to extend my gratitude to Dr. Martín Pereira-Fariña for co-supervising my doctoral project for about two years. The completion of this thesis would have been impossible without their helpful advice over the years. I would like to extend my sincere thanks to all my fellow (ex-)colleagues from the Singular Research Center for Information Technology (Centro Singular de Investigación en Tecnoloxías Intelixentes) of the University of Santiago de Compostela whose company made me feel greatly integrated to the workplace and far beyond it. I also gratefully acknowledge the assistance of the staff of the University of Santiago de Compostela whose guidance was particularly helpful when resolving various administrative issues that emerged along the way. Special thanks to the funding agencies that provided financial support for my doctoral project. Namely, I would like to express my gratitude to the Spanish State Research Agency (Agencia Estatal de Investigación) which funded, in its entirety, my predoctoral contract (ref.: FPI2019- 090153) as well as three nationwide-level projects (ADHERE - U: RTI2018-099646-B-I00; XAI4SOC: PID2021-123152OB-C21; DeepR3.gal: TED2021-130295B-C33) where I participated as a research assistant. In addition, I would like to extend thanks to the Department of Education and University Management (Consellería de Educación e Ordenación Universitaria) and the European Regional Development Fund (ERDF), which funded the project of excellence “Explica-IA” (ED431F 2018/02) that I had great pleasure of co-working on. Importantly, the present piece of research is, in part, a product of active international collaboration. I would thus like to recognise the effort of members of the Laboratory of the New Ethos (Warsaw University of Technology, Poland) who played an important role in shaping up the methodological basis of the argumentative framework proposed in this thesis. I would particularly like to thank Prof. Dr. Katarzyna Budzynska, my secondment supervisor, and Dr. Marcin Koszowy for their active cooperation on the final part of my doctoral project as well as numerous outdoor activities.
ILIA STEPIN afirma que as explicacións deben satisfacer unha serie de propiedades para que sexan eficaces. En primeiro lugar, as explicacións efectivas deben ser contrastivas, é dicir, non só explican por que a decisión ou predición automatizada dada é o caso, senón que tamén dan razóns polas que non son aplicables os resultados alternativos. Ademais, deben ser selectivas, é dicir, deben incluír só un número suficientemente pequeno das causas ou factores máis relevantes que conducen á decisión ou predición dada. Por último, pero non menos importante, as explicacións considéranse sociais, é dicir, son produto da interacción entre o explicador (o axente que explica o fenómeno dado) e o explicado (o destinatario da explicación). Como resultado desta hipótese, as explicacións automatizadas terán a máxima utilidade se se modelan de acordo cos requisitos descritos anteriormente. De acordo coa normativa vixente da Escola de Doutoramento Internacional da Universidade de Santiago de Compostela, a presente tese preséntase en forma de compendio de publicacións. En xeral, a tese divídese en nove capítulos. O Capítulo 1 introduce o tema da tese. No Capítulo 2 exponse a hipótese probada e os obxectivos xerais e específicos da tese. No Capítulo 3 descríbese a metodoloxía aplicada para acadar os obxectivos da tese. O capítulo 4 ofrece unha discusión xeral sobre as propiedades de explicación modeladas e as ferramentas de xeración de explicacións desenvolvidas nesta tese. O capítulo 5 recolle as principais contribucións que xurdiron neste proxecto. Os capítulos 6-8 están relacionados coa metodoloxía e as conclusións recollidas nos artigos de revistas que forman o núcleo desta tese. O capítulo 9 extrae as principais conclusións da presente tese. A continuación, resumimos capítulos 6-8 con máis detalle. Nesta tese, levouse a cabo o deseño, implementación e validación dun novo marco de xeración de explicacións que cumpre todos os requisitos xerais de explicación mencionados anteriormente (é dicir, as explicacións xeradas son contrastivas, selectivas e sociais). Defendemos o uso de modelos interpretables baseados en regras que se poidan utilizar (e cuxas predicións poidan ser explicadas posteriormente) de forma independente ou como representates para explicar as predicións de algoritmos de “caixa negra”. De calquera xeito, o marco proposto serve para explicar o resultado dun clasificador interpretable (por exemplo, unha árbore de decisións ou un sistema de clasificación difuso baseado en regras). No contexto de XAI, a predición dun clasificador pode explicarse (non necesariamente de forma contrastiva) en termos dos trazos máis característicos da instancia que conduciron á predición dada. En diante, referirémonos a tales explicacións como factuais. Para que as explicacións resultantes sexan contrastivas, buscamos modelar explicacións complementarias ás factuais, é dicir, que opoñen explícitamente o resultado da clasificación realmente predito a resultados alternativos hipotéticos. Noutras palabras, a predición do clasificador dado explícase non só en función das características que son máis relevantes para a predición, senón tamén en termos de clasificacións non preditas. Ademais, tales explicacións poden suxerir cambios mínimos nos valores das características para que o resultado previsto cambie da forma desexada. En XAI, 2
Resumo estas explicacións tamén se denominan comúnmente como contrafactuais (CF). Nótese que contrafactuais refírense a exemplos xa observados no pasado, mentres que transfactuais refírense a exemplos sintéticos aínda non vistos pero que se espera que se observen no futuro. Non obstante, para simplificar a notación, remitirémonos ás explicacións CF no resto desta tese, sen importar se os cambios suxeridos involucran a creación de exemplos sintéticos. Para fins ilustrativos, consideremos un escenario bancario común. Unha cliente dun banco, unha muller de mediana idade cuxos ingresos son 4.000 euros mensuais, solicita un préstamo a un ano de 30.000 euros. Adestrado para prever se as solicitudes de préstamo deben ser aceptadas ou rexeitadas, o sistema de clasificación bancaria suxire que o funcionario bancario debe rexeitar a solicitude da cliente. Para proporcionarlle ao cliente a recomendación máis relevante sobre como se podería modificar a decisión, o sistema presenta a seguinte suxestión de CF: “A solicitude de préstamo da cliente sería aprobada se os seus ingresos mensuais fosen polo menos 5.000 euros e tivese polo menos un préstamo activo menos”. Como se desprende do exemplo anterior, as explicacións de CF no contexto de problemas de clasificación son inherentemente contrastivas, xa que se opoñen de xeito explícito a diferentes resultados de clasificación. As explicacións contrastivas e, máis específicamente, CF son estudadas dende hai moito tempo nunha ampla gama de ciencias. Por exemplo, dise que forman parte integrante do razoamento humano. Arguméntase tamén que os contrafactuais representan o nivel máis alto de causalidade. Ademais, pódense xerar para calquera clasificador. Por estes motivos, as explicacións CF chamaron a atención de numerosos investigadores da XAI nos últimos anos. Convencionalmente, os contrafactuais considéranse explicacións agnósticas do modelo, post-hoc e locais. Son locais porque explican o comportamento do sistema a partir das súas predicións individuais. Sábese que os contrafactuais explican as predicións de forma post-hoc, xa que se xeran despois de obter a saída do sistema. Notablemente, esta familia de explicacións é coñecida, en xeral, por ser independente do modelo, xa que os correspondentes métodos de xeración de explicacións están deseñados para operar só na entrada dada e na saída prevista do sistema sen acceder necesariamente aos elementos internos do sistema. Non obstante, a diversidade dos métodos de xeración de explicacións contrastivas e CF recentemente emerxentes mostra que non se limitan necesariamente a esta definición convencional. No Capítulo 6, revisamos as teorías existentes sobre a explicación contrastiva e CF dunha ampla gama de ciencias. Ademais, analizamos os marcos computacionais de última xeración deseñados para a xeración dos dous tipos de explicación mencionados anteriormente. Ademais, inspeccionamos o grao de sinerxía entre os enfoques teóricos da explicación contrastiva e CF e as súas contrapartes computacionais de última xeración. Cabe destacar que as explicacións CF posúen unha serie de propiedades importantes que se poden utilizar como medidas de utilidade da explicación (validez, proximidade, accionabilidade, diversidade, por citar algunhas). Nesta tese, centrámonos na modelización de CFs que 3
ILIA STEPIN se presume que son suficientes co seguinte subconxunto de propiedades. En primeiro lugar, os CF deben ser válidos, é dicir, deben levar a predicións correctas correspondentes ao resultado alternativo desexado. En segundo lugar, espérase que unha explicación CF inclúa só un conxunto de cambios mínimos nos pares característica-valor da instancia para que cambie a clasificación prevista. De feito, o axente explicado está interesado en recibir a explicación CF máis relevante para a instancia que se está a considerar. Tendo en conta o exemplo bancario anterior, mentres que un CF que indica que os ingresos mensuais deberían ser de 6.000 euros ou máis seguirán sendo válidos neste escenario, o usuario final está interesado en manter este valor o máis próximo posible aos seus ingresos reais, polo que a explicación CF sería preferible indicar que os ingresos mensuais deberían ser de 5.500 euros (sempre que os dous CF sexan válidos). En terceiro lugar, espérase que unha explicación CF sexa accionable, é dicir, só as características que se poidan modificar de xeito viable forman parte da explicación. De feito, se o sistema suxire que se reduza a idade do cliente, o CF correspondente é inútil, aínda que todos os demais cambios se poidan facer con éxito. Por último, espérase que os contrafactuais sexan diversos, é dicir, que abrangan varios CF válidos distintos (un punto único ou agrupados en conxuntos) que teñan un poder explicativo equivalente e, idealmente, que se atopen dispersos nas rexións de datos que cobren. De feito, o usuario final pode atopar explicacións máis satisfactorias que conteñan non só valores específicos das características CF dadas, senón intervalos de tales valores. A presentación de CFs tan diversas pode aumentar a flexibilidade do comportamento do usuario, xa que o destinatario da explicación ten a posibilidade a escoller o escenario a seguir que mellor lle conveña. A diversidade de explicacións de CF baseadas en regras pode manifestarse de varias maneiras. Por unha banda, as características de explicación pódense representar numericamente, en forma de intervalos (por exemplo, “5.500 ≤renda ≤ 6.000”). Por outra banda, pódense ofrecer no seu lugar as correspondentes descricións textuais (por exemplo, “a renda é alta”). En ambos casos, os CF xerados automaticamente inclúen un conxunto de datos CF que permiten ao usuario final escoller o valor alternativo máis axeitado para as funcións dadas entre o intervalo de valores suxerido. Non obstante, non está claro se tales etiquetas lingüísticas (é dicir, “alta” do exemplo anterior) fan as características explicativas correspondentes máis comprensibles ou fáciles de usar e, polo tanto, a explicación xeral máis efectiva. Para facer as explicacións selectivas, confiamos no uso de elementos internos de modelos baseados en regras. Algúns destes algoritmos de clasificación interpretables por deseño agregan información sobre as características que son máis relevantes para a predición dada. Por exemplo, as árbores de decisión conteñen os valores de características máis relevantes no camiño de decisión. Polo tanto, a explicación fáctual pódese reconstruír resumindo a información agregada no camiño desde a raíz ata o nodo folla previsto. Non obstante, a xeración de explicacións CF efectivas para familias específicas de algoritmos interpretables, como árbores de decisión ou 4
Resumo sistemas de clasificación baseados en regras difusas, seguen sendo pouco estudadas. Ademais, estes clasificadores baseados en lóxica difusa ofrecen ferramentas que, por deseño, permiten aos desenvolvedores constituír explicacións textuais equivalentes utilizando o repertorio de termos lingüísticos. Así, potenciar árbores de decisión (difusas) e clasificadores baseados en regras difusas con novos métodos de xeración de explicacións CF permítenos mellorar aínda máis o seu potencial explicativo. No Capítulo 7, propoñemos tres algoritmos de xeración de explicacións CF que producen explicacións textuais factuais e CF para clasificadores baseados en regras preseleccionadas. En primeiro lugar, deseñamos un algoritmo (en diante, denomínase XOR) que xera explicacións CF ordenando as representacións vectorizadas das regras CF de acordo co seu grao de relevancia para a instancia de proba. Supoñemos que tales explicacións baseadas en regras levan á xeración de explicacións CF válidas. Posteriormente, introducimos unha variante alternativa do mesmo algoritmo (en diante, denomínase EUC) que relaciona os vectores de funcións de pertenza difusa coas regras CF mediante a medición da distancia euclidiana entre todos os pares de tales vectores. Para ambos algoritmos, propoñemos ademais o mecanismo de aproximación lingüística, é dicir, un método para asociar intervalos de características numéricas a termos lingüísticos. Esta extensión permítenos xerar diversas explicacións automáticas equivalentes en formato numérico ou puramente textual. A pesar de que os dous algoritmos introducidos anteriormente ofrecen explicacións facilmente interpretables, son específicos do modelo, é dicir, requiren acceso aos elementos internos do modelo e non se poden aplicar directamente a calquera clasificador. Non obstante, os algoritmos de xeración de explicacións CF independentes do modelo son capaces de explicar universalmente a saída de calquera clasificador tanto de xeito factual como contrafactual. Así, tamén propoñemos un algoritmo de xeración de explicacións CF xenética independente do modelo (en diante, denomínase GEN) que produce explicacións CF optimizando a poboación inicial de forma iterativa ata que se identifique o único punto de datos máis próximo á instancia de proba. En conxunto, ambos grupos de algoritmos de xeración de explicacións CF (é dicir, aqueles específicos do modelo e os agnósticos do modelo) poden usarse de forma complementaria entre si, especialmente se as explicacións resultantes veñen en diferentes formatos. Co fin de comparar a utilidade das explicacións específicas do modelo baseadas en coñecementos imprecisos fronte a outras que apuntan a puntos de datos específicos, realizamos dous estudos de avaliación humana (é dicir, Survey GM eSurvey TS) onde comparamos a eficacia das explicacións resultantes para as instancias de proba preseleccionadas. En ambos estudos, adestramos sistemas de inferencia difusa de Mamdani que fan predicións utilizando o algoritmo FURIA e xeran as explicacións correspondentes utilizando todos os algoritmos propostos (é dicir, XOR, EUC e GEN). Adestramos aos clasificadores nun conxunto de datos de clasificación de tipos de cervexa para xerar posteriormente explicacións lingüísticas para os estímulos 5
ILIA STEPIN da enquisa preditos correctamente. Cabe sinalar que todas as explicacións xeradas supoñense accionables debido á estrutura do conxunto de datos utilizado nos experimentos: en calquera caso, sempre é posible modificar os valores das características dentro dos intervalos de valores suxeridos. Survey GM está deseñado para que o usuario poida avaliar a calidade da explicación automatizada en base ás catro máximas de Grice (cantidade, calidade, relevancia e forma), que transformamos en cinco aspectos explicativos (informatividade, fiabilidade, precisión, relevancia e lexibilidade) para que os participantes do estudo comprendan máis facilmente a súa tarefa de avaliación. Durante o estudo, os participantes avaliaron cada aspecto das tres explicacións (unha por cada método de explicación proposto) empregando unha escala Likert de 7 puntos. Pola súa banda, Survey TS é unha variante simplificada de 5 puntos baseada na escala Likert empregada en Survey GM onde se avalía unha única explicación en termos de fiabilidade e satisfacción xeral. Ambas enquisas leváronse a cabo para un público obxectivo que tiña suficiente coñecemento do dominio e un alto nivel de experiencia. A investigación e os protocolos experimentais foron aprobados polo comité ético da Universidade de Santiago de Compostela. Mentres que se propuxeron un gran número de métricas computables automaticamente para estimar a calidade das explicacións automatizadas, estas adoitan servir para avaliar a calidade desde o punto de vista algorítmico (por exemplo, a distancia xeométrica á instancia de proba). Non obstante, as métricas automáticas que estiman aspectos da percepción do usuario seguen sendo escasas. Para abordar este problema, propoñemos a métrica da complexidade da explicación percibida, é dicir, unha estimación do complexa que parece ser unha explicación desde o punto de vista do usuario ao lela. A nova métrica proposta está inspirada no Gunning Fog Index (un indicador da facilidade de entender o texto por parte do público destinatario). En particular, baséase en dous factores que están presentes nas explicacións textuais baseadas en regras xeradas mediante os métodos de xeración de explicacións propostos: a lonxitude da explicación e a relación agregada do número de termos lingüísticos de todas as características utilizadas na explicación. Os resultados dos estudos de avaliación humana realizados mostran que a complexidade da explicación percibida ten unha correlación positiva moderada coa informatividade estimada polo usuario e unha forte correlación negativa coa relevancia e a lexibilidade estimadas polo usuario, mentres que as puntuacións métricas propostas non se correlacionan coa fiabilidade ou a precisión. Polo tanto, pódese concluír que a métrica proposta comprende tres dos aspectos de explicación mencionados anteriormente para usuarios que teñan coñecementos e experiencia suficientes no dominio. Ademais, usar a puntuación de complexidade da explicación percibida pode ser útil para diminuír os custos de avaliación humana (para o público obxectivo), xa que se pode calcular para substituír (en parte) os estudos de usuarios correspondentes. Compre sinalar que os algoritmos propostos no Capítulo 7 limítanse á xeración de explicacións contrastivas selectivas non sociais, é dicir, carecen de calquera interacción directa co 6
Resumo usuario final e só representan os datos CF máis relevantes desde o punto de vista algorítmico. Nestas configuracións, o usuario final ten que tomar unha decisión sobre a fiabilidade das explicacións automatizadas ofrecidas sobre a base dunha única información. Se a explicación non se considera o suficientemente fiable ou satisfactoria, o usuario final pode querer descartala aínda que sexa válida. Polo tanto, é indispensable que o usuario final explore o espazo de explicación se o considera necesario. Para facer sociais as explicacións resultantes, modelamos a interacción entre o sistema e o usuario en forma de diálogo explicativo argumentativo onde o usuario é capaz de discutir sobre pezas de explicación específicas e, deste xeito, explorar o espazo de explicación ata que poida tomar unha decisión informada sobre a predición do sistema. Para iso, ampliamos os nosos algoritmos de xeración de explicacións cun módulo de xeración de diálogos explicativos. Deseñado como un axente conversacional, o modelo de diálogo resultante garante unha comunicación dialóxica interactiva entre o sistema e o usuario onde este está habilitado para solicitar e procesar as explicacións necesarias ofrecidas de forma comprensible e humana. Ademais, esta extensión mellora o marco de xeración de explicacións proposto coa opción de formar explicacións interactivas dinámicas en contraste coas xenéricas estáticas. No Capítulo 8 propoñemos o denominado “xogo de diálogo explicativo”, un modelo formal de diálogo explicativo baseado no enfoque correspondente á modelización do diálogo a partir da teoría da argumentación. O modelo de diálogo está deseñado de forma descendente, é dicir, baséase nun protocolo de diálogo predefinido que contén catro posibles solicitudes de usuarios (as de explicación factual ou CF, detalle, aclaración e explicación alternativa), con varias posibles respostas do sistema asociadas a distintas solicitudes do usuario. Ademais, xeneralizamos o modelo de diálogo explicativo na forma dunha gramática de diálogo sen contexto para facelo universalmente aplicable á saída de calquera sistema de clasificación baseado en regras mellorado cun explicador que é capaz de producir explicacións textuais baseadas en regras. Validamos o modelo de diálogo resultante realizando un estudo de avaliación humana mediante tres casos de uso: clasificación da posición do xogador de baloncesto, clasificación do tipo da cervexa e clasificación da enfermidade da tiroide. Nestes escenarios de diálogo de busca de información, unha das partes do diálogo (neste caso, o usuario) é inicialmente informada sobre os datos que se están procesando e despois pretende dar sentido á información dada pola outra parte (neste caso, a predición do sistema). Nos tres casos de uso, os diálogos pretenden explicar a predición dun único sistema para unha instancia de datos previamente seleccionada e clasificada correctamente. Como neste experimento se aborda o aspecto comunicativo da xeración de explicacións, utilízanse árbores de decisión nítidas como clasificadores, xunto co método XOR empregado para xerar as explicacións correspondentes, co fin de garantir a transparencia dos resultados experimentais. Para analizar as transcricións de diálogos recollidas, aplicamos técnicas de minería de pro- 7
ILIA STEPIN cesos que tratan as instancias de diálogo explicativo como fíos de proceso. En particular, realizamos a denominada comprobación de conformidade para relacionar o protocolo de diálogo e o corpus de diálogos explicativos realmente rexistrados. Concluimos que os participantes no estudo fan un uso activo de todos os tipos de solicitudes ofertadas. Polo tanto, o procedemento de verificación da conformidade confirma a utilidade do modelo de diálogo proposto inicialmente na súa totalidade. Ademais, rexistramos un gran número de solicitudes de explicacións de CF alternativas (segunda e terceira mellor clasificadas polo sistema) para case todas as clases de CF en todos os casos de uso. Esta observación apunta ademais á necesidade de presentar diversos CFs múltiples en diálogos de busca de información. Cabe destacar que todos os traballos de revistas e software que constitúen a base da presente tese están a disposición do público. Finalmente, realizamos observacións finais e esbozamos direccións para o traballo futuro no Capítulo 9. 8
Summary Artificial Intelligence (AI) plays an increasingly important role in a large number of daily life activities. AI applications are found in numerous products, from banking to health care to manufacturing to education. However, the fast progress of the present-day AI raises concerns related to its interpretability and explainability. On the one hand, AI models are rapidly becoming overly complex for a general audience to understand the nature of the decision-making systems that they make part of. This may undermine trust in automated decisions produced by such systems and increases reluctance to use them. On the other hand, shifting from hard-coded rulebased to data-driven machine learning (ML)-based AI algorithms has resulted in the nature of such algorithms being concealed even from their developers. In order to demystify such “black-box” algorithms to both lay users and domain experts, researchers from numerous fields of science called for making the present-day AI explainable. This resulted in countless research projects forming the basis of the recently emerged eXplainable AI (XAI) community. In line with scientific aspirations, the ubiquitous use of AI has led to major changes in legal regulation, which are reflected in, for example, the European Union’s (EU) General Data Protection Regulation (GDPR) or the recently proposed Artificial Intelligence Act (AIA) which was voted for by the EU Parliament in June 2023, being in the final stage before becoming law and coming into force in each member state. Poor explanatory capacities of “black-box” AI models have motivated discussions on a favourable use of so-called interpretable models, instead. Hereinafter, the concept of interpretable models refers to the family of algorithms that grant access to their human-comprehensive internals. Indeed, more interpretable but (possibly) less accurate ML models may appear to be more efficiently applicable to solving various challenging problems than more robust but less transparent algorithms, especially in cases of high-stakes decisions. Nevertheless, such explanation generation-related sub-tasks as, for example, evaluation and communication remain being demanding tasks even for interpretable models. A large body of interdisciplinary research on the nature of explanation claims that explanations should satisfy a number of properties for them to be effective. First, effective explanations are claimed to be contrastive, i.e. they do not only explain why the given automated decision or prediction is the case but also give reasons why alternative outcomes are not applicable. In
ILIA STEPIN addition, explanations should be selected, i.e. they should include only an adequately small number of the most relevant causes or factors that lead to the given decision or prediction. Last but not least, explanations are deemed social, i.e. they are a product of interaction between the explainer (the agent that explains the given phenomenon) and the explainee (the recipient of the explanation). As a result, automated explanations are hypothesised to have maximal utility if modelled in accordance with the requirements outlined above. In accordance with the actual regulations of the international doctoral school of the University of Santiago de Compostela, this thesis is presented in form of a compendium of publications. Overall, it is divided into nine chapters. Chapter 1 introduces the topic of the thesis. Chapter 2 states the hypothesis tested and the general and specific objectives of the thesis. Chapter 3 describes the methodology applied to reach the thesis objectives. Chapter 4 provides the reader with a general discussion on the explanation properties modelled and the explanation generation tools developed in this thesis. Chapter 5 lists main contributions that emerged as part of the doctoral project. Chapters 6-8 relate to the methodology and findings reported in the journal papers that form the core of this thesis. Chapter 9 draws main conclusions from the present thesis. Let us now summarise Chapters 6-8 in more detail. We design, implement, and validate a novel explanation generation framework whose output explanations are claimed to meet all the aforementioned general requirements to explanation (i.e. being contrastive, selected, and social). We advocate the use of interpretable rule-based models that can be used (and whose predictions can be subsequently explained) independently or as proxies to explain predictions of “black-box” algorithms. Either way, the proposed framework serves the purpose of explaining the outcome of a given rule-based interpretable classifier (e.g., a decision tree or a fuzzy rule-based classification system). In the context of XAI, a classifier’s prediction can (not necessarily contrastively) be explained in terms of the most characteristic features of the test instance that led to the given prediction. Hereinafter, we refer to such explanations as factual. To make the output explanations contrastive, we pay particular attention to modelling explanations that are complementary to factual ones, i.e. they explicitly oppose the actually predicted classification outcome to hypothetical alternative outcomes. In other words, the given classifier’s prediction is explained not only in terms of the features that are the most relevant to the prediction but also in terms of non-predicted classifications. Further, such explanations can suggest minimal changes in feature values so that the predicted outcome changes in a desired way. In XAI, these are also commonly referred to as the so-called counterfactual (CF) explanations. Notice that, counterfactuals refer to examples already observed in the past while transfactuals refer to synthetic examples not seen yet but expected to be observed in the future. Anyway, for simplicity of notation, we will refer to CF explanations in the rest of this thesis, no matter if suggested changes involve the creation of synthetic examples. 10
Summary For illustrative purposes, let us consider a common banking scenario. A client of a bank, a middle-aged woman whose income equals €4.000 per month, is applying for a one-year loan of €30.000. Trained to predict whether loan applications should be accepted or rejected, the banking classification system suggests that the bank officer should decline the client’s request. To provide the client with the most relevant recommendation of how the decision can be changed, the system outputs the following CF suggestion: “The client’s loan application would be approved if her monthly income were at least €5.000 and if she had at least one active loan less.” As follows from the example above, CF explanations in the context of classification problems are inherently contrastive, as they explicitly oppose different classification outcomes. Contrastive and, more narrowly, CF explanations have long been studied in a wide range of sciences. For instance, they are claimed to make an integrative part of human reasoning. Further, counterfactuals are argued to represent the topmost level of causation. In addition, they can be generated for any classifier under consideration. For these reasons, they have attracted attention of numerous XAI researchers in recent years. Conventionally, counterfactuals are considered local post-hoc model-agnostic explanations. They are local because they explain the system’s behaviour on the basis of its individual predictions. Counterfactuals are known to explain predictions in a post-hoc manner, as they are generated after the system’s output has been obtained. Remarkably, this family of explanations is, in general, known to be model-agnostic, since the corresponding explanation generation methods are designed to operate only on the given input and predicted output of the system without necessarily accessing the system’s internals. However, the diversity of the newly emerging contrastive and CF explanation generation methods shows that they are not necessarily limited to this conventional definition. In Chapter 6, we review existing theories of contrastive and CF explanation from a wide range of sciences. Further, we analyse the state-of-the-art computational frameworks designed for generation of the two aforementioned kinds of explanation. In addition, we therein inspect the degree of synergy between the theoretical approaches to contrastive and CF explanation and their state-of-the-art computational counterparts. Noteworthy, CF explanations possess a number of important properties that can be used as measures of explanation utility (validity, proximity, actionability, diversity, to name a few). In this thesis, we focus on modelling CFs that are hypothesised to suffice the following subset of such properties. First, CFs must be valid, i.e. they must lead to correct predictions corresponding to the desired alternative outcome. Second, a CF explanation is expected to include only a set of minimal changes to the test instance feature-value pairs for the predicted classification to change. Indeed, the explainee is interested in receiving the piece of CF explanation that is the most relevant to the test instance under consideration. Considering the banking example above, whereas the CF stating that the monthly income should be €6.000 or more will still be valid in this scenario, the end user is interested in keeping this value as close as possible to her actual income, 11
ILIA STEPIN In this regard, the property of contrastiveness is at the core of the so-called counterfactual explanations (or counterfactuals, or CFs, for short), a sub-group of contrastive explanations that suggest minimal changes to the input feature values so that the output changes in the desired way [33]. Contrastive by nature, CFs are found to be inherent to human reasoning [4] and can therefore greatly facilitate explanation processing by end users [5]. For these reasons, explaining predictions counterfactually has become among key explainability issues, especially when explaining “black-box” models [21]. Given an evident lack of transparency in the reasoning of many complex AI models (e.g. neural networks), the use of possibly less accurate or robust but more interpretable models has been actively argued for [32]. In light of this, we explore the potential of rule-based classification systems to provide their end users with automated explanations for their predictions. Indeed, the potential of interpretable models for explanations is left largely underexplored [19]. Enhancing rule-based classification systems with effective methods of CF explanation generation allows them to become self-explanatory while providing their end users with contrastive selected explanations. Further, such self-explanatory rule-based interpretable classifiers can then be used as part of more complex explainers to address the issue of explanability of “black-box” models. In order to preserve the state-of-the-art levels of performance while gaining explainability, “black-box” models can be enhanced with explanation generation modules that make use of (possibly, surrogate) interpretable models [9]. Such interpretable models (e.g., decision trees or fuzzy rule-based classification systems) [2] have shown to effectively explain “black-box” models in a post-hoc manner when, for example, trained on a local neighbourhood around the test instance [41]. In this regard, they can serve as a proxy to approximate given single “blackbox”-based predictions. As the need for explaining decisions made by AI-based systems is recognised legally, various researchers are urging for making a step forward towards responsible, human-centric AI [6]. Whereas several automatic metrics have been designed to estimate the quality of automated CF explanations with respect to their computational aspects [26], human evaluation remains among the key challenges for truly effective CF explanation generation [43]. Indeed, only a limited number of state-of-the-art CF explanation algorithms have undergone assessment by potential beneficiaries of such explanations [16]. Human evaluation of automated explanations is closely connected with the social aspect of explanation. It is often addressed in XAI by means of engaging the end user in explanatory dialogue with the system [42]. Further, insights from humanities and social sciences (e.g., argumentation) allow us to propose explanatory dialogue models that rely on a consolidated body of knowledge about human reasoning and connect it to that of an AI-based agent. In fact, argumentation makes an integrative part of certain explanation theories and therefore appears to be a suitable methodological fit to bridge the gap between the explainer (the explanation 18
Chapter 1. Introduction generation module) and the explainee (the end user). Despite specific methodological differences, argumentation and explanation are found to greatly completement each other [3]. For example, some theories of explanation conceptualise explanations as arguments [12]. Whereas argumentation theories provide a diverse repertoire of frameworks that is capable of generating explanations for automatic predictions in a wide range of tasks [46], we aim to explore its potential as a communication channel between the end user and the system to enhance the previously designed framework for contrastive-counterfactual selected explanations for interpretable rule-based classification systems with a social dimension. 1.2 THESIS STRUCTURE The present thesis contains nine chapters. The remainder of the thesis is structured as follows. Chapter 2 states the hypothesis tested in this thesis as well as the general and specific objectives. As we list the objectives of the thesis, we refer the reader to the publications where the objectives were reached. Chapter 3 describes the general methodology applied throughout the thesis and describes specific tools that were used in order to reach the thesis objectives. Chapter 4 provides the reader with a general discussion on explanation properties in the context of XAI and analyses in detail the strengths and weaknesses of their modelling in this thesis. Chapter 5 lists the contributions of this thesis, i.e., the software developed to reach the thesis objectives and all the publications that emerged during the doctoral project. Chapter 6 provides the reader with the background information on contrastive and CF explanations. In addition, it examines theoretical foundations thereof, the related state-of-the-art computational frameworks, and inspects the degree of synergy between the former and the latter. Chapter 7 introduces three algorithms for CF explanation generation (namely, XOR, EUC, and GEN) used to explain predictions of an FRBCS. In addition to discussing technicalities of the aforementioned algorithms, it evaluates the algorithms via two human evaluation studies. Further, it proposes a novel metric of perceived explanation complexity (PEC) that aims to facilitate evaluation of automatically generated explanations. Chapter 8 proposes an argumentative framework for communication of automatically generated rule-based explanations. In particular, it formalises explanatory dialogue in form of the so-called “dialogue game” and describes in detail the corresponding dialogue protocol. Further, it additionally represents the protocol in form of context-free dialogue grammar to make the protocol universally applicable to other explainer-classifier pairs that are capable of generating textual rule-based explanations. Last but not least, it reports the results of a human evaluation experiment that serves the purpose of validation of the proposed explanatory dialogue model. 19
ILIA STEPIN Finally, Chapter 9 presents main conclusions derived from the results of the doctoral project and outlines prospective directions for future work that are relevant to the problems of CF explanation generation, communication, and evaluation. 20
2 Hypothesis and objectives In this thesis, we develop an explanation generation framework for interpretable (i.e., “whitebox”) classifiers (e.g., decision trees) and semi-interpretable (i.e., “grey-box”) rule-based classification systems (e.g., fuzzy inference systems). Despite the fact that such models provide predictions that can be easily interpreted factually, their CF potential remains understudied. We formulate the main hypothesis tested in the present thesis as follows: “By modelling explanations satisfying the properties specifically relevant for XAI and enhancing them with dialogic interactive facilities, we can convey both factual and CF explanations that are appealing for a good number of users in different application domains”. The general objective of the present doctoral thesis is to advance state-of-the-art XAI technologies for (1) automatic generation of factual and CF explanations for interpretable rule-based classifiers and (2) effective and comprehensive communication of such explanations. The implemented explanation generation framework is expected to output explanations that satisfy the aforementioned requirements to effective explanations (i.e. being contrastive, selected, and social). More precisely, the following specific objectives are considered to achieve the overall goal: O1. Design, implement, and validate a framework for factual and CF explanation generation applied to given pretrained (semi-)interpretable rule-based classifiers. This objective has been successfully reached in the following publications: • Ilia Stepin, Jose M. Alonso, Alejandro Catala, and Martín Pereira-Fariña. “A Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable Artificial Intelligence”. IEEE Access, vol. 9, pp. 11974–12001, 2021. DOI: 10.1109/ACCESS.2021.3051315; • Ilia Stepin, Jose M. Alonso-Moral, Alejandro Catala, Martín Pereira-Fariña. “An empirical study on how humans appreciate automated counterfactual explanations which embrace imprecise information”. Information Sciences, vol. 618, pp. 379– 399, 2022. DOI: 10.1016/j.ins.2022.10.098. O2. Design, implement, and validate a conversational agent endowed with credibility via natural language processing and argumentation technologies in the context of XAI in order
ILIA STEPIN to communicate and customise automatically generated explanations. This objective has been successfully reached in the following publication: • Ilia Stepin, Katarzyna Budzynska, Alejandro Catala, Martín Pereira-Fariña, Jose M. Alonso-Moral. “Information-seeking dialogue for explainable artificial intelligence: Modelling and analytics”. Argument and Computation, in press, 2023. DOI: 10.3233/AAC-220011. O3. Develop a human evaluation framework for the purpose of validation of the algorithms designed to achieve O1 and O2. This evaluation framework serves the purpose of estimating various aspects of automatically generated explanations and that of assessing the quality of the process of communication of such explanations, respectively. This objective has been successfully reached in the publications addressing both O1 and O2. 22
3 Methodology The methodology employed in this thesis bases on the iterative development approach. In order to achieve the objectives listed in Section 2, the work on each of them presupposes the following consecutive steps: 1. Requirement specification: defining the research objective while considering possible limitations of the corresponding state-of-the-art AI techniques and software. Special attention is paid to aspects of automatic explanation generation and communication (i.e., properties of explanation, requirements to the communicative aspects of the language used for modelling argumentative explanatory dialogue, use cases, target audience, etc.); 2. Literature review: a bibliographic study that serves to identify the state-of-the-art conceptual, theoretical, and computational frameworks designed to address the goal-specific problems. This study includes an analysis of the advantages and disadvantages of the identified methods; 3. Implementation: design and development of conceptual models and algorithms aimed at producing advances with respect to the state-of-the-art. Theoretical contributions are followed by software implementations to be validated empirically at the next step; 4. Validation: the process of verification of the software as well as revision of the algorithm or model implemented if the experimental results obtained are not satisfactory; 5. Integration: once validated, the new algorithms or prototypes are integrated with those validated previously so that they make part of the unified framework. In order to achieve objective O1, we first perform a systematic literature review (SLR) of the state-of-the-art methods of contrastive and CF explanation generation in the context of XAI following the guidelines for performing SLRs in software engineering [17, 18]. In particular, we formulate a series of research questions to be addressed, design a search strategy, and extract and synthesise data in accordance with the predefined inclusion and exclusion criteria. In addition, we perform the so-called “snowballing” procedure, i.e. a revision of the bibliography lists of the previously collected studies, following the corresponding guidelines [45]. Subsequently, based
ILIA STEPIN on the insights from our SLR, we develop a conceptual framework for factual and CF explanation generation for interpretable rule-based classifiers that outputs contrastive selected explanations. The framework includes one model-agnostic and two model-specific explanation generation algorithms. To generate automatic explanations in natural language, we adapt one of the most commonly used natural language generation (NLG) pipelines [7, 31]. We opted for using a template-based NLG approach instead of an end-to-end neural NLG approach to maximise the fidelity to the data of the generated narratives versus their naturalness. Accordingly, the use of pre-trained large language models for NLG falls outside the scope of this thesis. Consequently, we evaluate the proposed methods in a series of human evaluation experiments contributing to reaching objective O3, as we adhere to the human evaluation guidelines designed specifically for XAI [13]. To accomplish objective O2, we develop an argumentative conversational agent relying on the “dialogue game”-based theoretical approach to argumentative dialogue modelling [30]. More specifically, we design a set of original requests and responses that constitute a newly proposed explanatory dialogue protocol. Similarly to O1, we evaluate the designed argumentative framework by performing a human evaluation study and analyse the collected dialogue transcripts. In addition, we employ concepts from process mining to perform the so-called “conformance checking” treating instances of the collected explanatory dialogues as processes [27]. To reach objective O3, we design a software framework relying on human evaluation guidelines for XAI [13] and implement it as a stand-alone application. Its flexible structure allows us to adapt it to the needs of specific experiments carried out as part of the present thesis. As a result, we use it to evaluate all the CF explanation generation and communication methods proposed in this thesis. In order to see the connection between the research questions posed and the tools developed to answer them (as well as the publications addressing them), we kindly refer the reader to Fig. 5.1 from Chapter 5 which lists the main contributions of this thesis. 24
4 General discussion In this thesis, we explore the potential of interpretable rule-based classification systems to generate effective automated explanations in accordance with recent requirements to the quality of such explanations. Recall that effective explanations are claimed to be contrastive, selected, and social [23]. In this section, we discuss peculiarities of modelling automated explanations ensuring encapsulation of these key properties in the context of the explanation generation framework for interpretable rule-based classifiers developed in this thesis. In particular, we first inspect computational aspects of contrastive and CF explanations for classification problems. Then, we discuss how one-shot selected CF explanations can be generated for predictions made by interpretable rule-based classification systems. Finally, we inspect how a social layer can be added on top of the previously proposed explanation framework. More specifically, we focus on explanatory dialogue modelling as a means of explanation communication. We thus examine (1) the most relevant aspects of formal explanatory dialogue modelling for textual rule-based explanations and (2) a general architecture of the corresponding argumentative conversational agent. As a result, we propose an explanation generation framework whose output embraces all the aforementioned properties of explanations for interpretable rule-based classifiers (see Fig. 4.1). The general discussion on enhancing automated explanations with the aforementioned properties is structured in the following manner. Section 4.1 discusses computational aspects of CF explanations and their connection with the family of contrastive explanations as defined in the literature. Section 4.2 examines aspects of one-shot, selected CF explanation generation for interpretable rule-based classification systems. Section 4.3 explores how insights from argumentation theory can enhance the social aspect of automated explanations in the context of explanatory dialogue systems. 4.1 MAKING EXPLANATIONS CONTRASTIVE(-COUNTERFACTUAL) Explanations are argued to be contrastive, as the facts that they explain are (sometimes, implicitly) opposed to pieces of information related to a contrast case [23]. They can be defined as answers to the contrastive why-question (i.e., “Why Prather than Q?”) where Pis the fact being explained and Qis an alternative non-observed foil. Then, Pis said to be explained
ILIA STEPIN Figure 4.1: Explanation properties that the proposed argumentative conversational agent embraces. contrastively if a causal difference between Pand not-Q is verified (the so-called “difference condition”) [20]. Ensuring that automatically generated explanations are contrastive is claimed to greatly contribute to their readability, as contrastive explanations are said to prune the space of all causal factors aiding finer-grained understanding [15]. Theories of contrastive explanation largely rely on causal accounts of explanation. However, a number of existing explanation generation algorithms offer purely non-causal explanations [25]. Further, the notion of contrastive explanation is found to greatly overlap with that of CF explanation in the XAI community, despite certain methodological differences. Thus, causal accounts of contrastive explanation imply that CF explanations can be deemed contrastive so long as they respect the difference condition [24]. Nevertheless, CF explanations are imposed a number of additional constraints that make them a sub-group of contrastive explanations if causality is not being considered. In the context of classification problems, contrastive and/or CF explanations differentiate the given piece of factual information (e.g., “Your loan application has been rejected because your monthly income is too low”) from some CF information (e.g., “Your loan application would have been approved if you had at least one active loan less”). Altogether, a combination of factual and CF explanations enables the end user to construct a mental representation of the AI-based agent’s reasoning for all possible outcomes. CF explanations are a powerful tool for explaining predictions or decisions made by AI- based agents. In addition to factual explanations, they provide complementary information to 26
Chapter 4. General discussion explain the AI-based agent’s reasoning, so that a combination of factual and CF explanations can explain all possible predictions of the AI-based agent irrespective of whether they took place or not. CF explanations are said to be post-hoc (i.e., they explain predictions of pretrained models) and local (i.e., they are designed to explain individual predictions). In addition, CFs can be model-agnostic (i.e., the corresponding explanation generation method operates only on the input feature values and the output prediction of the classifier) or model-specific (i.e., the explainer also has access to the classifier’s internals). Automated CF explanations are claimed to have a number of properties (see [8, 26] for an exhaustive list thereof): •Validity. A CF explanation is said to be valid iff the corresponding CF leads to the desired CF prediction; •Proximity. A CF is said to be proximate iff the distance between the test instance and the CF data point that the given CF explanation is related to is as small as possible; •Sparsity. A CF is said to be sparse iff it contains the minimal number of features in comparison to other valid CFs; •Diversity. CFs are said to be diverse iff they form a set of valid CF data points that are at the same time maximally different from each other so that the explainee is offered a number of legit suggestions on how to change the given prediction to the alternative desired one. Whereas the quantitative metrics defined to measure the properties listed above are commonly used for evaluating CFs [43], some may be incompatible with others. For example, there exists a trade-off between proximity and diversity: no CF explanation generation method is claimed to maximise both due to their divergent treatment of CFs with respect to the distance from the test instance [26]. In this case, other factors (e.g., the target audience or the application domain) may become crucial to estimate the quality of automated CFs. In addition, researchers distinguish various other properties that automated CF explanations are desired to have. These include actionability (i.e., the mere ability to change the feature values as the given CF explanation suggests), causality (i.e., establishing causal relations between the features with respect to a given causal model and ensuring that such relations are maintained in the given CF), and fairness (i.e., the given CF explanation is unbiased with respect to specific protected features, e.g., gender or race). Indeed, theoretical accounts of contrastive and CF explanations are found to largely differ from each other in terms of their relation to the causal aspect of explanation. A vast majority of theoretical academic studies appeal to the causal nature of contrastive and/or CF explanation [36]. However, modelling CFs that possess these properties falls outside the scope of this thesis and is left for future work. 27
ILIA STEPIN finds to be the most relevant for the given (factual or CF) class may not coincide with user expectations or preferences, leading to decreased utility of such an explanation. It is therefore of paramount importance to enable the end user to inquire alternative explanations so that the user can prune the explanation space until she is fully satisfied with the information accumulated along the explanatory dialogue. To ensure that the desired aforementioned advantages of user-system interaction are achieved fully, argumentation theory methods can be used to model a communication channel between the user and the explainer. In this thesis, we design explanatory user-system dialogue applying the so-called “dialogue game” approach to argumentation [22]. This mechanism allows us to (1) formalise and integrate request types described below, (2) enable the end user to iteratively explore the explanation space by arguing over the explanations offered previously, (3) personalise explanations giving to the end user full freedom to request only necessary and sufficient information about the dataset, prediction, or explanation components. The corresponding dialogue protocol establishes a typology of user’s requests, explainer’s responses, and transitions between the dialogue states. Thus, the set of the proposed dialogue requests includes the following categories: •Factual and CF explanation requests. These include why and why-not questions to the explainer for factual and CF classes, respectively. •Detailisation requests. These tackle the switch from purely linguistic values of specific features that make part of the given piece of explanation to their numerical counterparts. •Clarification requests. These requests are meant to question definitions of specific features that make part of the given piece of explanation. •Alternative explanation requests. These requests are designed to enable the user to explore the explanation space. Notably, they are made unavailable for factual explanations in the case of DTs, as alternative decision paths could be erroneous with respect to the actual prediction and do not adequately explain the classifier’s reasoning for the given test instance. The proposed formal model of explanatory dialogue has been implemented in form of a task-oriented dialogue system1. The dialogue system’s main tasks are to (1) communicate to the end user explanations generated automatically by an explainer and (2) handle follow-up user requests concerning explanation-related details. Adapting a classic pipeline for task-oriented 1The source code is made publicly available at https://gitlab.citius.usc.es/ilia.stepin/ fcfexpgen, branches “dialgame” and “dialgame_nlu”. 34
Chapter 4. General discussion Figure 4.2: The dialogue system pipeline making use of the proposed argumentative explanatory model. spoken dialogue systems [44] to (textual) explanatory dialogue modelling, we designed an argumentative conversational agent that handles user requests sequentially in accordance with the pipeline of dialogue system components that comprises the following four components: •The natural language understanding (NLU) module. It serves two purposes: (1) to identify user’s intent and (2) recognise all entities that the request has (if any). •The dialogue state tracker (DST). It defines the state that the dialogue is currently in based on the user’s intent recognised by the NLU module. •The dialogue policy manager (DPM). It selects the most appropriate response among those available at the given dialogue state passed by DST. •The NLG module. It generates well-formed, grammatical system’s responses based on the information received from DPM. Altogether, DST and DPM are said to constitute the dialogue manager. The argumentative dialogue protocol serves as the basis of the dialogue manager, as it tracks the state of the dialogue and defines all possible transitions among dialogue states. Fig. 4.2 illustrates the pipeline of the components of the implemented conversational agent. Let us consider a beer style classification problem to illustrate request processing. Given a pretrained beer style classifier, the user passes the characteristics of a specific beer (e.g., colour, bitterness, and strength) to the classification system and obtains the classifier’s prediction (e.g., “This beer is Blanche”). Then, the user seeks an explanation for the given prediction and submits the corresponding request to the dialogue system (e.g., “Why is this beer Blanche?”). At 35
ILIA STEPIN first, the NLU module analyses the user’s request to (1) identify what category it belongs to and (2) recognise all entities that the request contains. In our example, the NLU module first attempts to solve the user intent classification problem where it assigns probabilities to each category of user requests, e.g., (“why-explain”, 0.95). The same procedure is applied to entity recognition. The intent and entity rankings are passed on to DST. Then, DST selects the most probable intent (in this case, “why-explain”) and switches the dialogue to the corresponding state (in our example, that of a factual explanation request). Given the dialogue state, DPM seeks the most adequate response among the possible options. In accordance with the dialogue protocol, the system is allowed to respond to a factual explanation request (i.e., “why-explain”) by either offering a factual explanation to the end user (i.e., “explain-f ”) if it is able to generate it or refusing to offer it, otherwise (i.e., “no-explain-f ”). In our example, the system finds that the decision path responsible for the given prediction can be summarised in terms of two features (i.e., colour and bitterness) with the corresponding values (i.e., black and high, respectively). The explanatory feature-value pairs are then passed on to the NLG module, which generates a well-formed, grammatical utterance. Once the NLG module returns the system’s utterance (in this example, “The beer is Blanche because its colour is black and bitterness is high.”), it is presented to the end user. Similarly to the factual explanation request from the example above, the explanatory dialogue system follows the aforementioned pipeline to handle all other types of user requests defined in the argumentative dialogue protocol (i.e., CF explanation, detailisation, clarification, and alternative explanation requests). In each case, DPM makes calls to external modules if necessary (e.g., the explainer when generating explanations or the knowledge base containing domain knowledge when processing clarification requests). The proposed explanatory dialogue protocol provides a transparent means of explanation communication in information-seeking settings. Being transparent, the proposed model can be aligned with regulatory requirements to automatic explanation generation. Furthermore, it is shown to take into account user preferences, as it favours diversity of the output explanations and allows the user to further question all variable components of such explanations. The proposed argumentative dialogue model can be potentially used as a tool for assessing effectiveness of CF explanations generated by other rule-based CF explanation generation algorithms. To facilitate adaptation to different rule-based CF explainers, the proposed dialogue model is further conceptualised using the formalism of dialogue grammars. Notably, the proposed dialogue protocol also allows us to quantitatively estimate the necessity in diversity of qualitative CF explanations. Recall that the proposed dialogue protocol represents the explanation space for each (possibly, CF) class as a list of explanations ranked by relevance, as measured by the explainer. Given a corpus of explanatory dialogues collected, we can compare the explainer-measured relevance to the demand in the given explanations based on 36
Chapter 4. General discussion the data collected from actual users. Statistics of alternative explanation requests for each class can be useful for this purpose. On the one hand, a large number of alternative explanations asked for may be a signal of little utility of the initially offered explanations. On the other hand, the empirical data from end users may not adequately reflect their satisfaction with the explanation space in its entirety, as end users are at all times exposed to only a part of the explanation space unless they sequentially request all the explanations that the system can offer to them. In this regard, a metric of similarity between the set of explanations (at least, potentially) generated and that of actually requested may be useful for assessing automatically the user satisfaction with the explanation space, in general. Finally, the concept of diversity of CF explanations regarded in terms of a set of single alternative explanations allows us to question the nature of the mere definition of a CF explanation. Indeed, automated CFs generated following the conventional definition search minimally different feature-value pairs that ensure a different classification. However, if there is a strong tendency to disregard minimally different CF data points that form the basis of a CF explanation and it is shown to be consistent for different explainers and audiences, it may be timely to reconsider the definition of a CF explanation or empower it with the property of human-centricity that goes beyond existing automatically computable metrics. Whereas the piece of work presented in this thesis only makes the first step in this direction, it can serve an inspiring source of ideas for elaboration on the user-centric prospects of automated CF explanations. 37
5 Contributions The work on the present thesis resulted in (1) several pieces of software developed to reach the thesis objectives and (2) various publications that emerged as a result of the studies carried out during the doctoral project. Section 5.1 details the software produced as part of the thesis. Section 5.2 lists all the publications that this thesis bases upon. 5.1 SOFTWARE The following pieces of software have been developed in order to achieve the thesis objectives: C1: FCFExpGen1– a framework for factual and CF explanation generation for interpretable rule-based classification systems. The proposed framework includes the following three algorithms: XOR: a model-specific algorithm that selects candidate CF rules and ranks them by relevance to the test instance using the eXclusive-OR (XOR) function. The rule claimed to be the most relevant represents a set of CF data points that are minimally different from the test instance in terms of their features. This CF set forms the basis of the output CF explanation; EUC: a model-specific variant of the XOR algorithm that utilises Euclidean distance as a metric of relevance of the candidate CF rule to the test instance (i.e., it measures proximity of the vector representation of the CF rule to the test instance in the selected n-dimensional space); GEN: a model-agnostic genetic algorithm that solves an optimisation problem looking for the closest single CF data point w.r.t. the test instance under consideration. C2: DialGame2– an argumentative conversational agent for communication of automatically generated textual rule-based factual and CF explanations. Noteworthy, the dialogue man- 1https://gitlab.citius.usc.es/ilia.stepin/fcfexpgen (branch “xor_euc_gen”) 2https://gitlab.citius.usc.es/ilia.stepin/fcfexpgen (branches “dialgame” and “dialgame_nlu”)
ILIA STEPIN ager (i.e., an integrative part of the conversational agent) implements a dialogue protocol basing on the argumentation theory-based technique referred to as “dialogue game” [30]. C3: SurveyGenerator3– a web-tool for carrying out human evaluation experiments designed for assessing various explanation aspects as well as the quality of explanation communication. Examples of the experiments carried out include: –Survey GM: the explanation evaluation survey that enables the end-user to rate a series of distinct explanations for the same test instance in terms of informativeness, trustworthiness, accuracy, relevance, and readability; –Survey TS: a simplified version of Survey GM which welcomes the end-user to rate a single explanation for the given test instance in terms of trustworthiness and satisfaction. It is worth noting that all the source code, the data used in the human evaluation experiments and the corresponding experimental results are made publicly available and can be reached at a public Gitlab repository. 5.2 PUBLICATIONS The work on the present doctoral thesis has resulted in three journal papers, three papers presented at international conferences and included in conference proceedings (both main and other tracks), and one book chapter. Namely, the following journal publications cover the algorithms developed and evaluated within the doctoral project [36, 37, 38]: • Ilia Stepin, Jose M. Alonso, Alejandro Catala, Martín Pereira-Fariña. “A survey of contrastive and counterfactual explanation generation methods for explainable artificial intelligence”. IEEE Access, vol. 9, pp. 11974–12001, 2021. DOI: 10.1109/ACCESS.2021.3051315; • Ilia Stepin, Jose M. Alonso-Moral, Alejandro Catala, Martín Pereira-Fariña. “An empirical study on how humans appreciate automated counterfactual explanations which embrace imprecise information”. Information Sciences, vol. 618, pp. 379–399, 2022. DOI: 10.1016/j.ins.2022.10.098; • Ilia Stepin, Katarzyna Budzynska, Alejandro Catala, Martín Pereira-Fariña, Jose M. Alonso- Moral. “Information-seeking dialogue for explainable artificial intelligence: Modelling and analytics”. Argument and Computation, in press. DOI: 10.3233/AAC-220011. 3https://gitlab.citius.usc.es/jose.alonso/surveygenerator 40
Chapter 5. Contributions Further, the work in progress has been presented at several international conferences (including main tracks, workshops, and doctoral consortia), which resulted in the following publications [34, 35, 39]: • Ilia Stepin, Alejandro Catala, Jose M. Alonso, Martín Pereira-Fariña. “Paving the way towards counterfactual generation in argumentative conversational agents”. In Proceedings of the 1st Workshop on Interactive Natural Language Technology for Explainable Artificial Intelligence (NL4XAI) collated with the Conference on International Natural Language Generation (INLG), pp. 20-25, Tokyo (Japan), 2019. DOI: 10.18653/v1/W19- 8405; • Ilia Stepin, Jose M. Alonso, Alejandro Catala, M. Pereira-Fariña. “Generation and evaluation of factual and counterfactual explanations for decision trees and fuzzy rule-based classifiers”. In Proceedings of the IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), Glasgow (UK), 2020. DOI: 10.1109/FUZZ48607.2020.9177629; • Ilia Stepin. “Argumentation-based interactive factual and counterfactual explanation generation”. In Proceedings of the 1st Doctoral Consortium at the European Conference on Artificial Intelligence (DC-ECAI 2020), pp. 61-62, Santiago de Compostela (Spain), 2020. Notably, further experiments concerning specific technicalities of the XOR algorithm for automated factual and CF explanation generation can be found in the following book chapter [40]: • Ilia Stepin, Alejandro Catala, Martín Pereira-Fariña, Jose M. Alonso. “Factual and counterfactual explanation of fuzzy information granules”. In: Pedrycz, W., Chen, SM. (eds) Interpretable Artificial Intelligence: A Perspective of Granular Computing. Studies in Computational Intelligence, vol 937. Springer, Cham. DOI: 10.1007/978-3-030-64949- 4_6 In addition, a proof of concept of the argumentative framework for factual and CF explanation communication was presented orally at the European Conference on Argumentation in Rome (Italy) on September 30, 2022. The full paper is currently under review, pending to be finally published by College Publications in their “Studies in Logic and Argumentation” book series4in 2023. Besides, one of the model-specific interpretable fuzzy rule-based explanation generation algorithms (i.e., EUC) is presented in an immersive article entitled “How to build selfexplaining fuzzy systems: From interpretability to explainability” and submitted to the special issue “Artificial Intelligence eXplained” (AI-X) of the IEEE Computational Intelligence Magazine. At the moment of writing, the manuscript is undergoing a second review round. 4https://www.collegepublications.co.uk/logic/sla/ 41
ILIA STEPIN Fig. 5.1 visualises the main contributions (i.e., software and journal papers) listed in Sections 5.1-5.2, respectively. Figure 5.1: Main research questions and answers (regarding techniques, software and publications). 42
6 State-of-the-art methods of contrastive and counterfactual explanation generation As discussed in Chapter 4, the problem of contrastive and CF explanation generation is among the most trending XAI topics [21]. The recent rise of attention to contrastive and CF explanation encourages a discussion about the nature of these concepts and their (mis-)use in XAI. Whereas XAI is a newly emerged research field, the notion of explanation, in general, and its contrastive and CF sub-types, in particular, have long been discussed in social sciences. It is therefore important to know how theoretical concepts related to contrastive and CF explanation can guide XAI researchers in developing effective explanation generation methods. In this chapter, we (1) inspect theoretical foundations of the concepts of contrastive and CF explanation, (2) explore technical aspects of the state-of-the-art computational frameworks of generation of contrastive and/or CF explanations, and (3) discuss how theoretically grounded the computational frameworks are. We show that a majority of state-of-the-art contrastive and CF explanation generation frameworks only loosely (if at all) rely on theoretical models of explanation from social sciences. Instead, they mainly represent data-driven solutions to optimisation problems oftentimes neglecting a wide body of knowledge about the nature of explanation accumulated over the centuries. As a consequence, it sometimes leads to terminological confusion and makes us strive for standardisation of the explanation-related notions used across sub-fields of XAI. The results from this chapter are published in the following paper [36]: Ilia Stepina, Jose M. Alonsoa, Alejandro Catalaa, Martín Pereira-Fariñab. “A survey of contrastive and counterfactual explanation generation methods for explainable artificial intelligence”. In IEEE Access (Open Access), vol. 9, pp. 11974-12001, ISSN: 2169-3536, 2021. DOI: 10.1109/ACCESS.2021.3051315 aCentro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela, Rúa de Jenaro de la Fuente Domínguez, s/n, 15782 Santiago de Compostela, Spain
ILIA STEPIN aCentro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela, Rúa de Jenaro de la Fuente Domínguez, s/n, 15782 Santiago de Compostela, A Coruña, Spain bLaboratory of The New Ethos, Warsaw University of Technology, plac Politechniki 1, 00-661, Warsaw, Poland cDepartamento de Electrónica e Computación, Universidade de Santiago de Compostela, Rúa Lope Gómez de Marzoa, s/n, 15782 Santiago de Compostela, A Coruña, Spain dDepartamento de Filosofía e Antropoloxía, Universidade de Santiago de Compostela, Plaza de Mazarelos s/n, 15705 Santiago de Compostela, A Coruña, Spain Scientific production indicators: At the moment of writing, the scientific production indicators for the year of publication (i.e., 2023) are unavailable for Argument and Computation, the journal where Chapter 8 was published. In the year immediately preceding that of publication (i.e., 2022), the journal had a CiteScore index of 3.6 (calculated by Scopus on 05 May, 2023) and an impact factor of 1.4 (2022 Journal Citation Reports). In addition, it had the following positions in the categories listed below: • Scopus: Q1 (rank #95/1078) in Linguistics and Language (the 91st percentile), Q2 (rank #57/172) in Computational Mathematics (the 67th percentile), Q2 (rank #366/792) in Computer Science Applications (the 53rd percentile), Q3 (rank #159/301) in Artificial Intelligence (the 47th percentile); • JCR: Q4 (rank #155/192) in Computer Science and Artificial Intelligence. As of 29 June 2023, the publication has not been cited yet. Personal authorship statement: In accordance with the Contributor Roles Taxonomy (CRediT), the personal authorship contribution comprises the following roles: methodology, software, validation, investigation, data curation, writing - original draft, visualization. Publishing rights: The journal paper where the results of this Chapter are published is licensed under a Creative Commons Attribution-NonCommercial 4.0 License. The individual or entity exercising the licensed rights (hereinafter, “the user”) is free to copy and redistribute the material in any medium or format under the following terms1: 1For more information, see https://creativecommons.org/licenses/by-nc/4.0/ 50
Chapter 8. Argumentative explanation communication for rule-based classification systems •Attribution: The user must give appropriate credit, provide a link to the license, and indicate if changes were made. The user may do so in any reasonable manner, but not in any way that suggests the licensor endorses the user or the user’s use. •NonCommercial: The user may not use the material for commercial purposes. 51
9 Conclusion In this thesis, we addressed the problem of explanation generation and communication for interpretable rule-based systems. More specifically, we focused on the task of generation of interactive CF explanations, which meet theoretically grounded requirements to quality explanations for XAI. To address this challenge, we first performed a literature review of theoretical foundations of the family of contrastive and CF explanations and the state-of-the-art methods of their automatic generation, which resulted in a two-level taxonomy of contrastive and CF explanations. Taking into consideration the insights from the review, we then designed, implemented, and validated a computational framework for generating factual and CF explanations associated to interpretable rule-based classification systems. It includes one model-agnostic and two model-specific algorithms, all of which offer human-comprehensive explanations in natural language. Further, we enhanced the framework with an argumentative dialogue generation module, which allows for interactive explanations in agreement with end user’s needs. All in all, the generated explanations have been shown to be contrastive, selected, and social. In what follows, we summarise main lessons learned from the work carried out during this doctoral project. Section 9.1 encapsulates our concluding remarks for each piece of research reported in Chapters 6-8. Section 9.2 outlines prospective directions for future work. 9.1 CONCLUDING REMARKS The literature review on contrastive and CF explanations revealed several gaps in the inspected sub-field of XAI. First, the state-of-the-art computational frameworks of contrastive and CF explanation generation are scarcely grounded on explanation theories from social sciences. This is, in part, due to the fact that the existing theories and computational frameworks mainly address distinct aspects of explanation generation. Thus, theories of contrastive explanation often discuss products of the explanatory process in terms of cause-and-effect relationships whereas a large number of computational contrastive and/or CF explanation generation methods focus on non-causal (e.g., spacial) relations between the test point whose prediction is to be explained and potential CF data points. Second, it turns out that the terms “contrastive” and “counterfactual” are often used interchangeably in the XAI community despite certain methodological
ILIA STEPIN differences. This observation calls for standartisation of the terminology used in the field. In order to unify the terminology (where applicable), we suggest that the term “contfactual” be promoted for contrastive-CF explanations. We believe that this term adequately encompasses the aspect of contrastiveness in CF explanations and vice versa. Third, the state-of-the-art computational frameworks have been observed to greatly lack human evaluation support. In fact, a vast majority of contrastive and CF explanation generation methods have only been evaluated using data-driven automatically computed metrics. Nevertheless, human evaluation studies are indispensable for shifting towards human-centric AI despite being expensive and difficult to design. In order to reduce the gap between the automatic data-driven and human evaluation-based metrics, we proposed the metric of perceived explanation complexity (PEC), i.e. a measure of how complex the given piece of textual explanation seems to be for the end user to process it. For the target audience (in this case, users who have a high degree of expertise in XAI or related fields), human evaluation experiments have shown that the computed PEC scores correlate with informativeness, relevance, and readability of automated generations. Thus, the proposed metric can effectively replace human evaluation experiments measuring the aforementioned explanation aspects for domain experts or highly qualified specialists. Overall, the quality of the explanations generated by all the proposed algorithms was positively evaluated in the human evaluation studies that we carried out in this thesis. In terms of all the assessed explanation aspects (i.e. informativeness, trustworthiness, accuracy, relevance, readability, and satisfaction), the resulting scores were above average for all the proposed CF explanation generation methods. However, none of these methods has been found to consistently outperform the others in all the explanation aspects. In this regard, we conclude that the methods modelling imprecise knowledge (e.g., XOR and EUC) and those making use of precise feature values (e.g., GEN) are best used complementarily to each other to satisfy the needs of a wider audience of users. Whereas such complementary information can be aggregated in a single piece of text, the social aspect of explanation remains unaddressed in case of one-shot explanations. To overcome this issue, we proposed an argumentative dialogue protocol to model information-seeking explanatory dialogues and developed a corresponding conversational agent. The human evaluation results of the dialogue protocol validation prove the necessity for all the proposed types of requests and responses for effective explanation communication. Further, a large number of requests for alternative CF explanations testify that the most relevant CF explanations from the algorithmic point of view may oftentimes not seem optimal from the user’s point of view. Whereas further comparative studies are necessary to analyse explanatory power of distinct explanation generation algorithms from the cognitive point of view, it can be concluded at this stage that the best-ranked CFs (i.e., most relevant or minimally different CFs from the test instance) may have to be combined with one or more alternatives given a set of multiple candidate CFs. 54
Chapter 9. Conclusion Striving for enabling the end user to play a decisive role in the process of explanation communication, we believe that it is indispensable to further emphasise the social aspect of automated explanations, e.g., by designing additional metrics of the quality of explanation communication similarly to the PEC score proposed in this thesis. In light of the statements made above, we conclude that the explanation generation framework proposed in this thesis offers factual and CF explanations that turn out to be appealing to end users. Thus, textual explanations generated in both modalities (those modelling imprecise knowledge and those outputting specific numerical values or intervals thereof) received higherthan-average estimates from the target audiences in all the experiments carried out. In addition, end users appreciated their interaction with the argumentative conversational agent which was carefully designed to communicate automated factual and CF explanations. 9.2 FUTURE WORK The research results presented in this thesis indicate several directions for future work. From the theoretical point of view, the proposed framework can be extended to introduce causal relations between the predicted data and related features. In fact, this could further bridge the gap between theoretical and computational paradigms of contrastive and CF explanation generation and therefore appears highly desirable in light of the results obtained from the performed literature review. From the algorithmic point of view, the proposed explanation generation framework should be further extended with a surrogation approach to handle other types of classifiers (including those non-interpretable and non-rule-based). Nevertheless, in the absence of access to the internals of the classifier and/or the feature space, it may be essential to transform the proposed model-specific explanation generation algorithms into their model-agnostic equivalents. Further, different settings may require changes in the dialogue protocol that models the communication process between the explainer and the user for the sake of deeper customisation. In addition, further extension is required to guarantee properties of CF explanations not addressed in this thesis. For example, the presented algorithms do not allow to assess straightforwardly how actionable the generated CFs are. Hence, enhancing output CF explanations with other desired CF properties is believed to further increase effectiveness of such explanations. Importantly, it seems impossible to achieve the state of human-centric AI without formalising and modelling ethical relations on the basis of the data being processed. Thus, bias mitigation, yet another highly relevant line of research in the XAI community, is another algorithmic challenge to address. Finally, it is of our particular interest to further adapt the designed human evaluation framework for future experiments on explanation, trustworthiness, and satisfaction. Altogether, the prospective extensions of the work presented in this thesis are believed to have great potential for moving forward from XAI to Trustworthy AI. 55
Ethical considerations The experiments whose results are reported in the present thesis were designed to involve human evaluation. Therefore, a prior permission to carry them out had been obtained from the Ethics Committee of the University of Santiago de Compostela (a copy of the corresponding certificate is attached below). All the information collected from the human evaluation study participants was in agreement with the European Union’s General Data Protection Regulation (GDPR). Further, human evaluation was based solely on non-personal or anonymous data. In addition, all the participants gave informed consent confirming the following: • the participant reached the age of majority; • participation in the study was completely voluntary; • participation in the study could be terminated at any time; • participant’s anonymous responses would be used for research purposes in accordance with the GDPR. None of the experimental results could anyhow be used against the study participants. Risks of misuse of the collected results were minimal. In light of an increasing use of AI-based applications for automated text and/or image generation, it is important to mention that no piece of text or pictures from the present thesis or any of the published or accepted articles was generated using any generative AI-based applications (e.g., ChatGPT).
Bibliography [1] S. Ali, T. Abuhmed, S. El-Sappagh, K. Muhammad, Jose M. Alonso-Moral, R. Confalonieri, R. Guidotti, J. Del Ser, N. Díaz-Rodríguez, and F. Herrera. Explainable artificial intelligence (XAI): What we know and what is left to attain trustworthy artificial intelligence. Information Fusion, in press, 2023. DOI: 10.1016/j.inffus.2023.101805. [2] J. M. Alonso, C. Castiello, L. Magdalena, and C. Mencar. Explainable Fuzzy Systems - Paving the Way from Interpretable Fuzzy Systems to Explainable AI Systems, volume 970. Springer International Publishing, 2021. DOI: 10.1007/978-3-030-71098-9. [3] F. Bex and D. Walton. Combining explanation and argumentation in dialogue. Argument & Computation, 7(1):55–68, 2016. DOI: 10.3233/AAC-160001. [4] R. M. J. Byrne. Counterfactual thought. Annual Review of Psychology, 67:135–157, 2016. DOI: 10.1146/annurev-psych-122414-033249. [5] R. M. J. Byrne. Counterfactuals in explainable artificial intelligence (XAI): Evidence from human reasoning. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), pages 6276–6282, 2019. DOI: 10.24963/ijcai.2019/876. [6] V. Dignum. Responsible artificial intelligence: how to develop and use AI in a responsible way. Artificial Intelligence: Foundations, Theory, and Algorithms. Springer, Cham, 2019. DOI: 10.1007/978-3-030-30371-6. [7] A. Gatt and E. Krahmer. Survey of the state of the art in natural language generation: Core tasks, applications and evaluation. Journal of Artificial Intelligence Research, 61:65–170, 2018. DOI: 10.1613/jair.5477. [8] R. Guidotti. Counterfactual explanations and how to find them: literature review and benchmarking. Data Mining and Knowledge Discovery, pages 1–55, 2022. DOI: 10.1007/s10618-022-00831-6. [9] R. Guidotti, A. Monreale, F. Giannotti, D. Pedreschi, S. Ruggieri, and F. Turini. Factual and Counterfactual Explanations for Black Box Decision Making. IEEE Intelligent Systems, 34(6):14–23, 2019. DOI: 10.1109/MIS.2019.2957223.
List of Acronyms AI artificial intelligence AIA artificial intelligence act CF counterfactual DPM dialogue policy manager DST dialogue state tracker DT decision tree FRBCS fuzzy rule-based classification system FURIA fuzzy unordered rule induction algorithm GDPR general data protection regulation ML machine learning NLU natural language understanding NLG natural language generation XAI explainable artificial intelligence
ILIA STEPIN APPENDIX. PUBLISHED OR ACCEPTED ARTICLES This appendix contains three journal papers that form the basis of the present thesis. All of them are Open Access publications, making them publicly available to the interested reader. In this appendix, they appear in the following order: • Ilia Stepin, Jose M. Alonso, Alejandro Catala, Martín Pereira-Fariña. “A survey of contrastive and counterfactual explanation generation methods for explainable artificial intelligence”. In IEEE Access, vol. 9, pp. 11974-12001, ISSN: 2169-3536. IEEE Inc. (Open Access), 2021. DOI: 10.1109/ACCESS.2021.3051315 • Ilia Stepin, Jose M. Alonso-Moral, Alejandro Catala, Martín Pereira-Fariña. “An empirical study on how humans appreciate automated counterfactual explanations which embrace imprecise information”. In Information Sciences, vol. 618, pp. 379-399. ISSN: 0020-0255. Elsevier (Open Access), 2022. DOI: 10.1016/j.ins.2022.10.098 • Ilia Stepin, Katarzyna Budzynska, Alejandro Catala, Martín Pereira-Fariña, Jose M. Alonso- Moral. “Information-seeking dialogue for explainable artificial intelligence: Modelling and analytics”. In Argument and Computation, ISSN print: 1946-2166; ISSN online: 1946-2174, in press. IOS Press (Open Access). DOI: 10.3233/AAC-220011 68
Received December 16, 2020, accepted January 3, 2021, date of publication January 13, 2021, date of current version January 22, 2021. Digital Object Identifier 10.1109/ACCESS.2021.3051315 A Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable Artificial Intelligence ILIA STEPIN 1, JOSE M. ALONSO 1, (Member, IEEE), ALEJANDRO CATALA 1, AND MARTÍN PEREIRA-FARIÑA 2 1Centro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain 2Departamento de Filosofía e Antropoloxía, Universidade de Santiago de Compostela, 15705 Santiago de Compostela, Spain Corresponding author: Ilia Stepin ([email protected]) This work was supported in part by the Spanish Ministry of Science, Innovation and Universities under Grant RTI2018-099646-B-I00 and Grant RED2018-102641-T, in part by the Galician Ministry of Education, University and Professional Training under Grant ED431F 2018/02, Grant ED431C 2018/29, Grant ED431G/08, and Grant ED431G2019/04; and in part by the European Regional Development Fund (ERDF/FEDER Program). ABSTRACT A number of algorithms in the field of artificial intelligence offer poorly interpretable decisions. To disclose the reasoning behind such algorithms, their output can be explained by means of so-called evidence-based (or factual) explanations. Alternatively, contrastive and counterfactual explanations justify why the output of the algorithms is not any different and how it could be changed, respectively. It is of crucial importance to bridge the gap between theoretical approaches to contrastive and counterfactual explanation and the corresponding computational frameworks. In this work we conduct a systematic literature review which provides readers with a thorough and reproducible analysis of the interdisciplinary research field under study. We first examine theoretical foundations of contrastive and counterfactual accounts of explanation. Then, we report the state-of-the-art computational frameworks for contrastive and counterfactual explanation generation. In addition, we analyze how grounded such frameworks are on the insights from the inspected theoretical approaches. As a result, we highlight a variety of properties of the approaches under study and reveal a number of shortcomings thereof. Moreover, we define a taxonomy regarding both theoretical and practical approaches to contrastive and counterfactual explanation. INDEX TERMS Computational intelligence, contrastive explanations, counterfactuals, explainable artificial intelligence, systematic literature review. I. INTRODUCTION In the last few decades, the field of Artificial Intelligence (AI) has witnessed major changes. As available computational resources have grown significantly, AI algorithms are attracting a significant amount of attention in industry and research [1]. While a great number of such algorithms present strikingly accurate decisions, their decision-making apparatus is frequently left unclear to users of such applications. In particular, a number of Machine Learning (ML)-based algorithms are often perceived as ‘‘black-box’’ algorithms because they are overloaded with millions of hardly interpretable parameters to be optimized at the training stage. This fact makes the algorithm’s output hard to explain. A lack of The associate editor coordinating the review of this manuscript and approving it for publication was Francesco Piccialli. the ability to explain such automatic decisions undermines users’ trust and hence decreases usability of such systems [2]. Furthermore, it prevents users from a responsible exploitation of their decisions [3]. In addition, many of the existing eXplainable AI (XAI1) methods provide summaries of automatically made predictions rather than true explanations [4]. As a result, the need to motivate automatic decisions with a clear explanation of why the algorithm outputs a particular decision has made the XAI research field grow quickly [5]. Since the number of high-stakes AI applications found in daily life increases, the requirements to their explana- 1XAI stands for eXplainable Artificial Intelligence. This acronym was made popular by the USA Defense Advanced Research Projects Agency when launching to the research community the challenge of designing self-explanatory AI systems (https://www.darpa.mil/program/ explainable-artificial-intelligence). 11974 This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License. For more information, see https://creativecommons.org/licenses/by-nc-nd/4.0/ VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI tory capacity increase accordingly. This also provokes the introduction of regulations and laws concerned with explanation requirements for AI-based applications. For instance, the need for explaining reasoning mechanisms behind such applications is now legally regulated in the European Union by means of the General Data Protection Regulation.2 According to these legal provisions, the data subject must be provided with ‘‘meaningful information about the logic involved’’ in the automatic decision making process, which is commonly referred to as the ‘‘right to explanation’’ [6]. Thus, an AI application is expected not only to provide accurate decisions but also to justify them in a comprehensive manner to end-users. The goal of approaching human-centric AI has led towards a deeper research on the nature of explanation. However, no agreement about a definition of explanation has been reached despite the fact that explanation has called a significant amount of attention in, e.g., philosophy of science [7], [8]. In its most general form, explanation is normally treated as ‘‘an answer to the question of why something is the case’’ [9]. In the context of AI, it often bases on judgments about why a certain outcome is predicted by an AI algorithm and hypotheses about causes with respect to given effects [10]. The need of generating more human-like explanations has attracted AI researchers’ attention to particular properties of explanation as well as its sub-types [11]. Thus, it appears particularly challenging to explain a given algorithm’s output in terms of reasonable yet non-occurring alternatives given a possibly infinite set of such options. Furthermore, this can be enhanced with the ability of suggesting relevant changes in the input so that the algorithm outputs a different decision. Given a rising interest towards these types of explanation (referred to as contrastive and counterfactual, respectively) within the XAI community, it is of crucial importance to review the existing theoretical accounts of contrastive and counterfactual explanation as well as state-of-the-art computational frameworks for automatic generation thereof. Thus, the aim of this study is to fulfill the next three objectives: (1) to scrutinize theoretical works on the contrastive and counterfactual accounts of explanation; (2) to summarize state-of-the-art methods in the field of automatic explanation generation thereof; and (3) to discuss a degree of synergy between the revised theories and their related up-to-date implementations. The rest of the manuscript is organized as follows. Section II introduces the notions of contrastive and counterfactual explanation as well as their main application areas. Section III presents the terminology used throughout the review, poses the research questions, and describes the methodology employed to address the given questions. Section IV presents the main findings collected within the present survey and the emerging taxonomy thereof. Section V discusses peculiarities of the existing theoretical and compu- 2https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX: 02016R0679-20160504 tational frameworks of contrastive and counterfactual explanation. Finally, we conclude in Section VI. II. BACKGROUND A. CONTRASTIVE EXPLANATION Findings on explanation accumulated in humanities and social sciences show that it is intrinsically contrastive [11]. The property of contrastiveness presupposes that an explanation answers the given why-question regarding the cause of the event in question (‘‘Why did Phappen?’’) in terms of hypothesized non-occurring alternatives (‘‘Why did Phappen rather than Q?’’) [12]. Thus, supporters of the pragmatic approach to explanation argue that it is exactly the ability to distinguish the answer to an explanatory question from a set of contrastive hypothesized alternatives that provides the explainee with sufficiently comprehensive information on the reasoning behind the question [13]. This approach is also claimed to set a minimum criterion that an explanation must fulfill: it must favor the probability of the observed event P to all the hypothetical alternatives (Q1,Q2,...,Qn) [14]. Contrastive explanation is among influential topics in cognitive science [15]–[17]. Thus, contrastive explanations are claimed to be inherent to human cognition [16]. Indeed, we are used to question those decisions that we once made, especially if such decisions or coinciding circumstances resulted in tragic events [18]. In addition, contrastive reasoning forms the basis of abductive inference [19], i.e., the process of inferring certain facts that render some observation plausible [20]. In other words, a given observation can be explained on the basis of the most likely among a pool of competing hypotheses [21]. B. COUNTERFACTUAL EXPLANATION Given the property of contrastiveness, it is possible to imagine explanatory alternatives to how things would stand if a different decision had been made at some point. They can serve to explain potential consequences of such contrastive non-taken alternative decisions. In this case, the mind is assumed to construct and compare mental representations of an actually happened event and that of some event alternative to it [22]. Cognitive scientists refer to such mental representations of alternatives to past events as counterfactuals (‘‘contrary-to- fact’’) [15]. The process of ‘‘thinking about past possibilities and past or present impossibilities’’ is therefore called counterfactual thinking [23]. Alternatively, the combination of imagining an alternative scenario in relation to the one that actually happened and the exploration of its consequences is referred to as counterfactual reasoning [24]. In addition, counterfactual reasoning is claimed to be a key mechanism for explaining adaptive behavior in a changing environment [25], [26]. Counterfactuals describe events or states of the world that did not occur and implicitly or explicitly contradict factual world knowledge [27]. Formulated in natural language, counterfactuals are usually presented in the form of conditional VOLUME 9, 2021 11975
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI statements. Broadly speaking, they contain: (1) an antecedent describing an outcome alternative to an actual event; (2) a consequent describing (a set of) consequences, had the antecedent been the case; and (3) a binary counterfactual dependency relation between them. Thus, Grahne defines a counterfactual to be a conditional statement where the antecedent ‘‘can contradict the current state of affairs, or our current knowledge thereof’’ [28]. However, despite a general agreement on structural properties of counterfactuals, existing interpretations of counterfactual conditionals still compete. As such, further constraints imposed on their structure differ depending on the approach adopted. According to Ginsberg [29], a counterfactual is a conditional statement of the form ‘‘If P, then Q’’ where Pis ‘‘expected to be false’’. Aumann limits a counterfactual to be a conditional with a false antecedent only [30]. In contrast, Spohn argues that both the antecedent and the consequent of a counterfactual must be false [31]. All in all, counterfactual conditional statements are claimed to enable people to produce utterances that are factually false yet truthful irrespective of the interpretation adopted [32]. A line of research devoted to modeling human counterfactual reasoning has been thoroughly investigated in computer science. Thus, counterfactual reasoning in computer science is defined as the process of evaluating conditional claims about alternative possibilities and their consequences [33]. It is argued to be valid arising from antecedents that are true in a hypothetical model but false in reality [34]. In this setting, the truth of a counterfactually inferred statement is resolved by: (1) modeling a situation where the smallest possible change in features of the actual world (as set in the antecedent) leads to a different (possibly, desired) state of things (the so-called ‘‘closest’’ or ‘‘nearest’’ possible world); and (2) estimating what is true in that setting [35]. Moreover, counterfactuality is among the most fundamental concepts in theories of causation [36], [37]. Indeed, counterfactuals are argued to represent a causal relation between the event happened in reality and its imaginary counterpart. A counterfactual definition of a cause of an arbitrary event traces back to Hume [38]. According to him, a cause is an object (antecedent) that justifies the existence of another object (consequent) which it is followed by: ‘‘If the first object had not been, then the second never had existed’’. Therefore, once a causal connection between the antecedent and the consequent is established, a counterfactual conditional can be generalized to be a conditional claim about an alternate possibility and its consequences of the form ‘‘If X were to occur, then Ywould (or might) occur’’ [33]. Similarly, Kment applies a similarity-based approach between possible worlds to formulate a general account of counterfactuals [39] driven by a non-epistemic interpretation of explanation (i.e., factors that serve as reasons for some fact to obtain are responsible for that fact). The conditional structure of counterfactual statements gave rise to a probabilistic account of such statements. Thus, Pearl extended the definition of the causal counterfactual to estimate the probability of the truth of the consequent caused by the antecedent (‘‘a probability statement about the truth of y, had xbeen true, when it is known that yhad been false when xwas false’’) [37]. This approach to counterfactuals motivated a number of experiments on the existence of the relation between counterfactuals and conditional probability. In support of this assumption, Over et al. [40] showed the existence of connection between counterfactuals and conditional probability, as they experimented with probability judgments about counterfactuals. Thus, they proposed that the subjective probability of the counterfactual at the present time is the same as the conditional probability P(y|x) at some earlier time. Twenty-six subjects were asked to estimate the probability of truth of thirty-two counterfactual conditionals with both affirmative and negative antecedents and consequents. Their findings point to a strong correlation between the probability of the counterfactual conditional and causal strength judgments. On a similar note, Edgington regarded counterfactual judgments as uncertain conditional statements and therefore evaluated them by estimating their conditional probability given some endorsing event [41]. C. DISTINCTION BETWEEN CONTRASTIVE AND COUNTERFACTUAL EXPLANATION It is important to note that some researchers tend to either collapse or intentionally distinguish contrastive reasoning from counterfactual reasoning despite their conceptual similarity. For instance, Lombrozo treated counterfactual and contrastive explanations as equivalent assuming hypothesized events non-occurred in reality to be ‘‘counterfactual cases’’ where a subset of these cases forms a contrastive explanation [10]. In contrast, McGill and Klein distinguished contrastive reasoning from its counterfactual counterpart [42]. According to them, contrastive reasoning is concerned with situations where different target situations are analyzed (‘‘What made the difference between the employee who failed and the employees who did not fail?’’). On the other hand, counterfactual reasoning is claimed to deal with cases where the antecedent is altered to account for changes in the outcome (‘‘Would the employee have failed had she not been a woman?’’). Alternatively, Fang et al. [43] referred to contrastive reasoning as a procedure operating on ‘‘butstatements’’, as in ‘‘all cars are polluting, but hybrid cars are not polluting’’, which serves a principally different explanation generation task in comparison with the other aforementioned approaches. D. CONTRASTIVE AND COUNTERFACTUAL EXPLANATION IN THE CONTEXT OF XAI The stochastic nature of predictions made by various AI algorithms is claimed to be among the main obstacles in reaching a true explanation [44]. Research on automatic contrastive and counterfactual explanation generation shows a number of considerable observations that help overcome this issue. Thus, empirical studies prove that incorporating contrastiveness improves the quality of explanations offered 11976 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI to the end-user [45]. Furthermore, contrastive explanations can be used to personalize human-machine interaction when a user is engaged in an explanatory dialogue with an AI application. Thus, they can be employed with the aim of adjusting the contents of the explanation for the algorithm’s output in accordance with the user’s preferences [46]. Finally, the ability to explain a decision contrastively is claimed to lead to responsible decision-making [47]. It is important to note that contrastive explanations point to the difference between the actual and a hypothetical decision. On the other hand, counterfactual explanations specify necessary minimal changes in the input so that a contrastive output is obtained. However, these terms are sometimes used interchangeably in the context of XAI [48], [49]. Various families of techniques have been proposed to generate contrastive and counterfactual explanations of AI algorithm output. In the context of XAI, an explanation for an automatic decision or prediction, treated as an observation, can be obtained abductively by attempting the search problem over the set of the known information concerning that observation [50]. Alternatively, counterfactual explanation is widely addressed in the paradigm of case-based reasoning, i.e., a family of problem solving methods based on appeals to precedent solutions. In this setting, generating the most suitable counterfactual may be viewed as a search problem where the most similar precedent is looked for among those making part of the case database [14]. Furthermore, Keane et al. argue that applying case-based reasoning techniques for generating counterfactuals increases their explanatory competence [51]. Counterfactual explanations are normally considered contrastive by nature and therefore present a source of valuable complementary information to a given automatic prediction [52]. For instance, a counterfactual explanation of an ML-based algorithm prediction may describe ‘‘the smallest change to the feature values that changes the prediction to a predefined output’’ [53]. An important advantage of counterfactual explanations over their non-counterfactual analogs is that they are devoid of any prerequisites to the data or model. Indeed, counterfactual explanations are dataagnostic as they can be based on the features of the neighbouring data examples extracted from the same training set and/or on the data generated synthetically around the data instance in question. In addition, counterfactual explanations are, in principle, model-agnostic, as they are suitable to explain the output of any black-box algorithm in a post-hoc manner. Whereas counterfactual explanation generation is concerned with a number of technical challenges, it also requires to take into account several ethical aspects. For instance, their use is expected to be safe (revealing model’s internals through counterfactuals may lead to model stealing) [54], fair (discriminatory explanations should be avoided) [55], actionable (suggested changes in the input should be feasible) [56], and accountable (ensuring responsibility for the explanations provided) [57]. III. METHODOLOGY The present survey has been undertaken as a systematic literature review following the guidelines by Kitchenham and Charters [58], Kitchenham et al. [59], and Wohlin [60]. The background notation necessary to follow the findings of the review is specified in Section III-A. In short, the study comprises three phases as established in the research method by Kitchenham and Charters [58]: (1) planning the review procedure; (2) conducting the review; and (3) reporting the results. During the first phase, three research questions (RQ1,RQ2, and RQ3) were specified (see Section III-B). Subsequently, we determined a search strategy to retrieve primary studies, i.e., we collected all the relevant publications investigating the research questions (see Section III-C). Then, we developed inclusion and exclusion criteria (see Section III-D) in order to select the studies relevant for this article. When the same publication was retrieved from multiple sources, all-but-one instances of the publication (duplicates) were discarded. In addition, we identified and added manually other relevant publications extracted from the bibliography lists of the previously selected manuscripts to ensure a maximum coverage of the related subject areas. It is worth noting that this additional procedure is informally known as snowballing [60]. Finally, we extracted and synthesized the data necessary to address the research questions (see Section III-E). A. PRELIMINARY TERMINOLOGY As has been shown in Section II, contrastive and counterfactual explanations presuppose a diverse nature across various application domains. Hence, let us now define the general terms used henceforth in this manuscript. As we are primarily concerned with explainability of AI algorithms, we define explanation in terms of the observed output of such an algorithm. Thus, we regard an explanation as a non-empty set of pieces of information justifying the given algorithm’s output for an input data instance. The explanation for the given output on the basis of the features of the input data instance is deemed as factual. An explanation opposing the actual outcome to one of possible other outcomes is considered to be contrastive (e.g., ‘‘The data instance is of class A and not B because ...’’). An explanation containing instructions on how the output could have been changed constitutes a counterfactual explanation (e.g., ‘‘The data instance would be of class B if ...’’). Explanations exhibiting patterns of both contrastive and counterfactual explanation are deemed to be contrastivecounterfactual explanations (e.g., ‘‘The data instance is of class A and not B because .... However, it would be of class B if ...’’). We distinguish between contrastive and counterfactual explanation throughout the rest of the manuscript if and only if only one of these two terms is used in the given primary study. In contrast, we unify the notions of counterfactual and contrastive explanation introducing the term ‘‘contfactual explanation’’ or ‘‘contfactual’’ to identify potential similarities and differences of both types of explanation VOLUME 9, 2021 11977
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI within a broader scope of literature. This term is used hereafter wherever both terms for contrastive and counterfactual explanation can be used interchangeably. The terms ‘‘contrastive explanation’’ and ‘‘counterfactual explanation’’ are only used when they are found in the corresponding study and cannot be used interchangeably in the given context. Notice that the term ‘‘contfactual explanation’’ is not equivalent to ‘‘contrastive-counterfactual explanation’’ but covers both independently used types of explanation as well as their fusion. A theoretical framework providing justification and a reasoning mechanism for obtaining a contfactual explanation is regarded as a theory of contfactual explanation. Altogether, we use the term contfactual explanation generation to refer to the process of automatic composition of contfactual explanations for a given output of an AI algorithm in the form of a complementary piece of information associated to a factual explanation. B. RESEARCH QUESTIONS In order to reach the three objectives of the study as formulated in Section I, the following three research questions were specified: •RQ1: How are contfactual explanations defined in the literature? •RQ2: What are the state-of-the-art methods of contfactual explanation generation? •RQ3: How grounded are the state-of-the-art contfactual explanation generation methods on the theoretical approaches to contfactual explanation? C. SEARCH STRATEGY We selected the digital libraries Scopus and Web of Science (WoS) to retrieve relevant publications from. These libraries do not only include research publications in computing but also index studies across all scientific fields, which allows for an objective analysis of the interdisciplinary literature relevant to the research questions posed. Subsequently, we performed six queries over the title, abstract, and author keywords in the aforementioned libraries (see the overall structure of the query pipeline in Fig. 1). It is worth noting that the proximity operator NEAR is used following the WoS notation whereas the equivalent proximity operator Wis used for the same queries in Scopus. The following search strings were used for querying the digital libraries: q1=counterfactual* W/3 expla* q2=contrastive* W/3 expla* q3=q1OR q2 q4=q3AND (defin* OR theor* OR infer* OR implic*) q5=q3AND (generat* OR implement* OR framework* OR develop* OR software* OR model* OR artificial intelligence OR AI) AND SUBJAREA(Computer Science OR Mathematics OR Engineering) q6=q4AND q5 FIGURE 1. A pipeline of the queries executed. The queries found in the dashed area are considered preparatory to those directly addressing the research questions. The search was performed on October 2nd , 2020. The search web tools of the selected digital libraries allow researchers to reproduce the original study. Furthermore, their use guarantees performing equivalent queries across both libraries. In order to capture all relevant publications, we only used the corresponding word-stems to allow for maximal diversity of the retrieved papers. For instance, the search item ‘‘expla*’’ was used to cover all publications containing such word-forms as ‘‘explanation’’, ‘‘explaining’’, ‘‘explanatory’’, and so on and so forth. Queries q1and q2embrace all the up-to-date publications containing mentions of counterfactual and contrastive explanation, respectively, found across all subject areas. In addition, we used a window span of three words (i.e., ‘‘NEAR/3’’) to ensure that the attributes ‘‘counterfactual’’ and ‘‘contrastive’’ relate to explanation. The resulting sets of publications were then unified (q3). Subsequently, the preprocessed collection of publications was split into two overlapping subsets aiming to distinguish the publications covering theoretical accounts of contfactual explanation with the aim of extracting the related definitions, theories (or their inferences or implications) (q4) and existing computational frameworks for contfactual explanation generation (q5). The terms ‘‘definition’’, ‘‘theory’’, ‘‘inference’’, and ‘‘implication’’ as well as their corresponding word-forms (q4) were expected to appropriately limit the pool of the unified set of publications with the aim of retrieving definitions as required for addressing RQ1. Similarly, we used the terms ‘‘generation’’, ‘‘implementation’’, ‘‘framework’’, ‘‘development’’, ‘‘software’’, ‘‘model’’, and their corresponding word-forms (q5) to retrieve publications concerning contfactual explanation generation frameworks. In addition, the terms ‘‘artificial intelligence’’ and ‘‘AI’’ were used to ensure retrieving relevant AI-related publications. Since RQ2addresses purely technical issues of 11978 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI state-of-the-art implementations of such tools, we further imposed an additional restriction on q5so that it would return only publications from such subject areas as computer science, mathematics, and engineering. Last but not least, the findings from q4and q5were merged to examine the connection between the existing theories of contfactual explanation and frameworks for automatic contfactual explanation generation in the context of XAI (q6). It is important to note that publications retrieved as a result of q4form an exhaustive set of papers addressing RQ1. Similarly, publications obtained as a result of q5address RQ2. Finally, the papers that q6returned address RQ3. D. INCLUSION AND EXCLUSION CRITERIA The publications retrieved during the initial search were subsequently inspected on the basis of the following inclusion and exclusion criteria. To address the epistemology of contfactual explanation, we filtered the retrieved publications to include in the collection of primary studies only those satisfying the following criteria: (1) a publication proposes a contrastive or counterfactual or contrastive-counterfactual approach to explanation or (2) it contains a clearly formulated definition of counterfactual or contrastive explanation referring to other publications in the corresponding field. In order to capture existing computational frameworks for contfactual explanation generation, we included publications that: (1) present a novel approach, method, or framework for contfactual explanation generation whose output can serve to explain the reasoning of an AI algorithm and (2) are found in such subject areas as computer science, mathematics, engineering as well as in their sub-fields. In contrast, we excluded duplicate reports of the same studies appeared in both Scopus and WoS. As for the publications related to RQ1, we also removed: (1) the studies whose contents did not introduce any contfactual theory of explanation or (2) those containing no formal or informal definition of contrastive or counterfactual or contrastivecounterfactual explanation. As for the publications related to RQ2, we discarded: (1) the publications which were not related to AI algorithms or applications as well as (2) those where the proposed framework did not provide any human-comprehensible contfactual explanations as output. E. DATA EXTRACTION AND SYNTHESIS Table 1shows the number of publications retrieved after each independent query, duplicates found among them in Scopus and WoS, as well as Candidate Primary Studies (CPS). Note that the numbers of duplicates indicated in Table 1refer only to within-query duplicates, i.e., the same publications retrieved from Scopus and WoS for the given single query. Recall that q4and q5exhaustively cover all the three research questions. Hence, the numbers of CPS are calculated as a sum of the publications retrieved after q4and q5. Furthermore, CPS are reduced by the number of publications addressing TABLE 1. Numbers of publications retrieved after each single query as well as those forming the pool of candidate primary studies. The numbers of publications making part of the primary studies are highlighted in bold. FIGURE 2. A flow diagram of the primary study selection on the basis of queries q4and q5(nis the number of publications at each stage). RQ3because they are found in both sets of publications collected for RQ1and RQ2and are therefore duplicates. Fig. 2displays the flow diagram of the primary study selection. A sum of 338 publications (207 from Scopus and 131 from WoS) made up the collection of CPS addressing the research questions. 107 within-query duplicates were identified and removed from further analysis. In addition, 29 more duplicates were excluded when merging the sets of publications retrieved after q4and q5. All in all, 136 duplicates were removed. The title, abstract, and author keywords of each candidate primary study were screened to discard the studies irrelevant to the research questions posed. As shown in Fig. 2, 75 publications were deemed irrelevant and filtered out at this stage. A deeper analysis of the remaining 127 publications VOLUME 9, 2021 11979
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI FIGURE 3. A taxonomy of contfactual explanation emerging from our systematic literature review. TABLE 2. The exhaustive list of all the primary studies in relation to each research question. enforced us to discard 33 studies which did not satisfy the inclusion criteria. Finally, 19 papers were added to the review upon inspecting the bibliography of the primary studies. As a result, 113 unique publications formed the exhaustive pool of primary studies. Table 2presents the list of primary studies selected for the review. Thus, a collection of 74 out of 113 (65.49%) original primary studies were found to formulate definitions for contfactual explanation and/or address theoretical accounts thereof (RQ1). In addition, 52 out of 113 (46.02%) publications describe frameworks (or extensions of other frameworks) for contfactual explanation generation (RQ2). Note that 13 out of 113 (11.50%) primary studies were found to address both RQ1and RQ2and therefore answer RQ3. The following data were extracted from each primary study: title, authors, year of publication, author keywords. In addition, all publications related to RQ1were read to analyze contfactual theories of explanation and, subsequently, extract the sought-for definitions of contfactual explanation. As RQ2concerned a broader number of technical characteristics of contfactual explanation frameworks, we additionally extracted the following information: (1) the problem that the retrieved framework aims to solve; (2) the method proposed for contfactual explanation generation; (3) the form of output explanation (for instance, textual or visual); and (4) the corresponding evaluation methods. Based on the data extracted from the primary studies, the publications were grouped and classified in accordance with the aforementioned criteria. IV. RESULTS Prior to answering the research questions, we carried out a bibliometric analysis over the results of the general independent queries on counterfactual and contrastive explanation (q1 and q2, respectively) as well as their union (q3). We report the results of the bibliometric analysis in Section IV-A. The findings related to the theoretical accounts of contfactual explanation (RQ1) are presented in Section IV-B. The analysis of the computational frameworks for contfactual explanation generation (RQ2) can be found in Section IV-C. Finally, the publications describing theoretically grounded computational frameworks (RQ3) are reported in Section IV-D. An emerging taxonomy of contfactual explanation frameworks is depicted in Fig. 3and forms the core of the results discussed in the rest of the manuscript. A. BIBLIOMETRIC ANALYSIS The bibliometric analysis over the queries q1,q2, and q3 allows us to obtain a big picture of the research area of contfactual explanation generation and spot its key characteristics. To illustrate the state of affairs within the field, we report annual scientific production and maps of author keywords revealing the main problem-specific notions. The reference 11980 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI TABLE 6. A classification of the contfactual explanation generators by AI problem. they emphasize the need for causal attribution, as ignoring causal relations may lead to generating unfeasible counterfactuals. Therefore, they suggest a hybrid framework for counterfactual explanation generation. •Hybrid contrastive-counterfactual explanation. Kuorikoski and Ylikoski point to the multifaceted nature of contrastive-counterfactual explanation. They argue ‘‘there exist constitutive and possibly formal counterfactual dependencies as well as combinations of these’’ [98]. Similarly, Pexton suggests a two-level hierarchy of explanation [108]: microphysical explanations are non-causal and form the lower-level of the hierarchy whereas manipulable causal explanations are placed at the higher-level. C. CONTFACTUAL EXPLANATIONS AS DEFINED IN AUTOMATIC GENERATION FRAMEWORKS (ANSWER TO RQ2) The analysis of the primary studies related to RQ2allows us to categorize the state-of-the-art contfactual explanation generation frameworks in accordance with the following criteria: (1) the problem the solution for which is to be explained (i.e., the AI problem); (2) the method employed to generate such an explanation (i.e., the explainability method); (3) the output representation of the explanation; and (4) the evaluation method thereof. 1) AI PROBLEM Contfactual explanations are used to justify automatic decisions obtained for a variety of AI-related problems. Table 6 provides the reader with a taxonomy of the state-of-the- art frameworks from the primary studies. It is derived from the considered publications in terms of the domain tasks that these frameworks are used for. As depicted in Fig. 9, most contfactual explanation generation frameworks deal with counterfactual explanation (31 out of 52 frameworks; 59.62%). In contrast, 17 out of 52 (32.69%) generate contrastive explanations. Only four studies (7.69%) fuse contrastive and counterfactual explanations. One of these studies [129] deals with both classification and regression. •Contfactuals for classification. A vast majority of state-of-the-art AI applications that generate contfactuals (42 out of 52; 80.77%) are used to explain the FIGURE 9. Numbers of frameworks grouped by AI problem with respect to the type of contfactual explanation generated. outcome of ML-based classifiers, i.e., algorithms that learn a mapping function f:X−→ Yfrom a training dataset of nlabeled examples X= {xi|1≤i≤n}to a discrete output variable (class) Y= {yj|1≤j≤m} where mis the number of classes. Indeed, contfactuals are particularly suitable for informing the end-user why a given data example is assigned a particular class label. Thus, the outlined classification-oriented frameworks are evaluated on classifiers based on logistic regression [55], [136], [153], [158], decision trees [46], [80], [122], [140], [150], [155], [159], gradient boosted decision trees [147], support vector machines [131], [138], [146], random forests [81], [86], [142]–[144], neural networks [6], [48], [49], [91], [129], [130], [133], [135], [139], [141], [145], [148], [151], or combinations of these [100], [105], [134], [152], [154], [160]. In three studies [67], [128], [137], the classifiers used in the experiments are not specified. •Contfactuals for regression. One of the classificationoriented frameworks [129] is extended to also handle the regression problem, i.e., learning a mapping function f from a training dataset Xto a continuous output variable Y. However, the continuous output is, in this case, subsequently converted to a lower-scale discrete value mapped to a textual description similar to that typical of a classification problem. The other frameworks addressing the regression problem aim to leverage gradientboosted decision trees [61] and indicate how large errors in regression tasks could be overcome [56]. •Contfactuals for knowledge engineering. The first of the considered frameworks (in chronological order) [93] offers explanations by reasoning abductively over the information extracted from a given knowledge base to answer a specific contrastive question. In this setting, an explanation is considered to be a consistent set of disjunctive literals for the explanation-seeking question. It is worth noting that the framework is not designed to provide explanations for ML algorithms. VOLUME 9, 2021 11987
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI TABLE 7. A classification of the contfactual explanation generators by explainability method. •Contfactuals for planning. Contfactual explanation generation appears highly relevant to sequential tasks in robotics such as automatic planning [94], [157], [161]. Moreover, some of the robotics-related frameworks found in reinforcement learning settings provide explanations for policies that a robot selects at a given time step [132], [156]. •Contfactuals for recommendation. Ghazimatin et al. propose a graph-based recommendation system in a counterfactual setup [84]. They obtain a counterfactual explanation by removing a minimal set of user actions so that the output recommendation changes. •Contfactuals for conflict resolution. Mosca et al. introduce an argumentation-based framework for social network management [149]. They use contrastive explanations to answer critical questions about agent actions in the context of multi-user privacy conflict. 2) EXPLAINABILITY METHOD All the frameworks generating contfactual explanations can be classified by their explainability method as either modelspecific or model-agnostic. The former type of implementations is meant for explaining decisions of particular AI algorithms. The frameworks of the latter type generate explanations irrespective of the nature of the underlying algorithm. Table 7presents the publications under study grouped in terms of the explainability method that they apply. The distribution of model-specific and model-agnostic explainability methods for generation of different types of contfactual explanation is shown in Fig. 10. Most frameworks deal with counterfactual model-agnostic methods. •Model-specific contfactual explanation generators. Several model-specific frameworks generate counterfactuals to explain the output of decision trees [46], [80], [122], [155]. For instance, Fernández et al. [80] present a recursive algorithm which extracts counterfactuals in the form of contrast-class decision tree nodes. The relevance of the generated counterfactuals is then measured by calculating a variant of the Gower distance. The proposed metric penalizes the number of feature changes when traversing the tree so that sparsity is promoted. Alternatively, Sokol and Flash rely on the Manhattan distance measuring leaf-to-leaf distance in the tree to retrieve the most relevant counterfactuals [46], [155]. Designed specifically for decision trees, their ‘‘Glass-box’’ frame- FIGURE 10. Numbers of frameworks grouped by explainability method with respect to the type of contfactual explanation generated. work is argued to be easily extendable to capture the output of other logical (rule-based) models. Aguilar- Palacios et al. generate contrastive explanations using gradient boosted decision trees to forecast promotional sales [61]. The researchers make use of the weighted Euclidean distance to present the forecast as a contrast to the neighbouring vectorized promotions. Stepin et al. retrieve counterfactuals from a rule matrix where each rule is encoded in terms of all possible feature values [122]. Subsequently, the generated counterfactuals are ranked using a XOR-based distance to find the most relevant counterfactual pertinent to the given contrast class. This method is further extended to generating counterfactuals for fuzzy decision trees. A number of frameworks address specific properties of counterfactuals. Thus, Ustun et al. tackle the problem of actionability, i.e., constraining the generated counterfactuals in such a manner that the imposed changes ‘‘do not alter immutable features’’ and that they ‘‘do not alter mutable features in an infeasible way’’ [158]. To approach this problem, a mixed integer programming method is employed. Russell et al. adopt a similar approach to encompass continuous and discrete variables as well as the combination of the two [153]. The main focus of the work is however placed on assessing coherence and diversity of generated counterfactuals. In order to guarantee the coherence of the counterfactual data example used for explanation, an integer programming-based method is proposed. In addition, the generated counterfactual explanations are claimed to be diverse, as diversity constraints are applied iteratively to a set of candidate counterfactuals. However, this framework is limited to: (1) explaining predictions of only linear classifiers and (2) a simple structure of the textual explanation template. A large number of frameworks are limited to explaining the output of particular models due to task-specific constraints. For instance, several explanation generators address computer vision tasks. Hendricks et al. bind 11988 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI an input visual image with a paired textual counterfactual explanation generated by a recurrent neural network [86]. In their framework, a number of candidate explanations (both image-relevant and non-relevant) are generated, paired, and ranked. The best counterfactual explanation is then selected to be the most class-specific to the counterfactual image while being the most relevant to the input image. Goyal et al. argue that their model is more faithful by design, as it generates visual explanations directly from ‘‘the target model based on the receptive field of the model’s neurons’’ [139]. Two model-specific frameworks are found in the context of video processing. For instance, Akula et al. [128] present an empirical study where input video frames are paired with the corresponding AND-OR graphs, i.e., compositional recursively defined graphbased knowledge representations capturing contextual information. The explanations based on such graphs are passed on to human subjects to evaluate the contrastive answers to the predefined questions. Alternatively, Kanehira et al. train a post-hoc explanatory model to justify a video classifier’s output [91]. A counterfactual explanation is, in this case, dependent on how likely a selected region in the given frame is classified positive and not negative, hence all such regions are scored and normalized. In accordance with the findings in the previous Section IV-C1, contfactual explanations have a great potential for automatic planning-related tasks. Most explanation generators meant for planning-based tasks are model-specific due to the problem- and approachspecific restrictions preventing them from being used for other AI challenges. For instance, Kim et al. employ a Bayesian probabilistic model for generating contrastive explanations [94]. Thus, the framework operates on a pair of plan traces defined in terms of linear temporal logic templates. The problem of obtaining contrastive explanations is designed as a Bayesian inference problem, with the posterior distribution to be maximized defined as the probability of a contrastive explanation given a set of positive and negative plan traces. Conversely, Sreedharan et al. consider the task of automatic analysis of counterfactual explanations in their ‘‘Hierarchical Expertise-Level Modeling’’ framework [156]. A robot provides a user with a plan for the next action to take. Then, the robot expects the user to respond with a set of foils. The robot’s task is then to convincingly refute the foils by offering a minimal explanation for why the foils are not acceptable under the given circumstances. In addition, Chakraborti et al. formulate the multi-model planning problem as a tuple consisting of the planner’s model of the problem and the corresponding human approximation thereof [132]. As plan explicability is reformulated in terms of its comprehensibility by an end-user. The robot’s model is adapted to the updates of human’s model of the problem. The problem of contrastive explanation generation for planning is also found to be framed in the reinforcement learning setting. For instance, Sukkerd et al. formulate the planning problem as the shortest stochastic path problem and develop the corresponding problem solver to obtain a contrastive explanation [157]. Hence, their objective is to find an optimal policy ‘‘that minimizes the expected cumulative cost of reaching a goal state over all closed policies’’. The explanation is believed to justify the rejection of the policies alternative to the optimal one. In addition, Zhao and Sukkerd explain an autonomous system’s behaviour modeling it as a Markov decision process [161]. Thus, a contrastive explanation is presented as a product of the analysis of the optimal policy at the next time step and an opposing policy on the basis of the objective values. •Model-agnostic contfactual explanation generators. A large number of model-agnostic frameworks treat contfactual explanation generation as an optimization problem in a post-hoc manner. Wachter et al. design a generic counterfactual explanation framework to find the closest point to the test data example [6]. Fixing the optimal set of weights of a trained classifier, the objective function minimizes the distance between the nearest data points of opposing classes. Note that counterfactual data points can be synthesized artificially. The researchers suggest the use of the Manhattan distance weighted by the inverse median absolute deviation to calculate the proximity of a counterfactual to the input data example. Another case of counterfactual explanation generation regarded as an optimization problem is the ‘‘Constrained Adversarial Examples’’ framework [148]. Adversarial examples that could serve as the basis for the counterfactual explanation of the output of deep learning models are searched for with the aim of minimizing the loss with respect to the attributes (features) between the original and counterfactual data examples. The researchers attempt to find the best counterfactual explanation by minimizing the number of attributes changed. Furthermore, the gradient direction is constrained to ensure the ethical adequacy of the explanation generated. Dandl et al. [134] formulate counterfactual search as a multiobjective optimization problem using a distance metric for mixed feature spaces aiming to obtain sparse and most plausible counterfactuals. Labaien et al. generate contrastive explanations for time-series data [141]. The explanation generation is considered a two-fold optimization problem of finding pertinent positives and negatives. Pawelczyk et al. make use of an autoencoder architecture for a pretrained classifier performing counterfactual search in the nearest neighbor style [151]. Model-agnostic frameworks are largely found to use decision trees as part of the reasoning mechanism instead of explaining their output. In contrast to the model-specific frameworks operating on decision tree output, Guidotti et al. employ decision trees as part VOLUME 9, 2021 11989
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI of reconstructing the reasoning behind any arbitrary classifier in a post-hoc fashion [140]. In their ‘‘Local Rule-based Explanation’’ framework, they generate a local neighbourhood for the given pre-classified data example using a genetic algorithm and subsequently train a decision tree on that newly obtained dataset to select a minimally distant foil within that local neighbourhood. Similarly, van der Waa et al. randomly sample or generate a data set in the neighbourhood local to the data point in question [159]. A decision tree is then trained to select the foil based on the minimum number of nodes between the original data point and the candidate foils. Furthermore, their ‘‘Foil- Trees’’ framework provides the methodological basis for perceptual-level contrastive explanation generation within the ‘‘Perceptual-Cognitive Explanation’’ framework [150]. Subsequently, the generated contrastive explanations are attributed to a specific group of users by means of ontology engineering at the cognitive level of the framework to make the explanations adaptive. In contrast, Martens and Provost argue that decision trees are an inadequate tool for representing, e.g., large documents [146]. Hence, they suggest the model-agnostic ‘‘Search for Explanations for Document Classification’’ algorithm for retrieving counterfactual explanations. However, it is only directly applicable to binary linear classifiers, whereas heuristics are proposed for nonlinear models. Several model-agnostic frameworks aim at measuring specific properties of contfactuals. Anjomshoae et al. [129] focus on contrastive explanations that maximize contextual importance and contextual utility. On the one hand, contextual importance measures the extent to which the input feature values affect the black-box algorithm’s output. On the other hand, contextual utility testifies how favorable the values of the selected features are for a given decision. Thus, the context-based values are calculated for each feature used by a black-box model observing the changes in the output as the input varies across the range of all possible input values. Being based on model-agnostic and problem-independent concepts, this framework is shown to be universally applicable to various classification and regression algorithms. However, the scalability of such an algorithm is limited to the use-cases operating on a small number of features. A similar limitation is observed due to possibly high variability of the input. Laugel et al. raise the issue of justification for counterfactual explanation [144]. They argue that a synthesized counterfactual data point must be connected to the training data. Counterfactuals are selected from a local neighbourhood circling around the test example with the radius of the distance to the closest correctly predicted data point of a contrast-class. The candidate counterfactuals are then clustered, as the initial local neighbourhood is updated to become a more extensive hyperspherical layer, until it can no longer be extended. Laugel et al. [100] enhance the work on justified counterfactual explanations. They argue that the distance from the test instance to a counterfactual does not sufficiently measure counterfactual’s relevance, as the counterfactual in question may appear disconnected from the ground-truth data. Thus, a counterfactual is deemed justified if it can be connected to an associated ground-truth data instance without crossing the decision boundary. Fernández et al. introduce the notion of counterfactual sets to enhance counterfactual diversity [81]. They explain random forest predictions by fusing different tree predictors so that the resulting counterfactual set contains the most relevant counterfactual. The other neighboring counterfactuals serve to diversify the output explanation. Mothilal et al. are also concerned with counterfactual diversity [105]. They design a loss function with a diversity metric over the generated counterfactuals to provide end-users with multiple relevant counterfactual explanations. Kusner et al. propose a causal model to assess the so-called ‘‘counterfactual fairness’’ [55]. It is worth noting that counterfactuals are presented in the form of conditional distributions and not structural equations despite the fact that the causal model employed follows Pearl’s formalism [37]. Similarly to the model-specific frameworks, numerous model-agnostic explanation generators are found to be task-specific. In computer vision-related classification tasks, Chang et al. find the smallest region in the image whose substitution would change the classifier’s prediction [133]. They employ a generative model to construct a saliency map while masking the other regions of the input image. Similarly, Dhurandhar et al. address an optimization problem over a perturbation variable to produce a contrastive explanation for the image classification task [135]. However, the proximity of the selected counterexample to the test point is, in this case, guaranteed by using an autoencoder. 3) OUTPUT REPRESENTATION The considered frameworks output contfactual explanations in several ways. Depending on the problem considered, contfactuals are presented in the form of: (1) intervals or specific values of the appropriate feature values whose alteration would have changed the output (i.e., numerical or featurebased output); (2) single- or multiple-sentence coherent text (i.e., linguistic output); (3) specific regions in the input image (i.e., visual output); or (4) a multi-modal combination of (some of) the above (see Table 8). As depicted in Fig. 11, most frameworks focus on numerical counterfactual output. •Numerical (feature-based) contfactual explanation. Numerical values (or intervals of values) associated to the most relevant features usually explain the behavior of AI algorithms. They can be represented as logical formulas [67], [93], [94] or in tabular form reflecting necessary changes to affect the decision [55], [56], [61], 11990 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI TABLE 8. A classification of the contfactual explanation frameworks by output representation. FIGURE 11. Numbers of frameworks grouped by output representation with respect to the type of contfactual explanation generated. [81], [105], [136], [147], [148], [151], [154], [158], [160]. They can be extracted from interpretable featurevalue pairs as a result of pruning in the search space [132]. In addition, they can replicate the internals of the classifier’s structure, e.g., in the form of decision tree nodes or rules [80], [140], [152]. •Linguistic contfactual explanation is a piece of grammatical single- or multiple-sentence text in natural language. Single-sentence textual explanations combine a textual description with explicitly stated numerical feature values [6]. Such explanations suggest featurevalue based instructions [48], [122], [131] or alternative actions for a possible output change [84], [149]. They also answer end-user’s inquiries with respect to the automatic decision in question [128], [159]. In contrast, multiple-sentence explanations provide end-users with specific details for the given decision [153], [156], [157], [161] or explain multiple decisions at once [146]. •Visual contfactual explanation. On the one hand, visual explanations for non-visual input data (i.e., datasets containing continuous or categorical featurevalue pairs) plot feature-value pair dependencies [134], [141], [142]. On the other hand, visual input data (i.e., images) are associated with saliency maps [133] or explained by contrastive patterns between the given data example and that of an opposing class in one iteration [130], [144] or a series thereof [49]; by depicting critical regions absent in the input data example that determine what lacks in the image to be classified differently [135]; or by visualizing spatial regions associated to data examples of opposed classes [139]. •Multi-modal contfactual explanation is a combination of numerical and/or linguistic and/or visual explanations. Multi-modal explanations are claimed to enhance human-robot interaction [46]. They are often selected to be the most appropriate where the problem addressed is concerned with pairing a computer vision problem with a natural language processing task such as object detection and language grounding [86]. In addition, visual-linguistic explanations identify counterfactuality in videos [91] and allow for dialogic interaction [150]. Such hybrid counterfactuals (in terms of their output representation) may as well complement each other while addressing the same task. For instance, Gomez et al. visualize the generated explanations in the form of bar plots combining them with explicitly stated numerical values [138]. Alternatively, Liu et al. combine feature importance bar plots with visual input and output [145]. While a textual explanation summarizes the degree of importance of the selected features, a visual explanation may present contextual in-method metrics that justify the classifier’s reasoning [129]. Contfactual explanations, as a mixture of tabular and visual output representations, appear also in an augmented reality framework [137]. Nevertheless, explanations of different modalities are not necessarily merged. To ensure the universality of the proposed approaches, specific feature values are presented for tasks with datasets containing only continuous features, i.e., where the same method is used to output images for a handwritten digit classification problem [143]. Finally, interaction with users can be enhanced by means of voice-based explanations combined with textual explanations [46], [155]. 4) EVALUATION METHOD Evaluation of generated contfactual explanations is an issue of main concern. Unfortunately, despite an increasingly expanding use of contfactual explanations, no uniform set of evaluation methods has been adopted so far. Hence, it is worth taking a look at evaluation methods from other generationoriented sub-areas of AI. For instance, it is common to distinguish between intrinsic and extrinsic evaluation methods in natural language generation [163]. Intrinsic evaluation implies assessing the performance of a natural language generation system (or its modules) as an isolated unit. In contrast, extrinsic (task-based) methods are designed to estimate how successfully the system performs with respect to an external task. In addition, Gatt and Krahmer make a distinction between ‘‘objective’’ (automatic, corpus-based) and ‘‘subjective’’ (human judgements) metrics [164]. Objective metrics include (but are not limited to) precision- and/or recalloriented scores, number of insertions/deletions/substitutions, VOLUME 9, 2021 11991
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI TABLE 9. A classification of the contfactual explanation generators by evaluation method. FIGURE 12. Numbers of frameworks grouped by evaluation method with respect to the type of contfactual explanation generated. etc. In turn, subjective metrics measure readability, accuracy, relevance of the generated text, as perceived by humans. Thanks to their methodological universality, they can be extrapolated to other (non-linguistic) modalities of generated explanations (e.g., numerical, visual, or multi-modal) and are therefore used to form the basis of the evaluation method classification in this review. It is worth noting that all the considered frameworks are evaluated by means of intrinsic (either subjective or objective) metrics. Hence, we only make a clear distinction between subjective and objective evaluation methods in this study (see Table 9for details). A distinction between the use of different types of evaluation metrics can be seen in Fig. 12. It is easy to appreciate how most frameworks deal with objective evaluation of counterfactuals. Let us give further details below, regarding the four groups of publications in Table 9. •No evaluation details provided. 17 out of 52 (32.69%) of the considered publications do not evaluate their frameworks, i.e., neither automatic metrics for contfactual explanation generation are suggested nor a human evaluation survey is presented in such publications. However, whereas certain publications do not provide any specific evaluation method, some do stress that human evaluation should be encouraged to estimate the quality and effectiveness of the generated counterfactuals [46], [150], [153]. •Subjective evaluation. The subjective methods include human preferences for certain types of contfactual explanation over others. For instance, Akula et al. show that that contrastive explanation-seeking questions are in general better answered by means of contfactual explanations [128]. They classify contrastive questions in the following 10 categories suggesting the template questions for arbitrary objects x,x1, and x2(all being of some class X) and y,y1, and y2(all being of some other class Y): –WH-X: ‘‘Why xrather than not x?’’; –WH-X-NOT-Y: ‘‘Why xrather than y?’’; –WH-X1-NOT-X2: ‘‘Why x1rather than x2?’’; –WH-NOT-Y: ‘‘Why not y?’’; –NOT-X: ‘‘Is it xrather than not x?’’; –NOT-X1-BUT-X2: ‘‘Is it x1rather than x2?’’; –NOT-X-BUT-Y: ‘‘Is it xrather than y?’’; –DO-X-NOT-Y: ‘‘What if it is xrather than y?’’; –DO-NOT-X: ‘‘What if it is not x?’’. –DO-X1-NOT-X2: ‘‘What if it is x1and not x2?’’ It is worth noting that 6 out of 10 question types (WH-NOT-Y, NOT-X, NOT-X1-BUT-X2, NOT-X-BUT- Y, DO-NOT-X, and DO-X1-NOT-X2) matched with automatically generated contfactuals are shown to be highly preferred to factual explanations. In addition, Ferrario et al. propose an augmented realitybased setting to favor interactivity and facilitate explaining ML algorithm output to non-experts [137]. However, the two aforementioned studies [128], [137] lack an evaluation of the quality of the generated contfactual explanations themselves. Lucic et al. asked 75 subjects to judge interpretability, actionability, and trustworthiness of the generated contfactual explanations [56]. They concluded contfactual explanations are highly interpretable and actionable. In addition, they help users understand why the model makes large errors while solving a regression problem but do not support users’ trust in the model’s output. In addition, Hendricks et al. provide results of human evaluation for the generated explanations [86]. However, these only include evaluations for factual explanations and are therefore excluded from the taxonomy group being discussed. •Objective evaluation. A majority of the researchers propose objective (automatic) methods for evaluating automatically generated contfactuals. A number of the frameworks are evaluated by means of accuracy-based metrics [61], [91], [94], [159]. Kanehira et al. propose one accuracy-based evaluation metric for visual and linguistic explanations each: negative class accuracy and concept accuracy, respectively [91]. Negative class accuracy estimates the quality of the visual explanation as the ratio of the probability of the contrast class after the image region in question is masked out. In turn, concept accuracy estimates how compatible the output linguistic explanation is to its visual counterpart. It is calculated as the intersection over union between a given region 11992 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI and all bounding boxes in the image. Kim et al. define the domain-specific accuracy for the automatic planning problem of unique contrastive explanations as a sum of the number of traces in a positive set of traces satisfying a contraint for a contrastive explanation and those of the negative set where the constraint is unsatisfied over all plan traces [94]. The consistency of output explanations is otherwise shown by measuring their accuracy on the basis of mean error, mean absolute error, and mean absolute percentage error [61]. An extensive number of evaluation methods are found to be strictly task- or approach-specific. Hendricks et al. measure word detection (i.e., which words are not image relevant by holding out one word at a time from the sentence to determine the least relevant word in the explanation) and word correction (i.e., a number of replacements of the foiled word with words from a set of target words) [86]. Similarly, Martens and Provost estimate explanation complexity by calculating the average number of words in the shortest explanation and problem complexity according to the overall number of generated explanations [146]. Fernández et al. evaluate the relevance of the generated counterfactuals measuring a Gower distance-based metric in comparison with the number of feature changes and minimum distance (in terms of decision tree nodes) between the leaves in the given decision tree classifier [80]. Van der Waa et al. evaluate the generated explanations by means of such modelspecific metrics as the average length of the explanation in terms of decision tree nodes and the F1-score of the foil-tree on the test set compared to the model’s output [159]. Kusner et al. estimate counterfactual fairness on the basis of the density of the predicted data for their causal models [55]. Laugel et al. claim that understandability of the generated explanations can be estimated by means of their sparsity defined as the number of nonzero coordinates of the explanation vector [143]. Moore et al. measure the number of solutions, the distances to the nearest training set data points, and the transferability of the generated counterfactuals to other datasets and classifiers [148]. Sreedharan et al. calculate the number of predicates that are used to generate the model lattice [156]. Similarly, Chakraborti et al. calculate the number of nodes in the search space remaining after pruning [132]. Goyal et al. report how often the discriminative regions lie inside the test data example segmentations as well as relevant specific key regions [139]. Labaien et al. calculate the number of changes to switch from the original to the selected contrastive sample following the dataset constraints [141]. To estimate faithfulness of the generated counterfactuals, Pawelczyk et al. suggest calculating the so-called degree of difficulty of a counterfactual suggestion to measure how costly it is to achieve the state of the given suggestion [151]. Aiming to provide realistic counterfactuals, Sharma et al. introduce the counterfactual explanation robustness-based score defined as the expected distance between the input instances and their corresponding counterfactuals [154]. In addition, the generated counterfactuals are inspected in terms of fairness which is calculated as the expected distance between the input and a counterfactual over distinct values for a specified feature set. Merrick and Taly evaluate output explanations in terms of mean feature attributions to show the importance of relevant references [147]. Gomez et al. evaluate counterfactuals in terms of data distribution, feature importance, as well as possible and actionable changes to the input [138]. Dandl et al. use the hypervolume indicator metric to estimate the quality of the estimated Pareto front during counterfactual search [134]. In addition, Chang et al. measure the weakly supervised localization error for an image detection task – the intersection-over-union ratio over 0.5 with any of the ground truth bounding boxes and the saliency metric, i.e., ‘‘the log ratio between the bounding box area and the in-class classifier probability after upscaling’’ [133]. Several metrics can be extended to be applied to other approaches. Lash et al. estimate how much the probability of a given prediction reduces given a feature perturbation as determined by a contrastive explanation [142]. Dhurandhar et al. employ the concept of pertinent positives (i.e., ‘‘factors whose presence is minimally sufficient in justifying the final classification’’ [135]) and pertinent negatives (i.e., ‘‘factors whose absence is necessary in asserting the final classification’’) to evaluate factuals and counterfactuals, respectively, for a given classification task. Both types of evaluation methods highlight the features supporting evidence as formulated in the contrastive explanation on the basis of the values that a perturbation variable takes on. Fernández et al. evaluate counterfactuals in terms of the average of the pairwise distances based on the feature type and the percentage of valid counterfactuals [81]. Mothilal et al. stress that counterfactuals should be evaluated in terms of validity (i.e., whether a generated counterfactual really leads to a different outcome), proximity (i.e., feature-wise distance between the original and counterfactual samples), sparsity (i.e., number of features differing in the original and counterfactual samples), and diversity (i.e., feature-wise distance between each pair of counterfactuals) [105]. Similarly, Stepin et al. calculate factual and counterfactual explanation length to estimate conciseness of the generated explanations [122]. They also compute the number of counterfactuals and their best minimal distance to the factual explanation to assess the relevance of counterfactuals. Rajapaksha et al. consider coverage (as an indicator of representativeness of a rule for a given dataset), confidence (i.e., the percentage of instances in the dataset which contain the consequent and antecedent together over the number of instances which only contain the antecedent), lift (i.e., an association between antecedent VOLUME 9, 2021 11993
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI and consequent), leverage (i.e., the observed frequency between the antecedent and consequent), and the number of features in explanation for evaluating their framework against other rule-based methods [152]. Also, White and Garcez reintroduce fidelity to the underlying classifier on the basis of distance to the decision boundary [160]. In addition, some of the model-agnostic frameworks [140], [159] allow for measuring how well the output of black boxes (i.e., actual output to be explained) and grey boxes (i.e. interpretable intermediate predictors) mimic the local neighbourhood (i.e., fidelity) and the data example to be explained (i.e., hit). Laugel et al. measure how justified counterfactuals are by averaging a binary score (one if the explanation is justified following the proposed definition, zero otherwise) over all the generated explanations [100], [144]. It is worth noting that the run-time of explanation generation algorithms is reported in addition to the evaluation metrics for several frameworks [132], [139], [146], [152], [156], [159]. •Hybrid evaluation. Two frameworks are evaluated in terms of both automatic metrics and human judgments. Ghazimatin et al. calculate explanation length to discuss comprehensiveness of explanations as well as estimate their usefulness and credibility by surveying 500 subjects [84]. In addition, Le et al. compute fidelity, conciseness, information gain, and influence [48]. Automatic metrics are complemented with a user study on intuitiveness, friendliness, comprehensibility, and understandability of generated explanations. D. LINKS BETWEEN THEORETICAL AND PRACTICAL CONTRIBUTIONS TO CONTFACTUAL EXPLANATION GENERATION (ANSWER TO RQ3) We find that only few of the existing computational frameworks are grounded on theories of contfactual explanation. Indeed, only 13 out of 113 studies (11.50%) were present in both of the pools of primary studies related to RQ1and RQ2. Table 10 summarizes the characteristics of such theoretically grounded contfactual explanation generation frameworks. Moreover, only 3 out of the 13 (23.08%) studies interpret the insights from the theoretical foundations to propose their own contfactual explanation definition for problem-oriented purposes. Kean states that ‘‘explanation in artificial intelligence is based on the inference of deduction’’ [93]. He complements a deductive evidence-based explanation with a redefined abductive contrastive explanation drawing parallels to the ‘‘inference to the best explanation’’ [103]. He models Lipton’s theoretical framework distinguishing two types of contrastive explanation: non-preclusive (i.e., non-restrictive) and preclusive. The key aspect distinguishing the two types of contrastive explanation is in regard to how a model explains the contrast given an explanation-seeking question. Thus, a non-preclusive contrastive explanation is ‘‘irrelevant to the model of explaining the contrast’’ being ‘‘necessary in the model of explaining the question’’. On the contrary, a preclusive contrastive explanation is assumed to be restricted by a negated model of the contrast. Aguilar-Palacios et al. [61] refer to Lipton’s definition of contrastive explanation [12]. Referring to Pearl [37], Bertossi redefines causal explanation in the context of XAI to be ‘‘a set of feature values for the entity under classification that is most responsible for the outcome’’ [67]. The rest of works redefine contfactual explanation on the basis of the problem-specific constraints without explicitly referring to the theoretical foundations described in Section IV-B. Driven by the task of automatic planning, Kim et al. define a contrastive explanation to be a constraint satisfied by a specific set plan traces [94]. Fernández et al. define a counterfactual to be a set of feature changes that turn the given data example to be classified differently [80]. Whereas this definition is applicable to the classification problem in general, the applicability of the framework is restricted to decision trees only. Similarly, Hendricks et al. explain visual concepts for the image classification task on the basis of the so-called counterfactual evidence (i.e., an attribute discriminative enough for another class of objects in the image absent in the given image) [86]. Ghazimatin et al. [84] define a counterfactual on the basis of their model’s internal structure: an explanation is deemed counterfactual if after removing the edges from the recommendation graph, the user receives a different top-ranked recommendation. In addition, Kanehira et al. only specify the linguistic form of a counterfactual explanation without defining it explicitly [91]. Finally, there is a number of marginal interpretations of contfactuals among the RQ3-related studies. Laugel et al. denote a counterfactual as a specific data instance that changes the algorithm’s prediction [100]. Poyiadzi et al. denote a counterfactual to be ‘‘the new state of the object’’ [49]. Nevertheless, the most commonly acceptable definition of a contfactual in the observed RQ3-related studies states that a contfactual explanation is a set of minimal feature modifications that makes the model change the prediction [81], [105], [122]. V. DISCUSSION The findings show that a large body of research has been elaborated on theoretical accounts of contrastive, counterfactual, and contrastive-counterfactual explanation. In addition, the topic has recently attracted attention from researchers in XAI (see Fig. 13). Thus, 50 out of the 52 considered state-of-the-art contfactual explanation generation frameworks (96.15%; RQ2) have been developed from 2017 to 2020. The results of the study, in relation to RQ1, show that a majority of the considered theoretical accounts of contfactual explanation (49 out of 74; 66.22%) speculate on the causal nature of explanation. However, whereas most researchers in philosophy of science have mainly used the concept of counterfactuality to explain causal relations between entities in question, causal inference is poorly addressed in the 11994 VOLUME 9, 2021
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI TABLE 10. A summary of characteristics of theoretically grounded computational frameworks for contfactual explanation generation. CT stands for contrastive explanation, CF means counterfactual explanation, and CT-CF is contrastive-counterfactual explanation. FIGURE 13. Numbers of theoretical and computational contfactual explanation generation frameworks grouped by year of publication. For illustrative purposes, only the studies published from January 1990 to September 2020 are displayed. pool of publications concerning computational frameworks of contfactual explanation generation. Kean directly refers to a causal account of contrastive explanation to address the problem of abductive reasoning [93]. In addition, Lucic et al. [56] explicitly specify that their method is based on previous work on philosophical accounts of contrastive explanation [12] as well as on causal attribution [165], [166]. Kusner et al. [55] make use of causal inference models and the corresponding tools provided by Pearl [37]. They assess how discriminatory the generated counterfactual explanations are for the given classification task output. On the other hand, Bertossi redefines the concept of causal explanation [67]. Following a causal account of contfactual explanation, Fernández et al. introduce weakly causal irreducible counterfactual explanation [136]. As most of the current ML-tasks are centered around singling meaningful patterns out from unstructured data, establishing causal relations appears to be among the AI problems that are yet to attract global attention. This partly explains why most of the modern contfactual explanation generators focus on feature perturbation when searching for the most relevant contfactuals and not establishing causal relations between them. At the same time, the other computational frameworks are primarily non-causal. Furthermore, a strikingly low number of such frameworks appear to be rooted in theoretical accounts of explanation due to an imbalance in favor of causality-oriented theoretical accounts. However, the amount of publications for RQ3may be somewhat misleading, as contfactual explanations are often redefined without specifically referring to theoretical contributors in explanation. This is VOLUME 9, 2021 11995
I. Stepin et al.: Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable AI hypothesized to be due to the problem-specific necessities ignored in previous theoretical works from different branches of science. For instance, Laugel et al. emphasize that the minimal perturbations required to change the predicted class of a given observation enable a user to understand which features locally impact the prediction and therefore how it can be changed [144]. In this interpretation, counterfactual explanations are conceptually most similar to counterfactuals as defined by Lewis [36]. Indeed, while the concept of ‘‘the closest possible world’’ does not always appear in the related publications, it turns out to be implicitly wired in almost all works. Instead, other considered frameworks do not appeal directly to any of the theoretical accounts of explanation addressed in RQ1. Hence, it is worthwhile taking a look at how contrastive and counterfactual explanations are redefined in the frameworks not appearing in Section IV-D. In general, a consensus among researchers has been observed on how contfactual explanations are defined irrespective of the theoretical framework proposed by individual authors. In humanities and social sciences, a major difference in various contfactual theories of explanation is observed to concern the causal nature of explanation and its extrapolation to non-causal cases. In computer science and AI, notions of counterfactual and contrastive explanation are found to be the most dissimilar when applied to non-overlapping problems. Nevertheless, the corresponding line of research in AI makes little use of the rich theoretical background accumulated by now. While some rule-based approaches used in expert systems are justified theoretically (e.g., see [93]), newly emerging tasks present novel challenges for theorists and call for updating the theories developed so far. More precisely, ML-specific contfactual explanations are designed to answer the question: ‘‘Why was the outcome Y observed instead of Y0?’’ [148]. Anjomshoae et al. define finding a contrastive explanation as ‘‘contrasting instance against the instance of interest’’ [129]. Fernández et al. specify that a counterfactual is generally regarded as a hypothetical instance similar to an example whose explanation is of interest but with a different predicted class [80]. Also, counterfactual explanations ‘‘show a difference in a particular scenario that causes an algorithm to change its mind’’ [155]. As a majority of the considered frameworks are designed for tackling classification problems, contfactual explanations operate on the notion of a contrast-class (e.g., see [155]) answering the question: ‘‘How is the prediction altered when the observation changes, given a classifier and an observation?’’ Furthermore, these changes are normally expected to be minimal [143]. However, certain application domains as well as the selection of a classifier require researchers to redefine contfactuals imposing task-dependent constraints, which makes it nearly impossible to connect them to any of the existing theories of contfactual explanation. For instance, Martens and Provost define a contrastive explanation for a document classification task to be a minimal set of words such that removing all words within this set from the document changes the predicted class from the class of interest [146]. In addition, Guidotti et al. reformulate a counterfactual to be a set of split conditions of a decision tree describing the minimal number of changes in the feature values of a test example [140]. In image classification, it is found necessary to detect specific regions in the given test image. For this type of tasks, the contrastive explanationseeking question is formulated as follows: ‘‘Which parts of the image, if they were not seen by the classifier, would most change its decision? or which inputs, when replaced by an uninformative reference value, maximally change the classifier output?’’ [133], [139]. Similarly, Dhurandhar et al. ask what should be minimally and necessarily present and absent in the given image to justify its classification [135]. Alternatively, counterfactuals are viewed as ‘‘solutions that are guaranteed to map back onto the underlying data structure’’ [153]. Redefined contrastive explanations are also found in the domain of robotics and automatic planning. According to Sukkerd et al., a contrastive explanation answers the question why a generated behavior is optimal with respect to the planning objectives of an autonomous system [157]. Alternatively, contrastive explanations are used to answer why-not questions about the system’s behavior in which the consequences of the counterfactuals in question are pointed out [161]. In addition, the nature of the explanation-seeking questions for computational frameworks deserves further discussion. Sokol and Flash distinguish three types of counterfactual explanations: (1) a plain counterfactual (‘‘Why?’’) generated as the shortest possible class-contrastive counterfactual; (2) a counterfactual explanation not conditioned on the indicated feature(s) (‘‘Why despite?’’); and (3) a (partially) fixed counterfactual explanation (‘‘Why given?’’) which is conditioned on a predetermined set of features [46]. Hilton proposes different types of contrastive questions such as: (1) ‘‘Why X rather than not X?’’; (2) ‘‘Why X rather than the default value for X?’’ and (3) ‘‘Why X rather than Y?’’ [166]. Following this distinction, Akula et al. extend this set of contrastive questions to formulate ten contrastive question types for counterfactual explanation generation [128] (see Section IV- C4). Alternatively, only linguistic templates for such explanations are defined without any theoretical grounding in accordance with any accounts described in Section IV-B. For instance, Sokol and Flash define a counterfactual explanation to be a piece of text following the template: ‘‘The prediction is hpredictioni. Had a small subset of features been different hfoili, the prediction would have been hcounterfactual predictioniinstead’’ [155]. Remarkably, a wide range of frameworks favor automatic evaluation methods. Thus, they rarely place the end-user in the center of the explanation evaluation process. However, we find an increasing number of interactive frameworks that attempt not only to present the automatically generated explanations to the end-user but also interact with him or her [46], [150], [155]. Promoting interactivity (e.g., by engaging the end-user to participate in an explanatory dialogue with the 11996 VOLUME 9, 2021
having them explained, legal regulations concerning data processing are becoming widely adopted, e.g. the General Data Protection Regulation (GDPR) in the European Union [33]. Moreover, a new European regulation on AI is in progress and highlights the importance of preserving the European values by promoting trustworthy and responsible human-centric AI [9,34]. The gap between obscurity of automatic decisions and their explainability can be overcome by using interpretable models [37]. Among all AI tools, such soft computing techniques as fuzzy sets and systems have been shown to be not only interpretable but also explainable [3]. Thus, two key advantages are distinguished when relating the properties of interpretability and explainability of fuzzy systems. First, their transparent (i.e., interpretable) structure allows for making unambiguous inferences of why the given output was produced. Second, the use of linguistic variables and rules enables such systems to be explainable, i.e., to produce comprehensible explanations in natural language. Nevertheless, the ability to demonstrate evidence on why specific output is produced (i.e., explain the factual output) may not be sufficient to display the underlying reasoning to the end user. Therefore, a factual explanation may need to be complemented with an explanation of why some other output was not produced. Opposed to factual explanations justifying the given prediction, counterfactual (CF) explanations (or counterfactuals) inform the end user about minimally different alterations to the input features for the outcome to change [41]. In the context of classification problems, CF explanations are typically designed as answers to the template question ‘‘Why was Ppredicted rather than Q?” where Pis the output (factual) class and Qis a non-predicted hypothesized alternative CF class [29]. CF explanation generation is often regarded as an optimization problem in search of the data point of another class which represents the closest data point alternative to the test instance in an n-dimensional Euclidean space [46]. In the context of fuzzy sets and systems, however, such minimal changes may be described not only by means of a continuous variable representing numerical feature values (which we call ‘‘quantitative CFs” in this paper) but also by a discrete linguistic variable whose values are linguistic terms (which we refer to as ‘‘qualitative CFs” in this paper). In the former case, distinctive (numerical) features point to specific values, which are minimally different from those the test instance has, that should be set for the outcome to change. In the latter case, linguistic terms represent sets of suitable CF feature values in form of text and conceal the underlying numerical intervals. The difference in end user’s perception of these types of CF explanations remains unclear [45]. On the one hand, it may be affected by peculiarities of the structure of explanation, such as the number of explanatory features or explanation length. On the other hand, user’s perception may be influenced by a degree of precision of the explanation content. Thus, qualitative CFs may be regarded as pieces of imprecise information which can facilitate understanding of the communicated explanation but may, however, be underinformative or even misleading to the end user. Conversely, quantitative CFs specify finegrained changes to values of features. Last but not least, existing metrics for measuring quality of CF explanations (e.g., validity, proximity, diversity, among others) are strongly related to the data used for explanation generation [31]. However, those metrics ignore perceptual skills of the explanation’s recipient and may not be sufficient for assessing the overall explanation effectiveness. In order to make another step towards human-centric AI, it therefore appears necessary to propose novel means of capturing and assessing human perception of explanations. As part of previous work [41], we introduced a method for generating qualitative CF explanations applied to decision trees (DT). Then, we generalized this method to fuzzy information granules [43]. In this paper, our contribution is fourfold. First, we extend our previous work with a generalized Euclidean distance-based metric for CF explanation generation which better grasps membership function values. Second, we propose a novel genetic-based quantitative CF explanation generation method. Third, we define a new metric for assessing the complexity of automated explanations. Fourth, we carefully validate both qualitative and quantitative CF explanations via human evaluation in agreement with the best known practices for fair and sound evaluation of Natural Language Generation (NLG) and analyze the findings in terms of explanation complexity as expected to be perceived by the end user. The rest of the manuscript is structured as follows. Section 2presents a brief overview of existing methods for quantitative and qualitative CF explanation generation. Section 3introduces our methods for generating CF explanations associated to fuzzy rule-based classification systems (FRBCS). Section 4describes the key characteristics of the experimental design for subsequent human evaluation studies. Section 5goes in detail with the analysis of the data collected in two evaluation surveys. Section 6discusses the findings and offers suggestions on how they can be exploited. Finally, we outline directions for future work and conclude in Section 7. 2. Related work CF explanation generation has in recent years attracted increasing attention from researchers in the AI field. As CFs oppose actual and potential outcomes, they are most widely used to explain the output of various classifiers, from linear machine learning models to deep neural networks [42]. Further, they are extensively found across different application domains. For example, CFs are found applicable in healthcare where they, e.g., serve to provide a patient with a bigger picture of the risk of developing diabetic retinopathy [26] or in banking where CFs suggest recommendations on necessary changes to have a loan application approved if previously rejected [16].In addition, CF explanations are as well extensively used in robotics (e.g., in planning – to justify the choice of a robot over other feasible but unfavored possible solutions [44]). Despite numerous potential application domains, the use of CFs is advised to be controled due to possible malicious implications. As I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 380
such, they have been misused or misinterpreted (what may lead to data breaches) in cases of, e.g. password masking or evoting [20]. Other privacy concerns include inferring sensitive patterns of the training data or manipulations with the revealed internals of the model [40]. In the context of qualitative CFs, a number of generation methods output CF sets to support diversity. For example, Sokol and Flash inspect the internal structure of DTs in their ‘‘Glass-Box” framework for generating CF sets [40]. Thus, the authors retrieve CF sets from the decision paths ranking them by their leaf-to-leaf distance to the actual prediction. On a similar note, Stepin et al. generate set-based (i.e., qualitative) CFs from either crisp or fuzzy DTs [41] but also regarding fuzzy information granules [43] while introducing an extra-linguistic layer to approximate numerical intervals or membership function values, respectively, using predefined linguistic terms. Whereas the aforementioned methods are model-specific, i.e., they only allow for explaining counterfactually the given output of the DT itself, DT-based approaches are also used for model-agnostic methods. In their LOcal Rule-based Explanation (LORE) method, Guidotti et al. employ a genetic algorithm to first synthesize a local neighborhood around the test instance which is subsequently used to train a DT and generate CF sets [17]. The collection of CF sets is then reconstructed from the decision paths. Then, the minimally different CF set is selected on the basis of the (minimal) number of Boolean split conditions of the DT that the given CF path does not satisfy. Maaroof et al. extend LORE to fuzzy logic-based applications by proposing Contextualised LORE for Fuzzy attributes (C-LORE-F) [26]. Alternatively to LORE, the researchers formulate a local neighborhood generation approach for solving the uniform cost search problem. Potential neighbors are generated by applying iterative changes over a single feature taking into account intersections between two corresponding fuzzy sets. Further, the authors propose to induce the rules instead of building up a DT using the Dominance-based Rough Set Approach (DRSA) where the decision rules take into consideration the preference directions of the input variables. In addition, Fernández et al. extract CF sets from a random forest classifier by partly fusing individual tree predictors [12]. Further, their Random Forest Optimal Counterfactual Set Extractor (RF-OCSE) prunes the search space of candidate CFs using the minimum observable approach to filter out CFs whose distance to the test instance exceeds the best up-to-now distance. On the other hand, quantitative (i.e., single-point-output) CF explanation generation methods address the optimization problem searching for an individual data point found to be minimally different from the test point under consideration in accordance with the selected distance function, e.g., Manhattan distance weighted by the inverse median absolute deviation [46]. Similarly, Moore et al. use a differentiable model on the basis of a gradient-based method over the cross entropy loss function to identify a single minimally distant CF data point [30]. Alternatively, genetic algorithms are also frequently used to generate CFs [39]. Model-agnostic genetic algorithms are used not only to generate a local neighborhood but also to identify a specific optimal CF data point. In addition to the standard genetic algorithm, Lash et al. apply local search to non-mutated children so that the best solution is preserved for the next generation [24]. Sharma et al. propose another approach called Counterfactual Explanations for Robustness, Transparency, Interpretability, and Fairness of Artificial Intelligence (CERTIFAI) where a genetic algorithm based on natural selection, mutation, and crossover appeals to user feedback (regarding feature mutation, feature range specification, and enquiries for a specific number of explanations) [39]. Whereas these user constraints allow for generating actionable human-centric explanations, imposing too severe restrictions may overreduce the search space resulting in generating null explanations. In addition, Schleich et al. make use of a complete search space in their GeCo framework [38]. Thus, the authors present a customizable genetic algorithm enhanced with two optimization techniques to reduce memory costs and running time. The compressed d-representation of the input features reduces the memory storage required for mutation-related calculations whereas the so-called partial evaluation optimizes the evaluation of the classifier, as static components of the classifier can be pre-evaluated using an equivalent sub-model of the same classifier [38]. Finally, both qualitative and quantitative generation methods are primarily evaluated with automatically computable metrics (e.g., fidelity, validity, proximity, or diversity) [12,17,31]. Unfortunately, empirical studies involving human evaluation for assessing the goodness of automated CFs are scarcely found in the literature. Baaj and Poli show that explanations based on the use of linguistic terms appear rather satisfactory and convincing despite being overly repetitive for a general audience [5]. Wang and Yin state that CFs increase understanding for users who have sufficient domain knowledge but fail to calibrate trust in the model [47]. Further, Lucic et al. demonstrate that CFs help users understand why a model makes large errors [25]. Olson et al. show that CFs can be also effective for non-expert users in the identification of flawed agents [32].In addition, Woodcock et al. stress that lay users trust CFs only if the information gap in the existing domain knowledge between them and expert users is not significant, specifically in the healthcare domain [48]. Nevertheless, unlike our work, none of the aforementioned studies contrasts the output of single-point-output quantitative generation methods and setbased qualitative ones. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 381
3. Explanation generation methods 3.1. Notation The methods proposed in this study address a multi-class classification problem, i.e., learning a mapping function h:X!Yfrom a dataset X¼x i fgj n i¼1 containing nlabeled instances to a discrete output variable (class) Y¼y j j m j¼1 where mis the number of classes. The dataset is characterized by the set of pnumerical 1 features F¼fk fgj p k¼1, which are mapped to the corresponding linguistic variables. By definition [49], each feature is a tuple fk¼Lfk;Tfk L;Ufk;Gfk;Mfk DE ;8fk2Fwhere Lfkis the name of the feature fk;Tfk L¼tfk l no js l¼1is the set of linguistic terms defined in the universe of discourse Ufk;Gfkand Mfkbeing syntactic and semantic rules, respectively. Let VT¼STfk L;8fk2Fdenote the set of all linguistic terms. In our experiments (see Sections 4 and 5), we aim to explain (both factually and counterfactually) the output of an FRBCS [23] which is defined by the following components: a knowledge base containing a set of input and output variables and a rule base which represents a set R¼r i w i ðÞ fgj jRj i¼1 of weighted fuzzy rules of the form r i w i ðÞ:IFL f 1 ist f 1 1 AND ... L f k ist f k k ... AND ... hi THENyISy i , where r i 2R;w i 20;1½is the rule weight (i.e., the higher w i the more relevant r i ), t f k k 2T f k L ;f k 2F;y i 2Y; a fuzzy processing structure containing fuzzification and defuzzification interfaces as well as a fuzzy reasoning mechanism. Given an input vector x¼x 1 ;...;x p and a rule r i 2R, its activation degree a i is computed as a i (x) ¼ l t f1 1 x 1 ðÞ... l t fk k x k ðÞ... l t fp p x p , being l t fk k x k ðÞthe membership degree of the value x k for the linguistic term t k associated to feature f k , and is a t-norm such as minimum or product. Any rule r i can be denoted as a tuple r i w i ðÞ¼AC i ;cq i hiwhere AC i is an antecedent (i.e., a non-empty set of feature-value pairs) and cq i is a consequent (i.e., a class label). The output class y FAC 2Ypredicted by an FRBCS is said to be the factual explanation class. All the rules from the rule base that lead to the predicted outcome form a set of factual explanation rules R FAC ¼S r j 2R r j jcq j ¼y FAC , being R FAC #R. Similarly, all the non-predicted classes form a set of CF classes, with a collection of the corresponding rules mapped to each of them: R CF ¼S r j 2R r j jcq r j ¼y CF no ;Y CF ¼y CF jy CF 2Yny FAC fg. Given an FRBCS s, a data instance x2X, and the classification output y FAC predicted by s, each class y j 2Yis associated with a single explanation of why xis classified in the given way. Hence, there exists only one factual explanation E FAC sð, x,y FAC Þ. In addition, there is a non-empty set of CF explanations E CF sð,x,Y CF Þ¼ S y CF 2Y CF E CF sð,x,y CF Þfor each non-predicted class y CF 2Y CF . Throughout the manuscript, we assume that the output is explained in its entirety if the corresponding explanation contains a factual explanation specifying why the given decision is made as well as jYj1 CF explanations indicating why all the alternative classification options are discarded. Therefore, a (full) explanation for a data instance x2Xis assumed to contain one factual explanation and a non-empty set of CF explanations: Es ð,x,YÞ¼E FAC s ð,x,y FAC Þ[E CF s ð,x,Y CF Þ. Accordingly, explanation generation methods aim to produce (1) a factual explanation for the test instance and (2) the most relevant CF explanations for all the CF classes. 3.2. Factual explanation generation We design the process of explanation generation to include three main stages (text planning, sentence planning, and surface text realization) as in the NLG pipeline proposed by Reiter and Dale [35]. We selected this NLG pipeline because it is by far the most commonly used in the scientific community [14]. It is worth noting that we apply the same NLG pipeline no matter if we consider either factual or CF explanations: Text planning, where the information to be conveyed in the text is identified (content determination), as well as some order and general structure of the text is planned. In the case of CF explanations, content determination relies on relevance estimation (as described in the next section). Sentence planning, which includes grouping of messages when needed (sentence aggregation) and decisions about the words/expressions to be used (referring expression generation and/or lexicalization). This stage is crucial to avoid repetitions and make the output text more natural. 1 The use of categorical features is out of the scope of this work. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 382
Surface text realization, which consists of generating a syntactically, morphologically, and orthographically correct text. This last stage is implemented using a pool of templates dynamically instantiated, populated and mixed with a Python wrapper of the SimpleNLG library [6]. Specifically, the factual explanation generation process presupposes the following steps: factual explanation rule selection, linguistic approximation of the feature values used in the antecedent (optionally), and linguistic realization. First, the factual explanation rule is selected from all the rules whose consequent is the predicted class. To do so, we calculate the product of the activation degree a j of each rule r j 2Rand its associated rule weight w j , s.t. argmax w j a j , i.e., the factual explanation rule has the maximum product of the activation degree a j and rule weight w j . Second, if the rules are semantically grounded, i.e., if they use meaningful strong fuzzy partitions (SFP), the feature values in the factual explanation are readily available and mapped to the corresponding linguistic terms (e.g., ‘‘IF Color IS Pale AND Strength IS Standard THEN Beer style IS Blanche” where Pale and Standard are expert-defined linguistic terms). Otherwise, i.e., if only local semantics are available (e.g., ‘‘IF Color IS MF0 AND Strength IS MF1 THEN Beer-style IS Blanche” where MF0 and MF1 are two membership functions with local semantics), linguistic approximation is necessary to generate a meaningful explanation. Notice that the mechanism of linguistic approximation is also used for qualitative CF explanation generation and will be described in detail in the next section. Finally, once the relevant pieces of information are identified, linguistic realization is performed. 3.3. Qualitative counterfactual explanation generation In this section, we introduce a new method for generating qualitative CF explanations (hereinafter denoted as EUC). This method can be regarded as an extension of our previously proposed method (hereinafter denoted as XOR)[43]. The EUC method aims to be more sensitive than XOR to variations in membership functions. Despite certain methodological differences, both methods form a pipeline containing the following steps to be described in detail below (see Fig. 1): CF rule representation, relevance estimation, linguistic approximation (optional in terms of the local/global semantics attached to the FRBCS), and textual explanation generation. CF rule representation. First of all, the test instance (as well as all the CF candidates) must be represented in a compatible form. Both EUC and XOR methods reason over the information retrieved from the rule base. Multiple candidates form CF sets which are labeled in accordance with the selected linguistic terms for the given features. Thus, we regard CF sets as collections of data instances covered by the rules leading to the desired CF class. In this sense, there exist as many potential CFs as there are rules that lead to the desired CF class. For a given FRBCS, a test instance x2Xcan be represented as a vector x¼x 1jV T j ¼ l x t i ðÞj jV T j i¼1 hi of membership function values of each linguistic variable. Similarly, each CF rule can be regarded in terms of the membership function values that the linguistic variables take on. Therefore, each CF rule r CF 2R CF is vectorized over V T for compatibility purposes so that the collection of such vectorized rules makes up a rule-term matrix M jR CF jjV T j where the i-th row corresponds to a CF rule and the j-th column corresponds to the given linguistic term t j 2V T . Hence, the rule-term matrix is populated with such membership values as functions of a given linguistic term M ij ¼ l x t ij . Fig. 1. CF explanation generation pipeline. The shadowed building blocks influence the surface realization of the output explanation. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 383
It is worth noting that the XOR method additionally binarizes both the test instance vector and the rule-term matrix, at the cost of information loss because of the test instance and rule vectors being approximated. Instead, the EUC method represents the original information without further approximation. This is claimed to better capture fuzzy variable ambiguity and avoid potential information loss. Relevance estimation. Given vector representations of the candidate CF rules, it becomes essential to identify the CF set that is minimally different from (and therefore most relevant to) the test instance. Whereas XOR calculates relevance by minimizing the number of different bits, EUC relates each vectorized CF rule to the test instance vector in a jV T j-dimensional space and measures CF relevance as the Euclidean distance dbetween pairs of vectors x;r CF i ;16i6jR CF j; being r CF i ¼M i; the vector associated to row iin matrix M, i.e., the vector which corresponds to CF rule i. d XOR x;r CF i ¼P j jx j r j CFi j jV T j 20;1 ½ ; d EUC x;r CF i ¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi P j x j r j CF i 2 r20;1 ½Þ . where x j and r j CF i are the j-th elements in vectors xand r CF i , respectively. The candidate CF rules are then ranked in accordance with the given distance metric. Subsequently, we include the minimally distant (or most relevant) CF rules for each CF class in the pool E CF of the resulting CF explanations for the given test instance x. If multiple CF rules are equally minimally distant from x, such rules are deemed equally explanatory. In this case, the most relevant CF is selected randomly. Representing the test instance and CF rules in a Euclidean jV T j-dimensional space is hypothesized to better capture fuzzy-specific properties of an FRBCS. For example, the Euclidean distance appears more sensitive to changes in membership function values. The number of unique values that the XOR-based distance can take on is limited by jV T j. In consequence, several CF rules may result in having the same relevance score while being distinct in the number of features or their labeling. On the contrary, EUC provides a more flexible and diverse measure of relevance of different CF rules and therefore gives a better insight into the fuzzy system’s behavior. Linguistic approximation. If the linguistic terms are not based on a SFP and therefore not semantically grounded, the selected CF rule must be enhanced with an additional linguistic layer so that the output explanation is meaningful to the end user. Once the CF rules are ranked by relevance and the most relevant CF is identified, it must therefore be linguistically approximated. To do so, each fuzzy set corresponding to the linguistic term of the selected CF rule is mapped to the gold standard annotations. Note that this mapping is actionable if the a -cut is applied to such a fuzzy set given some threshold value d. To illustrate the process of linguistic approximation, consider a fuzzy set FS characterized by a trapezoidal membership function and three linguistic terms (T¼t 1 ;t 2 ;t 3 fg ) which are candidates to be associated with FS (see Fig. 2 for details). Given some cut-off threshold value d 1 , the fuzzy set FS can be projected to an interval of numerical values L¼ v d 1 ; v d 2 .In addition, each linguistic term t i 2Tcan be projected to an interval t id1 16i6jTjðÞ. Then, the interval Lcan be compared with the intervals t id1 using the Jaccard Similarity Index [13]: 8Lt f a 2V T :St id1 ;LðÞ¼ t id1 \L t id1 [L20;1½;ð1Þ Fig. 2. Illustrative example of the linguistic approximation mechanism. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 384
where t id1 is the numerical interval closer to the linguistic term t i , and Lis the numerical interval associated to the selected a - cut. As follows from Fig. 2,St 3d1 ;LðÞ>St 2d1 ;LðÞ>St 1d1 ;LðÞ. Hence, the feature f j characterized by fuzzy set FS is verbalized as ‘‘f j is t 3 ” in this case. Note that the threshold value dfor the a -cut serves as a hyperparameter. The previously proposed XOR method uses heuristics to specify dmanually. Instead, both qualitative CF generation methods now use major voting in order to reduce possible approximation error. Thus, given some small enough step, we inspect all the approximated linguistic terms over the cut-off interval 0;1½for each term in the given CF rule and assign a confidence score to each term t i as follows: ct i ðÞ¼ #t i 1þ 1 step , being #t i the number of times t i is the winner. For each feature f j involved in the classification and considered in the output explanation, we apply major voting to identify which linguistic term is covered by the widest range of the inspected approximations using the approximation confidence score ct i ðÞas a reference, so that the selected linguistic term is t j 2V T jargmax c t j . Considering the example in Fig. 2, let step be 0.01. We therefore perform n¼1þ1=0:01 ¼101 linguistic approximations. Suppose that the term under consideration is mapped to the set of linguistic terms as indicated in Table 1. Approximation confidence scores are calculated for all the competing linguistic terms. Since we aim to use the most frequently found term among all the considered threshold values, the linguistic term that has the highest score (in this case, t 3 ) is selected for the output explanation. It is worth noting that in this illustrative example, the selected linguistic term is the same as the one selected when considering only d 1 . However, in the general case they may be different. Therefore, it is recommended to follow the major voting approach instead of relying only on a single dvalue selected heuristically. As only two building blocks (relevance estimation and linguistic approximation) influence the output explanation (see the shadowed blocks in Fig. 1), XOR and EUC generate CFs following one of the three scenarios below: the two methods select the same rule to be the most relevant, the approximation algorithm gets the same semantically grounded linguistic terms; the two methods select two different CF rules (e.g., ‘‘IF f 1 IS MF 0 and f 2 IS MF 0 THEN y cf ” and ‘‘IF f 1 IS MF 1 and f 2 IS MF 1 THEN y cf ”) which nevertheless generate identical CF explanations due to a large enough overlap between the corresponding fuzzy sets. This scenario is possible when all the features used in both rules are identical and their non-semantically grounded values overlap to a large enough extent; the two methods select two different CF rules (e.g., ‘‘IF f 1 IS MF 0 and f 2 IS MF 0 THEN y cf ” and ‘‘IF f 1 IS MF 2 and f 3 IS MF 4 THEN y cf ”) where feature values are approximated to different linguistic terms. Textual explanation realization. At the last stage, the selected factual and CF pieces of information are converted to explanations in natural language while applying the NLG pipeline introduced in the previous section. It is worth noting that the text and sentence planning along with text realization for a factual explanation follow the structure of the corresponding winner rule from the rule base. Thus, a factual explanation is assumed to include a subordinate clause of cause (e.g., ‘‘The data instance xis of class y f because f 1 is v 1 and f 2 is v 2 ”), which lists the features and the corresponding values or linguistic terms that influenced the actual decision. On the other hand, a CF explanation is verbalized in natural language as a complex conditional sentence that adopts the structure of the rule, e.g., ‘‘xwould be of class y cf if f 1 were v 2 and f 3 were v 4 ” for the given CF class y cf . Implementation details. The XOR and EUC methods are implemented as open source software in Python and are made publicly available at a Gitlab repository 2 . 3.4. Quantitative counterfactual explanation generation In this section, we present a new method for CF explanation generation which is grounded in evolutionary and bioinspired computation algorithms for explainable AI [11]. More precisely, we have implemented a Genetic Algorithm (hereafter denoted as GEN) which takes as the starting point the genetic fuzzy tuning approach previously proposed by Alonso et al. [4]. Indeed, the original algorithm was first introduced by Cordon and Herrera [7] and later adapted to explainable SFP tuning in [4]. GEN manages a population Pwith Nindividuals which evolve in ggenerations. The given test instance xis used for building the first individual of the population. Each individual is associated to a real-coded chromosome which is made up of p Table 1 Approximation confidence score calculation. Term dApproximation confidence t 1 [0.0, 0.3) 30/101 = 0.297 t 2 [0.3, 0.5) 20/101 = 0.198 t 3 [0.5, 1.0] 51/101 = 0.505 2 https://gitlab.citius.usc.es/ilia.stepin/fcfexpgen (branch ‘‘xor_euc_gen”) I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 385
genes, with each gene representing one of the features in F. Since all the features are numerical, gene i21;p½encodes the double value associated to feature i. The rest of the population is generated randomly. Thus, a random value is assigned to each gene iwithin its variation interval which is determined by the numerical range associated to feature i. The pseudocode of the developed algorithm is as follows (see the GEN shadowed block in Fig. 1): 1. Initialize the generation counter, g¼0, and evaluate the initial population, P 0ðÞ . Evaluating a population means computing Fitness for each individual in the population. Here, Fitness is computed as the Euclidean distance between the data instance ^ xassociated to the current chromosome and the original test instance x, if the inferred output is in agreement with the target CF class. Otherwise, Fitness equals the maximum distance which comes out from the Euclidean distance between the two vectors representing the extreme values (min/max) for the variation intervals associated to each feature. Hence, the smaller Fitness, the better. 2. while g<MaxGener and Fitness PStopThres and Nbest 6NrepThres g:¼gþ1 Select P ðgÞ from P ðg1Þ Crosso v er P ðgÞ Mutate P ðgÞ Elitist selection P ðg1Þ E v aluate P ðgÞ end while The procedure ends either when the maximum number of generations (MaxGener) is reached, or Fitness is under the predefined threshold (StopThres), or the number of consecutive generations for which the best fitness value remains the same (Nbest) is greater than the predefined threshold (NrepThres). On the one hand, MaxGener should be defined empirically in terms of the complexity of the dataset under consideration. It must be large enough to guarantee that GEN converges to a good enough solution. On the other hand, StopThres and NrepThres are threshold values to speed up the procedure, so that the algorithm stops before MaxGener is reached in case Fitness is small enough or becomes constant for a large enough number of generations. For each generation, the following steps are repeated: The selection of P gðÞ from P g1ðÞ is made as a deterministic tournament selection procedure. Each individual in the new population, P gðÞ , is chosen from the previous one, P g1ðÞ , after making a tournament that involves TS individuals randomly selected from P g1ðÞ . The best individual is selected in any tournament. The selection pressure can be adjusted by changing TS 6N. The larger TS, the smaller the chance of weak individuals to be selected. For example, if TS ¼N, then all the individuals in P gðÞ are equal to the best one in P g1ðÞ , what is unsatisfactory from the point of view of diversity in the population. The BLX a crossover operator [10] is applied to P g ðÞ . The parents, i.e., the selected chromosomes in the current population, are crossed over in pairs. Each pair of parents, dad ¼d 1 ;;d p and mom ¼m 1 ;;m p , is replaced in the new population by two offsprings, O d ¼o d1 ;;o dp and O m ¼o m1 ;;o mp , where o dj and o mj are random values from the intervals [min dj ;max dj ] and [min mj ;max mj ], respectively. I j =[I l j ;I u j ] is the variation interval of gene j. According to the taxonomy for the crossover operator presented by [21], a ¼0:3 is a suitable value for letting BLX a exploit the nature of real coding as follows: min dj =maximum I l j ;d j a jd j m j j max dj =minimum d j þ a jd j m j j;I u j min mj =maximum I l j ;m j a jm j d j j max mj =minimum m j þ a jm j d j j;I u j A uniform mutation operator is considered. The value of the selected gene is changed by another one generated randomly within its variation interval. The elitist selection ensures perpetuating the best individual from the given generation to the next one. If the best individual, B i in P g1ðÞ , is not included in P gðÞ , then the worst individual in P gðÞ is replaced by B i . I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 386
Once GEN ends, we have identified a new data instance ^ xthat is assumed to minimally change the original test instance x while making the FRBCS infer the desired CF output 3 . Then, it is time for generating the related CF explanation in natural language. To do so, we once again apply the NLG pipeline described previously. First of all, we compute the percentage of modification Dj¼100 ^ xjxj Ijassociated to each feature jto go from xto ^ x. The text which describes Djis as follows: xjis [slightly] increased jdecreased; where increased appears if Dj>0. On the contrary, decreased is used if Dj<0. In addition, the linguistic modifier slightly appears only in case of small modifications, i.e., only if 0:96Dj65, which means the percentage of modification is smaller or equal than 5%. Notice that nothing is said about feature jif Dj<0:9. In this case, we consider the feature jto remain the same assuming that such a small change (less than 0.9%) does not have sufficient explanatory power for the recipient of the explanation. This assumption is made heuristically in accordance with our previous experience with designing NLG systems while keeping in mind the limited processing capability of human beings [28]. As a result, the generated textual explanations are shorter and easier to process while referring only to relevant changes. Afterwards, at the sentence planning stage, for the sake of simplicity and naturalness, we aggregate those pieces of information associated to different features which are affected by the same type of modification (e.g., ‘‘f 1 and f 2 are slightly increased” replaces to ‘‘f 1 is slightly increased and f 2 is slightly increased”). We also apply lexicalization for each feature to be described in a fully meaningful way. Therefore, increased and decreased are replaced by more meaningful terms (e.g., strength is bigger or color is darker). Finally, text realization is done again using the following template and the SimpleNLG library with the aim of ensuring syntactically, morphologically and orthographically correct final text: ‘‘[Output Class Name] would be [CF Class Name] if [Name of the most Relevant Feature j ]were [linguistic description of D j ] (new data value) [AND...]”. Notice that the new values for the features associated with the most relevant changes are given in brackets. Implementation details. The GEN method is implemented as a piece of open source software in Python and is made publicly available at a Gitlab repository 4 . It is also integrated with the open source software GUAJE 5 which is devoted to facilitating the design of explainable fuzzy systems [3]. The following GEN parameters are considered when generating the quantitative CF explanations under evaluation in the rest of the paper: population length (N¼30), tournament size (TS ¼2), mutation probability (mprob ¼0:1), crossover probability (cprob ¼0:8), a-crossover (a¼0:3), MaxGener = 1000, StopThres =0,NrepThres = 30. The interested reader is kindly referred to Appendix A for further details about how such parameters were selected. 4. Evaluation design In this section, we specify some of the key features that subsequent human evaluation studies rely upon. Section 4.1 introduces the dataset and FRBCS whose classifications are explained. Then, Section 4.2 presents a novel metric for measuring the complexity of automated explanations. 4.1. Dataset and fuzzy inference system The experiments have been carried out using the BEER dataset 6 . It contains characteristics of 400 instances of beer each of which belongs to one of 8 classes (Blanche, Lager, Pilsner, IPA, Stout, Barleywine, Porter, or Belgian Strong Ale). All data instances are described in terms of three features: color, strength, and bitterness. The corresponding linguistic terms and their ranges of values are displayed in Table 2. It is worth noting that all linguistic terms are commonsense and fully meaningful because they were provided by expert brewers. In our experiments, we generate explanations for an FRBCS associated with the Fuzzy Unordered Rule Induction Algorithm (FURIA) [22]. The min–max inference mechanism [27] is applied so that both conjunction (AND) and implication (THEN) are implemented by the t-norm minimum, and the output accumulation is done by the t-conorm maximum. All membership functions are trapezoidal. All rule weights are set to the default value of 1. In addition, it is necessary to apply linguistic approximation as part of the explanation generation pipeline because FURIA rules are endowed only with local semantics. It is worth noting that such a linguistic approximation makes use of meaningful SFP-based linguistic terms as well as their combinations. Thus, explanations may contain combinations of adjacent terms (e.g., ‘‘Feature 1 is Term 1 or Term 2 ”) with the aim of enhancing further their explanatory capacity. Fig. 3 illustrates the SFP associated to color. In this work, we use the same FRBCS that was previously designed and evaluated in [43] with 10-fold cross-validation, achieving 95.5% of correctly classified instances and F1-score equals 0.954 (see the confusion matrix in Table 3 for further details). Notice that, with the aim of avoiding generation of misleading explanations and mainly because the present work focuses on the intended human evaluation, the misclassified test instances are excluded from further analysis in the rest of this manuscript. Whereas explaining misclassification is a challenging problem, it falls outside the scope of this work. 3 Due to the well-known random heuristic nature of genetic algorithms, they avoid stacking in a local minimum but they can not always guarantee the convergence to the global minimum. Anyway, as shown in Appendix A, GEN succeeds to be effective in the search of ‘‘sub-optimal” solutions which are expected to be close enough to the optimal one. 4 https://gitlab.citius.usc.es/ilia.stepin/fcfexpgen (branch ‘‘xor_euc_gen”) 5 https://gitlab.citius.usc.es/jose.alonso/guaje/ 6 The BEER dataset is publicly available athttps://dx.doi.org/10.13140/RG.2.2.20313.67680 I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 387
4.2. Perceived explanation complexity The use of explanations in natural language poses the problem of adequate estimation of explanation complexity. For example, it remains unclear whether the use of adjacent linguistic terms in an explanation (e.g., ‘‘...if color were pale or straw”) increases or decreases understandability (and therefore effectiveness and usability) of such an explanation. As the starting inspiring point for our proposal of automatic calculation of explanation complexity, we refer to existing readability tests in linguistics, which estimate how easily a text can be read by the intended audience. More precisely, the well-known Gunning Fog Index [19] is the weighted average of the normalized sentence length and the percentage of complex words in the text. Similarly, an estimate of complexity of a feature-based linguistic explanation (as perceived by the end user) may rely on the explanation length as well as on the number of features and linguistic terms used in the explanation. Fig. 3. Interpretation of SFP-based linguistic terms associated to Color. Table 3 FURIA confusion matrix. UC stands for Unclassified instances. Predicted class Observed class BLA LAG PIL IPA STO BAR POR BSA UC Blanche (BLA) 50 Lager (LAG) 48 1 1 Pilsner (PIL) 1 49 IPA 1 43 5 1 Stout (STO) 50 Barleywine (BAR) 5 43 1 1 Porter (POR) 1 1 47 1 Belgian Strong Ale (BSA) 1 1 1 47 Table 2 Numerical intervals associated to each SFP-based linguistic term. Feature Linguistic term Range of values Color Pale [0.0, 3.0] Straw [3.0, 7.5] Amber [7.5, 19.0] Brown [19.0, 29.0] Black [29.0, 45.0] Bitterness Low [7.0, 21.0] Low-medium [21.0, 32.5] Medium–high [32.5, 47.5] High [47.5, 250.0] Strength Session [0.035, 0.052] Standard [0.052, 0.067] High [0.067, 0.090] Very high [0.090, 0.136] I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 388
In light of the above, we formally define the perceived explanation complexity (PEC) of an automated explanation eas follows: PEC eðÞ¼kmin l eðÞ; r ðÞ r þ1kðÞ 1 jFjX F e i¼1 t f i jT f i L jð2Þ where k20;1½is the weight regularizing the impact of the explanation length and number of features and terms used in the explanation, leðÞis the explanation length in characters, r is a normalization hyperparameter over the explanation length, jFj is the total number of features in the dataset, F e is the number of unique features used in the given explanation, t f i is the number of terms associated with the i-th feature used in the explanation, jT f i L jis the power of the set of linguistic terms of the i-th feature. In the case of the qualitative methods XOR and EUC, the basic linguistic terms to take into account are those already described in Table 2. However, in order to guarantee a fair comparison between quantitative and qualitative CF explanations, it is necessary to linguistically represent numerical feature value changes suggested by the quantitative method GEN. The sets of linguistic terms associated to each feature by the GEN method are the following: T L ColorðÞ¼ fdarker, slightly darker, lighter, slightly lighterg. T L BitternessðÞ¼ fsmaller, slightly smaller, bigger, slightly biggerg. T L Strength ðÞ ¼fsmaller, slightly smaller, bigger, slightly biggerg. To illustrate computation of PEC(e), let us consider the following example: given a data instance, k¼0:5 and r ¼150, we have three alternative CF explanations with their corresponding complexity scores. XOR: ‘‘Beer style would be Stout if color were black.” PEC eðÞ¼0:5 46 150 þ0:5 1 3 1 5 ¼0:153 þ0:033 ¼0:186 EUC: ‘‘Beer style would be Stout if bitterness were low or low-medium, color were black, and strength were standard or high or very high.” PEC eðÞ¼0:5 130 150 þ0:5 1 3 2 4 þ 1 5 þ 3 4 ¼0:433 þ0:242 ¼0:675 GEN: ‘‘Beer style would be Stout if color were bigger (30.501) and strength were smaller (0.078).” PEC eðÞ¼0:5 90 150 þ0:5 1 3 1 4 þ 1 4 ¼0:300 þ0:083 ¼0:383 Noteworthy, it always holds that PEC eðÞ20;1½.PEC(e) is null only if the explanation is empty and the associated weight k¼1. On the contrary, the highest value of PEC(e) is obtained when the explanation length is equal to the normalization hyperparameter r or all the dataset features and all the linguistic terms are included in the explanation. However, both of these special cases are of no interest, as the empty explanation has got null explanatory power whereas explanation including all the possible categories of features is clearly misleading. 5. Human evaluation The human evaluation study consisted of two online questionnaires that allowed us to assess how the metric PEC is related to different explanation aspects. Section 5.1 presents the instruments and design of the first questionnaire (hereinafter referred to as Survey GM because the items to rate are associated to the so-called Gricean Maxims [15] as we will show below) as well as the analysis of collected data and the discussion of main results. In the light of lessons learned from this survey, we developed a subsequent one (hereinafter referred to as Survey TS because the focus is on assessing Trustworthiness and Satisfaction of the given explanations) whose experimental design and main discoveries are described in Section 5.2.In both surveys, all the subjects participated voluntarily and anonymously. This research obtained ethics approval from the University Ethics committee. 5.1. Survey GM: Evaluating CF explanations in terms of Gricean Maxims 5.1.1. Experimental settings The first experiment was designed as a within-subject study. In order to perform a comparative analysis of qualitative and quantitative CF explanations, we considered only those test instances for which the qualitative methods (XOR and EUC) generated distinct explanations (thus avoiding misleading repetitions). Since the BEER dataset has 8 classes, given a test instance we have 1 factual class and 7 alternative CF classes. Because the FURIA rules were trained and evaluated with 10-fold cross-validation, the 400 data instances in the BEER dataset were split 10 times into training set (90%) and test set (10%). As a result, we built 10 sets of FURIA rules. They were used to make predictions for all test instances in each fold (see details in Table 4). Then, we filtered out unclassified and misclassified test instances with the aim of avoiding the inclusion of void or misleading explanations to be evaluated in the survey. Notewor- I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 389
survey is costly as it requires higher cognitive load and more time from the participants. Therefore, calculating PEC automatically allows the survey designer to set up and deploy a shorter questionnaire and thus easier to fill (Survey TS). It is worth noting that PEC strongly correlates with several explanation aspects but does so in different directions, so an FRBCS designer is advised to carefully select the method of explanation generation based on the peculiarities of the application domain and/ or intended audience. All in all, the insights from this work are expected to advance methods of generation and evaluation for various explanation approaches. As such, they are expected to be helpful for designing future human evaluation surveys in the area of explainable AI. Moreover, as part of future work, we will go deeper with selecting and fusing CF explanations with the aim of customizing them for users having different profiles in different application scenarios. Further research is therefore necessary: (1) to extend the proposed CF explanation generation methods beyond numerical features; (2) to better assess the impact of the PEC hyperparameters ( r and k); and (3) to better understand the connection between complexity and trustworthiness of automated explanations. Notice that, the conclusions derived from the current study are only applicable to the target population under consideration. As part of future work, for the sake of generalization, we intend to design and carry out other similar experiments with a larger and wider panel of respondents, including non-expert lay users. Finally, we plan to use PEC as one of the criteria to optimize when designing explainable multi-objective evolutionary fuzzy systems. Funding Ilia Stepin is an FPI researcher (grant PRE2019-090153). Jose M. Alonso-Moral is a Ramon y Cajal researcher (grant RYC- 2016–19802). This work was supported by the Spanish Ministry of Science and Innovation (grants RTI2018-099646-B-I00, PID2021-123152OB-C21, and TED2021-130295B-C33) and the Galician Ministry of Culture, Education, Professional Training and University (grants ED431F2018/02, ED431G2019/04, and ED431C2022/19). All the grants were co-funded by the European Regional Development Fund (ERDF/FEDER program). Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Appendix A In addition to the human evaluation study on the automatically generated CFs, we performed three independent experiments on the genetic algorithm hyperparameter fine-tuning. In particular, we estimated the impact of the following hyperparameters associated to the GEN method: (i) the size of the population, (ii) the crossover probability and the corresponding alpha value, and (iii) the mutation probability. All the experiments were run for the five survey stimuli where both the predicted classes and the CF classes were known. The experimental results were assessed in terms of the best achieved fitness scores. Fig. 4 summarizes the impact of the population size (10, 20, 30, 40, 50). It can be observed that the default population size (30) provides good results, on average, for all the test instances under consideration. Fig. 4. An empirical assessment of the impact of the population size in the GEN method. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 396
Fig. 5 shows the results of the experiment on the crossover probability values (0.7, 0.8, 0.9), considering different a values (0.2, 0.3, 0.4). In short, the combination of the crossover probability (0.8) and a ¼0:3 yields the best results for the considered CF data points. Fig. 6 illustrates the impact of the selected mutation probability values (0.05, 0.1, 0.15, 0.2). It can be seen that doubling the default mutation probability value may result in worsened performance of the algorithm. To sum it up, the analysis carried out allows us to conclude that the selected hyperparameter values do not only agree with the guidelines found in the literature (e.g., [21]) but also prove to be effective in the given experiments and can indeed be recommended for future use. All the detailed calculations as well as additional plots and the source code for replicating this experimental analysis can be found in our Gitlab repository:https://gitlab.citius.usc.es/ilia.stepin/fcfexpgen (branch ‘‘xor_euc_gen”). Fig. 5. An empirical assessment of the impact of the crossover hyperparameters (the crossover probability and the acrossover operator) in the GEN method. Fig. 6. An empirical assessment of the impact of the mutation probability in the GEN method. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 397
References [1] A. Abdul, J. Vermeulen, D. Wang, B.Y. Lim, and M. Kankanhalli. Trends and trajectories for explainable, accountable and intelligible systems: An HCI research agenda. In Proceedings of the Conference on Human Factors in Computing Systems (CHI), pages 1–18, Montreal QC, Canada, 2018. Association for Computing Machinery. https://doi.org/10.1145/3173574.3174156. [2] A. Adadi, M. Berrada, Peeking inside the black-box: A survey on explainable artificial intelligence (XAI), IEEE Access 6 (2018) 52138–52160, https://doi. org/10.1109/ACCESS.2018.2870052. [3] J.M. Alonso, C. Castiello, L. Magdalena, C. Mencar, Explainable Fuzzy Systems - Paving the Way from Interpretable Fuzzy Systems to Explainable AI Systems, volume 970, Springer International Publishing (2021), https://doi.org/10.1007/978-3-030-71098-9. [4] J.M. Alonso, O. Cordón, S. Guillaume, and L. Magdalena. Highly interpretable linguistic knowledge bases optimization: Genetic tuning versus soliswetts. Looking for a good interpretability-accuracy trade-off. In Proceedings of the IEEE International Conference on Fuzzy Systems, pages 901–906, London, UK, 2007. https://doi.org/10.1109/FUZZY.2007.4295485. [5] I. Baaj and J.-P. Poli. Natural language generation of explanations of fuzzy inference decisions. In Proceedings of the IEEE International Conference on Fuzzy Systems, pages 1–6, New Orleans, LA, USA, 2019. https://doi.org/10.1109/FUZZ-IEEE.2019.8858994. [6] A. Cascallar-Fuentes, A. Ramos-Soto, A. Bugarín, Adapting SimpleNLG to Galician Language, in: In Proceedings of the International Conference on Natural Language Generation, Association for Computational Linguistics (ACL), 2018, https://doi.org/10.18653/v1/W18-6507. [7] O. Cordón, F. Herrera, A Three-Stage Evolutionary Process for Learning Descriptive and Approximate Fuzzy Logic Controller Knowledge Bases from Examples, International Journal of Approximate Reasoning 17 (4) (1997) 369–407, https://doi.org/10.1016/S0888-613X(96)00133-8. [8] R. Dale, E. Reiter, Computational Interpretations of the Gricean Maxims in the Generation of Referring Expressions, Cognitive science 19 (2) (1995) 233–263, https://doi.org/10.1016/0364-0213(95)90018-7. [9] V. Dignum, Responsible Artificial Intelligence: How to Develop and Use AI in a Responsible Way. Artificial Intelligence: Foundations, Theory, and Algorithms, Springer, Cham, 2019, https://doi.org/10.1007/978-3-030-30371-6. [10] L.J. Eshelman and J.D. Schaffer. Real-Coded Genetic Algorithms and Interval-Schemata. In L. Darrell Whitley, editor, Foundations of Genetic Algorithms, volume 2 of Foundations of Genetic Algorithms, pages 187–202. Elsevier, 1993. https://doi.org/10.1016/B978-0-08-094832-4.50018-0. [11] A. Fernandez, F. Herrera, O. Cordon, M. Jose del Jesus, F. Marcelloni, Evolutionary Fuzzy Systems for Explainable Artificial Intelligence: Why, When, What for, and Where to?, IEEE Computational Intelligence Magazine 14 (1) (2019) 69–81, https://doi.org/10.1109/MCI.2018.2881645. [12] R.R. Fernández, I.M. de Diego, V. Aceña, A. Fernández-Isabel, J.M. Moguerza, Random forest explainability using counterfactual sets, Information Fusion 63 (2020) 196–207, https://doi.org/10.1016/j.inffus.2020.07.001. [13] S. Fletcher, M.Z. Islam, Comparing sets of patterns with the Jaccard index, Australasian Journal of Information Systems 22 (2018), https://doi.org/ 10.3127/ajis.v22i0.1538. [14] A. Gatt, E. Krahmer, Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation, Journal of Artificial Intelligence Research 61 (2018) 65–170, https://doi.org/10.1613/jair.5477. [15] H.P. Grice, Logic and Conversation, in: P. Cole, J.L. Morgan (Eds.), Syntax and Semantics: Speech Acts, Academic Press, 1975, pp. 41–58, https://doi.org/ 10.1163/9789004368811_003. [16] R. Guidotti, Counterfactual explanations and how to find them: literature review and benchmarking, Data Mining and Knowledge Discovery (2022) 1– 55, https://doi.org/10.1007/s10618-022-00831-6. [17] R. Guidotti, A. Monreale, F. Giannotti, D. Pedreschi, S. Ruggieri, F. Turini, Factual and Counterfactual Explanations for Black Box Decision Making, IEEE Intelligent Systems 34 (6) (2019) 14–23, https://doi.org/10.1109/MIS.2019.2957223. [18] D. Gunning, E. Vorm, J.Y. Wang, M. Turek, DARPA’s explainable AI (XAI) program: A retrospective, Applied AI Letters 2 (4) (2021), https://doi.org/ 10.1002/ail2.61 e61. [19] R. Gunning, Technique of clear writing, McGraw-Hill, 1968. [20] C. Herley and W. Pieters. If You Were Attacked, You’d Be Sorry: Counterfactuals as Security Arguments. In Proceedings of the 2015 New Security Paradigms Workshop, NSPW ’15, pages 112–123, New York, NY, USA, 2015. Association for Computing Machinery. https://doi.org/10.1145/2841113. 2841122. [21] F. Herrera, M. Lozano, A.M. Sánchez, A Taxonomy for the Crossover Operator for Real-Coded Genetic algorithms: An Experimental Study, International Journal of Intelligent Systems 18 (3) (2003) 309–338, https://doi.org/10.1002/int.10091. [22] J. Hühn, E. Hüllermeier, FURIA: an algorithm for unordered fuzzy rule induction, Data Mining and Knowledge Discovery 19 (3) (2009) 293–319, https:// doi.org/10.1007/s10618-009-0131-8. [23] H. Ishibuchi, T. Nakashima, M. Nii, Classification and modeling with linguistic information granules: Advanced approaches to linguistic Data Mining, Springer Science & Business Media (2004), https://doi.org/10.1007/b138232. [24] M. Lash, Q. Lin, N. Street, J. Robinson, and J. Ohlmann. Generalized Inverse Classification. In Proceedings of the International Conference on Data Mining (SDM), pages 162–170. Society for Industrial and Applied Mathematics, 2017. https://doi.org/10.1137/1.9781611974973.19. [25] A. Lucic, H. Haned, and M. de Rijke. Why Does My Model Fail? Contrastive Local Explanations for Retail Forecasting. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, pages 90–98, Barcelona, Spain, 2020. Association for Computing Machinery. https://doi.org/10.1145/3351095.3372824. [26] N. Maaroof, A. Moreno, A. Valls, M. Jabreel, M. Szelag, A Comparative Study of Two Rule-Based Explanation Methods for Diabetic Retinopathy Risk Assessment, Applied Sciences 12 (7) (2022) 1–18, https://doi.org/10.3390/app12073358. [27] E.H. Mamdani, Application of Fuzzy Logic to Approximate Reasoning Using Linguistic Systems, IEEE Transactions on Computers 26 (12) (1977) 1182– 1191, https://doi.org/10.1109/TC.1977.1674779. [28] G.A. Miller, The magical number seven, plus or minus two: Some limits on our capacity for processing information, Psychological Review 63 (2) (1956) 81–97, https://doi.org/10.1037/0033-295x.101.2.343. [29] T. Miller, Explanation in artificial intelligence: Insights from the social sciences, Artificial Intelligence 267 (2019) 1–38, https://doi.org/10.1016/j. artint.2018.07.007. [30] J. Moore, N. Hammerla, and C. Watkins. Explaining deep learning models with constrained adversarial examples. In Proceedings of the Pacific Rim International Conference on Artificial Intelligence (PRICAI), pages 43–56. Springer, 2019. https://doi.org/10.1007/978-3-030-29908-8_4. [31] R.K. Mothilal, A. Sharma, and C. Tan. Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, pages 607–617, Barcelona, Spain, 2020. Association for Computing Machinery. https://doi.org/10.1145/3351095.3372850. [32] M.L. Olson, R. Khanna, L. Neal, F. Li, W.-K. Wong, Counterfactual state explanations for reinforcement learning agents via generative deep learning, Artificial Intelligence 295 (2021) 1–29, https://doi.org/10.1016/j.artint.2021.103455. [33] Parliament and Council of the European Union. General Data Protection Regulation (GDPR), 2016. URL:http://data.europa.eu/eli/reg/2016/679/oj. [34] Parliament and Council of the European Union. A European Approach to Artificial Intelligence, 2022. URL:https://digital-strategy.ec.europa.eu/en/ policies/european-approach-artificial-intelligence. [35] E. Reiter, R. Dale, Building Natural Language Generation Systems, in: Studies in Natural Language Processing, Cambridge University Press, 2000, https://doi.org/10.1017/CBO9780511519857. [36] M.T. Ribeiro, S. Singh, and C. Guestrin. ”Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the International Conference on Knowledge Discovery and Data Mining (SIGKDD), pages 1135–1144, San Francisco, California, USA, 2016. Association for Computing Machinery. https://doi.org/10.1145/2939672.2939778. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 398
[37] C. Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nature Machine Intelligence 1 (5) (2019) 206–215, https://doi.org/10.1038/s42256-019-0048-x. [38] M. Schleich, Z. Geng, Y. Zhang, and D. Suciu. GeCo: Quality Counterfactual Explanations in Real Time. In Proceedings of the Very Large Data Bases (VLDB) Endowment, volume 14(9), pages 1681–1693, 2021. https://doi.org/10.14778/3461535.3461555. [39] S. Sharma, J. Henderson, J. Ghosh. CERTIFAI, A Common Framework to Provide Explanations and Analyse the Fairness and Robustness of Black- Box Models, Association for Computing Machinery, 2020, pp. 166–172, https://doi.org/10.1145/3375627.3375812. [40] K. Sokol and P. Flach. One Explanation Does Not Fit All: The Promise of Interactive Explanations for Machine Learning Transparency. KI – Künstliche Intelligenz, 2020. https://doi.org/10.1007/s13218-020-00637-y. [41] I. Stepin, J.M. Alonso, A. Catala, and M. Pereira-Fariña. Generation and evaluation of factual and counterfactual explanations for decision trees and fuzzy rule-based classifiers. In Proceedings of the IEEE World Congress on Computational Intelligence (WCCI), Glasgow, UK, 2020. https://doi.org/10.1109/ FUZZ48607.2020.9177629. [42] I. Stepin, J.M. Alonso, A. Catala, M. Pereira-Fariña, A Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable Artificial Intelligence, IEEE Access 9 (2021) 11974–12001, https://doi.org/10.1109/ACCESS.2021.3051315. [43] I. Stepin, A. Catala, M. Pereira-Fariña, J.M. Alonso, Factual and Counterfactual Explanation of Fuzzy Information Granules, in: Interpretable Artificial Intelligence: A Perspective of Granular Computing, Springer International Publishing, 2021, pp. 153–185, https://doi.org/10.1007/978-3-030-64949- 4_6. [44] R. Sukkerd, R. Simmons, D. Garlan, Toward Explainable Multi-Objective Probabilistic Planning, in: IEEE/ACM 4th International Workshop on Software Engineering for Smart Cyber-Physical Systems (SEsCPS), Gothenburg, Sweden, 2018, pp. 19–25. [45] S. Verma, J. Dickerson, and K. Hines. Counterfactual explanations for machine learning: A review. In Proceedings of the Machine Learning: Retrospectives, Surveys and meta-Analyses (ML-RSA) Workshop at the Conference on Neural Information Processing Systems (NeurIPS), 2020. [46] S. Wachter, B. Mittelstadt, C. Russell, Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR, Harvard Journal of Law & Technology 31 (2) (2018) 841–887, https://doi.org/10.2139/ssrn.3063289. [47] X. Wang, M. Yin, Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making, in: 26th International Conference on Intelligent User Interfaces, IUI ’21, Association for Computing Machinery, 2021, pp. 318–328, https://doi.org/10.1145/ 3397481.3450650. [48] C. Woodcock, B. Mittelstadt, D. Busbridge, G. Blank, et al, The Impact of Explanations on Layperson Trust in Artificial Intelligence-Driven Symptom Checker Apps: Experimental Study, Journal of Medical Internet Research 23 (11) (2021), https://doi.org/10.2196/29386 e29386. [49] L.A. Zadeh, The concept of a linguistic variable and its application to approximate reasoning, Information Sciences 8 (3) (1975) 199–249, https://doi. org/10.1016/0020-0255(75)90036-5. I. Stepin, J.M. Alonso-Moral, A. Catala et al. Information Sciences 618 (2022) 379–399 399
CORRECTED PROOF Argument & Computation -1 (2023) 1–59 1 DOI 10.3233/AAC-220011 IOS Press Information-seeking dialogue for explainable artificial intelligence: Modelling and analytics Ilia Stepin a,c,∗, Katarzyna Budzynska b, Alejandro Catala a,c, Martín Pereira-Fariña dand Jose M. Alonso-Moral a,c aCentro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela, Rúa de Jenaro de la Fuente Domínguez s/n, 15782 Santiago de Compostela, A Coruña, Spain E-mails: [email protected],alejandr[email protected],[email protected] bLaboratory of The New Ethos, Warsaw University of Technology, plac Politechniki 1, 00-661, Warsaw, Poland E-mail: katarzyna.b[email protected] cDepartamento de Electrónica e Computación, Universidade de Santiago de Compostela, Rúa Lope Gómez de Marzoa, s/n, 15782 Santiago de Compostela, A Coruña, Spain dDepartamento de Filosofía e Antropoloxía, Universidade de Santiago de Compostela, Plaza de Mazarelos s/n, 15705 Santiago de Compostela, A Coruña, Spain E-mail: martin.pereir[email protected] Abstract. Explainable artificial intelligence has become a vitally important research field aiming, among other tasks, to justify predictions made by intelligent classifiers automatically learned from data. Importantly, efficiency of automated explanations may be undermined if the end user does not have sufficient domain knowledge or lacks information about the data used for training. To address the issue of effective explanation communication, we propose a novel information-seeking explanatory dialogue game following the most recent requirements to automatically generated explanations. Further, we generalise our dialogue model in form of an explanatory dialogue grammar which makes it applicable to interpretable rule-based classifiers that are enhanced with the capability to provide textual explanations. Finally, we carry out an exploratory user study to validate the corresponding dialogue protocol and analyse the experimental results using insights from process mining and argument analytics. A high number of requests for alternative explanations testifies the need for ensuring diversity in the context of automated explanations. Keywords: Explainable Artificial Intelligence, information-seeking dialogue game, explanation locutions, counterfactual explanation, process mining analytics, argument analytics 1. Introduction Explainability in the context of Artificial Intelligence (AI) has long attracted attention of researchers from computer science [57] and argumentation [21]. The first explanation generation methods turned up *Corresponding author. Tel.: +34 8818 16394; E-mail: [email protected]. 1946-2166 © 2023 – The authors. Published by IOS Press. This is an Open Access article distributed under the terms of the Creative Commons Attribution-NonCommercial License (CC BY-NC 4.0).
CORRECTED PROOF 2I. Stepin et al. / Information-seeking dialogue for XAI in the 1980s along with the so-called Expert Systems [74]. More precisely, the first explainers addressed the challenge of explaining the output of expert systems and logic programs [7], which eventually led to the emergence of the research field that we now call Computational Argumentation. Recent years have witnessed a new boost of interest in developing eXplainable AI (XAI), as novel machine learning (ML) algorithms produce highly accurate yet oftentimes poorly explainable predictions [1]. As defined at present, XAI aims to (1) generate explainable models preserving a high level of accuracy and (2) enable the end user, e.g., a client of a bank or a patient of a hospital, with the opportunity to understand, trust, and manage the given AI-based systems [2,29] (e.g., querying a bank loan management system to identify reasons for the loan application being rejected or a hospital information system to receive treatment-related recommendations). The obscure nature of the underlying reasoning of the state-of-the-art predictive algorithms has given way to the so-called “right to explanation” [80]. The corresponding legal regulations are being increasingly adopted worldwide [87]. For example, the European Union (EU)’s General Data Protection Regulation (GDPR) acknowledges the right of the user “not to be subject to a decision evaluating personal aspects relating to him or her which is based solely on automated processing and which produces adverse legal effects concerning, or significantly affects, him or her” [51]. In addition, current EU’s legal regulations in, for example, the financial domain require that algorithmic transparency be provided for automatic trading techniques (see the Directive 2014/65/EU on Markets in Financial Instruments, commonly known as MiFID II [52] for details). Being a controversial topic of primary importance for numerous stakeholders, its juridical basis is constantly updated. Thus, the newly proposed EU’s AI Act (AIA) [53] establishes a taxonomy of AI-based systems and requires that high-risk AI applications offer explanations for their decisions or recommendations to their end users. In order to mitigate algorithmic transparency issues of the state-of-the-art AI algorithms, a use of interpretable models is advised [59]. Interpretable rule-based models (such as, e.g., decision trees (DT) or decision rules) are known to provide user-friendly explanations [47]. Remarkably, DTs can be used as part of more complex model-agnostic explainers that are able to justify predictions of other arbitrary classifiers if they are, for example, trained on a local synthetically generated neighbourhood around the test instance [28]. Despite the fact that only few XAI frameworks offer explanations in natural language [12], DTs have also been shown to be a powerful tool for communicating textual explanations to end users, e.g., by engaging the user in an explanatory dialogue [70,79]. Explanations are claimed to have to necessarily be embedded in a dialogical interaction so that the end user is able to challenge the aspects of an explanation that have not been understood [63]. For illustrative purposes, let us consider a beer style classification problem (see Table 1for details). Given a number of predefined classes (i.e., beer styles) and an instance of beer, the classifier (System) makes a prediction on what beer style the test instance is (move m1). The end user (User) engages in a dialogue with the System to inquire an explanation for the given prediction (m2). Then, the System provides the User with a factual explanation based on the most distinctive features, pertinent to the test instance, that led to the given prediction (m3). Then, given the factual explanation, the User may want to clarify specific features (e.g., bitterness) that the offered explanation contains (m4). In response, the System offers a definition for the requested feature (m5). Alternatively, the User may inquire the System to explain why the given test instance is not predicted to be of another class (m6,m10). The pieces of such counterfactual explanations offered in response suggest minimal changes to the input so that its classification changes in the desired manner (m7,m11). Given the initial explanation, the User may want to specify a more precise description of the given feature (m8). Then, the System may present detailed information over the corresponding numerical interval (subject to availability) over which the feature is defined (m9). Further,
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 3 Table 1 An illustrative explanatory dialogue Move Agent Utterance m1System The present beer is Blanche. m2User Could you explain why you think so? m3System It is Blanche because colour is pale and bitterness is low. m4User What do you mean by bitterness? m5System Bitterness is a harsh, acrid taste that is one of the four basic taste sensations. m6User But is the beer not stout? m7System It would be Stout if colour were black and bitterness were high. m8User Could you specify how high bitterness is defined? m9System Sure! I define bitterness as high because it lies in the range from 47 to 250. m10 User But why is it not Porter? m11 System It would be Porter if colour were brown. m12 User I am not quite satisfied with your explanation. Could you offer me another one? m13 System Sure! It would be of class Porter if colour were brown and strength were session. m14 User Ok, now I trust your prediction. m15 System Thank you for your trust in me. Bye! the User may disagree with the explanation offered and argue over it (m12). The System should then offer an alternative explanation that would satisfy the User’s needs (m13). When the User is sufficiently informed about the reasons that led to the given prediction, he or she makes an informed decision on whether the System’s prediction should be trusted or not (m14). The explanatory dialogue ends with the System’s farewell locution (m15). As follows from Table 1, we consider two types of explanations: factual and counterfactual. Assuming knowledge of the feature space, factual explanations (illustrated with move m3in Table 1) aim to explain the given classifier’s prediction in terms of the most relevant feature values that led to that prediction. On the contrary, counterfactuals (illustrated with moves m7,m11,andm13 in Table 1) are post-hoc examplebased explanations that suggest a minimal change in feature values to those of the given data instance so that the system’s prediction changes as desired [71]. This paper introduces an explanatory dialogue game for communicating factual and counterfactual explanations for interpretable rule-based classifiers. We assume that the classifier is associated with an explainer that is capable of providing textual (rule-based) explanations. Based on the dialogue typology proposed by Walton and Krabbe [82], we model the information-seeking type of explanatory dialogue equipping it with a specific collection of locutions tailored for the aforementioned types of explanation that the user may ask the system. As a starting point, we consider the typology of dialogue moves proposed by Budzynska et al. [9]. In our work, we extend this typology of dialogue moves with a repertoire of locutions allowing for communication of factual and counterfactual explanations to enable the end user to interactively explore the explanation space. Then, we propose a context-free dialogue grammar to generalise the formal structure of the resulting dialogue model. Despite an empirically shown strong need in both factual and counterfactual explanations [41] and at least a hundred of counterfactual explanation generation methods proposed by now in the context of XAI, less than a third of these methods are evaluated in user studies [37]. To address this issue, we subsequently perform a pilot user study to evaluate the proposed dialogue model. Moreover, we analyse the collected dialogue transcripts treating instances of explanatory dialogue as processes using the state-of-the-art techniques from process mining and argument analytics [43].
CORRECTED PROOF 4I. Stepin et al. / Information-seeking dialogue for XAI As a result, we bridge the gap between ML practitioners and the argumentation community by making the following contributions: •we model information-seeking explanatory dialogue based on the fundamental notions from the argumentation theory and apply the dialogue model in the context of XAI; •we propose a set of original dialogue locution types that are found specifically suitable for effective communication of factual and counterfactual explanations; •we demonstrate the explanatory utility of the proposed dialogue protocol via a human evaluation study based on three use cases for an interpretable rule-based classifier leaving open-source implementations of the dialogue game and the human evaluation toolkit available for public use; •we suggest formal means for extending the proposed protocol to make it applicable to modelling dialogic human-machine interaction for classification tasks in other applications. The rest of the manuscript is structured as follows. Section 2introduces the classification problem formally and outlines the common properties of explanations claimed to be essential for explaining solutions to such a problem. In addition, we subsequently discuss possible discrepancy between automatically generated explanations and user-preferred explanations. Section 3defines an explanatory dialogue game as an interface between an explanation generation module and the end user. Section 4introduces essential process mining concepts and shows how we apply them to explanatory dialogue analysis. Section 5presents the experimental settings of the human evaluation study carried out to assess the utility of the proposed dialogue protocol. Section 6reports the experimental results obtained from the human evaluation study. Section 7discusses the dialogue model validation results. Section 8presents an overview of related work regarding formal explanatory dialogue models as well as recent argumentation-based techniques for explanatory dialogue modelling. Finally, we outline prospective directions for future work and conclude in Section 9. 2. Preliminaries In this section, we first outline a definition of the classification problem and assumptions about the nature of classifiers and explainers that we are driven by (see Section 2.1 for details). Then, we formally define essential explanation-related concepts that we utilise throughout the manuscript in Section 2.2. Finally, we draw reader’s attention to possible discrepancies between the user-preferred explanations and those offered to him or her by the explainer in Section 2.3. 2.1. The classification problem As outlined in Section 1, we focus on communicating to the end user automated explanations for the output of an interpretable rule-based ML classifier. Figure 1depicts a general architecture of the modelled explanation communication process. The System is assumed to include, at least, the following core components: an interpretable rule-based classifier, an explainer, a knowledge base, and a dataset that the classifier is trained on. The User starts the communication process by sending a classification request for a specific test instance to the System in form of the test instance’s characteristics (i.e., features). The classifier is pretrained on a given dataset X={xi}|n i=1containing nlabelled instances to learn a mapping function c:X−→ Ywhere Y={yj}|m j=1is a discrete output variable (class), mbeing the number of classes.
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 5 Fig. 1. A schema of the modelled system-user explanation communication process. This paper focuses on designing an explanatory dialogue game for communication of factual and counterfactual explanations for interpretable rule-based classifiers (the shaded block). In this work, we assume knowledge of the feature space: the dataset is said to contain linearly scaled numerical features. In addition, all the numerical feature values are said to be mapped to the corresponding feature-dependent linguistic variable [86]. Therefore, each data instance xi∈X=Fi,y iis associated to class yi∈Yand defined over the set of p3-tuple features Fi={fk,vk,tk}|p k=1where each feature fkis assigned to the corresponding numerical value vkand linguistic term tk(e.g., age, 20, young). The values of the linguistic variables (i.e., the so-called linguistic terms) may be defined by an expert. In this case, they are mapped to expert knowledge-based numerical intervals covering all the values of the corresponding feature. Otherwise, the linguistic variable is assigned to a set of textual values and mapped to equal-size numerical intervals. In this respect, the set of textual values that the linguistic variable can take on is of arbitrary cardinality. The classifier predicts the class label ˆyfor the given test instance xtest =Ftest,y teston the basis of the learned mapping function c. The test instance classification is predicted correctly if the predicted class label and the actual test instance class label are the same (i.e., ˆy=ytest). Otherwise, the predicted class is deemed wrong (i.e., ˆy= ytest). Altogether, the interpretable rule-based classifier and the explainer are said to form an explainable classifier. Once the classifier outputs a prediction, the associated explainer attempts to generate an explanation in natural language for that prediction. Upon request, the explanation is passed to the User via the explanatory dialogue game, which serves as a communication channel between the explainable classifier and the User. During their intercourse, the User is assumed to be able to submit further explanation-related requests and receive responses processed by the dialogue game module whereas the dialogue game module can query the explainer for further explanation-related information. 2.2. Explanation to the classification The upsurging need for explaining a classifier’s output is raising interest in the mere nature of the explanation. For instance, social sciences testify that explanations are expected to be contrastive, selected,
CORRECTED PROOF 6I. Stepin et al. / Information-seeking dialogue for XAI and social [45]. First, the property of contrastiveness implies establishing a relation not only between the cause and effect of the phenomenon under consideration but also another relation between the cause and a given non-observed effect (i.e., another alternative effect). Second, explanations are as well argued to be selected, i.e., only the most relevant causes should make part of a specific explanation. Third, explanations are claimed to be social, i.e., they are a product of interaction between the explainer and the explainee. Contrastiveness plays an important role when explaining a solution to the classification problem, as different classes are opposed to the others on the basis of distinctive feature values. Further, contrastiveness is inherent to counterfactual (CF) explanations (or counterfactuals, for short). In the context of XAI, counterfactuals suggest minimal changes in feature-value pairs for a different outcome to be obtained [71]. CFs are said to be post-hoc (i.e., they are generated for pretrained classifiers) and local (i.e., they explain the classifier’s output w.r.t. a specific test instance) [27]. CFs may be (1) model-agnostic if they operate only on the given input (i.e., a test instance) and output (i.e., a prediction) of the classifier or (2) model-specific if they utilise the internals of the classifier to explain the given output [47,71]. CF explanations are claimed to have a number of desired properties against which CF explanation methods can be evaluated [27]. For example, CFs should be valid (i.e., CFs should truly lead to the desired hypothetical outcome), proximate (i.e., CFs should suggest only minimal changes to the test instance w.r.t. the selected distance metric), sparse (i.e., CFs should minimise the number of features whose values are to be changed), actionable (i.e., CFs should suggest feasible changes), and diverse (i.e., CFs should offer multiple alternatives). An exhaustive list of such properties can be found in recent surveys on CF explanation generation and evaluation [27,49,78]). A large number of explanation generation methods are evaluated using automatically computable metrics that assess the aforementioned properties of CF explanations [49]. However, such metrics oftentimes do not take into consideration user feedback at all. Whereas considering the social factor may not be necessary when, e.g., measuring validity, estimating CF diversity may have to directly involve capturing effects of the interaction between the system and the user. Indeed, CF explanations suggesting minimal changes in feature values may not always be equally appreciated by end users. Given a variety of potential CFs, different users may prefer distinct CFs for the same hypothetical output. Further, the social aspect of explanation becomes crucially important when two alternative automatically generated pieces of explanation are deemed equally explanatory (e.g., when the distances from the test instance to two or more closest CF data points are the same or when two CF sets have the same coverage). As the state-of-the-art AI technologies are shifting towards being user-centric [83], it appears indispensable to enhance existing explanation generation modules with a system-user communication interface that would allow end users to produce such inquiries for alternative CFs in the course of an explanatory dialogue, even if the user is not aware of the dataset-related peculiarities. Various state-of-the-art CF explanation generation frameworks are known to offer diverse CFs ([15,17,35,49,60,62,75,85], among others). However, the format of such CFs raises several important concerns. First, most of such frameworks lack any interaction with end users leaving the users without further guidance when interpreting the generated explanations. Second, some explainers output a set of distinct CFs altogether [49,60]. In these settings, the Grice’s maxim of quantity [25] may be violated, as only a subset of the offered explanations can be sufficient for the end user. Third, a large number of diverse CF explanation generation frameworks provide their output in tabular form [15,17,35,49,62,75]. Whereas natural language generation tools can be used to transform tabular data into text, a taxonomy of necessary explanation-related requests and responses remains missing. To address these issues, we propose a transparent explanatory dialogue model for diverse factual and counterfactual explanation
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 13 On the one hand, the set of requests from the user to the system REQ={REQ-explanation( ˆy),reqdetailisation( ˆy,e,),req-clarification( ˆy,e,),req-alternative( ˆy,e)} consists of the following items:2 •REQ-explanation( ˆy): the set of user requests for explanation for system’s prediction ˆy; •req-detailisation( ˆy,e,): the user request for further details on feature (i.e., the corresponding numerical intervals) that makes part of a high-level (either factual or CF) explanation efor prediction ˆy; •req-clarification( ˆy,e,): the user request for clarification of the meaning of a specific feature that makes part of (either factual of CF and either high-level or low-level) explanation efor prediction ˆy; •req-alternative( ˆy,e): the user request for an alternative (either factual or CF and either high-level or low-level) explanation provided that the user is not satisfied with the previously offered explanation efor system’s prediction ˆy. Further, the set of user explanation requests REQ-explanation( ˆy)consists of the following possible locutions: •req-why( ˆy): the user request for a factual explanation for the system’s prediction ˆy; •req-why-not( ˆy,y): the user request for a CF explanation concerning the CF class y∈Y\{ˆy}for prediction ˆy(i.e., to specify why some CF class ywas not predicted instead of ˆy). On the other hand, the set of responses (replies) that the system sends back to the user REP={REP- explanation( ˆy),REP-detailisation( ˆy,e,),REP-clarification( ˆy,e,),REP-alternative( ˆy,e)} mirrors the set of user requests: •REP-explanation( ˆy): the set of system responses in an attempt to explain prediction ˆy; •REP-detailisation( ˆy,e,): the set of system responses in an attempt to provide details (i.e., numerical intervals) with respect to feature of explanation efor system’s prediction ˆy; •REP-clarification( ˆy,e,): the set of system responses in an attempt to clarify feature making part of (either factual or CF) explanation efor prediction ˆy; •REP-alternative( ˆy,e): the set of system responses in an attempt to provide the user with an explanation alternative to the previously offered (either factual or CF and either high-level or low-level) explanation efor prediction ˆy. In addition, the set of replies to requests for (initial, non-alternative) explanation REP-explanation( ˆy) consists of the following items: •rep-why( ˆy): the system attempts to factually explain the prediction ˆyon the basis of the known features that led to that decision and offers a factual explanation if it is able to, or refuses to offer it, otherwise; •rep-why-not( ˆy,y): the system attempts to provide the user with a CF explanation for prediction ˆy for the given CF class yor refuses to offer it, otherwise. The set of replies to detailisation requests REP-detailisation( ˆy,e,)consists of the following items: •rep-detailisation( ˆy,e,): the system provides the numerical intervals over which the corresponding linguistic term of the requested explanation feature making part of explanation eis defined; 2Sets of requests are denoted using uppercase letters (as in, e.g., REQ-explanation) whereas single instances of requests are denoted using only lowercase letters (as in, e.g., req-detailisation).
CORRECTED PROOF 14 I. Stepin et al. / Information-seeking dialogue for XAI •rep-no-detailisation( ˆy,e,): the system refuses to provide numerical intervals on the requested feature’s linguistic term in explanation e, e.g. due to their unavailability. The set of replies to clarification requests REP-clarification( ˆy,e,)consists of the following items: •rep-clarification( ˆy,e,): the system provides the user with a definition of the requested feature making part of explanation efor prediction ˆyretrieving it from the knowledge base; •rep-no-clarification( ˆy,e,): the system refuses to clarify the requested feature making part of explanation efor prediction ˆydue to, e.g., its absence in the knowledge base. The set of replies to alternative explanation requests REP-alternative( ˆy,e)consists of the following items: •rep-alternative( ˆy,e): the system recognises the fact that the user is not satisfied with the offered (factual or CF) explanation efor prediction ˆy, seeks the most relevant alternative to it, generates and offers an alternative explanation to the user; •rep-no-alternative( ˆy,e): the system recognises the fact that the user is not satisfied with the offered (factual or CF) explanation efor prediction ˆy, seeks the most relevant alternative to it, but is unable to generate it. 4) Dialogue protocol. An explanatory dialogue between the system and the user is modelled following the rules specified in the dialogue protocol. The protocol determines turntaking rules, the rules governing user’s and system’s allowed moves at each stage of the explanatory dialogue, and the termination states of the dialogue. Thus, the locution types above are directly mapped to the speech acts produced by the system and the user as specified in the dialogue protocol. All of the aforementioned protocol rules are specified in Appendix B. 5) Knowledge store. Let Kbe the knowledge store which accumulates user’s knowledge w.r.t. explanations requested during his or her interaction with the system. Knowledge store Kis initialised to be an empty set: K=∅. When the system generates a factual or CF explanation (locutions explain-f (ˆy, E,ef)and explain-cf (ˆy,E,y,ecf ), as specified in the dialogue protocol), the corresponding piece of explanation is added to the knowledge store: K=K∪ef(ˆy) or K=K∪ecf (ˆy,y), respectively. The same applies to alternative explanations of either kind (locutions alter-f (ˆy,E,ef,e f)and alter-cf (ˆy,E, y,ecf ,e cf )). 6) Explanation store. Let Ebe the explanation store which tracks the current state of the explaineepreferred explanation throughout the dialogue. Explanation store Eis initialised to be an empty set: E=∅. Similarly to the knowledge store, a factual or CF explanation is added to the explanation store once generated: E=E∪ef(ˆy) or E=E∪ecf (ˆy,y), respectively. If the user finds the offered factual or CF explanation not satisfactory enough and asks for an alternative explanation (locutions whyalternative(ˆy,E,ef)and why-not-alternative(ˆy,E,y,ecf ), respectively), the corresponding explanation is removed from the explanation store: E=E\ef(ˆy) or E=E\ecf (ˆy,y), respectively. Noteworthy, the user cannot request an alternative explanation to any explanation non-offered previously. Further, the user can only submit explanation-related requests (detailisation, clarification, alternative) for the piece of explanation being processed. The resulting explainee-preferred explanation is the union of all the pieces of explanation found in the explanation store when a terminal dialogue state is reached. 7) Detailisation store. Let DET be the store that contains the features of the currently processed high-level explanation for which further details can be requested. DET is initialised to be empty, as the explanatory dialogue starts: DET =∅. The user can submit a detailisation request to the system only if a high-level (either factual or CF) explanation e=eh f|eh cf is being processed. Recall that for
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 15 each feature of the currently processed high-level explanation e, the feature is defined in terms of a linguistic variable mapped to the corresponding linguistic terms. When a new piece of high-level explanation is offered to the end user, DET is reinitialised with the set of features that the currently processed explanation contains: DET ={},∀∈e. The user can ask the system to provide him or her with the numerical intervals for the linguistic term of the given explanation feature only once during a sub-dialogue concerning a specific piece of explanation. Thus, the corresponding feature is eliminated from the detailisation store once the system has generated a response θ(locution elaborate(ˆy,E[,y], e,,θ)): DET =DET \{}.IfDET =∅, it is prohibited for the user to submit a detailisation request (locution what-details(ˆy,E[,y], e,)). When the user makes the final decision w.r.t. the system’s claim (i.e., either accepts or rejects it), the detailisation store is nullified: DET =∅. 8) Clarification store. Let CLAR be the clarification store that contains the explanation features whose meaning can be clarified. Similarly to the detailisation store, CLAR is initialised to be empty: CLAR =∅. When a new piece of explanation is offered, CLAR is populated with all the features that the explanation being processed e=eh f|eh cf |el f|el cf contains: CLAR ={},∀∈e. Noteworthy, the definitions for all the features that the dataset contains are precollected, mapped to one another by an expert or retrieved from a dictionary, and stored in the knowledge base. The user can ask to clarify a specific feature from the clarification store only once during a sub-dialogue concerning a specific piece of explanation. Then, the corresponding feature is eliminated from the clarification store after the system’s response υ(locution clarify(ˆy,E[,y], e,, υ)): CLAR =CLAR \{}.IfCLAR =∅, it is prohibited for the end user to submit a clarification request (locution what-is(ˆy,E[,y], e,)). When the user makes the final decision w.r.t. the system’s claim (i.e., either accepts or rejects it), the clarification store is nullified: CLAR =∅. 9) CF class store. Let CFS be the CF class store that contains all CF classes. It is initialised upon the successful execution of the factual explanation request (locution explain-f (ˆy,E,ef))sothatCFS =Y\{ˆy}for some prediction ˆy∈Y. The user is allowed to request a CF explanation for each class from CFS only once (locution why-not-explain(ˆy,E,y)). In addition, the user is allowed to ask for a (series of) alternative CF explanation(-s) for the same CF class (locution why-not-alternative(ˆy,E,y, ecf )as many times as there are alternative CFs for that class. Once a CF explanation is requested for some CF class y, it is eliminated from the CFS store: CFS =CFS \{y}. When the user makes the final decision w.r.t. the system’s claim (i.e., either accepts or rejects it), the CF class store is nullified: CFS =∅. 10) Knowledge Base. The knowledge base contains the dataset-related domain knowledge including a specification of all the dataset features (e.g., linguistic terms, the corresponding intervals, and definitions of all the features that the dataset contains). 3.2. Illustrative example Having introduced the proposed formalism for explanatory information-seeking dialogue modelling, let us now illustrate it taking the previously considered example for reference (see Table 1for details). Thus, we are considering the beer style classification problem for the beer dataset that contains the following classes: Ybeer ={Blanche, Lager, Pilsner, IPA, Barleywine, Stout, Porter, Belgian strong ale}. Table 2outlines the states of the detailisation, clarification, and CF class stores of the example explanatory dialogue after each dialogue move. Table 3outlines the states of the knowledge and explanation stores for the same example dialogue. Initially, the system claims that some instance of beer is of class Blanche (move m1). All the stores that make part of the dialogue model (K, E, DET, CLAR, CFS) are initialised to be empty. At the next
CORRECTED PROOF 16 I. Stepin et al. / Information-seeking dialogue for XAI Table 2 A move-by-move formal description of the stores governing the example of explanatory dialogue from Table 1 Move Locution DET CLAR CFS m1claim (ˆy,E)∅∅∅ m2why-explain (ˆy,E)∅∅∅ m3explain-f (ˆy,E,ef){colour, bitterness}{colour, bitterness}{Lager, Pilsner, IPA, Barleywine, Stout, Porter, Belgian strong ale} m4what-is(ˆy,E,ef,) {colour, bitterness}{colour, bitterness}{Lager, Pilsner, IPA, Barleywine, Stout, Porter, Belgian strong ale} m5clarify (ˆy,E,ef,,υ) {colour, bitterness}{colour}{Lager, Pilsner, IPA, Barleywine, Stout, Porter, Belgian strong ale} m6why-not-explain (ˆy,E,y){colour, bitterness}{colour}{Lager, Pilsner, IPA, Barleywine, Stout, Porter, Belgian strong ale} m7explain-cf (ˆy,E,y,ecf ){colour, bitterness}{colour, bitterness}{Lager, Pilsner, IPA, Barleywine, Porter, Belgian strong ale} m8what-details (ˆy,E,ecf ,) {colour, bitterness}{colour, bitterness}{Lager, Pilsner, IPA, Barleywine, Porter, Belgian strong ale} m9elaborate (ˆy,E,ecf ,,θ) {colour}{colour, bitterness}{Lager, Pilsner, IPA, Barleywine, Porter, Belgian strong ale} m10 why-not-explain(ˆy,E,y){colour}{colour, bitterness}{Lager, Pilsner, IPA, Barleywine, Porter, Belgian strong ale} m11 explain-cf (ˆy,E,y,ecf ){colour}{colour}{Lager, Pilsner, IPA, Barleywine, Belgian strong ale} m12 why-not-alternative(ˆy,E,y,ecf ){colour}{colour}{Lager, Pilsner, IPA, Barleywine, Belgian strong ale} m13 alter-cf (ˆy,E,y,ecf ,e cf ){colour, strength}{colour, strength}{Lager, Pilsner, IPA, Barleywine, Belgian strong ale} m14 accept-u (ˆy,E)∅∅∅ m15 accept-s (ˆy,E)∅∅∅ step, the user requests a factual explanation for the given prediction (m2). The system provides the user with a factual explanation (m3). As the factual explanation is generated, both DET and CLAR stores are populated with the corresponding features (colour and bitterness). Further, the piece of factual explanation ef(ˆy=Blanche)is placed to both the knowledge store and the explanation store. In addition, the CF store CFS is populated with all the CF classes. At the next stage, the user asks the system to clarify the notion of bitterness (m4) and receives the corresponding definition from the system (m5). As the clarification request for a given feature can only be submitted once while processing a specific piece of explanation, bitterness is then eliminated from the CLAR store. Once the factual explanation is offered, the user may commit to the factual explanation offered and inquire a CF explanation for some CF class. In the present example, the user seeks, at this stage, to know why the classifier did not predict the given beer to be Stout (m6). Then, the classifier presents the most relevant piece of CF explanation for this CF class in accordance with its ranking (m7). The CF explanation ecf (y=Stout)is then added to both the knowledge and explanation stores, whereas the class Stout is removed from the CFS store. Then, the DET and CLAR stores are updated with the features that the newly offered CF explanation contains. As the user requires more detailed information on bitterness (m8), the system retrieves the requested numerical interval over which the value of bitterness is defined to be high (m9). The feature bitterness is then removed from the DET store. Then, the user proceeds to request a CF explanation for class Porter (m10). Similarly to the previously offered explanations, DET and CLAR are updated accordingly, as the most relevant piece (from explainer’s point of view) of CF
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 17 Table 3 An example explanatory dialogue schema Block Move Utterance K E Cm1System: The test instance is of class y.∅∅ Em2User: Could you explain why you think so? ∅∅ m3System: It is of class ybecause feature1is term1.{ef(ˆy)}{ef(ˆy)} m4User: What do you mean by feature1?{ef(ˆy)}{ef(ˆy)} m5System: feature1is definition for feature1.{ef(ˆy)}{ef(ˆy)} m6User: But why is it not of class y?{ef(ˆy)}{ef(ˆy)} m7System: It would be of class yif feature1{ef(ˆy),ecf (y)}{ef(ˆy),ecf (y)} were term2and feature2were term3. m8User: Could you specify how feature1is defined? {ef(ˆy),ecf (y)}{ef(ˆy),ecf (y)} m9System: feature1is defined to be term2because {ef(ˆy),ecf (y)}{ef(ˆy),ecf (y)} it is found in the interval [term2min,term2max]. m10 User: But why is the test instance not of class y?{ef(ˆy),ecf (y)}{ef(ˆy),ecf (y)} m11 System: It would be of class y if feature1{ef(ˆy),ecf (y), ecf (y)}{ef(ˆy),ecf (y), ecf (y)} were term3and feature3were term3. m12 User: I am not quite satisfied with your explanation. {ef(ˆy),ecf (y), ecf (y)}{ef(ˆy),ecf (y)} Could you offer me another one? m13 System: Sure!Itwouldbeofclassy if... {ef(ˆy),ecf (y), ecf (y), e cf (y)}{ef(ˆy),ecf (y), e cf (y)} Tm14 User: Okay, I trust your prediction. {ef(ˆy),ecf (y), ecf (y), e cf (y)}{ef(ˆy),ecf (y), e cf (y)} m15 System: Thank you for your trust in me. Bye! {ef(ˆy),ecf (y), ecf (y), e cf (y)}{ef(ˆy),ecf (y), e cf (y)} In the left-hand side column (“Block”), C stands for claim, E – for explanation, T – for termination). explanation is generated and offered for the class Porter (m11). Then, the class Porter is excluded from the CFS store whereas the newly offered CF explanation is added to the knowledge and explanation stores. However, as the user is left dissatisfied or not convinced enough with the offered explanation, he or she inquires an alternative explanation to the previously offered one (m12). Then, the latest offered explanation is removed from the explanation store. Subsequently, if the next best ranked alternative can be offered, it is added to the explanation store (m13). The DET and CLAR stores are then updated accordingly. Having processed the presented explanations in their entirety, the user makes an informed decision that the classifier’s prediction can be accepted (m14). The system terminates the dialogue outputting a farewell locution (m15). Table 3generalises the presented example of explanatory dialogue for any dataset where features, linguistic terms, and classes serve as dataset-specific variables. It is possible to generalise any explanatory dialogue modelled in accordance with the proposed framework using the suggested template utterances. Noteworthy, three main building blocks of such explanatory dialogue (C – claim, E – explanation, and T – termination) can be distinguished. Figure 5presents the corresponding (partial, for illustrative purposes) parse tree of such a generalised explanatory dialogue. 3.3. Explanatory dialogue grammar (EDG) As follows from the example of dialogue presented in Section 3.2, the proposed dialogue model has a hierarchical structure with respect to its main building blocks. This observation allows us to reflect the modular composition of explanatory dialogue (following our model) in a context-free dialogue grammar. As the transitions between the states of the dialogue are finite and predefined, the use of the correspond-
CORRECTED PROOF 18 I. Stepin et al. / Information-seeking dialogue for XAI Fig. 5. A parse tree of the example of explanatory dialogue. Shaded nodes are non-terminals corresponding to specific speech acts. The subtrees in the dashed regions represent dialogue moves.
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 19 ing EDG allows us to (1) generate any explanatory dialogue that is valid in accordance with the dialogue protocol restrictions and (2) parse any actually valid explanatory dialogue or make a conclusion that the present explanatory dialogue is invalid with respect to the dialogue model constraints. Further, a grammar-based dialogue model can take into account modifications in the dialogue protocol if those are deemed necessary. In light of the above, we define an EDG following Chomsky’s definition of a context-free grammar as a tuple G=T,N,P,Swhere Tis the set of terminals, Nis the set of non-terminals, Pis the set of production rules (productions), and Sis the start token. In our model, Tcorresponds to a sentence actually uttered by each participant in the course of a dialogue. Nencompasses the internal building blocks of the dialogue as well as the speech acts involved (see the shaded nodes in Fig. 5for details). Thus, any explanatory dialogue is said to have three main building blocks (those corresponding to the nonterminals CLAIM, EXPLANATION, TERMINATION). In accordance with current legal requirements to explanation for AI, the block EXPLANATION enables the user to exercise the right to explanation and is made optional. All the non-terminals produced from the non-terminal EXPLANATION are designed in accordance with the predefined requests and responses (see Section 3.1 for details). In addition, P is composed in accordance with the dialogue protocol settings (see Appendix Bfor details). Note that productions can be subdivided in two groups, i.e., dataset-independent and dataset-specific productions. Dataset-independent production rules form the core of the proposed explanatory dialogue model and can be used in any application domain so long as it meets the settings of the classification problem as described in Section 2.1. The dataset-independent rules valid for the illustrative example of an explanatory dialogue are outlined in Appendix C. In turn, dataset-specific rules follow the structure of the given dataset and they are restricted by the information provided by the given interpretable rule-based classifier and the corresponding knowledge base. Finally, the start token Sis known to always be the non-terminal DIALOGUE node, i.e., the root node in the tree depicted in Fig. 5. 4. Process mining for dialogue analytics The proposed model of explanatory dialogue is designed in a top-down manner, which signals certain shortcomings. Thus, the dialogue protocol bases on the assumption that the taxonomy of requests and responses proposed in Section 3inspired by findings from the literature exhaustively covers user’s needs and system’s abilities when engaged in an explanatory dialogue. However, in the absence of any empirical evaluation, such assumptions may result being purely speculative. For example, specific requests may be utilised to a very limited extent or even not utilised at all. Alternatively, there may exist requests that are not included in the original model, which may nevertheless be considered essential for humanmachine interaction by the explainees. Either way, modifications to the model should be grounded on the data obtained from the end users. As such data-driven conclusions on the utility of the top-down dialogue model can only be made upon empirical evaluation, a user study is necessary to validate the proposed model. In addition to analysis of free-form user feedback, evaluation of a dialogue model can be automated by inspecting dialogue patterns in the collected dialogue transcripts. In these settings, dialogues can be treated as iterative processes whose key patterns allow us to discern strengths and weaknesses of the dialogue model. To analyse dialogues as processes, we propose a use of process mining techniques. Process mining is the subfield of data science that aims to provide tools for discovering insights into operational processes and thus supports process improvements [76]. Following the process mining terminology [50], an instance of a process (i.e., a specific explanatory dialogue) is denoted as a trace τ.
CORRECTED PROOF 20 I. Stepin et al. / Information-seeking dialogue for XAI Table 4 An example of an event log (the activities in bold are those produced by the system; the user-produced activities are those in italics) Case Activity Start End Dialogue1claim 2022-06-09 11:54:12 2022-06-09 11:54:12 Dialogue1why-explain 2022-06-09 11:54:12 2022-06-09 11:54:21 Dialogue1explain-f 2022-06-09 11:54:21 2022-06-09 11:54:22 Dialogue1what-details 2022-06-09 11:54:22 2022-06-09 11:54:42 Dialogue1elaborate 2022-06-09 11:54:42 2022-06-09 11:54:42 Dialogue1why-not-explain 2022-06-09 11:54:42 2022-06-09 11:55:58 Dialogue1explain-cf 2022-06-09 11:55:58 2022-06-09 11:56:00 Dialogue1what-details 2022-06-09 11:56:00 2022-06-09 11:56:32 Dialogue1elaborate 2022-06-09 11:56:32 2022-06-09 11:56:33 Dialogue1accept-u 2022-06-09 11:56:33 2022-06-09 11:57:28 Dialogue1accept-s 2022-06-09 11:57:28 2022-06-09 11:57:28 Dialogue2claim 2022-06-15 17:03:34 2022-06-15 17:03:34 Dialogue2why-explain 2022-06-15 17:03:34 2022-06-15 17:04:22 Dialogue2explain-f 2022-06-15 17:04:22 2022-06-15 17:04:23 Dialogue2what-is 2022-06-15 17:04:23 2022-06-15 17:04:50 Dialogue2clarify 2022-06-15 17:04:50 2022-06-15 17:04:50 Dialogue2why-not-explain 2022-06-15 17:04:50 2022-06-15 17:05:38 Dialogue2explain-cf 2022-06-15 17:05:38 2022-06-15 17:05:40 Dialogue2why-not-alternative 2022-06-15 17:05:40 2022-06-15 17:06:12 Dialogue2alter-cf 2022-06-15 17:06:12 2022-06-15 17:06:13 Dialogue2what-details 2022-06-15 17:06:13 2022-06-15 17:06:59 Dialogue2elaborate 2022-06-15 17:06:59 2022-06-15 17:07:00 Dialogue2reject-u 2022-06-15 17:07:00 2022-06-15 17:07:49 Dialogue2reject-s 2022-06-15 17:07:49 2022-06-15 17:07:49 Subsequently, each trace consists of the set of activities A(in this case, locutions). In turn, a specific instance (realisation) of an activity α∈A(i.e., a dialogue move) is referred to as an event ε. Altogether, a collection of explanatory dialogues makes up the so-called event log. An example of an event log basing on a collection of explanatory dialogues is depicted in Table 4. It contains two traces (i.e., Dialogue1and Dialogue2) that represent instances of the recorded explanatory dialogues between (possibly, different) user(-s) and the given system (i.e., an interpretable rulebased classifier). In total, the process model contains 22 events each of which is essentially a specific dialogue move paired with the corresponding locution. Figure 6illustrates the corresponding process model graph. The visual representation of the process model facilitates detection of the activity patterns (i.e., subprocesses characterising common parts of distinct dialogues) taking place in the collection of dialogues. A dialogue protocol can be represented as a finite state machine whose nodes are the locutions modelled, edges being legitimate transitions between different states of the dialogue (e.g., from a request to all possible responses). In terms of process mining, one can represent the dialogue protocol as the so-called process model – a directed graph M=N,Ewhere the set of nodes N⊆A∪{Start,End}is composed of the process activities and the set of edges E⊆N×Nrepresents (possibly, causal) relations between pairs of activities where Start and End are, respectively, the start and end time of execution of the corresponding activity.
CORRECTED PROOF I. Stepin et al. / Information-seeking dialogue for XAI 21 Fig. 6. The graphical view of the process model corresponding to the example Dialogue1in Table 4. To analyse the actually recorded dialogues quantitatively, we suggest that the so-called conformance checking procedure be applied. In process mining, conformance checking is applied to relate the events in the actually registered processes and the process model in order to identify commonalities and discrepancies between the former and the latter. In the case of evaluating the proposed dialogue game, all the moves made by both dialogue game players follow the previously defined dialogue protocol. Hence, no deviation from the protocol can be observed. Instead, conformance checking allows us to highlight the most (and the least) frequent dialogue patterns in the event log and evaluate it against the process model (i.e., the dialogue protocol). Conformance checking can lead to obtaining data-driven knowledge of the least frequently submitted requests and/or dialogue state transitions, which can be used to modify the originally proposed dialogue protocol in order to increase its quality. To sum it up, the proposed dialogue model can be evaluated in two complementary ways: qualitatively and quantitatively. On the one hand, qualitative free-form user feedback (e.g., in the form of a post-experiment survey) can point to missing requests or transitions between existing requests in the dialogue protocol. On the other hand, the least frequent dialogue patterns may signal their futility for explanatory purposes of the dialogue model. In process mining, a frequency threshold value can, for example, be set to subsequently optimise the process model by removing the least observed model patterns. Similarly, the least frequent requests or responses may be removed from the dialogue protocol if the empirically grounded threshold value is available and set prior to evaluation. As a result, process mining is shown to serve as a methodological basis for quantitative evaluation of the proposed dialogue model. In combination with free-form user feedback for qualitative evaluation of the dialogue protocol, process mining is able to provide us with further insights w.r.t. the quality of a dialogue model. 5. Experimental settings In order to evaluate the proposed model of explanatory dialogue following the aforementioned evaluation framework, we carried out an exploratory user study. In the remainder of this section, we describe the setup of the human evaluation study. Thus, Section 5.1 describes the datasets used as the basis for training the classifiers for the study. Section 5.2 outlines technicalities of the explanation generation method used in the given experiment. Section 5.3 outlines the distinctive characteristics of the classifiers trained on the aforementioned datasets. Section 5.4 discusses the stimuli selection as well as the design of the dialogue system used in the experiment. 5.1. Datasets In our study, we used the following three datasets: basketball player position [3], beer style [13], and thyroid disease diagnosis [19]. All three datasets serve to solve a multiclass classification problem in three different application domains. First, the basketball players position dataset presupposes
CORRECTED PROOF 22 I. Stepin et al. / Information-seeking dialogue for XAI five classes related to the following player positions: Ybasketball ={point-guard, shooting-guard, smallforward, power-forward, center}. Second, the beer style dataset (as was used in the illustrative example in Section 3.2) categorises instances of beer to belong to one of the following eight classes: Ybeer ={Blanche, Lager, Pilsner, IPA, Barleywine, Stout, Porter, Belgian strong ale}. Third, the thyroid disease dataset presupposes the following four potential labels: Ythyroid ={no hypothyroid, primary hypothyroid, compensated hypothyroid, secondary hypothyroid}. To guarantee consistent and comparable results, only numerical continuous features were used for training the corresponding classifiers. Further, all the features were mapped to linguistic terms as follows. The beer style dataset was annotated by an expert brewer, therefore it contains original feature-value partitions. The features from the other datasets were split in three uniform intervals of equal length, each of which was mapped to the following linguistic terms: low, medium, high(except for the feature height, which is described with 5 linguistic terms, in the basketball player position dataset). Table 5 summarises information on the features from all the datasets as well as the corresponding linguistic terms, with the numerical intervals attached. 5.2. Explanation generation method To evaluate the dialogue game proposed in this paper as a communication interface between the system and the user, we generate multiple factual and CF explanations using the XOR method [72]. This explanation generation method operates on the rule base (i.e., a set of decision paths to each class) of a rule-based interpretable classifier (e.g., a fuzzy rule-based classification system or a decision tree DT where branches are first transformed into a list of rules). All automatic explanations follow the structure of the decision path (in the case of the factual explanation) or the minimally different decision path leading to the given CF class (in the case of the CF explanation). The following pipeline of four steps constitutes the explanation generation process: (1) Rule vectorisation. Each rule found in the rule base is represented as a (binary, in the case of the XOR method) vector of all possible feature-value pairs. In the case of a DT, the values of the vector are all the unique conditions (e.g., “bitterness ⩽10”) found in the set of DT nodes. (2) Relevance estimation. Once the rules are vectorised, a distance is calculated between vectors representing the decision path vector (responsible for the prediction) and each rule leading to the given (factual or CF) class. In the case of the XOR method, the exclusive-OR function calculates the distance between the vectors. The vectors are then ranked in accordance with the distances. The minimally distant rule is selected as a template for the output explanation following the conventional definition of a CF explanation. (3) Linguistic approximation. Each interval found in the selected rule is mapped to the predefined linguistic terms by measuring the similarity between the set of numerical values corresponding to this interval and each set of numerical values for the corresponding feature. The most similar linguistic term is selected for the given feature. (4) Surface realisation. The linguistically approximated rule is passed on to the surface realisation module that outputs a template-based grammatically correct high-level explanation. Similarly, the corresponding numerical intervals are used to generate a low-level explanation. For DTs, factual explanations are essentially the feature-value intervals aggregated along the decision path. This explanation generation method presupposes that alternative factual explanations cannot be generated because alternative decision paths leading to the same predicted class would not adequately
[Document text truncated for crawler view.]