scieee AI-readable full text Open interactive document viewer

A systematic review of human-centered explainability in reinforcement learning: transferring the RCC framework to support epistemic trustworthiness

Moll, Maximilian; Dorsch, John

Abstract

This paper presents a systematic review of explainable reinforcement learning methodologies with an emphasis on human-centered evaluation frameworks. Drawing from literature between 2017 and 2025, we apply and extend the Reasons, Confidence, and Counterfactuals (RCC) framework—originally designed for supervised learning—to reinforcement learning contexts. Our analysis reveals two predominant explanatory strategies: constructive, where explicit explanations are generated, and supportive, where users must infer reasoning from provided visual or textual cues. Our review also emphasizes human factor considerations, like task complexity, explanation formats, and evaluation methodologies. Particularly, for the latter, our analysis shows that improvement of the quality of decision is rarely measured.

Full text

Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness Maximilian Moll1and John Dorsch2 1Universit¨at der Bundeswehr M¨unchen, Werner-Heisenberg Weg 39, Neubiberg, 85579, Germany. 2Center for Environmental and Technology Ethics, Prague (CETE-P), Institute of Philosophy, Czech Academy of Sciences, Jilsk´a 352, Prague, 11000, Czech Republic. Contributing authors: [email protected];dorsc[email protected]; Abstract This paper presents a systematic review of explainable reinforcement learning methodologies with an emphasis on human-centered evaluation frameworks. Drawing from literature between 2017 and 2025, we apply and extend the Reasons, Confidence, and Counterfactuals (RCC) framework—originally designed for supervised learning—to reinforcement learning contexts. Our analysis reveals two predominant explanatory strategies: constructive, where explicit explanations are generated, and supportive, where users must infer reasoning from provided visual or textual cues. Our review also emphasizes human factor considerations, like task complexity, explanation formats, and evaluation methodologies. Particularly, for the latter our analysis shows that improvement of the quality of decision is rarely measured. Keywords: XAI, Reinforcement Learning, Decision Support, AI ethics, epistemic trust 1 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w 1 Introduction Explainable Reinforcement Learning (XRL) has emerged as a critical research domain due to the increasing desire to integrate reinforcement learning (RL) agents into human-intelligent systems. As RL agents become integral to high-stakes decisionmaking processes—ranging from healthcare (Yu et al. 2021) and finance (Hambly et al. 2023) to autonomous vehicles (Irshayyid et al. 2024) —it is imperative that their actions and recommendations are transparent, interpretable, and trustworthy. Explainability in RL facilitates user understanding, enables informed human oversight, and supports effective collaboration by clarifying the rationale behind an agent’s behavior. Particularly, in contexts demanding robust decision support, XRL systems help human operators critically evaluate automated recommendations, calibrate trust, and maintain agency. It should be noted that XRL is only one part of explainable AI (XAI), the attempt to more widely provide explanations in artificial intelligence systems. However, since RL introduces its own set of peculiarities to this challenging overarching domain, XAI in general will not be discussed. For a good introduction to the topic, the reader is referred to Liao and Varshney (2021). Prior research in explainability for supervised learning has demonstrated the importance of explanations aligned with human epistemic processes. The Reasons, Confidence, and Counterfactuals (RCC) framework (Dorsch and Moll 2025), originally distilled from findings in supervised learning contexts, emphasizes that effective explanations must clearly articulate why decisions are made, how confident the system is in its recommendations, and what alternative decisions could have been made. Applying and extending this RCC framework to RL contexts presents unique challenges due to sequential decision-making dynamics, long-term strategic reasoning, and inherent uncertainties. As the deployment of RL agents expands into sensitive domains, growing attention has turned toward the ethical dimensions of AI trust. A key distinction has emerged between moral and epistemic trustworthiness (Dorsch and Deroy 2024). While moral trust presupposes social accountability and the ability to engage in normative responsibility—capacities current RL systems do not and arguably cannot possess—epistemic trust is a more attainable goal. Epistemic trust concerns the credibility of an agent’s informational contributions: it centers on an agent’s reliability in decision-making and its ability to transparently communicate the basis for its decisions in ways that support human understanding and judgment (Alvarado 2023). In the context of reinforcement learning, achieving epistemic trustworthiness requires explainability mechanisms that clarify not only what an agent is doing, but why it is doing it. XRL thus plays a crucial role in enabling users to understand, scrutinize, and appropriately calibrate their reliance on AI agents. This paper systematically reviews human-centered explainability methods in RL, explicitly focusing on empirical studies involving human participants. Given the ultimate goal of XRL is to improve human-AI collaboration in decision-making processes, publications with human studies are being selected, to evaluate explanatory methods not only on their theoretical coherence but also on their effectiveness in actual human-centered decision-support applications. 2 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w 2 Publication Collection The literature was systematically collected following a structured search strategy. The search was conducted using three major academic databases—Google Scholar, IEEE Xplore, and Scopus—and focused on the period from 2017 to 2025. To ensure comprehensive coverage of the field, the search utilized the following keywords: “Explainable Reinforcement Learning”, “Interpretable Reinforcement Learning”, “XRL”, “Explainability in Reinforcement Learning”, “Explainable Reinforcement Learning Visualization”, and “Causal Explainable Reinforcement Learning”. Initially, a preliminary exclusion criterion was applied, which required papers to either have a minimum of five citations in at least one database or to clearly demonstrate methodological advancement over existing techniques through comparative analysis with established pipelines. This step was intended to prioritize studies with recognized impact or clear novelty. Approximately 60 papers met this initial criterion. Following the preliminary screening, a secondary, more detailed evaluation was conducted. Papers were specifically assessed for the presence of user studies, as this review aimed to emphasize human-centered explainability methods. Papers that lacked such empirical evaluation involving human participants were subsequently excluded. Finally, two papers exclusively addressing interpretability without explicit explanatory components, (Wang et al. 2019;Kohler et al. 2024) were excluded from further analysis. This decision was based on the study’s primary interest in methods that provide explanations accessible to decision makers. An overview over the final selection can be found in table 1. 3 RCC Framework and Its Application to Reinforcement Learning The RCC framework was originally developed in response to a series of empirical findings in supervised learning models used in AI decision support systems (AI-DSS) (Dorsch and Moll 2025). These findings centered on the widely adopted switching percentage paradigm, where users are asked whether they would change their own prediction to match the model’s. The percentage of users who switch serves as a measure of how epistemically trustworthy the explanation is—not just whether it is understood, but whether it informs action. A clear pattern emerged: the way information is communicated matters as much as the information itself. In some cases (Ribeiro et al. 2018), textual explanations outperformed visual ones—even when the visualizations were more detailed—while in others, graphical formats proved more effective, provided they aligned with users’ reasoning strategies. The key takeaway is that effective explanation design must be grounded in the problem context, not just in model architecture. Supporting human reasoning in medium to high-stakes environments requires understanding how users make decisions and tailoring explanations accordingly. The core idea behind RCC is to align explainability with core epistemic norms that govern how people justify beliefs and assess credibility. Rather than expose internal model mechanics—which often benefit developers more than users—RCC focuses on 3 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w Table 1 Overview over all selected publications listing the extend to which the discussed methology provides - or can be used to deduce - reaons, counterfactuals, and confidence: yes: ✓, no: ✗, partially: ∼. Also the mode for communicating the explanation and the evaluation measure in the human study are being listed. Publication Reasons Counterfactuals Confidence Mode Evaluation Measure Greydanus et al., 2018 ∼ ∼ ✗Visual Satisfaction Scale, Kind of Switching Huang et al., 2018 ✗ ✗ ✓Visual Satisfaction Scale, Taking Control Iyer et al., 2018 ∼ ∼ ✗Visual Explanation Construction, Action Prediction Waa et al., 2018 ✓ ✓ ✗Text Satisfaction Scale Tabrez et al., 2019 ✓✗ ✗ Text Satisfaction Scale Madumal et al., 2020 ✓ ✓ ✗Text, Diagram Satisfaction Scale, Action Prediction Madumal et al., 2020 ✓ ✓ ✗Text Satisfaction Scale, Task prediction Puri et al., 2020 ∼ ∼ ✗Visual Decision Accuracy Sequeira et al., 2020 ✓∼✓Visual Perception of Aptitude Silva et al., 2020 ∼ ∼ ∼ Diagram Satisfaction Scale, Task prediction Cruz et al., 2022 ✓ ✓ ✓ Text Satisfaction Scale McCalmone et al., 2022 ✓✗ ✗ Diagram Action Prediction Milani et al., 2023 ✓ ✓ ✗Text Satisfaction scale Amitai et al., 2024 ✗∼✗Visual Identification of trigger event Festor et al., 2024 ✗ ✗ ∼Text, Numbers Switching Percentage, Usefulness Takagi et al., 2024 NA NA NA Visual Specific the kinds of information that support epistemic evaluation: users expect agents to articulate reasons, calibrate their level of certainty, and consider plausible alternatives. These are not just abstract ideals but functional epistemic practices that help people determine whether to accept or reject a claim. RCC treats these practices as design principles for human-centered AI systems. In what follows, we extend this framework to RL systems, where decision-making is sequential, policy-driven, and reward-based, and where the need for clear and intuitive explanation remains just as urgent. 3.1 Reasons The first component of the RCC framework is the need for a system to explain why a particular decision makes sense. In RL, this cannot be answered at a single level. RL agents operate across multiple interacting structures—states, actions, policies, and reward functions—each of which can serve as a legitimate source of reasoning. A comprehensive explanation requires clarity about which level is being referenced and how it contributes to the agent’s behavior. •State-Level Reasons These explain why the current state warrants particular attention, typically by 4 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w highlighting salient features of the environment. For example: “Thermal signatures indicate minimal civilian presence,” or “Sensor data suggests high enemy concentration in the eastern corridor.” These reasons often align with how human operators scan for situational cues and support intuitive situational awareness. •Action-Level Reasons These justify why this action was chosen in this state. They are usually grounded in predicted action-values (e.g., Q-values) or learned rules. For instance: “Choosing a flanking maneuver here yields a 30% higher expected success rate than a frontal assault.” Action-level reasons help connect technical output to decision-maker intuitions, especially in dynamic or adversarial settings. •Policy-Level Reasons Sometimes, the explanation requires reference to a broader behavioral pattern. Policy-level reasons describe why an agent tends to act in a certain way across contexts—for example: “The system prioritizes low-detection routes in urban terrain when enemy drone presence exceeds 20%, based on prior mission success rates.” These explanations make visible the strategy guiding local choices. •Reward Function-Level Reasons At the deepest level, reasons may refer to what the agent is optimizing. For example: “This recommendation reflects a reward structure that penalizes delays more heavily than exposure risk, prioritizing rapid neutralization.” Making this level of reasoning transparent is essential in sensitive domains like defense, where human values must align with what the system treats as optimal. These layers are not mutually exclusive; in practice, meaningful explanations often combine them. What matters is that users can see not only what the agent is doing, but why, in a form that maps onto the structure of human reasoning. 3.2 Confidence The second component of RCC is confidence—the system’s ability to indicate how certain it is about a given recommendation. In human-AI collaboration, confidence plays a pivotal role in trust calibration. It helps users weigh the system’s advice, calibrate reliance on its outputs, assess how much deference is appropriate, and determine when to intervene. In supervised learning, confidence is often tied to prediction probabilities. In RL, it varies with algorithmic structure: •In value-based methods, confidence can be inferred from the Q-value gap—the difference between the highest and second-highest expected returns in a given state. A large gap signals strong preference; a narrow one may suggest indecision. However, Q-value gaps can also reflect saliency, not true lack of confidence. •In policy-gradient methods, confidence is more naturally read from the probability distribution over actions. A sharply peaked distribution implies decisiveness; a flatter one reflects uncertainty or exploration. Crucially, confidence must be communicated carefully. Studies show that presenting confidence scores too early can introduce anchoring effects, biasing users toward overreliance—even when they have relevant expertise (Ghai et al. 2021;Ma et al. 5 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w 2023). To mitigate this, confidence should ideally be revealed after the user has formed an initial judgment, supporting critical reflection rather than passive acceptance or deference. 3.3 Counterfactuals The third component of RCC is counterfactual reasoning—the system’s ability to justify why it did not choose a different action. This contrasts with reasons (why this?) by shifting the focus to “why not that?”, revealing how the system weighs competing alternatives and signaling the presence of a meaningful causal structure. In RL systems, counterfactuals often draw from the same sources as reasons—such as action-values or policy expectations—but with a comparative framing. For example: “A ground assault was not chosen because its success rate under current conditions is 25% lower than an airstrike.” Such contrasts help users evaluate whether the selected action is robust or overly sensitive to specific assumptions. The key design question is whether users can actively interrogate counterfactuals—e.g., “Why not action X?”—or must infer them indirectly. Here, the standard is more flexible than for reasons. Even passive exposure to the consequences of unchosen actions—via simulated rollouts, comparative Q-values, or hypothetical trajectories—can effectively support user understanding. Ultimately, the goal is to provide contrastive clarity: to illuminate not just why the system acted, but why other options were rejected. This supports transparency, improves strategic oversight, and aligns explanation with how humans naturally evaluate decisions. 4 RCC in RL literature In this section we review the collected literature along the three dimensions of the RCC-framework. We will analyze their presence as well as potential short comings and gaps. 4.1 Reasons Based on the literature analyzed, two very different approaches can be observed. One group of papers aims to provide explanations directly, which we will call constructive. The remaining group presents helpful information to the user, who is then tasked to find the reasoning themselves, which we will call supportive. We will now discuss both in more detail. Clear representatives of constructive approaches are Madumal et al. (2020b,a); Milani et al. (2023). Here, full explanations are being generated based on different forms of relevant information. Madumal et al. (2020b,a) rely on a handcrafted actioninfluence graph and neural networks to predict which significant future action will be made possible by the chosen action and the amount of return obtained. Milani et al. (2023) replace the action-influence graph by an automatically inferred Bayesian network and expands the distal information to also include significant states. The studies show that these explanations are well received by participants. 6 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w Waa et al. (2018) follow a similar approach. However, an abstracted state space is used to make it more user friendly. Similarly, rewards are being transformed into outcome descriptions. While the primary focus of the paper is the production of counterfactuals, they also provide an approach for generating explanations that highlight the rough policy to be followed over a number of steps, the kind of abstracted states to be visited, and the expected positive and negative outputs. Thus, causal dependency is not explicitly modeled, but derived through abstractions of states and rewards. On the supportive side, there is first and foremost a group of papers using variations of saliency visualization (Puri et al. 2019;Iyer et al. 2018;Greydanus et al. 2018). Restricted to deep RL algorithms with visual input, these methods emphasize salient regions of the state image. Based on these visual cues, users are expected to infer the reasons behind action selection. While narrowing the scope of consideration from the entire state to its most relevant aspects can aid interpretability, a substantial understanding of the environment is still required. This limitation points to a broader issue: even when visualizations are helpful, they do not guarantee that users correctly infer the agent’s underlying policy. Indeed, none of the three studies discussed above provides a mechanism to determine whether users genuinely arrive at the correct explanation or merely guess the right action—raising a general concern in the field that correct decisions may not reflect true understanding. The primary difference among the three approaches lies in how each algorithm determines which parts of the state are most significant. One illustrative example is Puri et al. (2019), who approach this challenge by analyzing changes in the Q-value when features of the state are removed. To ensure that feature selection remains both specific and relevant, they require that a feature significantly influences the Q-value of the action being explained, but not the Q-values of alternative actions. This typically results in a small set of highlighted features, reducing the cognitive burden on users attempting to infer the rationale behind the action. Their approach is modestly supported by a human study, which found a higher rate of action prediction compared to comparable methods. A related approach is taken by Iyer et al. (2018), who pursue simplification through object-based grouping rather than feature reduction. Instead of eliminating individual features, they train their network on object representations and highlight those requiring the smallest change to affect the Q-value of the selected action. The study confirmed that these object saliencies matched the predictive value of more detailed information, effectively compressing the state’s complex information into a more concise representation. However, once again, there is insufficient evidence to determine whether participants’ correct choices were grounded in an accurate grasp of the decision-making process. In contrast to this ambiguity, Greydanus et al. (2018) provide more compelling evidence that saliency can enhance epistemic understanding. They used pixel-based saliency to highlight aspects of the input relevant to the learned policy, and participants were capable of distinguishing between overfitted and properly trained agents. This is not only practically useful, but also represents a preliminary step toward demonstrating that participants can infer correct reasoning, not merely correct actions. 7 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w Another group distinctly recognizable on the supportive side deals with policy extraction (Mccalmone et al. 2022;Takagi et al. 2024;Silva et al. 2020). The fundamental observation here is that in complex environments, it is often not obvious what the policy actually is—a necessary prerequisite for understanding it. Mccalmone et al. (2022) adopt a graph-based representation of the policy, in which actions are represented as edges. To achieve the necessary simplifications for making the policy more approachable, states are abstracted into nodes using natural language descriptions provided by users or domain experts. This results in a high-level summary of the policy, which can then be used to support understanding of the rationale behind each action. This was supported by findings that non-expert users were able to correctly align given states with their abstractions and successfully predict subsequent actions. Building on the idea of state abstraction, Takagi et al. (2024) adopt a similar strategy but replace natural language summaries with a latent-space representation. States are encoded using a variational autoencoder (VAE), clustered in latent space, and then the median of each cluster is mapped back to the input space. These representative states are then arranged in temporal order to form an abstracted policy trajectory. User studies showed no significant difference between these abstracted trajectories and full trajectories when participants were asked to match them with observed behavior, suggesting that the method effectively simplifies the policy while preserving interpretability. However, the approach provides no method for evaluating the reasons extracted from the automatically generated state abstractions. Moreover, selected actions do not feature prominently in the representations, limiting their utility for understanding action-specific rationale. As an alternative to abstraction-based simplification, Silva et al. (2020) explore the use of decision trees as policy approximators trained directly, rather than extracting an interpretable policy post hoc. Specifically, they replace the neural network in Deep Q-Networks with a differentiable decision tree, which is then discretized after training. In terms of performance, this approach proves successful, achieving results comparable to neural networks in the evaluated environments. Moreover, user studies indicate that participants clearly preferred working with decision trees over neural networks when performing action prediction tasks. While the approach does not generate explicit reasons, the transparent sequence of threshold-based decisions may assist expert users in forming their own explanations through deliberate analysis—though this remains an open question for future research. A final set of studies explores explanation through example behaviors (Sequeira and Gervasio 2020;Huang et al. 2018;Amitai et al. 2024). The core challenge in this approach lies in selecting examples that are genuinely informative—presenting a large number of potentially irrelevant demonstrations is unlikely to aid understanding. Instead, these methods focus on identifying particularly critical moments in the agent’s behavior. Sequeira and Gervasio (2020), for instance, identify “interesting” situations along four dimensions. The first includes highly frequent states and states with very little prior experience—where the former must be handled reliably to ensure overall performance, and the latter warrant scrutiny due to their novelty. The second dimension 8 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w involves execution certainty, assessed via the breadth of the action-selection probability distribution, and will be discussed in more detail below. Finally, the third and fourth dimensions target actions that lead to local maxima or minima in the Qfunction—states that are particularly beneficial or adverse, respectively. In addition to individual states, the most likely sequences leading to local maxima are also considered informative. However, human studies suggest that effectively balancing these types of examples is non-trivial: prioritizing one dimension can present an incomplete picture, while distributing attention evenly across all of them may overwhelm users and create confusion. Huang et al. (2018) take a more targeted approach, assuming that the human expert is already highly familiar with the task domain and needs only to verify the agent’s behavior in a few critical states. This enables the expert to assess whether the agent’s policy can be trusted or whether manual intervention is warranted. The goal, then, is to identify an appropriate set of such critical states. The authors report that selecting these states based on value—that is, choosing states where the correct action yields a higher Q-value than the average—outperforms a policy-based strategy that uses action-selection entropy as a proxy for uncertainty. Their study further shows that this value-based approach can indeed enhance perceived trust in the agent. While this method focuses on verifying alignment with human expectations, it implicitly assumes that deviations from those expectations signal incorrect behavior. Nonetheless, it could be fruitfully combined with earlier methods to generate explanations selectively, concentrating on only the most critical situations. In contrast to methods that present users with pre-selected examples, Amitai et al. (2024) introduce a logic-based querying system that enables users to specify the kinds of situations they wish to examine. Queries are defined by properties of the start and end frames, along with additional constraints, and are used to retrieve relevant clips from a pre-generated library of agent behavior. The study demonstrates that even non-expert users were able to quickly gain proficiency with this querying approach, positioning it as a promising method for enabling more interactive and user-directed exploration of policy behavior. The literature reviewed above demonstrates that most approaches engage with the dimension of providing reasons, albeit in varied ways. However, in many cases, a substantial part of the explanatory burden falls on the user. This may involve supplying the underlying causal model (Madumal et al. 2020b), or inferring reasons from highlighted features of the state (Puri et al. 2019). In more minimal approaches, users are presented only with pre-selected or self-selected examples, and there is currently little research evaluating whether the reasons inferred from such examples are actually valid. A further complication arises in example-based methods, where reasons are often conflated with counterfactuals—that is, with demonstrations of what would have happened under alternative actions. In summary, while many of the supportive approaches show promise, they tend to function as isolated components within a larger explanatory ecosystem—one that has yet to be fully developed or systematically studied. 9 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w Iyer, R., Li, Y., Li, H., Lewis, M., Sundar, R., Sycara, K.: Transparency and explanation in deep reinforcement learning neural networks. In: AIES 2018 - Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pp. 144–150. Association for Computing Machinery, Inc, ??? (2018). https://doi.org/10.1145/3278721. 3278776 Kohler, H., Delfosse, Q., Akrour, R., Kersting, K., Preux, P., Darmstadt, T.: Interpretable and editable programmatic tree policies for reinforcement learning (2024) Liao, Q.V., Varshney, K.R.: Human-centered explainable ai (xai): From algorithms to user experiences. arXiv preprint arXiv:2110.10790 (2021) Mccalmone, J., Le, T., Alqahtani, S., Lee, D., Lee, D..: Caps: Comprehensible abstract policy summaries for explaining reinforcement learning agents; caps: Comprehensible abstract policy summaries for explaining reinforcement learning agents. In: Nt’l Conf. on Autonomous Agents and Multiagent Systems (AAMAS). Online, ??? (2022). https://github.com/mccajl/CAPS Ma, S., Lei, Y., Wang, X., Zheng, C., Shi, C., Yin, M., Ma, X.: Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision-making. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023) https://doi.org/10.1145/3544548.3581058 Milani, R., Moll, M., Leone, R.D., Pickl, S.: A bayesian network approach to explainable reinforcement learning with distal information. Sensors 23 (2023) https://doi. org/10.3390/s23042013 Madumal, P., Miller, T., Sonenberg, L., Vetere, F.: Distal explanations for explainable reinforcement learning agents (2020) https://doi.org/10.48550/arXiv.2001.10284 Madumal, P., Miller, T., Sonenberg, L., Victoria, F.V.: Explainable reinforcement learning through a causal lens. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 2493–2500 (2020). https://doi.org/10.1609/aaai.v34i03.5631 . www.aaai.org Puri, N., Verma, S., Gupta, P., Kayastha, D., Deshmukh, S., Krishnamurthy, B., Singh, S.: Explain your move: Understanding agent actions using specific and relevant feature attribution. In: International Conference on Learning Representations (ICLR (2019). http://arxiv.org/abs/1912.12191 Ribeiro, M.T., Singh, S., Guestrin, C.: Anchors: High-precision model-agnostic explanations. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018) Sequeira, P., Gervasio, M.: Interestingness elements for explainable reinforcement learning: Understanding agents’ capabilities and limitations. Artificial Intelligence 16 Author’s Accepted Manuscript. This version has been accepted for publication in Human-Intelligent Systems Integration, but has not been copyedited or typeset. Please cite as: Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w 288 (2020) https://doi.org/10.1016/j.artint.2020.103367 Silva, A., Killian, T., Jimenez, I.R., Son, S.-H., Gombolay, M.: Optimization methods for interpretable differentiable decision trees in reinforcement learning. In: Proceedings of the 23rdInternational Conference on Artificial Intelligence and Statistics (AISTATS) (2020) Tabrez, A., Agrawal, S., Hayes, B.: Explanation-based reward coaching to improve human performance via reinforcement learning. In: 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE Press, ??? (2019) Takagi, Y., Tabalba, R., Kirshenbaum, N., Leigh, J.: Abstracted trajectory visualization for explainability in reinforcement learning. In: Proceedings - 2024 IEEE Conference on Artificial Intelligence, CAI 2024, pp. 75–82. Institute of Electrical and Electronics Engineers Inc., ??? (2024). https://doi.org/10.1109/CAI59869.2024. 00023 Waa, J.V.D., Diggelen, J.V., Bosch, K.V.D., Neerincx, M.: Contrastive explanations for reinforcement learning in terms of expected consequences (2018) Wang, J., Gou, L., Shen, H.W., Yang, H.: Dqnviz: A visual analytics approach to understand deep q-networks. IEEE Transactions on Visualization and Computer Graphics 25, 288–298 (2019) https://doi.org/10.1109/TVCG.2018.2864504 Yu, C., Liu, J., Nemati, S., Yin, G.: Reinforcement learning in healthcare: A survey. ACM Computing Surveys (CSUR) 55(1), 1–36 (2021) 17