scieee AI-readable full text Open interactive document viewer

A User Study on Explainable Online Reinforcement Learning for Adaptive Systems

Metzger, Andreas; Laufer, Jan; Pohl, Klaus

Abstract

Online reinforcement learning (RL) is increasingly used for realizing adaptive systems in the presence of design time uncertaintybecause Online RL can leverage data only available at run time. With Deep RL gaining interest, the learned knowledge is no longerrepresented explicitly, but hidden in the parameterization of the underlying artificial neural network. For a human, it thus becomespractically impossible to understand the decision making of Deep RL, which makes it difficult for (1) software engineers to performdebugging, (2) system providers to comply with relevant legal frameworks, and (3) system users to build trust. The explainable RLtechnique XRL-DINE, introduced in earlier work, provides insights into why certain decisions were made at important time steps.Here, we perform an empirical user study concerning XRL-DINE involving 73 software engineers split into treatment and controlgroup. The treatment group is given access to XRL-DINE, while the control group is not. We analyze (1) the participants’ performancein answering concrete questions related to the decision making of Deep RL, (2) the participants’ self-assessed confidence in giving theright answers, (3) the perceived usefulness and ease of use of XRL-DINE, and (4) the concrete usage of the XRL-DINE dashboard.

Full text

A User Study on Explainable Online Reinforcement Learning for Adaptive Systems ANDREAS METZGER ∗ ,paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Germany JAN LAUFER,paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Germany FELIX FEIT,paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Germany KLAUS POHL,paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Germany Online reinforcement learning (RL) is increasingly used for realizing adaptive systems in the presence of design time uncertainty because Online RL can leverage data only available at run time. With Deep RL gaining interest, the learned knowledge is no longer represented explicitly, but hidden in the parameterization of the underlying artificial neural network. For a human, it thus becomes practically impossible to understand the decision making of Deep RL, which makes it difficult for (1) software engineers to perform debugging, (2) system providers to comply with relevant legal frameworks, and (3) system users to build trust. The explainable RL technique XRL-DINE, introduced in earlier work, provides insights into why certain decisions were made at important time steps. Here, we perform an empirical user study concerning XRL-DINE involving 73 software engineers split into treatment and control group. The treatment group is given access to XRL-DINE, while the control group is not. We analyze (1) the participants’ performance in answering concrete questions related to the decision making of Deep RL, (2) the participants’ self-assessed confidence in giving the right answers, (3) the perceived usefulness and ease of use of XRL-DINE, and (4) the concrete usage of the XRL-DINE dashboard. CCS Concepts: •Software and its engineering →Designing software;•Computing methodologies →Machine learning. Additional Key Words and Phrases: adaptive system, machine learning, reinforcement learning, explainability, interpretability, debugging ACM Reference Format: Andreas Metzger, Jan Laufer, Felix Feit, and Klaus Pohl. 2024. A User Study on Explainable Online Reinforcement Learning for Adaptive Systems. 1, 1 (May 2024), 44 pages. https://doi.org/XXXXXXX.XXXXXXX 1 INTRODUCTION An adaptive system (a.k.a. self-adaptive system) can modify its own structure and behavior at run time based on its perception of the environment, of itself and of its requirements [ 25 , 56 , 65 ]. Examples of adaptive systems include elastic cloud systems [7], intelligent IoT systems [1], and proactive process management systems [34]. One key element of an adaptive system is its adaptation logic that encodes when and how the system should adapt itself. When developing the adaptation logic, developers face the challenge of design time uncertainty [ 4 , 38 ]. On the ∗Corresponding author. Authors’ Contact Information: Andreas Metzger, [email protected], paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Essen, Germany; Jan Laufer, [email protected]due.de, paluno (Ruhr Institute for Software Technology), University of DuisburgEssen, Essen, Germany; Felix Feit, [email protected], paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Essen, Germany; Klaus Pohl, [email protected]due.de, paluno (Ruhr Institute for Software Technology), University of Duisburg-Essen, Essen, Germany. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. ©2024 ACM. Manuscript submitted to ACM Manuscript submitted to ACM 1 2 Metzger et al. one hand, developers have to define when the system should adapt. To do so, they have to anticipate all potential environment states. However, this is infeasible in most cases due to incomplete information at design time. As an example, the concrete services that may be dynamically bound during the execution of a service orchestration and their quality-of-service characteristics are typically not known at design time. On the other hand, developers have to define how the system should adapt itself. To do so, they need to know the precise effect an adaptation has. However, the precise effect may not be known at design time. As an example, while developers may know in principle that enabling more features will negatively influence system performance, exactly determining the performance impact is more challenging. A large-scale survey about self-adaptation in industry performed by Weyns et al . [ 67 ] indicates that optimal design and design complexity together with design time uncertainty are the most frequently observed difficulties in designing adaptive systems in practice. 1.1 Online Reinforcement Learning for Adaptive Systems Online reinforcement learning (Online RL) is an emerging approach to realize adaptive systems in the presence of design time uncertainty [ 38 , 48 , 50 ]. Online RL means that RL is employed at run time. RL aims to learn an optimal action-selection policy, which is used to decide on which action (here: adaptation) 𝐴 to execute in any given environment state 𝑆 [ 60 ]. Using Online RL, adaptive systems can learn from actual operational data and thereby leverage information only available at run time. Based on run time monitoring data, RL receives a numeric reward 𝑅 for executing an adaptation. The reward expresses how suitable that adaptation was in the short term. The goal of Online RL is to maximize the cumulative reward received. Online RL thus helps to learn which adaptation 𝐴 to execute when faced with environment state 𝑆. Recent research on using RL for realizing adaptive systems leverages Deep RL algorithms, which represent their action-selection policy as an artificial neural network [ 19 , 35 , 62 ]. Benefits of Deep RL include that environment states are not limited to elements of finite or discrete sets, and that the used artificial neural networks can generalize well over unseen environment states. Some Deep RL algorithms can even capture concept drift in adaptive systems without the need to explicitly introduce mechanisms to observe such drift [36,48]. 1.2 Explainable Online Reinforcement Learning One key downside of Deep RL is that the learned action-selection policy is not represented explicitly, but is hidden in the parameterization of the artificial neural network. For humans, the decision-making of Deep RL thus essentially appears as a black box, as it is practically impossible to relate this parameterization to concrete RL decisions [ 51 ]. We thus require techniques to explain and interpret the internal workings of Deep RL and how its decisions are made [ 20 , 39 , 41 ]. Explaining the Deep RL decisions in the context of adaptive systems addresses different needs, including the following: • Explainability can help software engineers to debug the reward function by helping them to understand why Deep RL took certain decisions [ 17 ]. The successful application of Deep RL depends on how well the learning problem, and in particular the reward function, is defined [ 12 ]. Debugging of Deep RL is especially relevant for adaptive systems, because Online RL does not completely eliminate manual development effort. In particular, software engineers need to explicitly define a reward function, which quantifies the feedback to the RL algorithm. Such a reward function may be derived from a utility function that balances the various, often conflicting, goal dimensions. Getting such a utility function "right" – in a sense that it accurately reflects the trade-off among different goal dimensions – is a challenge [ 12 ]. Additionally, recent findings show that modeling the reward Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 3 function too closely on reality can slow down learning [ 36 ]. As a consequence, defining the reward function introduces a potential source for human error. • Explainability can aid system and service providers in regulatory compliance [ 41 ]. For example, in the EU, systems must comply with the relevant legal frameworks, such as the General Data Protection Regulation and the forthcoming AI Act. • Explanations facilitate users to build trust. They can understand how the system arrived at its results and thus can accept its results or not [41]. To provide insights into Online Deep RL’s decision-making for adaptive systems, we introduced the XRL-DINE technique [ 16 ]. XRL-DINE enhances and combines two existing explainable RL (XRL) techniques from the machine learning literature: Reward Decomposition [ 23 ] and Interestingness Elements [ 58 ]. Reward Decomposition uses a suitable decomposition of the reward function into sub-functions to explain the short-term goal orientation of RL, thereby providing contrastive explanations. Reward composition is especially helpful for the typical problem of adapting a system while taking into account multiple quality goals. Each of these quality goals could then be expressed as a reward sub-function. These explanations help to understand which goal an adaptation chosen by RL contributes to. Interestingness Elements collect and evaluate metrics at run time to identify relevant moments of interaction between the system and its environment. In particular when RL decisions are taken at run time, which is the case for Online RL for adaptive systems, monitoring all explanations to identify relevant ones introduces cognitive overhead. Interestingness Elements thereby facilitate selecting relevant actions. 1.3 Problem Statement and Contributions To assess the applicability and potential usefulness of XRL-DINE, we performed an initial evaluation in our previous work [ 16 ]. First, we prototypically implemented XRL-DINE using a state-of-the-art Deep RL algorithm, serving as proof-of-concept. Second, we demonstrated the potential usefulness of XRL-DINE by applying it to an adaptive web application exemplar [ 43 ]. Third, we measured indicators for the potential reduction in cognitive load required to interpret explanations depending on how XRL-DINE was configured. This initial evaluation indicated that XRL-DINE may be used by developers to gain insights and spot errors in the decision-making of RL-based adaptive systems. However, our initial evaluation did not directly take into account the "human factor" [13], meaning that we did not involve actual developers in our evaluation process. Our main new contribution compared to our earlier work in [ 16 ] is an empirical user study considering the "human factor" for evaluating XRL-DINE. We follow the human-grounded evaluation approach proposed in [ 13 ], which means that we conduct human-subject experiments involving simplified task that maintain the essence of the actual real-world application. In this user study, we involved 73 software engineers from academia and industry. We perform a comparative study, where the treatment group is given access to XRL-DINE, while the control group is not. In particular, we evaluate XRL-DINE considering the following main research questions: • RQ1 - What is the participants’ performance in answering concrete questions related to the decision-making of Deep RL? •RQ2 - How does the participants’ self-assessed confidence relate to the participants’ performance? •RQ3 - How do the participants in the treatment group perceive the usefulness and ease of use of XRL-DINE? •RQ4 - How did the participants in the treatment group use XRL-DINE? Manuscript submitted to ACM 4 Metzger et al. 1.4 Supplementary Material To facilitate reproducibility and replicability of our research, we provide supplementary material in an online repository at https://git.uni-due.de/rl4sas/xrl-dine. 1.5 Paper Organization Section 2provides relevant foundations. Section 3conceptually and formally describes the XRL-DINE technique. Section 4describes the proof-of-concept implementation of XRL-DINE. Section 5introduces the adaptive system exemplar we use as basis for our user study and gives a concrete example for applying XRL-DINE. Section 6describes the design of our user study, Section 7presents its execution, and Section 8presents and discusses its results. Section 9 discusses potential enhancements of XRL-DINE. Section 10 discusses related work. 2 FOUNDATIONS Below, we provide relevant foundations on online RL for adaptive systems, as well as explainable machine learning. 2.1 Online RL for Adaptive Systems 2.1.1 Reinforcement Learning (RL). In general, RL aims at learning an optimal action-selection policy 𝜋 of a so called RL agent by interacting with the RL agent’s environment, typically at discrete time steps. Upon executing an action 𝐴∈ A in state 𝑆∈ S at time step 𝑡 , the environment transitions to the next state 𝑆′ and awards a specific numeric reward 𝑅∈R based on a reward function R(𝑆, 𝐴)=𝑅 . The action-selection policy maps states S to a probability distribution over the action space, i.e., set of possible actions A . Formally, 𝜋 : S × A → [ 0 , 1 ] , giving the probability of taking action 𝐴 in state 𝑆 , i.e., 𝜋=P(𝐴|𝑆) . An optimal policy 𝜋 is a policy that optimizes the cumulative reward received [ 60 ]. XRL-DINE provides insights into the decision-making of a particular kind of Deep RL algorithms, so-called valuebased Deep RL algorithms – depicted in Fig. 1. In value-based RL, the action-selection policy 𝜋 depends on a learned action-value function, also called 𝑄 function, 𝑄(𝑆, 𝐴) . The action-value function gives the expected cumulative reward when executing action 𝐴 in state 𝑆 . Value-based Deep RL uses an artificial neural network to approximate 𝑄(𝑆, 𝐴) . During a policy update, the weights of the artificial neural network are updated using a trajectory of past actions, states and rewards [60]. Policy  Action A State S Reward R Action Selection Next state S’ RL Agent Action-value function Q(S,A) Policy Update Environment Fig. 1. Model of Value-based RL (adapted from [38]) Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 5 The actual policy 𝜋 in value-based Deep RL is realized by combining 𝑄(𝑆, 𝐴) with an action-selection function. The action-selection function determines for each state 𝑆 whether the RL agent should perform exploitation or exploration. During exploitation, the RL agent uses the learned knowledge to choose the best action based on 𝑄(𝑆, 𝐴) . During exploration the RL agent seeks to execute new actions. Obviously, actions should be selected that have shown to be effective (exploitation). However, to discover such actions in the first place, actions that were not selected before should be selected (exploration). Value-based Deep RL thus requires determining how to balance exploitation and exploration [ 48 ]. One typical solution is the 𝜖 -greedy mechanism. The action-selection function randomly chooses an action 𝐴∈ A with probability 𝜖 (exploration), and chooses the the action with the highest expected cumulative reward as determined by 𝑄(𝑆, 𝐴) with probability 1 −𝜖 (exploitation). To facilitate convergence of the learning process, one may implement a mechanism that decreases 𝜖over time (𝜖-decay), thereby reducing the amount of exploration. Based on the previous foundations, we illustrate the problem of explainability when using Deep RL. Fig. 2illustrates this problem using the classical RL toy example of cliff walk. In the cliff walk example, RL shall learn the optimal (i.e., shortest) path from start (S) to goal (G) while not falling off the cliff. Accordingly, falling off the cliff is penalized with a negative reward of 𝑅=− 100, while each step along the path is only penalized with 𝑅=− 1to encourage learning the shortest path. A classical RL algorithm to solve this problem is tabular Q-Learning, where the action-value function 𝑄 is represented explicitly in a table as shown on the left hand side of Fig. 2. Each cell gives the expected reward when performing action 𝐴 in state 𝑠∈𝑆 . This table clearly indicates that RL has learned to avoid falling off the cliff (actions that move the agent off the cliff have the lowest expected rewards) and to reach the goal in an optimal way (actions along the shortest path have the highest rewards). In comparison, the right hand side of Fig. 2shows the weights of the artificial neural network for Deep RL. State S Action A Learned Knowledge in Classical RL (here: Tabular Q-Learning) 24 25 26 27 28 29 30 31 32 33 34 35 UP -13,36 -12,57 -11,73 -10,74 -9,95 -8,91 -7,99 -6,98 -5,95 -4,92 -3,93 -2,98 RIGHT -12,00 -11,00 -10,00 -9,00 -8,00 -7,00 -6,00 -5,00 -4,00 -3,00 -2,00 -1,93 LEFT -12,99 -13,00 -11,98 -10,95 -9,99 -8,89 -7,97 -6,98 -5,94 -4,81 -3,94 -2,98 DOWN -13,95 -112,18 -112,80 -111,49 -112,13 -112,68 -112,91 -112,42 -111,81 -110,62 -112,86 -1,00 Action A State S: Learned Knowledge in Deep RL Falling off the cliff Reaching the Goal (G) Actions A = {UP, DOWN, LEFT, RIGHT} Reward States S= {0, …, 47} 011 24 35 Cliff Walk: Learn how to reach G from S, while not falling off the cliff Highest rewards Fig. 2. Illustration of how learned knowledge is represented for cliff walk example from [60]. 2.1.2 Adaptive systems. An adaptive system is capable of modifying its own structure and behavior at run-time based on its perception of changes in its environment, its requirements, and in the system itself [ 56 ]. Most software systems Manuscript submitted to ACM 6 Metzger et al. Adaptation Logic Analyze Monitor Execute Plan Knowledge Fig. 3. MAPE-K Model (adapted from [25]) exhibit some degree of adaptivity; e.g., they can respond to certain kinds of exceptions and error situations. To capture the key characteristics of the types of adaptive systems that we address, we use an emerging definition from software engineering research, which provides an architectural view-point [ 65 ]: An adaptive system can be structured into two main conceptual elements: the system logic and the adaptation logic. To define the conceptual elements of the adaptation logic, we follow the well-established MAPE-K reference model for adaptive systems depicted in Fig. 3. MAPE-K structures the adaptation logic into four main conceptual activities that rely on a common knowledge base [ 25 , 68 ]. These activities monitor the system and its environment, analyze monitored data to determine adaptation needs, plan adaptations, and execute these adaptations at run time. 2.1.3 Online RL for adaptive systems. Online RL applies RL during system operation, where actions have an effect on the live system, resulting in reward signals based on actual monitoring data [ 38 , 48 ]. Fig. 4depicts how the elements of value-based RL are integrated with MAPE-K to facilitate learning effective adaptations at run time. In the integrated model, action selection of RL takes the place of the analyze and plan activities of MAPE-K. The learned action-value function 𝑄(𝑆, 𝐴) takes the place of the adaptive system’s knowledge base. At run time, the policy is used by action selection to select an adaptation 𝐴 based on the current state 𝑆 determined by monitoring. Action Policy  Adaptation Logic realized using Value-based RL Execute Q(S,A) (Knowledge) Monitor Action Selection (Analyze+Plan) Policy Update Adaptation A State S Reward R Next state S’ Fig. 4. Integrated Model (adapted from [38]) Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 7 selection determines whether there is a need for an adaptation (given the current state) and plans (i.e., selects) the respective adaptation to execute. 2.2 Explainable Machine Learning Explainability is becoming increasingly important with the wide-spread use of deep learning for decision making [ 20 , 39 , 57 ]. While deep learning outperforms traditional machine learning (ML) models in terms of accuracy, one major drawback is that deep learning models essentially appear as a black box to its users and developers. This means that users and developers are not able to interpret a black-box ML model’s outcome, i.e., its decision, prediction, or action. Using such black-box models for decision-making entails significant risks [55]. Below, we provide definitions of the concepts "explanation", "explainability", and "interpretability" to serve as foundations for the remainder of this paper. Note, that there is not yet a consistent understanding and use of these terms. As an example, explainability and interpretability are two related but distinct concepts, but often used interchangeably in the ML literature. Explanation. The term "explanation" in explainable machine learning may refer to two concepts [ 39 ]: ( 𝑖 ) The process of explaining, and (𝑖𝑖) the result of this explanation process. Explanations may either be local or global [ 59 ]. A local explanation provides the causes for a concrete outcome of an ML model [ 63 ], thereby facilitating the understanding of the reasons for a specific decision [ 20 ]. A global explanation provides an understanding of the overall logic of the ML model, thereby allowing to follow the model’s entire reasoning leading to all the different possible outcomes [20]. We focus on local explanations. Explanation as a process involves (a) a cognitive process that determines the causes for an ML model’s outcome, and (b) a knowledge transfer process from the explainer to an explainee. Explainability. Explainabilty refers to the capability of being able to provide explanations. Explainability thus supports the aforementioned cognitive process. In other words, explainability characterizes the degree to which the cause for an ML model’s decision can be understood. Interpretability. A stronger concept than explainability is interpretability. Interpretability refers to the ease with which a human can understand the internal workings of an ML model [ 63 ]. Ideally, an interpretable ML model is one that is easy to understand and analyze, even for people who are not experts in the field of ML. Typical examples for interpretable ML models are decision trees or linear regression. Interpretability facilitates explainability, as understanding the internal workings of an ML model helps identify the causes for an ML model outcome. Both interpretability and explainability strongly depend on the explainee’s prior knowledge, experience and limitations [ 54 ]. For example, while a random forest (i.e., an ensemble of 𝑚 decision trees) is interpretable in principle, in practice one faces the explainee’s cognitive limitations, if, e.g., 𝑚is large or the number of input features is high. 3 THE XRL-DINE TECHNIQUE XRL-DINE provides insights into the decision making of Online Deep RL for adaptive systems. The main information provided by XRL-DINE to explainees are so called Decomposed Interestingness Elements (DINEs). DINEs allow explainees to understand the causes for RL’s decisions and thereby provide explanations why an adaptive system performed its adaptations. We first explain the two explainable ML techniques that XRL-DINE leverages as baselines: Reward Decomposition [ 23 ] and Interestingness Elements [ 58 ]. We then explain how these baseline techniques are adapted to deliver different types of Manuscript submitted to ACM 8 Metzger et al. DINEs focusing on different aspects of the decision-making process, and conclude with explaining the hyper-parameters of XRL-DINE. 3.1 Baseline Techniques and Limitations XRL-DINE combines Reward Decomposition [ 23 ] and Interestingness Elements [ 58 ] in such a way as to balance their respective limitations as follows. Reward Decomposition. Originally proposed to improve learning performance, Reward Decomposition was exploited in [ 23 ] for the sake of explainability. Reward Decomposition splits the reward function R(𝑆, 𝐴) of the RL agent into 𝑘 sub-functions R1(𝑆, 𝐴), . . . , R𝑘(𝑆, 𝐴) , called reward channels, which reflect a different aspect of the learning goal. For each of the reward sub-functions R𝑖(𝑆, 𝐴) a separate RL agent, which we call sub-agent, is trained, and which thus learns its own value-function 𝑄𝑖(𝑆, 𝐴) . To select a concrete action 𝐴 in state 𝑆 , an aggregated value-function 𝑄(𝑆, 𝐴) is computed by accumulating the action-values for each of the actions proposed by the different reward channels, i.e., ∀𝐴∈ A :𝑄(𝑆, 𝐴)=∑︁ 𝑖=1,...,𝑘 𝑄𝑖(𝑆, 𝐴) The resulting aggregated value-function 𝑄(𝑆, 𝐴) is then used for action selection, while trade-offs in decision making made by the composed agent become observable via the individual reward channels. When applied to adaptive systems, Reward Decomposition is especially helpful for the typical problem of adapting a system while taking into account multiple quality goals. Each of these quality goals could then be expressed as a reward sub-function. Explanations derived from reward decomposition help to understand which goal a chosen adaptation contributes to. However, no indication for the explanation’s relevance is provided, but instead it requires manually selecting the relevant time steps for which an RL decision should be explained. In particular when RL decisions are taken at run time, which is the case for Online RL for self-adaptive systems, observing all time steps and their explanations to identify the relevant ones can introduce a significant cognitive overhead for explainees. Interestingness Elements. The Interestingness Elements technique aims to facilitate the understanding of where the RL agent’s capabilities and limitations lie [ 58 ]. This is achieved by extracting interesting agent-environment interactions along a trajectory of time steps. XRL-DINE leverages two specific kinds of Interestingness Elements: "(un)certain executions" and "minima and maxima". (Un)certain executions: This Interestingness Element categorizes RL actions into whether the RL agent is certain or uncertain in its decision. The general intuition is that if an RL agent is uncertain in its decision in state 𝑆 , the RL agent will often perform many different actions when repeatedly faced with state 𝑆 , while if an RL agent is certain the same or similar actions are always performed. Or phrased differently, an RL agent is considered to be certain in its decision in state 𝑆if it is easy to predict the RL agent’s action 𝐴. To determine the uncertainty of a decision for a state 𝑆 , the evenness of the probability distribution over actions 𝐴∈ A is calculated. The probability distribution is approximated as: ˆ 𝜋(𝑆, 𝐴)=𝑛(𝑆, 𝐴)/𝑛(𝑆) Considering the observed trajectory of time steps, 𝑛(𝑆) is the number of times the RL agent was faced with state 𝑆 , while 𝑛(𝑆, 𝐴) is the number of times it executed action 𝐴 after observing 𝑆 . The evenness 𝑒(𝑆) and thus uncertainty for state 𝑆is then computed as: 𝑒(𝑆)=−∑︁ 𝐴∈A ˆ 𝜋(𝑆, 𝐴)ln ˆ 𝜋(𝑆, 𝐴)/ln |A| Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 9 An evenness of 𝑒(𝑆)=1indicates maximum uncertainty, while 𝑒(𝑆)close to zero indicates certainty. Minima and maxima: This Interestingness Element identifies where RL decisions led to a maximum or minimum reward, thereby helping to identify favorable and adverse situations for the RL agent. Maxima elements help understand how well the RL agent may perform with respect to its learning goal. Minima elements help understand how well the RL agent handles difficult situations. To determine minima and maxima, an estimate of the value function 𝑉(𝑆) is employed. 𝑉(𝑆) indicates the maximum value (expected reward) that can be achieved when taking a decision in state 𝑆 and is computed from the action-value function 𝑄(𝑆, 𝐴)as follows: 𝑉(𝑆)=max𝐴∈ A𝑄(𝑆, 𝐴) States with local minima Smin and local maxima Smax are computed as follows: Smin ={𝑆∈ S :∀𝑆′∈ S′:𝑉(𝑆) ≤ 𝑉(𝑆′)};Smax ={𝑆∈ S :∀𝑆′∈ S′:𝑉(𝑆) ≥ 𝑉(𝑆′)} S′is computed as follows: S′={𝑆′∈ S :∃𝐴∈ A :ˆ P(𝑆′|𝑆, 𝐴)>0} ˆ P(𝑆′|𝑆, 𝐴)is the probability of observing 𝑆′when executing action 𝐴in state 𝑆and is estimated by ˆ P=𝑛(𝑆, 𝐴, 𝑆′)/𝑛(𝑆, 𝐴) Considering the observed trajectory of time steps, 𝑛(𝑆, 𝐴, 𝑆′) is the number of times state 𝑆′ was visited after the RL agent executing action 𝐴in state 𝑆. When applied to adaptive systems, Interestingness Elements thus help reveal the RL agent’s confidence in its decisions, and thereby can can aid in debugging the reward function. As mentioned above, RL does not completely eliminate manual development effort as it requires to define a suitable reward function that helps the adaptive system to learn which adaptation to perform in which state. Accordingly, rewards quantify whether a chosen adaptation was a good one or not. If the RL agent is uncertain in a given state or only achieves rather low maxima this may point to a problem with how the reward function was defined; e.g., it may be that one needs to provide stronger rewards to guide the learning process towards more effective adaptations [36,69]. 3.2 “Reward Channel Dominance” DINE Building on Reward Decomposition, “Reward Channel Dominance” DINEs provide the information of how each subagent would influence each possible action 𝐴∈ A in the given state 𝑆 . “Reward Channel Dominance” DINEs thereby help understand why a concrete adaptation was chosen in a given state, and why not another possible adaptation was chosen. We introduce two types of “Reward Channel Dominance” DINEs: Absolute Reward Channel Dominance follows the original approach of Reward Decomposition and gives the actionvalues 𝑄𝑘(𝑆, 𝐴) for each action 𝐴∈ A of each reward channel 𝑘 for a given state 𝑆∈ S . Recall that the action values give the expected cumulative reward when executing action 𝐴 in state 𝑆 (see Section 2.1) and thereby determine which adaptation is chosen during exploitation in state 𝑆 . Fig. 5(a) shows an example of how this type of DINE is visualized in the XRL-DINE dashboard. The action values of the different reward channels are stacked for each possible action 𝐴∈ A, with the height of the bar indicating the aggregated action-value. Since each of the sub-agents has its own reward function 𝑅𝑘 and thus receives different kinds of rewards, the action values of the different sub-agents may have different ranges (e.g., including even negative values). This makes comparing Manuscript submitted to ACM 16 Metzger et al. Action 2 Action 2Action 4Action 4Action 4Action 2Action 2Action 1Action 3 Action 4Action 4Action 2Action 2Action 1Action 1 … Action 1 Action 2 Action 1 Action 1 4 3 Action 1 Action 2 Action 3 Action 4 Reward Channel 1 Reward Channel 2 Reward Channel 3 Action 1 Action 2 Action 3 Action 4 Reward Channel 1 Reward Channel 2 Reward Channel 3 Absolute Reward Channel Dominance Relative Reward Channel Dominance Action 1 Action 2 Action 3 Action 4 Action 5 Action 1 Action 2 Action 3 Action 4 Action 5 5 Reward Channel 1 Reward Channel 2 Reward Channel 3 … Maximum Minimum … 1 2 Fig. 10. XRL-DINE dashboard (5) Reward Channel Dominance for a selected time step is displayed in a stacked column chart. Each column represents a possible adaptation. The reward channels are shown in the same colors as for the other DINEs. 5 ADAPTIVE SYSTEM EXEMPLAR We apply XRL-DINE to a concrete adaptive system exemplar to serve as basis for our user study. We chose SWIM as an exemplar, which is one of the adaptive system exemplars provided by the SEAMS community [ 43 ]. SWIM simulates an adaptive multi-tier web application. It closely replicates the real-time behavior of an actual web application, while allowing to speed up the simulation to cover longer periods of real time. SWIM instantiates the auto-scaling problem in which the objective is to provision resources to satisfy conflicting business goals. Specifically, adaptation logic must be implemented that maximizes a given utility function, despite varying system load. SWIM has different monitoring metrics to determine the state 𝑆 of the system. These monitoring metrics include (1) the request arrival rate (i.e., “workload”), (2) the average throughput, and (3) response time. As all three environment variables are continuous, SWIM’s state space is continuous. Therefore, tabular RL solutions (see Section 2.1) cannot be directly applied to the exemplar. SWIM can be adapted via the following different adaptation actions 𝐴 : (1) additional web servers can be added / removed, resulting in the load being distributed across more / fewer servers; (2) the proportion of requests for which optional, computationally intensive content is generated (e.g., via recommendation engines) can be modified by setting a so called dimmer value. While both types of adaptations have an impact on user satisfaction (due to their influence on throughput and response time), adaptations of type (1) have an impact on costs (due to the costs of more/fewer servers), and adaptations of type (2) have an impact on revenue (due to recommendations leading to potential further purchases in the webshop). Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 17 5.1 Decomposed Reward Function To apply XRL-DINE, we took a reward function from the literature [ 42 ] and split it into three reward channels to form the following decomposed reward function: 𝑅=𝑎·𝑅user_satisf.+𝑏·𝑅revenue +𝑐·𝑅costs The weights were selected experimentally according to two criteria. First, no reward channel should dominate the other two reward channels to such an extent that the decisions of the other sub-agents have no influence on the choice of actions. The second, subordinate criterion is to conform as closely as possible to the original utility function. Specifically, the parameters were chosen such that User Satisfaction has the highest influence ( 𝑎= 4), Revenue has the second highest influence (𝑏=2), and Costs has the lowest influence (𝑐=1). The three reward sub-functions are defined as follows: 𝑅user_satisf. =             0.5 : 𝑥≤0.02 −0.5−𝑥−1 20 :𝑥≥1 0.5−𝑥−0.02 0.98 :otherwise In 𝑅user_satisf. , the perceived user satisfaction depends on the average latency 𝑥 . This utility function is provided as part of the SWIM exemplar. 𝑅revenue =𝜏·𝑎·(𝑑·𝑅𝑂+ (1−𝑑) · 𝑅𝑀) In 𝑅revenue , 𝜏 is the length of the time interval between two consecutive time steps, 𝑎 is the average arrival rate of requests and 𝑑 is the current dimmer value. 𝑅𝑀 is the reward obtained when processing a request without optional content. 𝑅𝑂 is the reward obtained when processing a request with optional content. The term 𝑑·𝑅𝑂+ ( 1 −𝑑) · 𝑅𝑀 thus represents the average reward controlled by the dimmer value for each request. 𝑅costs =−(𝜏·𝑐·𝑠) In 𝑅costs , 𝑠 indicates the number of servers currently in use. The parameter 𝑐 models the cost of using one server. This means, 𝑅costs is higher the fewer servers are used. Finally, all actions that cause a change in the system state (i.e., all actions except “No Adaptation”) are penalized by a reward of − 0 . 1to account for increased computational overhead and to incentivize the aggregated RL agent to learn a less jerky policy. 5.2 Selected Scenario To keep the cognitive load on the participants manageable and to keep the duration of the user study within reasonable limits (see the validity risk discussion in Section 8.6), we opted to select a single concrete scenario. We selected this concrete scenario based on two criteria: First, we wanted to keep the number of time steps relatively small, while still covering different possible types of decisions. In the end, we decided to limit the scenario to 21 consecutive time steps. Keeping the number of time steps low had the additional benefit that the XRL-DINE hyper-parameters only have a small effect on the number of Manuscript submitted to ACM 18 Metzger et al. DINEs shown. We thereby could keep the default values for these hyper-parameters as defined in the prototypical implementation, i.e., 𝜌= 0 . 3and 𝜙= 0 . 1, and thus did not have to consider them as independent variables in our study. Second, we wanted to understand how XRL-DINE helps to explain the decision-making of an already trained RL agent (i.e., an agent, where we mainly see exploitation and not random exploration; see Section 2.1). We thus start from a trained RL agent, for which we have experimentally measured convergence for a set of training data (a selected workload trace) by observing when average rewards stabilized. We then apply this trained RL agent to a new workload trace and randomly select 21 consecutive time steps towards the beginning of this new trace. Our chosen scenario3begins at time step 𝑇+1=22,575. 5.3 Applying XRL-DINE to the Selected Scenario Fig. 11 shows the contents of the main part of the XRL-DINE dashboard when applied to this scenario. The figure shows how the RL agent responds to an increase in user requests. This increase is reflected in the red curve in the XRL-DINE dashboard. This curve represents the normalized duration between incoming requests. The lower this value, the higher the request rate (i.e., the request rate is the inverse of the normalized duration between incoming requests). At time step 𝑇+ 2, there is a short peak in the request rate, which then decreases again. Between time steps 𝑇+ 4and 𝑇+ 6, the request rate increases again and then more or less remains at this higher level for the remainder of the example. The number of active servers is ten at time step 𝑇+ 1, but is lowered by the RL agent down to eight servers by time step 𝑇+5. The dimmer value is at 0.1 at time step 𝑇+1and is reduced by the RL agent to 0.0 by time step 𝑇+7. Increase Dimmer Remove Server Decrease Dimmer Remove Server Decrease Dimmer Decrease Dimmer Add Server Decrease Dimmer Add Server Add Server Add Server Add Server Add Server Add Server Remove Server Add Server Remove Server Remove Server Remove Server Remove Server Remove Server Add Server Add Server Add Server Add Server Add Server Active Servers Basic Response Time Basic Throughput Current Dimmer Is Booting Opt Response Time Request Arrival Mean Request Arrival Moving Mean Request Arrival Moving Variance Utilization Minimum Minimum Minimum Minimum Minimum Maximum Maximum Maximum Revenue Costs User Satisfaction No AdaptationNo AdaptationNo AdaptationNo AdaptationNo AdaptationNo AdaptationNo Adaptation 15 Time Step T + 1 T + 2 T + 3 T + 4 Avg Response Time Opt Throughput T + 5 T + 6 T + 7 T + 8 T + 9 T + 10 T + 11 T + 12 T + 13 REWARD CHANNEL: Fig. 11. XRL-DINE dashboard for scenario shown starting from time step 𝑇+1=22,575 3The interactive XRL-DINE dashboard for this scenario is available via https://git.uni-due.de/rl4sas/xrl-dine Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 19 The curve of the User Satisfaction reward channel depicts an increase between time steps 𝑇+ 1and 𝑇+ 3, which decreases again due to the shutdown of two servers after time step 𝑇+ 3. By lowering the dimmer value in time steps 𝑇+ 4and 𝑇+ 6and due to a more or less stable request rate, the reward for the User Satisfaction reward channel increases and stabilizes at a high level from time step 𝑇+ 6onward. Lowering the dimmer value also causes the reward for the Revenue channel to decrease and stabilize at a lower level. In contrast, the reward for Costs increases, since two less servers need to be operated. By choosing No Adaptation between time steps 𝑇+ 7and 𝑇+ 13, all reward channels receive a relative boost, as the − 0 . 1penalty for choosing any action except No Adaptation is no longer received. In summary, the aggregated RL agent’s strategy is to respond to an increase in the request rate by lowering the dimmer, thus trading a gain in User Satisfaction and Costs for a loss of Revenue. Between time steps 𝑇+ 2and 𝑇+ 13, the Revenue sub-agent decides to activate more servers. Yet, this decision is never taken by the aggregated RL agent until time step 𝑇+ 13, where the action chosen by the RL agent then is Add Server. The Costs sub-agent, in contrast, regularly wants to turn off more servers in the second half of the example, when the request rate drops again. This is in direct contradiction to the decision of the Revenue sub-agent. To better understand the aggregated agent’s internal decision making before the aggregated decision changes to "Add Server", we look at the “Reward Channel Dominance” DINEs for time step 𝑇+ 12, shown in Fig. 12. This DINE shows how the agent arrived at the decision to perform Add Server Action Value Absolute Reward Channel Dominance Revenue Running Costs User Satisfaction No Operation Increase Dimmer Decrease Dimmer Add Server Remove Server -20 -10 0 10 20 30 40 Relative Action Value Relative Reward Channel Dominance Revenue Add ServerIncrease Dimmer Remove Server 0 1 2 3 No Adaptation Costs Decrease Dimmer User Satisfaction Fig. 12. "Reward Channel Dominance" DINE for time step 𝑇+12 The action taken at time step 𝑇+ 12 is the last action before actually adding another server. In addition to the Revenue sub-agent, the Costs sub-agent also proposes an alternate action at this time step. The “Relative Reward Channel Dominance” DINE shows that the sub-agent for User Satisfaction has the greatest influence on the selected action No Adaptation. The other two sub-agents each show visibly less reward channel dominance. This imbalance suggests a strongly biased action selection. It can also be seen that the chosen action is closely followed by the second-best action Add Server. The alternative action Remove Server proposed by the Costs sub-agent is significantly worse, with visibly less reward compared to No Adaptation. In the example, the aggregated agent decides to sacrifice Revenue for higher User Satisfaction and lower Costs. Meanwhile, the sub-agent for Revenue suggests alternative actions. This is something one may expect when having knowledge of the domain. These suggestions all relate to the Add Server action. Adding more servers results in a lower average server utilization. This lower utilization would allow the dimmer to be raised again without compromising the dominant User Satisfaction reward. Thus, from a domain perspective, these alternative actions make sense. The proposed alternative actions of the Costs sub-agent also make sense from domain point of view, since removing servers intuitively leads to lower server costs. Manuscript submitted to ACM 20 Metzger et al. Summarizing, DINEs in this example facilitate explicitly representing the RL agent’s internal decision making. From a domain point of view, one can think of three reasonable strategies for responding to the unanticipated increase in request rate shown in the example: (1) more servers can be added to ensure a high level of Revenue and User Satisfaction but at the expense of Costs, (2) the dimmer value can be lowered to ensure a high level of User Satisfaction and Costs at the expense of Revenue, or (3) no action at all can be taken at the expense of User Satisfaction. In the chosen example, the agent follows the second strategy. When XRL-DINE is used for debugging, software engineers can evaluate whether this strategy is what they expected, and if not may change the reward function such as to prefer the alternative strategy suggested by the Revenue sub-agent. Of course, there may also be other causes for buggy behavior than a wrong reward function. For example, the state space may have been captured insufficiently (missing key state variables or including too many irrelevant ones), or the adaptation space may not have been completely defined (missing relevant adaptations). 6 USER STUDY DESIGN This section provides the research questions to be answered by our user study and the study design for each of these research questions. 6.1 Research Questions RQ1 - What is the participants’ performance in answering concrete questions related to the decision-making of Deep RL? The assessment of the quality of explanations is an active field of research [ 57 ]. Different techniques for assessing the quality of explanations exist, which include asking for a summary of the given explanation in someone’s own words, observing humans’ behavior when performing distinct tasks, or measuring the human task performance when using explanations. In our user study we are interested in measuring task performance by asking participants to answer dedicated questions related to the decision making of Online Deep RL. In particular, we compare task performance of a treatment group, which is given access to XRL-DINE, and a control group, which is given no access. We measure task performance in terms of effectiveness (rate of correctly solved tasks) and efficiency (number of correctly solved tasks per time). Accordingly, RQ1 entails the following sub-questions: •RQ1a - What is the participants’ effectiveness and how does it differ between the treatment and control group? •RQ1b - What is the participants’ efficiency and how does it differ between the treatment and control group? RQ2 - How does the participants’ self-assessed confidence relate to the participants’ performance? The participants are asked to self-assess their confidence in having correctly completed each task, thereby providing a complementary view to task performance measured in RQ1. Self-assessment measures the perceived understanding of participants [ 8 ], which can indicate how confident they are in their gained knowledge (here: about the decision making of the Online Deep RL agent). Self-assessed confidence may help identify areas of uncertainty in the explanations provided, as made evident by lower task performance. Accordingly, RQ2 entails the following sub-questions: • RQ2a - What is the participants’ self-assessed confidence and how does it differ between the treatment and control group? •RQ2b - How is the correlation between the participants’ self-assessed confidence and their effectiveness? RQ3 - How do the participants in the treatment group perceive the usefulness and ease of use of XRL-DINE? We are interested in how the participants perceive the usefulness and ease of use of XRL-DINE, which may provide indicators for the potential acceptance of XRL-DINE in practice and may exhibit opportunities for improvement. We Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 21 employ the Technology Acceptance Model (TAM) to measure these indicators [ 9 ]. TAM focuses on the technology users’ perception and on how likely they are to adopt and continue using the technology at hand. RQ4 - How did the participants in the treatment group use XRL-DINE? In this part of the study, we ask participants which part(s) of the XRL-DINE dashboard they used to solve a task. This helps analyze whether certain parts (and thus specific DINEs) are more popular than others for solving a task. Also, we ask the participants to identify problems in using the XRL-DINE dashboard. Together, this provides potential directions for future enhancements. Accordingly, RQ4 entails the following sub-questions: •RQ4a - Which XRL-DINE dashboard parts did the participants use? •RQ4b - What problems did the participants report when using the XRL-DINE dashboard? While RQ1 and RQ2 ask in how far XRL-DINE facilitates the study participants’ understanding concerning the decision making of the Online Deep RL agent, RQ3 and RQ4 ask about the usage and acceptance of XRL-DINE. 6.2 Overall Study Design The overall aim of our user study is to analyze in how far XRL-DINE may help software engineers to understand the decision making of an adaptive system realized using Deep Online RL. To this end, our user study follows the approach of a human-grounded evaluation [ 13 ], which means that we conduct human-subject experiments involving simplified tasks that maintain the essence of the actual real-world application. Like in the user study of Sequeira and Gervasio [ 58 ], we use an exemplar (here: SWIM) to evaluate the output of the respective explainable ML technique (here: XRL-DINE). We set up the user study in the form of an online questionnaire (for details see Section 7), because this facilitates repeatability (i.e., each participant receives always the same form of questions), simplicity (i.e., no need for filling and scanning paper-based results), and scalability (i.e., easy to share and distribute to additional participants) [ 27 ]. Moreover, by using an online questionnaire, the order of possible answers can be chosen randomly for each participant reducing the risk of response bias and order effects [ 27 ]. Another advantage is that the participants can participate in the study independently from each other and without being supervised. Finally, an online questionnaire allows us to provide the same information to each participant. Below, we introduce the design for each of the research questions and provide the questions that were posed in the online questionnaire together with our hypotheses. 6.3 RQ1 Design (Task Performance) RQ1 is about how software engineers perform when answering concrete questions concerning the decision making of Online Deep RL. We split the participants into two groups. The treatment group is using XRL-DINE, including all its DINEs, while the control group does not have access to any of the DINEs, but is only given the basic information about the Online Deep RL agent (such as its states, actions and rewards). Such a comparative user study appears to be unfair in principle. For the participants in the control group it is very challenging or even impossible to fully understand the decision making of the Online Deep RL agent (e.g., see Section 2.1). They may be lucky in doing some educated guessing (based on the basic information received), but have no realistic chance to solve tasks that entail the underlying decision-making policy of the Deep Online RL. Still, such a comparative study provides one essential insight. It contextualizes the results of the treatment group against a baseline, i.e., the results of the control group. Thereby, we can assess how “good” the results of the treatment group and thus the use of XRL-DINE are. Manuscript submitted to ACM 22 Metzger et al. Since our user study is executed in the context of a concrete adaptive system (SWIM), domain knowledge may help participants (of the control group and also the treatment group) to find the correct answers via educated guessing. To reduce the ability for educated guessing, one may remove all domain concepts from the information provided to the participants; e.g., by changing labels like User Satisfaction into XYZ. However, a study without domain context does not conform to the principle of human-grounded evaluation [13] and may also introduce new problems, such as participants having problems with handling the abstract labels. We thus decided to keep the domain concepts, and accept the risk that participants – of both groups – may do some educated guessing. Following the software engineering literature (e.g., see [ 6 , 49 ]), we measure human task performance in terms of effectiveness (RQ1a) and efficiency (R1b) as follows: effectiveness = Number of correctly performed tasks Number of all tasks =Í𝑐𝑖 𝑛 efficiency = Number of correctly performed tasks Time for performing all tasks [Minutes] =Í𝑐𝑖 Í𝑡𝑖 Here, 𝑐𝑖 is equal to 1 if the 𝑖 -th task ( 𝑖= 1 , . . . , 𝑛 ) was performed correctly, and 0 otherwise, while 𝑡𝑖 is the time spent (in minutes) for performing the 𝑖-th task. Questions. We ask the participants to perform eight tasks clustered into three task groups as shown in Fig. 13. The aim of each task is to answer a concrete question concerning the decision making of the Online Deep RL agent. The tasks cover different aspects of RL’s decision making. The participants had to solve the tasks and give a single-choice answer from a list of between three and seven answer choices. Task Group 1 Why is the adaptive system in a specific state? Task 1: How many adaptations were executed by the agent? [5] Task 2: Which action did the agent choose at timestep 10? [6] Task Group 2 Why did the RL agent make its decision (instead of an alternative decision)? Task 3: How often is the agent uncertain when making a decision? [6] Task 4: At timestamp 8 the agent is uncertain when making its decision. Why did the agent choose the adaptation „Decrease Dimmer“ instead of „Add Server“? [6] Task 5: At timestamp 15 the agent is uncertain when making its decision. Why did the agent choose not to adapt the system instead of „Add Server“? [7] Task Group 3 Which goals are pursued by the RL agent? Task 6: At timestamp 13 the agent is uncertain when making its decision. Which goal would the agent achieve when selecting „Remove Server“ instead of performing „No Operation“? [4] Task 7: At timestamp 9 the agent is uncertain when making its decision. Which goal would the agent achieve when selecting „Add Server“ instead of performing „Remove Server“? [4] Task 8: What is the main goal that the agent wants to achieve? [3] Fig. 13. Tasks for measuring participants’ effectiveness (RQ1a) and efficiency (RQ1b); number of answer choices are given in square brackets. Task Group 1: Tasks in this group are about determining why the adaptive system is in a specific state. In particular, participants have to differentiate between actions in general (i.e., every decision made by the Online Deep RL agent, which also includes “No Adaptation”; e.g., see Fig. 11) and actual adaptations (i.e., RL decisions that lead to a change of state). Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 23 Task Group 2: Tasks in this group are about determining why the Online Deep RL agent made a particular decision at a given time step and how certain the agent was about this decision compared with possible alternative decisions. In particular, participants had to grasp that the Deep Online RL agent may sometimes be rather uncertain, and in such situations identify possible alternative decisions. Task Group 3: Tasks in this group are about determining which goals are pursued by the Online Deep RL agent, and what different sub-goals may have been pursued when taking alternative decisions. Hypotheses. We hypothesize that the effectiveness (RQ1a) and efficiency (RQ1b) of the treatment group is significantly better than the effectiveness and efficiency of the control group. To determine whether there is a statically significant difference between the results of the treatment and control group [ 11 ], we use the Mann-Whitney U-Test, a nonparametric test, i.e., not assuming normal distribution of the results 4 . In addition to the 𝑝 value, which indicates how likely the observed difference between the treatment and control group is due to chance, we also provide the 𝑈 value, which indicates the size of the difference between the results. The smaller 𝑈 , the higher the probability that there is a difference. We also provide the effect size 𝑟 , which gives a normalized indication on how large the difference between the treatment and control group is (ranging from |𝑟|< 0 . 1indicating a small effect to |𝑟|= 0 . 5indicating a large effect). We adopt a significance level of 𝛼= 0 . 05 for all statistical tests. To execute the statistical tests we used DATAtab 5 , an online statistics calculator. DATAtab allows to import the raw user study data from Microsoft Excel. The raw user study data is part of the reproduction package provided (see Sec. 1.4). 6.4 RQ2 Design (Participants’ Confidence) We complement the measured task performance from RQ1 by asking participants to assess how confident they are in having completed the tasks correctly (similar to [ 8 ]). Such self-assessed confidence may indicate where the participants of the treatment group have problems with understanding the XRL-DINE explanations, and it can indicate where the participants of the control group have difficulties in understanding the decision making of Online Deep RL with the basic information they received. Questions. Following each concrete task from RQ1, the participants are asked how confident they are in having solved the previously given task correctly. To provide their response, participants are given a 4-point scale: Very unsure, A little unsure, Mostly confident, Very confident. We opted against a 5-point scale with a neutral answer to prevent central tendency bias. Hypotheses. We hypothesize that the self-assessed confidence of the treatment group is significantly higher than the self-assessed confidence of the control group (RQ2a), because the treatment group can resort to XRL-DINE for giving their answers, while the control group may have to rely on (educated) guessing. Also, we hypothesize that there is a significant positive correlation between self-assessed confidence and effectiveness, i.e., the rate of correct answers (RQ2b). To determine whether there is a statistically significant difference between the results of the treatment and control group, we use the Mann-Whitney U-Test like for RQ1. To determine the correlation between the self-assessed confidence and effectiveness, we use the nonparametric Spearman Rank Correlation coefficient, as we cannot assume normal distribution of the data (see RQ1 above). In addition to the correlation coefficient 𝑟 , we provide the 𝑝 value, which indicates how likely the observed correlation is due to chance. 4 Both numerical tests (i.e., Kolmogorov-Smirnov Test, Shapiro-Wilk Test, Anderson-Darling Test) and graphical tests (i.e., Histogram, Quantile-Quantile Plot) for normal distribution have shown that the results for both groups of participants are not normally distributed. 5https://datatab.net/statistics-calculator/ Manuscript submitted to ACM 24 Metzger et al. 6.5 RQ3 Design (Perceived Usefulness and Ease of Use) To measure the potential acceptance of XRL-DINE, we used the Technology Acceptance Model (TAM) [ 9 ]. TAM has been widely studied and emerged as a de-facto standard for analyzing technology acceptance [ 33 ]. Advantages of TAM are its low complexity, empirically validated measurement scales, as well as its robustness [ 26 , 44 ]. Research has also shown that the TAM can be used to evaluate software prototypes [ 10 , 29 , 44 ]. We use TAM to measure how the participants of the treatment groupd perceive the usefulness and ease of use of XRL-DINE. Perceived usefulness gives insights into how useful participants’ perceive XRL-DINE in supporting them to perform tasks related to the decision making of an Online Deep RL agent. Perceived ease of use provides an indicator about how usable, i.e., easy to use, XRL-DINE was perceived. Questions. Fig. 14 shows the TAM questions, which are based on the typical TAM questions but were slightly modified to fit the scope of the user study. Consistent with the TAM methodology, answers to each question are to be given on a 5-point-scale: extremely unlikely,quite unlikely,neither,quite likely, and extremely likely. Perceived Usefulness Faster Using the XRLDINE Dashboard in my job would enable me to accomplish tasks faster. Better Using the XRLDINE Dashboard would enable me to accomplish tasks better (e.g., with less errors). Increased Productivity Using the XRL-DINE Dashboard in my job would increase my job productivity. Increased Effectiveness Using the XRL-DINE Dashboard would increase my job effectiveness. Easier to do job Using the XRL-DINE Dashboard would make it easier to do my job. Useful Overall, I would find the XRLDINE Dashboard useful in my job. Perceived Ease of Use Easy to learn Learning to operate the XRLDINE Dashboard was easy for me. Easy to get results I found it easy to get the XRL-DINE Dashboard to do what I want it to do. Interaction clear How to interact with the XRLDINE Dashboard was clear and understandable. Visualization clear I find the visualization of the XRL-DINE Dashboard is clear and understandable. Easy to become skillful It was easy for me to become skillful at using the XRL-DINE Dashboard. Easy to use Overall, I find the XRL-DINE Dashboard easy to use. Perceived Usefulness Faster Using the XRLDINE Dashboard in my job would enable me to accomplish tasks faster. Better Using the XRLDINE Dashboard would enable me to accomplish tasks better (e.g., with less errors). Increased Productivity Using the XRL-DINE Dashboard in my job would increase my job productivity. Increased Effectiveness Using the XRL-DINE Dashboard would increase my job effectiveness. Easier to do job Using the XRL-DINE Dashboard would make it easier to do my job. Useful Overall, I would find the XRLDINE Dashboard useful in my job. Perceived Ease of Use Easy to learn Learning to operate the XRLDINE Dashboard was easy for me. Easy to get results I found it easy to get the XRL-DINE Dashboard to do what I want it to do. Interaction clear How to interact with the XRLDINE Dashboard was clear and understandable. Visualization clear I find the visualization of the XRL-DINE Dashboard is clear and understandable. Easy to become skillful It was easy for me to become skillful at using the XRL-DINE Dashboard. Easy to use Overall, I find the XRL-DINE Dashboard easy to use. Fig. 14. TAM questions 6.6 RQ4 Design (XRL-DINE Usage) To understand how participants of the treatment group use XRL-DINE, we analyze which parts of the dashboard they use for solving the tasks (RQ4a) and whether they face particular problems in using XRL-DINE (RQ4b). Questions. Concerning RQ4a, we ask "Which part(s) of the Dashboard helped you answer the previous question?" after the completion of each task from RQ1. Participants can give as answer one or several parts of the XRL-DINE dashboard or can also select “None” if they did not use any dashboard parts (e.g., because they performed educated guessing). Concerning RQ4b, we ask open-ended questions at the end of the online questionnaire in how far participants faced problems regarding the XRL-DINE dashboard. This may help spot problems related to user-interaction issues concerning how the XRL-DINE dashboard was designed. Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 25 7 USER STUDY EXECUTION Below, we introduce the technical setup of the user study, provide the overall structure of the online questionnaire, explain the user study procedure and provide descriptive statistics about the participants. 7.1 Technical Setup We selected the survey provider Limesurvey after comparing multiple survey providers 6 . In comparison to other survey tools, Limesurvey offers the possibility to measure the time the participants spent on each question, which we need to compute the participants’ efficiency (RQ1). The time spent for performing each task includes 1) reading the task, 2) finding the solution, and 3) submitting the solution. Before facing each task, a disclaimer text makes the participants aware of the upcoming task and time measurement. After finishing the task, the participants are informed that time measurement has stopped. The tasks of the user study were to be performed in the context of the SWIM exemplar (see Section 5). To this end, participants of the treatment group were provided with access to the interactive XRL-DINE dashboard instantiated for this exemplar. The control group were provided access to a reduced version of this dashboard, where all DINEs have been removed and only the chosen actions, states and rewards are shown. The reduced dashboard only consists of the parts (1)-(3) while part (4) and (5) have been removed (see Fig. 10). Both dashboard versions are part of the supplementary material (see https://git.uni-due.de/rl4sas/xrl-dine). 7.2 Online Questionnaire 7.2.1 Overall Structure. We have structured the online questionnaire into six main parts. In addition to parts that are directly related to the research questions from Section 6, we have added parts that provide additional information to the participants and also inquire additional background information from the participants (such as the participants’ experience and demographics). The control group receives the same questionnaire with the exception of the questions related to RQ3 and RQ4, because these questions concern XRL-DINE. Before participating in the study, participants are informed about the study purpose and procedure (see Section 7.3). They are also informed that their participation is voluntary, and they could stop participating at any time. Part 1. A welcoming-page introduces the study participants to the context of the study. The participants are asked to adopt the role of a software engineer who seeks to understand the decisions made by a Deep Online RL agent. After this introduction, the participants are asked to rate their experience in machine learning and reinforcement learning, as well as with self-adaptive systems, cloud computing and XRL-DINE. The answers are to be given on a 4-point scale: no experience,some experience,medium experience, and high experience. Part 2. The participants are introduced to the SWIM exemplar and XRL-DINE (treatment group only). The explanatory text presented in this part is available to the participants throughout the whole survey for reference. The participants are given the following pieces of information: •An explanation of the SWIM exemplar; •A list of the sub-goals, among which the RL agent should seek a trade-off; •A list of actions available to the RL agent together with examples of typical effects of these actions. 6https://www.limesurvey.org/ Manuscript submitted to ACM 32 Metzger et al. the treatment group participants not knowing the reward function, i.e., the treatment group can only guess the weights of the reward channels by analyzing the information provided by the XRL-DINE dashboard, which of course takes some time. However, this is just an assumption. We do not have any data from the user study to check or verify this assumption. Comparing Task Group 3 as a whole, no significant difference between the treatment and control group participants’ efficiency exists. The Mann-Whitney U Test across all tasks shows that the difference between the treatment and control group’s efficiency is not statistically significant, neither when considering all participants ( 𝑈= 445 , 𝑝 =. 397 ,𝑟 = 0 . 1), nor when only taking into account participants that solved all control questions correctly ( 𝑈= 127 , 5 , 𝑝 =. 886 ,𝑟 = 0 . 03). This means, the data does not allows us to draw a conclusion concerning our hypothesis that the efficiency of the treatment group is significantly better than the efficiency of the control group. Table 6. Results of the Mann-Whitney U Test comparing the efficiency of the treatment group and control group; significant results highlighted in bold. Task Group 1 Task Group 2 Task Group 3 Task 1 Task 2 Task 3 Task 4 Task 5 Task 6 Task 7 Task 8 U=264.5 U=295.5 U=120.5 U=246.5 U=278 U=229 U=273 U=240 p=.001 p=.007 p<.001 p=.001 p=.003 p<.001 p=.003 p=.001 r=0.40 r=0.32 r=0.59 r=0.42 r=0.37 r=0.43 r=0.37 r=0.40 U=247.5 U=141 U=488.5 p=.001 p<.001 p=.763 r=0.39 r=0.55 r=0.04 8.2 RQ2 Results (Self-Assessed Confidence) Below, we first present the participants’ self-assessed confidence for each task (RQ2a), and then analyze the correlation between the participants’ self-assessed confidence and their task performance (RQ2b). 8.2.1 Results for Self-Assessed Confidence (RQ2a). Table 7shows the mean results for the confidence self-assessment of the treatment and control group for each task. To quantify the confidence self-assessment of the treatment and control group, we mapped the confidence levels to numerical values (as shown in brackets after the confidence label). In the treatment group, 69% of the participants expressed that they are confident in solving the tasks correctly. Considering the numerical values, this maps to a mean confidence score of 2.96. When only taking into account participants that solved all control questions correctly, the confidence score rises to a mean value of 3.23. No clear trend can be identified for the control group (49% confident, 51% unsure). Their mean confidence score equals 2.53. Again, taking only participants into account that solved the control questions correctly, the confidence score is 2.74. Regarding Task Group 1, most participants of both treatment and control group assessed their confidence in solving the tasks correctly as “very confident”. In comparison, one can identify that the confidence level for participants of both treatment and control group decreases for tasks of Task Group 2 and 3. Especially for Task Group 2, participants of the control group assessed that they are unsure whether they solved the task correctly ( ≈ 75% unsure). More than 70% of the treatment group are still confident. Also for tasks from Task Group 3, the treatment group is more confident than the control group (63% treatment group vs. 40.6% control group). Table 8shows the results of Mann-Whitney U Tests comparing the self-assessed confidence of the treatment group and control group participants. As input for the Mann-Whitney U Tests we use the numerical values as already introduced. Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 33 Table 7. Confidence self-assessment for the treatment group and control group (shown in italics). Task Group 1 Task Group 2 Task Group 3 Task 1 Task 2 Task 3 Task 4 Task 5 Task 6 Task 7 Task 8 Mean Very unsure (1) 17% 2% 9% 15% 13% 9% 9% 33% 13% 9.5% 12.3% 17% 0% 0% 74% 47% 42% 47% 47% 5% 33% 0% 54.3% 33% A little unsure (2) 15% 19% 17% 19% 15% 19% 13% 28% 18% 17% 17% 20% 5% 0% 16% 21% 26% 16% 32% 32% 18% 2.5% 21% 26.6% Mostly confident (3) 24% 33% 30% 26% 30% 28% 28% 28% 28% 28.5% 28.6% 28% 11% 16% 5% 21% 16% 11% 5% 11% 12% 13.5% 14% 9% Very confident (4) 44% 46% 44% 41% 43% 44% 50% 11% 41% 45% 42.6% 35% 84% 84% 5% 11% 16% 26% 16% 53% 37% 84% 10.6% 31.6% The values either represent the participants’ confidence level for each task, the confidence level for the respective tasks in each task group, or the confidence level across all tasks. On the level of individual tasks, the participants’ confidence differs significantly for all tasks between the treatment and control group. However, based on the user study results we cannot identify a reason for why the participants of the treatment group report a significantly higher confidence for tasks 2, 4, 5, and 7 and why the participants of the control group report a significantly higher confidence for Task 1 and Task 8. When comparing the participants’ self-assessed confidence for each task group, a significant difference exists only for Task Group 2. Here, participants of the treatment group report a significantly higher confidence. However, the effect is small (𝑟=0.25). The Mann-Whitney U-Test across all tasks shows that the difference between the self-assessed confidence results of both treatment and control group was statistically significant, 𝑈= 330 , 𝑝 =. 022 ,𝑟 = 0 . 27. Taking into account participants that solved all control questions correctly, the difference is also significant and the effect size increases, 𝑈= 71 . 5 , 𝑝 =. 032 ,𝑟 = 0 . 38. Thereby, we can support our hypothesis that the self-assessed confidence of the treatment group is significantly higher than the self-assessed confidence of the control group, suggesting that the explanations of XRL-DINE foster self-assessed confidence in comparison to (educated) guessing. Table 8. Results of the Mann-Whitney U Test comparing the self-assessed confidence of the treatment group and control group; significant results highlighted in bold. Task Group 1 Task Group 2 Task Group 3 Treatment Group and Control Group DifferenceTask 1 Task 2 Task 3 Task 4 Task 5 Task 6 Task 7 Task 8 U=294 U=127.5 U=302 U=274 U=277.5 U=307.5 U=214.5 U=273 U=330 p=.022 r=0.27 p=.006 p<.001 p=.008 p=.003 p=.003 p=.01 p<.001 p=.003 r=0.36 r=0.59 r=0.35 r=0.36 r=0.36 r=0.32 r=0.46 r=0.37 U=365.5 U=344 U=371 p=.065 p=.035 p=.076 r=0.22 r=0.25 r=0.21 Manuscript submitted to ACM 34 Metzger et al. 8.2.2 Results for Correlation with Effectiveness (RQ2b). For the treatment group, we can observe a significant correlation between self-assessed confidence and effectiveness ( 𝑟= 0 . 53 , 𝑝 <. 001). Yet, for the control group there is no significant correlation between self-assessed confidence and effectiveness ( 𝑟= 0 . 17 , 𝑝 =. 499). Thereby, we can only partly support our hypothesis that there is a significant correlation between self-assessed confidence and task performance. The results suggest that while the explanations of XRL-DINE may facilitate the correct self-assessment of one’s own solution, the control group participants had problems in assessing the correctness of their answers due to the required educated guessing (e.g., several participants were falsely convinced that they had solved the task correctly). 8.3 RQ3 Results (Perceived Usefulness and Ease of Use) Fig. 16 shows the overall results for the perceived usefulness and ease of use as stacked bar charts. As can be seen, the participants rated their perception regarding the usefulness and ease of use of the XRL-DINE dashboard positively, i.e., 64% of all statements regarding the participants’ perceived usefulness was given a rating of Quite Likely or Extremely Likely. Only 15% perceive XRL-DINE as not being useful. Perceived ease of use came off slightly worse than perceived usefulness, but overall the positive ratings outweigh the negative ones here as well (53% positive vs. 28% negative). Perceived Usefulness 4% 11% 21% 49% 15% Perceived Ease of Use 0% 20% 40% 60% 80% 100% Extremely Unlikely Quite Unlikely Neither Quite Likely Extremely Likely 9% 19° *6 19^ 39% 14% Fig. 16. Overall results for TAM questions Fig. 17 shows a breakdown per individual TAM question as bar charts. Each bar provides the relative number of answers to each of the TAM questions. As can be seen, all questions were rated with at least 48% of positive votes. The best results were achieved for the statement “Using the XRL-DINE Dashboard would make it easier to do my job.” (72% positive ratings vs. 15% negative ratings). Yet, more than a quarter of all participants stated challenges when using XRL-DINE, i.e., issues with understanding the visualization and problems getting the dashboard to do what they wanted. Also, almost a third of all participants had problems to learn how to operate the XRL-DINE dashboard (31% negative ratings for the respective statement). Hence, although the perceived usability of XRL-DINE has been positive on average, there is potential for improvement, which discuss together with the feedback concerning the XRL-DINE usage for RQ4 below. 8.4 RQ4 Results (XRL-DINE Usage) Below, we first present which XRL-DINE dashboard the participants used (RQ4a), and then report what problems the participants mentioned when using the XRL-DINE dashboard (RQ4b). 8.4.1 Results for XLR-DINE Usage (RQ4a). Table 9shows the three most-used parts of the XRL-DINE dashboard as reported by the participants in the treatment group. Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 35 4% 4% 6% 2% 6% 4% 7% 6% 11% 7% 13% 7% 15% 19% 13% 7% 9% 6% 24% 20% 19% 19% 13% 22% 22% 20% 24% 28% 13% 19% 13% 22% 17% 26% 19% 19% 46% 50% 44% 50% 50% 52% 39% 41% 31% 37% 48% 39% 13% 7% 13% 13% 22% 20% 17% 11% 22% 11% 7% 13% Faster Better Increased Productivity Increased Effectiveness Easier to do job Useful Easy to become skillful Easy to use Perceived Usefulness and Perceived Ease of Use Extremely Unlikely Quite Unlikely Neither Quite Likely Extremely Likely Easy to learn Easy to get results Interaction clea Visualization clear Fig. 17. Results for individual TAM questions Table 9. Top three results for the Dashboard Part Usage of the Treatment Group Top1 Top2 Top3 Task 1 Trajectory of selected actions (89%) Uncertain Action DINE (11%) Reward Chart (6%) Task 2 Trajectory of selected actions (93%) State Progression (7%) Reward Chart (6%) Task 3 Uncertain Action DINE (85%) Trajectory of selected actions (28%) Reward Channel Extremum DINE (4%) Task 4 Contrastive Action DINE (46%) Uncertain Action DINE (28%) Trajectory of selected actions (20%) Task 5 Contrastive Action DINE (50%) Uncertain Action DINE (35%) Relative Reward Channel Dominance DINE (15%) Task 6 Contrastive Action DINE (46%) Uncertain Action DINE (31%) Relative Reward Channel Dominance DINE (24%) Task 7 Contrastive Action DINE (50%) Uncertain Action DINE (33%) Relative Reward Channel Dominance DINE (20%) Task 8 Reward Chart (52%) Uncertain Action DINE (26%) Reward Channel Extremum DINE (19%) For Task Group 1, the trajectory of selected actions is the most used dashboard part (Top1 for Tasks 1 and 2). In addition, participants reported they also used the Uncertain Action DINE, the State Progression and the Reward Chart. For Task Group 2, the participants state they mainly used the Uncertain Action DINE (Top1 for Task 3; Top2 for Tasks 4 and 5), followed by the Contrastive Action DINE (Top1 for Tasks 4 and 5). Besides these two DINES, the participants assessed they used the trajectory of selected actions as well as Reward Channel Extremum and Relative Reward Channel Dominance DINEs. For Task Group 3, results show participants used the Contrastive Action DINE (Top1 for Tasks 6 and 7) and the Reward Chart (Top1 for Task 8). Additionally, the Uncertain Action DINE was used for all three tasks from Task Group 3 (Top2). The results show that depending on the nature of task, participants choose different dashboard parts. They use Contrastive Action DINEs especially to solve tasks including “why”-questions (Tasks 4 and 5) and questions asking for a contrast (Tasks 6 and 7). To identify adaptations and actions, the trajectory of selected actions was used (Tasks 1 and 2). To determine whether the agent is uncertain when making a decision, the Uncertain Action DINE was chosen most often (Task 3). Finally, to identify the main goal of the agent, the Reward Chart was mostly used (Task 8). 8.4.2 Problems with XLR-DINE Dashboard (RQ4b). Below, we present relevant problems reported by the participants about the usage of XRL-DINE and selected suggestions on what kind of additional explanations would help the study participants to understand the decision making of the Online Deep RL agent. Manuscript submitted to ACM 36 Metzger et al. Only around a quarter of the participants stated they had problems with the XRL-DINE dashboard. The main problems mentioned were confusing graphs and lack of visual optimization. For example, one participant mentioned that “with all the lines in the state progression graph, [he is] not sure what information to use [...] as [he is] not sure how to interpret the data.” Another participant criticized that “some contrasting important interactions [are] yellow and some [are] red. The connection to the three goals is not immediately clear and the user might think yellow means danger and red means failure.” Suggestions for additional explanations included the following. One participant stated: “An explanation of the reinforcement learning agent’s action should include what action was chosen and for what reason. If possible, the reason should not only be technical, but also indicate the impact on the business goals. For example: The action “Add server” was chosen because the latency was higher than the defined value x. Above this value, a decrease in customer satisfaction is to be expected.” Another participant asked for a visualization of the states and agent interactions. Moreover, many participants asked for contrastive explanations, such as “If it chooses to lower the dimmer value, I want to know why it didn’t add another server, e.g.,[...]. Same goes for the other way around as well as similar decisions when scaling up/increasing the dimmer value.” 8.5 Discussion of Overall Results The results of the user study show that the study participants achieve significantly higher effectiveness and significantly higher self-assessed confidence across all tasks when having access to explanations given by XRL-DINE. Regarding the treatment and control group participants’ efficiency, there is no significant difference. When comparing the different task groups, the participants of the treatment group achieve significantly higher effectiveness in Task Group 2 (“why” questions) and Task Group 3 (“which goal” questions) compared to the control group participants. Regarding efficiency and self-assessed confidence, the treatment group participants achieve significant higher values for Task Group 2. We assume that having access to the explanations of XRL-DINE causes this difference. For Task Group 1, there was no significant difference in the effectiveness. However, the control group participants achieve a significantly higher efficiency compared to the treatment group. Theoretically, it should be possible to answer questions of Task Group 1 equally with the information provided by the XRL-DINE dashboard (treatment group) and the reduced dashboard used by the control group. A potential reason for the significantly higher efficiency of the control group participants might be that they required less time getting to know the reduced XRL-DINE dashboard functions compared to the full XRL-DINE dashboard functions used by the treatment group. Thus, the control group participants might have started solving the tasks earlier and therefore achieved a higher efficiency. However, we cannot verify this assumptions based on the data of the user study. All participants perceive XRL-DINE to be easy to use and useful. Participants suggested improvements in particular to the presentation of the explanations. We thus assume that explanations contribute positively to the ability of the participants to follow the RL agent’s decisions. Especially when answering why-questions (i.e., solving tasks from Task Group 2) the use of XRL-DINE’s explanations by the treatment group leads to significantly higher effectiveness, efficiency, and self-assessed confidence when answering compared to not having access to explanations (participants of the control group). 8.6 Validity Risks Our user study inherits the typical validity risks of user studies on task performance and assessing explanations (e.g., see [8,47]). We discuss them along the different types of risks below. Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 37 Internal Validity. There is a risk that participants may not seriously participate or may try to over-perform. To address this risk, we designed the study as an online questionnaire to be conducted within 30–45 minutes (e.g., by selecting a concrete scenario of manageable size), which Daun et al. [ 8 ] assume to be an adequate time to reduce the risk participants lose interest during participation. Also, we did not offer the option of pausing and returning to the questionnaire to avoid participants spend arbitrary long times for answering the questions. Since the participants were asked to participate directly by the second author, social pressure could be involved, which is associated with the risk of biased results. To minimise this risk, the study was conducted anonymously. In addition, the online questionnaire format allowed participants to participate in the study at a time and place of their choice. As a result, the authors of this paper are unable to track who actually participated and how well each person performed. Another risk is that the participants self-assessed experience (see Section 7.4) is affected by the the Dunning-Kruger effect [ 28 ], where participants with lower experience in a given area may overestimate their abilities. We took into account this uncertainty when establishing that the experiences of the treatment and control group is comparable, and also did not use the self-assessed experience for any further analyses which may have been affected by this uncertainty. Concerning the answer choices along a scale (e.g., in RQ2 and RQ3), we used verbal scale labels instead of numeric labels to provide a concise meaning of the choices. We also started with the negative choice first, as studies have shown that people tend to select the first option that fits within their range of opinion. By doing so, our study thus leads to more “conservative” results (cf. [37]). Construct Validity. We designed the survey in such a way that it provides participants of treatment and control group with a detailed introduction and thus equal grounding for answering the questions. As we explained in Section 7.2, the control group did not have access to XRL-DINE, but was given the reward function instead (which the treatment group did not see), thereby providing a more challenging baseline to compete against. One evident risk is the potential impact of the GUI design of the XRL-DINE dashboard and the given descriptive texts on the study results. Our study especially evaluates XRL-DINE including the given XRL-DINE dashboard as the way to transfer the explanations from the XRL-DINE engine to the human. Thus, we cannot distinguish between the effect of the explanations and the explanation transfer (i.e., visualization). This would require expanding our user study with evaluation approaches from human-computer interaction, for instance. The participants’ eight tasks consisted in answering different concrete questions. To reflect various insights into the decision-making of Deep RL, we covered different typical types of questions: "what/which", "why", and "how many". Still, these questions only represent a subset of possible questions. Another risk are the differences in size, demographics and experience of the treatment and control group. Controlling demographics by firstly collecting data around demographics and experience to secondly splitting the participants into two even groups obtaining a more even distribution between treatment and control group was not possible because the control group study was performed separately to address reviewers’ feedback. We therefore used statistical tests which support the comparison of results of unequal sample sizes [ 32 ]. Also, we carefully selected the participants of the control and treatment group such as to have comparable demographics and assessed that they have comparable experience levels as suggested by Doshi-Velez and Kim [ 13 ]. In addition, we made sure that the participants of the treatment and control group were disjoint to eliminate learning effects. Finally, we statistically measured significance of the comparative results to determine in how far one can draw conclusions from the differences in results. External Validity. To strengthen external validity, we used an actual adaptive system exemplar (SWIM) and trained the RL agent realizing SWIM’s adaptation logic using real-world workload traces. We tuned the hyperparameters of Manuscript submitted to ACM 38 Metzger et al. the RL algorithm experimentally, using educated guessing (e.g., comparable to [ 40 ]). We purposefully did not perform extensive, exhaustive hyperparameter tuning, e.g., using grid search, because our aim was not to improve or compare the performance of existing RL approaches, but to validate how XRL-DINE may be used to generate explanations for RL decisions. Still results are only for a single system and a selected scenario of 21 time steps with fixed values for 𝜙 and 𝜌 (see reasoning in Section 5.2), which thus limits generalizability. We used participants with various degrees and job types, providing a cross-section of some typical software engineering personnel. We could measure a positive effect of using XRL-DINE for all participants, however it appears that we could not measure a statistically significant effect across the different degrees and job types. If we were to generalize our findings for specific types of software engineers, an enlargement of the user study thus would be needed. 9 ENHANCEMENTS Below we discuss potential enhancements of XRL-DINE. Natural-language explanations. The literature distinguishes two major types of explanation formats [ 31 ]: (i) visual explanations, including graphical user interfaces, charts, data visualization, or heatmaps, and (ii) natural-language explanations. With the exception of the Contrastive Action DINE (see Section 3.4), XRL-DINE provides visual explanations. Compared with visual explanations, the benefits of natural-language explanations reported in the literature include (1) better understandability for people with diverse backgrounds as well as non-technical users, (2) increased user acceptance and trust, and (3) more efficient explanations [ 31 ]. One enhancement of XRL-DINE thus is to provide more of such natural-language explanations. A particular promising direction is to leverage the capabilities of modern AI chatbots (such as ChatGPT), which can provide natural-language answers to any natural-language question posed to them. However, as a downside of this flexibility, the underlying large language model may "hallucinate", i.e., generate nonsensical text unfaithful to the provided source input. This means AI chatbots may deliver explanations that do not faithfully explain the decision-making of RL, i.e., the explanations may exhibit low fidelity. In addition, AI chatbots may provide different explanations for the very same question asked, i.e., the explanations may exhibit low stability. Careful prompting and hyper-parameter tuning may be one direction to deliver high fidelity and stability of explanations [ 35 ]. Explainability in the presence of delayed rewards. Environment dynamics may delay the effects of adaptations. As an example, in the SWIM exemplar (introduced in Section 5) executing the Add Server action requires booting a new server, which takes time. As a result, the effect on latency and thus on rewards for executing this adaptation are delayed. In contrast, lowering the dimmer value has an immediate effect on latency and thus on rewards. Such timing-related differences may make interpreting DINEs more difficult. To address these timing-related differences, XRL-DINE may for instance be extended by the RUDDER approach for decomposition of delayed rewards [2]. Explaining long trajectories. While our user study focused on explaining a short trajectory of Deep RL decisions of 21 time steps, XRL-DINE in principle can be used to help explain long trajectories. To this end, the hyper-parameters 𝜌 and 𝜙 may be changed in such a way that depending how long the trajectory is, the number of generated DINEs remains manageable. As we showed in Section 3.6, the hyper-parameters allow tuning the rate of generated DINEs in a wide range. An interesting direction would be to auto-tune these hyper-parameters, having the user prescribe the number of DINEs to be generated. Also, for explaining such long trajectories, other forms of visualizations than what currently is offered in the XRL-DINE dashboard may be preferred; e.g., instead of showing "Reward Channel Extremum" DINEs as overlays to the received reward charts, a mere list of these DINE may be preferable. Explanations for decentralized adaptive systems. In a decentralized adaptive systems, the adaptation logic is decentralized across multiple systems [ 14 , 52 ]. As XRL-DINE is built to explain the decisions of a single RL agent, Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 39 XRL-DINE does not consider the decisions of other RL agents during explanations. Similarly to the aforementioned situations in which XRL-DINE may generate difficult to understand DINEs, the same may happen if XRL-DINE is directly applied to decentralized adaptive systems. Extending XRL-DINE depends on the fundamental, underlying approach of decentralization. On the one hand, such a decentralization may follow different patterns for how the MAPE-K elements are decentralized across systems [ 52 ]. On the other hand, there are different technical ways to decentralize the learning and decision making of Online RL, including Multi-Agent RL [ 45 ], hierarchical RL [ 5 ], and meta RL [70]. Explanations for policy-based RL to capture concept drifts. XRL-DINE works for value-based Deep RL, because XRL-DINE needs access to the learned action-value function 𝑄(𝑆, 𝐴) . As explained in Section 2.1, value-based RL typically relies on 𝜖 -greedy strategies to address the exploration-exploitation dilemma. This poses the challenge of when and how to increase 𝜖 again in order to capture concept drifts in the system environment. Concept drifts occur due to an evolution of the system environment that lead to a change of the effects of adaptations. As an example, if the physical machines that provide the virtual servers in the SWIM exemplar would be replaced by less powerful, but more energy efficient machines, this would impact on the effect of the Add Server action in terms of latency. Coping with such concept drift in value-based Deep RL would require to observe such concept drift and to increase the exploration rate 𝜖 to learn the changed effect of adaptations. In contrast, policy-based Deep RL [ 46 ] can naturally cope with such concept drifts. The fundamental idea of policy-based RL is to directly use and optimize a parameterized stochastic action-selection policy 𝜋𝜃(𝑆, 𝐴) in the form of a deep artificial neural network. Using stochastic action selection from this policy, policy-based RL captures concept drifts of the system environment without the need for software engineers to intervene [ 36 , 48 ]. However, this poses the question on how to generate DINEs from 𝜋𝜃(𝑆, 𝐴) instead from 𝑄(𝑆, 𝐴) . 10 RELATED WORK While generic explainable RL approaches are discussed in recent overview papers (such as [ 22 , 51 ]), these do not specifically address adaptive systems. We thus focus the following discussion on solutions that specifically provide explanations for adaptive systems, which can be clustered into the following main categories: Temporal graph models. This category of work uses temporal graph models as central artifact to derive explanations [ 18 , 61 ]. As proposed by García-Domínguez et al . [ 18 ], such a model may be used for so called forensic self-explanation. This means that the model may be queried via a dedicated query language. In addition, such a model may be used for so called live self-explanation. Here, one can submit queries to the running system and be presented with a live visualization. The underlying temporal model is kept up to date at run time (i.e., employed as a model at run time). The approach is comparable to the Interestingness Elements described in Section 2since explanations are generated based on execution traces. However, in contrast to XRL-DINE, interesting interactions must be extracted by manually writing queries using the provided query language. As follow-up work, suggestions for automating the selection of interesting interaction moments are proposed. Ullauri et al . [ 61 ] fully automated automate this by using complex-event-processing. While the aim of Ullauri et al . is to select interesting interactions to keep the size of the models at run time manageable, the aim of XRL-DINE is to reduce the cognitive load of developers. Compared to earlier work on temporal graph models for explainable adaptive systems, [ 61 ] stands out in providing explanations for RL decisions. While Ullauri et al . hint at the possibility of using model at run time queries to realize reward decomposition, it differs from XRL-DINE in that the combination of Interestingness Elements and Reward Decomposition is not considered. Manuscript submitted to ACM 40 Metzger et al. Goal models. This category of work uses goal-based models at run time [ 3 , 64 ]. Bencomo et al . [ 3 ] use higher-level system traces as explanations. Again, this can be considered similar to the idea of Interestingness Elements. Welsh et al . [ 64 ] employ a domain-specific language for providing explanations in terms of the satisficement of softgoals. In this regard, explanations are comparable to Reward Decomposition explanations introduced by Juozapaitis et al . [ 23 ] as described in Section 3in that the explanations refer to competing goal dimensions. Other than XRL-DINE, these techniques require making assumptions about the environment dynamics at design time, which can be a source for error due to design time uncertainty [ 66 ]. Also, in contrast to XRL-DINE, this category of work does not explicitly consider RL. Provenance graphs. Reynolds et al . [ 53 ] employs interaction data collected at run time to generate explanations in the form of provenance graphs. A provenance graph contains information and relationships that contributed to the existence of a piece of data. By keeping a history of different versions of the provenance graph, it is possible to determine at run time if and how the model has changed (using model versioning) and who has changed the model and why (using the provenance graph). Provenance graphs quickly can become too complex to be meaningfully interpreted by humans, thus a dedicated query language was introduced that allows extracting information of interest. Again, in contrast to XRL-DINE, this category of work does not consider RL. Anomaly detection. Ziesche et al . [ 71 ] suggest using machine learning to detect anomalous behavior of an adaptive system which may require an explanation. Again, this is similar to the idea of Interestingness Elements, which allows determining relevant points for explanation. In addition, they reduce the typically huge search space of possible reasons for such anomalous behavior by classifying the behavior into classes with similar reasons. Thereby, their approach can be considered a first step towards so called "self-explainable" systems that autonomously explain behavior that differs from anticipated behavior. Again, in contrast to XRL-DINE, this category of work does not explicitly consider RL. Explainability as a tactic. Li et al . [ 30 ] take a fundamentally different view on explainability. A formal framework is proposed in which explainability is not provided externally but is considered a concrete tactic of the adaptive system. In uncertain or difficult situations, the adaptive system can ask assistance from a human operator in making a decision rather than acting itself. Specifically, the overall system is modeled as a turn-based, stochastic multiplayer game in which three players participate. These players are (1) the actual self-adaptive system, (2) the environment, and (3) the operator. This game is then analyzed using a probabilistic model checker to determine when the involvement of a human operator is necessary. To prevent the operator from being permanently consulted, there is a cost to using this tactic that must be accounted for by the model checker. In this respect, this approach is similar to XRL-DINE, which aims to reduce the cognitive burden of the human. However, the motivation in XLR-DINE is the limited cognitive ability of the human, while the motivation for Li et al . is the time delay caused by involving an operator. Also, RL is not explicitly considered. Explainable online reinforcement learning. In our previous work [ 16 ], we introduced XRL-DINE providing detailed explanations of RL decisions at relevant points in time. XRL-DINE enhances and combines two existing explainable RL techniques from Juozapaitis et al . [ 23 ] and Sequeira and Gervasio [ 58 ]. The explainable RL techniques were proposed in isolation and not tailored to the needs of explaining adaptive systems. In [ 16 ], we thus proposed a combination and adaptation of these generic techniques. In contrast to our previous work, this paper adds the design, execution and analysis of a user study. Manuscript submitted to ACM A User Study on Explainable Online Reinforcement Learning for Adaptive Systems 41 11 CONCLUSION We introduced XRL-DINE, a technique that helps understanding the decision making of Online Deep RL for adaptive systems. We described the prototypical implementation of XRL-DINE using a state-of-the-art deep RL algorithm, serving as proof-of-concept. We used an adaptive systems exemplar to demonstrate the use of XRL-DINE, measure indicators for the cognitive load of using XRL-DINE. In particular, we performed a comparative user study involving 73 software engineers from academia and industry. Results show that XRL-DINE helps to correctly perform tasks (treatment group: 76% correctly executed tasks; control group: 45% correctly executed tasks) with a statistically significant difference between the treatment and control group. Moreover, these tasks can be performed in reasonable amount of time (on average 0 . 66 correct tasks per minute or in other words on average 01:15 minutes to give a correct answer). The analysis of the participants’ self-assessed confidence shows that 69% of the treatment group are confident that they solved a task correctly with a significant difference to the self-assessed confidence of the control group (only 49% confident). Additionally, XRL-DINE is perceived useful by the majority of participants and considered usable, with indications for future enhancements. ACKNOWLEDGMENTS We cordially thank the participants of our user study. Research leading to these results received funding from the EU’s Horizon 2020 and Horizon Europe R&I programmes under grant agreements 101070455 (DynaBIC) and 871493 (DataPorts). REFERENCES [1] Iván Alfonso, Kelly Garcés, Harold Castro, and Jordi Cabot. 2021. Self-adaptive architectures in IoT systems: a systematic literature review. J. Internet Serv. Appl. 12, 1 (2021), 14. [2] Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter. 2019. RUDDER: Return Decomposition for Delayed Rewards. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 13544–13555. [3] Nelly Bencomo, Kristopher Welsh, Pete Sawyer, and Jon Whittle. 2012. Self-Explanation in Adaptive Systems.. In 17th IEEE Intl Conf on Eng. of Complex Computer Systems, ICECCS 2012. [4] Radu Calinescu, Raffaela Mirandola, Diego Perez-Palacin, and Danny Weyns. 2020. Understanding Uncertainty in Self-adaptive Systems. In ACSOS 2020. IEEE, 242–251. [5] Mauro Caporuscio, Mirko D’Angelo, Vincenzo Grassi, and Raffaela Mirandola. 2016. Reinforcement Learning Techniques for Decentralized Self-adaptive Service Assembly. In Service-Oriented and Cloud Computing - 5th IFIP WG 2.14 European Conference, ESOCC 2016, Vienna, Austria, September 5-7, 2016, Proceedings (Lecture Notes in Computer Science, Vol. 9846), Marco Aiello, Einar Broch Johnsen, Schahram Dustdar, and Ilche Georgievski (Eds.). Springer, 53–68. [6] Mariano Ceccato, Alessandro Marchetto, Leonardo Mariani, Cu D. Nguyen, and Paolo Tonella. 2015. Do Automatically Generated Test Cases Make Debugging Easier? An Experimental Assessment of Debugging Effectiveness and Efficiency. ACM Trans. Softw. Eng. Methodol. 25, 1 (2015), 5:1–5:38. [7] Tao Chen, Rami Bahsoon, and Xin Yao. 2018. A Survey and Taxonomy of Self-Aware and Self-Adaptive Cloud Autoscaling Systems. ACM Comput. Surv. 51, 3 (2018), 61:1–61:40. [8] Marian Daun, Jennifer Brings, Patricia Aluko Obe, and Viktoria Stenkova. 2021. Reliability of self-rated experience and confidence as predictors for students’ performance in software engineering. Empir. Softw. Eng. 26, 4 (2021), 80. [9] Fred D Davis. 1989. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS quarterly (1989), 319–340. [10] Fred D Davis and Viswanath Venkatesh. 2004. Toward preprototype user acceptance testing of new information systems: implications for software project management. IEEE Transactions on Engineering management 51, 1 (2004), 31–46. [11] Francisco Gomes de Oliveira Neto, Richard Torkar, Robert Feldt, Lucas Gren, Carlo A Furia, and Ziwei Huang. 2019. Evolution of statistical analysis in empirical software engineering research: Current state and steps forward. Journal of Systems and Software 156 (2019), 246–267. [12] Daniel Dewey. 2014. Reinforcement Learning and the Reward Engineering Principle. In 2014 AAAI Spring Symposia, Stanford University, Palo Alto, California, USA, March 24-26, 2014. AAAI Press. Manuscript submitted to ACM