A machine learning framework for interpretable predictions in patient pathways: The case of predicting ICU admission for patients with symptoms of sepsis
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Zilker, Sandra; Weinzierl, Sven; Kraus, Mathias; Zschech, Patrick; Matzner, Martin Article — Published Version A machine learning framework for interpretable predictions in patient pathways: The case of predicting ICU admission for patients with symptoms of sepsis Health Care Management Science Provided in Cooperation with: Springer Nature Suggested Citation: Zilker, Sandra; Weinzierl, Sven; Kraus, Mathias; Zschech, Patrick; Matzner, Martin (2024) : A machine learning framework for interpretable predictions in patient pathways: The case of predicting ICU admission for patients with symptoms of sepsis, Health Care Management Science, ISSN 1572-9389, Springer US, New York, NY, Vol. 27, Iss. 2, pp. 136-167, https://doi.org/10.1007/s10729-024-09673-8 This Version is available at: https://hdl.handle.net/10419/315266 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Health Care Management Science (2024) 27:136–167 https://doi.org/10.1007/s10729-024-09673-8 A machine learning framework for interpretable predictions in patient pathways: The case of predicting ICU admission for patients with symptoms of sepsis Sandra Zilker1,2 ·Sven Weinzierl2·Mathias Kraus3·Patrick Zschech4·Martin Matzner2 Received: 8 February 2023 / Accepted: 13 April 2024 / Published online: 21 May 2024 © The Author(s) 2024 Abstract Proactive analysis of patient pathways helps healthcare providers anticipate treatment-related risks, identify outcomes, and allocate resources. Machine learning (ML) can leverage a patient’s complete health history to make informed decisions about future events. However, previous work has mostly relied on so-called black-box models, which are unintelligible to humans, making it difficult for clinicians to apply such models. Our work introduces PatWay-Net, an ML framework designed for interpretable predictions of admission to the intensive care unit (ICU) for patients with symptoms of sepsis. We propose a novel type of recurrent neural network and combine it with multi-layer perceptrons to process the patient pathways and produce predictive yet interpretable results. We demonstrate its utility through a comprehensive dashboard that visualizes patient health trajectories, predictive outcomes, and associated risks. Our evaluation includes both predictive performance – where PatWay-Net outperforms standard models such as decision trees, random forests, and gradient-boosted decision trees – and clinical utility, validated through structured interviews with clinicians. By providing improved predictive accuracy along with interpretable and actionable insights, PatWay-Net serves as a valuable tool for healthcare decision support in the critical case of patients with symptoms of sepsis. Keywords Patient pathway ·Process prediction ·Sepsis ·Interpretability ·Interpretable machine learning · Interpretation plots ·Deep learning Highlights •ThisarticleproposesPatWay-Net,anovel machine learningframeworkforpredictingcriticalpathwaysof patients with sepsis symptoms. Our framework retains patient pathway data in its natural form by combining non-linear multi-layer perceptrons (MLPs) for each static feature (i.e., static module) and an interpretable LSTM (iLSTM) cell for sequential features (i.e., sequential module). •Our results reveal that our approach outperforms commonly used interpretable machine learning models in our case, such as decision tree and logistic regression by 10.4% and 7.3% in terms of the area under receiver operating characteristic curve, respectively, and noninterpretablemodels,suchasrandomforestandXGBoost by 4.4% and 1.2%, respectively. BSandra Zilker [email protected] Extended author information available on the last page of the article •PatWay-Net provides decision support to clinicians and hospital management in predicting the pathway of a patient accurately while remaining interpretable and can, therefore, help to improve hospital resource management. •To enhance the model’s interpretability and utility for clinical decision-makers, we have developed a comprehensive dashboard that visualizes patient health trajectories, predictive outcomes, and associated risks, facilitating informed clinical and resource allocation decisions. •The clinical utility of our framework is supported by structured interviews with independent clinicians, confirming its interpretability and actionable insights for healthcare decision support. 1 Introduction As healthcare organizations face increasing demands and limited resources, the efficiency and compliance of health- 123 -
137 care processes are becoming increasingly important [1]. The pandemic has served as a stress test for these processes, revealing several weaknesses, such as gaps in resource allocation, inefficiencies in patient triage, and limitations in data-driven decision-making [2,3]. As a remedy, advanced decision support systems based on modern machine learning (ML)models canbeemployedtoimprovetheperformanceof healthcare processes and provide proactive insights for clinical decision-makers [3–5]. By using large amounts of data that are ubiquitously generated in today’s healthcare information systems, such models can learn non-trivial patterns from historical patient trajectories. A rich source of historical patient data is represented by so-called patient pathways, a timeline of each patient that describes the different departments, measurements, treatments, and transitions that a patient has gone through during aclinicalstay [6]. Thisinformationcanbeused to makeaccurate predictions about future health outcomes, informing the allocation of resources or the focus of medical professionals on specific patients [e.g., 6–9]. In this way, healthcare institutions can derive recommendations for managing and controlling patient pathways early and identify risks and issues before they emerge. Such recommendations are especially crucial in the context of sepsis symptoms, a complex and time-sensitive condition that demands rapid identification and intervention toimprovepatientoutcomes [10]. Byleveragingpatientpathway data, healthcare institutions can not only derive timely recommendations to manage and control disease progression but also identify risks and issues, such as early signs of sepsis, before they escalate [11]. Consequently, early detection and treatment of sepsis, facilitated by the analysis of patient pathways, can significantly reduce a patient’s deterioration. ML models represent a promising choice for predicting patientpathwaysas theycanrapidlyprocesslargeamounts of patient data and find latent patterns that help make informed decisions about patient outcomes. ML models come in various forms and facets. For critical applications, clinical decision-makers typically favor interpretable1ML models like decision trees, linear and logistic regression, and generalized additive models (GAMs) [e.g., 12–17]. They have the advantage of providing a clear understanding of how predictions are derived, which is crucial for making informed and accountabledecisions.Atthesametime,however,such interpretable models have the limitation that they cannot handle sequential data structures in their natural form, limiting their prediction capabilities for time-varying patient data. 1We make a strict distinction between the terms “interpretation” and “explanation”. Interpretation is derived from models designed to be intrinsically interpretable, whereas an explanation can be createdbyapplyingapost-hocanalyticalexplainable-artificial-intelligence approach to a black-box model (cf. Section 3). In contrast, there is an increasing interest in using more advanced and flexible models, such as bagged and boosted decision trees [e.g., 6,8,9,18] or deep neural networks (DNNs) [e.g., 19–21]. DNNs are of particular interest for predicting patient pathways because of their ability to automatically discover and learn complex patterns in highdimensional data [22]. This ability also allows them to capture hidden patterns in sequential data structures that are difficult to identify with traditional ML models. However, DNNs generally have the limitation that they lack model interpretability because their internal decision logic is not directly comprehensible by humans [4,23]. This renders them black boxes for model developers and decision-makers, which is why they are unsuitable for critical healthcare applications. To address the limitations of both research streams above, we propose PatWay-Net, an innovative ML framework that is designed for both high predictive accuracy and intrinsic interpretability in modeling pathways from patients with symptoms of sepsis. With this framework, we leverage the principle of interpretable ML models while harnessing the flexibility of a DNN architecture. More specifically, our contributions are as follows: •PatWay-Net is designed to constrain feature interactions, ensuring full model interpretability across the entire DNN architecture. •The architecture blends non-linear multi-layer perceptrons (MLPs) for static features with an interpretable LSTM (iLSTM) cell for sequential features, preserving the natural data structure of patient pathways. •AcomprehensivedashboardsupportsPatWay-Net’sapplicability by enabling clinical decision-makers to interpret PatWay-Net’s predictive outcomes and associated risks easily. •Structured interviews with independent medical experts rigorouslyvalidatePatWay-Net’s utilityand interpretability, attesting to its real-world healthcare applicability. We evaluate our proposed model using a real-life data set from an emergency department of a Dutch hospital, containing health records of patients with sepsis symptoms [11]. During their stay, patients go through different activities (e.g., changing departments, receiving medications) and develop different trajectories of severity, resulting in individual patient pathways. The data set contains a rich set of static and sequential features, such as socio-demographic data, blood measurements, medical treatments, and diagnoses,whichprovideavaluablebasisforpredictingthefuture behavior of individual pathways. Specifically, we use this information to predict whether a patient will be admitted to the intensive care unit (ICU), which constitutes a highly 123 A machine learning framework for interpretabl e…
138 Fig. 1 Illustration of underlying setting. Multiple tasks must be performed when a patient is transferred to a new department or receives a new treatment. Thus, early prediction of the various steps a patient goes through during their hospital stay leads to more efficient operations relevant prediction task for clinical professionals and administrative staff to support proactive resource allocation [9, 12,24]. By comparing different types of ML models, we show that PatWay-Net outperforms commonly used interpretable models, such as decision trees or logistic regression, and even non-interpretable models, such as random forest and XGBoost, in terms of area under the receiver operating characteristic curve (AUCROC) and F1-score. We then use PatWay-Net for interpreting both static and sequential features of the real-world setting to demonstrate its applicability for healthcare decision support. Our paper is organized as follows: Section 2motivates the task of predicting critical patient pathways from a clinical point of view. Section 3presents relevant background and related work. Section 4introduces our proposed ML framework for interpretable patient pathway prediction, PatWay-Net. Section 5outlines the evaluation and application results based on the real-life use case for predicting ICU admission for patients with symptoms of sepsis. Section 6 summarizes our work by drawing implications for research and practice, reflecting on limitations, and providing an outlook for future work. 2 Clinical relevance 2.1 Patient pathways and clinical decision support Healthcare processes are generally concerned with all activities related to diagnosing, treating, and preventing diseases to improve well-being [25]. This includes patient-related activities organized in patient pathways and administrative activities that support clinical tasks [26]. Patient pathways are directly linked to a patient’s diagnostic–therapeutic cycle and, therefore, do not constitute strictly standardized processes. However, accurate prediction of patient pathways is crucial for optimizing resource allocation, improving patient outcomes, and facilitating timely clinical interventions, thus making it an essential tool for enhancing healthcare efficiency and effectiveness. Figure 1illustrates a patient’s hospital stay at multiple departments. In each department, various tasks must be performed to ensure a safe and well-organized patient transition. In this example, the patient was transferred from the emergency room to the coronary care unit. Depending on the patient’s condition, the patient may be transferred to the ICU or the normal care unit (NCU). Therefore, both departments must be prepared for patients. By using a decision support system that accurately predicts the next station, resources for one of the departments can be saved. Technically, a patient visiting the hospital produces a patient pathway. A set of multiple patient pathways is then stored as an event log. Table 1presents an example representation of an event log. Here, one visit of a patient is represented by a patient pathway (ID = 1). In the beginning, the patient registered at the emergency room at 2024-02-20 12:11:01. Also, the gender of the patient is registered as male (M). In the next patient activity, the blood pressure is measured at 180. Later, medication is administered before the Table 1 Example event log with a single patient pathway following the scenario in Fig. 1 Patient pathway ID Patient activity Timestamp Blood pressure Gender 1 Emergency room registration 2024-02-20 12:11:01 - M 1 Measure blood pressure 2024-02-20 13:11:27 180 M 1 Give medication 2024-02-20 14:30:27 - M 1 Measure blood pressure 2024-02-20 15:45:55 195 M 1 ICU Admission 2024-02-20 16:12:02 - M 123 S.Zilker et.al.
139 blood pressure is measured for a second time at 195. The next activity then describes the patient being transferred to the ICU. AsshowninTable1,thepatientinformationinaneventlog is not structured to be easily processed by prediction models. Therefore, careful processing of static and sequential patient information is necessary to predict following patient activities accurately. In addition, experts generally have to ensure that the model learns meaningful patterns from the data, which constrains the model to be intrinsically interpretable. Both points are addressed in this work. In the development of decision support for healthcare applications, the involvement of medical experts is inevitable [27]. Their insights can ensure that the proposed approach aligns with the complexities and problems of clinical practice. In this work, a comprehensive dashboard serves as a translational interface, bridging the gap between high-level computational outputs and real-world clinical decisions. It provides a demonstration of a model’s potential for realworld applicability, ensuring that its capabilities are both understandable and useful to healthcare practitioners. 2.2 The case of sepsis Sepsis results from the body’s overwhelming response to an infection and can be life-threatening [28]. Therefore, sepsis is a time-sensitive issue that needs clinicians’ attention as early as possible to enable the best possible outcome for each patient[10]. Basedonthis, itiscriticaltopredictthisoutcome during an ongoing patient pathway to provide timely recommendationsforcontrollingthedisease’sprogression[11,29]. However, theimportanceofsepsisliesnotonlyintheurgency of its treatment but also in its complex and variable nature that can be detected in the resulting patient pathways [30]. While its symptoms and, thus, underlying medical indicators, can progress or change rapidly, treatment needs to be adapted dynamically which influences the patient pathway [29]. Ultimately, in the context of developing interpretable ML models for predicting patient pathways, the focus on patients with sepsis symptoms is crucial, given the imperativetoenhanceclinicaldecision-making, resource allocation, and ultimately, patient outcomes in this high-stakes domain. 3 Methodological background and related work ML models are increasingly being integrated into clinical applications to assist healthcare professionals in diagnosing diseases, predicting patient outcomes, and making treatment decisions [7–9]. While the predictive power of these models is often decisive, it is also essential that they provide comprehensible outputs due to the critical nature of healthcare decisions. Comprehensible outputs promote transparency, reduce the risk of unintended biases, and ensure the reliability of the model results, ultimately contributing to safer and more effective patient care [31–33]. From a methodological point of view, there are generally two distinct streams of research dealing with comprehending ML models. Table 2provides an overview of both streams with exemplary approaches, which can be further classified according to the type of input features they support. 3.1 Explainable machine learning The first stream of research refers to the concept of explainable ML. It promotes the use of flexible ML models with high predictive power, which subsequently require post-hoc explanation methods to convert their complex mathematical functions into easier-to-understand explanations [23,32]. CommonrepresentativesofflexibleMLmodelsforstaticfeatures are bagged and boosted decision trees such as random forest [42] and XGBoost [43]. Such models excel at handling static tabular data because they can capture complex interactions between features, allowing them to achieve high predictive performance [6,8,9,18]. In this work, we include both models as strong baseline approaches in our evaluation section. However, the construction of high-level interactions creates a lack of transparency because the individual feature effectsarenolonger understandable by humans and therefore require additional explanation methods. For sequential features, the field has increasingly focused on DNNs in recent years [22]. Their multi-layered network architecture allows them to automatically discover and learn complex patterns in high-dimensional data structures that are relevant for the prediction task [4,5]. Of particular interest are recurrent neural networks and long short-term memory (LSTM) networks because they can capture temporal patterns and therefore offer superior predictive performance compared to traditional approaches in dynamic and complex healthcare process environments [e.g., 19,20]. Furthermore, such network architectures have the advantage that they can be modified to capture static and sequential features simultaneously[e.g.,21,41]. Nevertheless, thenested,multilayeredstructureof DNNs also creates alackof transparency, because it is not directly observable what information in the input data drives the models to generate their prediction, rendering them black boxes for model users. In our work, we adopt the overall idea of an LSTM network [44] but propose a modification to ensure full model transparency. To turn the internal decision logic of black-box models into comprehensible results, the field of explainable ML has proposed a variety of post-hoc explanation methods [6,45]. Some of these methods are model-specific. That is, they are designed for specific types of models and derive explanations by examining internal model structures and parameters 123 A machine learning framework for interpretabl e…
140 Table 2 Positioning of our work with respect to related fields from a methodological perspective Explainable machine learning Interpretable machine learning Definition Refers to methods that aim to simplify (approximate) the decision logic of ML models that are not directly understandable to human users (known as black-box models). Refersto ML modelsthatare designed to beinherently understandable to human users. Main focus Encourages the use of flexible ML models with high predictive power that require post-hoc explanations to convert complex mathematical functions into a more understandable form for clinical model validation. Encourages the use of ML models that ensure a complete understanding and validation of the decision logic for fully transparent clinical decision support without the need for additional explanation methods. Static features Involves a scenario where a flexible black-box ML model is provided with a fixed-length feature vector (e.g., age, weight, vital signs of a patient), and the model’s response is analyzed after prediction using model-specific explanation methods such as layer-wise relevance propagation [e.g., 6] or model-agnostic explanation methods such as Shapley additive explanations [e.g., 18]. Interpretable ML models limit interactions between features to reduce complexity, allowing for comprehensive validation of the model’s performance. Typical interpretable ML models are linear models [e.g., 12,14,15], decision trees [e.g., 17], or generalized additive models [e.g., 13,16], as well as typical risk charts, such as the well-known simplified acute physiology score (SAPS) at the intensive care unit [34]. Sequential features A black-box sequential ML model is provided with temporalpatientdata(e.g., atrendofvitalsignsovera period), and the model’s response is analyzed post-hoc [e.g., 31, 35,36]. Typical sequential ML models with high predictive power are recurrent neural networks like long short-term memory (LSTM) networks [e.g., 19,20]. An interpretable ML model that processes temporal patient data with full transparency to medical professionals. Interpretable models allow for complete validation of model behavior. Examples include probabilistic finite automatons [e.g., 37], hidden Markov models [e.g., 38], and certain advances in neural networks [e.g., 39,40]. Static + sequential features A black-box model consists of two parts: One that can process static features and one that can process sequential features [e.g., 21,41]. The information about the patient from the two sources is then combined to compute the model output. Our research: A fully interpretable ML model that can process both static and time-varying patient data. (e.g., layer-wise relevance propagation for DNNs [6,36]). Other methods are model-agnostic and, therefore, broadly applicable to different ML models. One of the most widely used model-agnostic methods is Shapley additive explanations (SHAP) [46]. SHAP uses a game-theoretic approach to explain the output of any ML model. It has been applied, for example, to mortality prediction in ICUs [31] and to process prediction models based on general event logs [35]. An overview of existing post-hoc explanation methods is given by Loh et al. [32]. Overall, post-hoc explanation methods have the advantage of providing a high degree of flexibility while encouraging the use of models with high predictive performance.Furthermore, theycanleadtovaluableinsights, especially for exploratory analysis purposes [47]. However, post-hoc explanation methods must also be viewed with caution. They generally attempt to reconstruct the cause of a generated prediction by approximation. As a result, they can never fully explain the entire black-box model without losing information, which may lead to unreliable results. Similarly, explanations are provided only after a model’s prediction, making it impossible to fully validate the functioning of the model for all inputs before model deployment. This issue becomes particularly critical when the distribution of input data changes over time, and the model may need to handle input feature ranges that were not encountered during its training phase. Overall, such deficiencies can lead to misleading conclusions and potentially harmful results [33,48]. For this reason, we refrain from pursuing this general research stream in this paper. 3.2 Interpretable machine learning The second stream of research refers to the field of interpretable ML, which promotes the development of intrinsically interpretable models [23,49]. In this research stream, the structure of an ML model is constrained, such that the resulting model allows for a better understanding of how predictions are generated. Traditional representatives are linear models and decision trees, which are easy to comprehend and therefore often remain the preferred choice in critical healthcare applications [e.g., 12,14,15,17]. At the same time, however, they are generally too restricted to capture more complex relationships. A more advanced class of intrinsically interpretable ML modelsare GAMs[23,49].InGAMs,input featuresaremodeled independently in a non-linear way to generate univariate shape functions that can capture arbitrary patterns but remain fully interpretable. The resulting shape functions for each feature are summed up afterward to produce the final model output. Thus, GAMs include additive model constraints yet dropthelinearityconstraintofasimple logistic/linearregres- 123 S.Zilker et.al.
141 sion model. This structure is simply interpretable as it allows users to verify the importance of each feature. That is, the fitted shape functions directly reveal how each feature affects the predicted output without the need for additional explanation. In recent years, a wide variety of GAM variants have been proposed that can learn specific types of shape functions depending on the underlying learning procedure, for example, based on splines [50], decision trees [51,52], or even neural networks [53–55]. However, all of these approaches have in common that they primarily focus on processing static features and, therefore, cannot handle sequential data structures in their natural form [49]. As a consequence, their application in the healthcare domain is usually limited to preprocessed features in a static and aggregated form [e.g., 13,16]. In this work, we adopt the general idea of GAMs to capture non-linear effects of individual features and propagate this idea not only to static features but also to sequential features to obtain a powerful yet fully interpretable model. Apart from that, there are also interpretable ML models that are specifically designed to capture sequential patterns. Traditional approaches include probabilistic finite automatons [37] or hidden Markov models [38]. Such models have the drawback that they require explicit knowledge about the form of an underlying process model [56], which is challenging to discover or reconstruct from complex event data in dynamic healthcare environments [11,57]. Therefore, recent approaches increasingly pursue the idea of constraining the structure of DNN architectures to obtain models that can process sequential features in their natural form while remaining intrinsicallyinterpretable.Todate,however,littleworkexists in this area and current approaches often do not distinguish between sequential and static features [e.g., 39,40]. In summary, only a limited amount of approaches deal withthedevelopmentof intrinsicallyinterpretablemodelsfor transparent patient pathway prediction. In particular, it lacks an innovative approach that can capture non-linear relationships in the form of flexible shape functions for static as well as sequential patient features while providing comprehensible model outputs that visualize the different feature effects for transparent decision support. Likewise, to the best of our knowledge, none of the existing approaches can automatically detect and integrate (sequential) feature interactions to controlthemodel’sflexibilityforimproved predictiveperformance. As a remedy, we propose PatWay-Net, a novel ML framework that combines all these aspects within a single approach. 4 PatWay-Net This section describes PatWay-Net, an interpretable ML framework building on a DNN model with an architecture that transfers the ideas of GAMs into a novel, intrinsically interpretable LSTM module for sequential features, and intrinsically interpretable MLPs, for static features.2We apply this proposed DNN architecture of PatWay-Net to the problem of patient pathway prediction but want to emphasize that our proposed architecture is universal and can be applied to a variety of problem sets that combine sequential and static data (see also Appendix Dfor evaluations on other use cases). In the following, we first describe the underlying problem of patient pathway prediction (Section 4.1), before mathematically describing the architecture (Section 4.2) and the training process (Section 4.3) of PatWay-Net’s DNN model. Subsequently, we describe the different interpretation plots that can be derived from the intrinsically interpretable architectural design of PatWay-Net’s DNN model (Section 4.4). 4.1 Problem statement An ML model f∈Fshould map patient pathways to a target of interest, with Fdenoting the so-called hypothesis space. Patient pathways comprise two sets of information, one set describes static information about the patient, and one set describes dynamic or sequential information about the patient. Definition 1 (Patient Pathways) Mathematically, the information that describes a set of patients can be expressed as a tuple Xstatic,Xseq,(1) where Xstatic ∈Rs×qis the static patient data, and Xseq ∈ Rs×T×pis the sequential patient data. The dimension s denotes the number of patient pathways,qindicates the number of static variables that describe a patient (e.g., one-time diagnoses or gender), and Tand pdescribe the number of time steps that we recorded for the sequential information and the number of features tracked in each time step, respectively. A single patient’s patient pathway iis denoted by the static information X(i) static and the sequential data X(i) seq. The objective of this work is to find a prediction model f∈Fthat maps the static information Xstatic and sequential information Xseq about patients to target outcomes y= (y1,...,ys), that is f:Xstatic,Xseq→y.(2) The target outcomes ycan thereby represent various patient activities in the future, such as ICU admission. 2For reproducibility, all developed and used material can be found here: https://github.com/fau-is/patway-net 123 A machine learning framework for interpretabl e…
142 A timely prediction of the future occurrence of an activity is crucial, as it can prevent the worsening of the patient’s condition and initiate successful treatment by medical experts. Therefore, a prediction model should not only make predictions once the full patient pathway is present but should make predictions already at earlier stages, that is, with less information included in the patient pathways. Thus, we define the patient pathway prefix in the following. Definition 2 (Patient Pathway Prefix) Given patient pathway iwith static information X(i) static ∈Rqand sequential information X(i) seq ∈RT×p, the patient pathway prefix of length t∗ is defined as a tuple X(i) static,X(i) seq[: t∗],(3) where X(i) seq[: t∗]∈Rt∗×pdenotes the first t∗time steps of the patient’s sequential information. 4.2 Architecture of the DNN model The proposed interpretable architecture of PatWay-Net is shown in Fig. 2. It contains a static, a sequential, and a connection module. While the first two modules naturally model the event log data, the connection module maps the outputs of these modules onto predictions of patient activities (in our case ICU admission). 4.2.1 Static module The static module resembles a GAM [50], yet combines the underlying idea with the power of DNNs [54]. By making this architectural choice, we allow our proposed model to remain fully transparent. That is, the effect of each input feature on the model output can be fully assessed after training the model. This is achieved by mapping the input features separately to output values (i.e., there are no interactions between input features). This separation naturally constrains this DNN but, on the other hand, allows the visual inspection of the effect each static feature has on the network’s output. Consequently, although our proposed model is derived from the field of DNNs, we make careful choices about our architecture to allow for a fully transparent white-box model (in contrast to the black-box behavior of general DNNs). Mathematically, for qstatic input features Xstatic[1], ...,Xstatic[q], the static module maps the input features to outputs o1,...,oqthrough ol=fl MLP(Xstatic[l]), with l∈{1,...,q},(4) where each fl MLP denotes a neural network and ol∈Rindicates a single scalar. The neural networks of the architecture are trained in individual sub-modules so that the weights of the different neural networks are trained independently from each other (cf. the boxes around the neural networks of the static module in Fig. 2). With this architecture, we can later compute the outputs olfor various input values for each neural network fl MLP and, thereby, visually inspect the effect that the input has on the output. 4.2.2 Sequential module The sequential module extends the previous idea of our static module to a sequential setting. For this, we propose a novel interpretable LSTM (iLSTM) layer to encode the values of each sequential feature Xseq[j]into a vector hj t∈Rm, with j∈{1,2,...,p}, where mdenotes the hidden size for a single sequential feature in the iLSTM cell. To ensure intrinsic interpretability of the iLSTM, each sequential feature has its corridor throughout the gates and state vectors of the origi- Fig. 2 Illustration of the architecture of PatWay-Net’s DNN model consisting of a sequential, a static, and a connection module. Here, two static and two sequential features are shown, which run through their modules and are then connected 123 S.Zilker et.al.
143 nal LSTM [44], without the possibility to interact with any other feature (similar to the previous static module, in which each static feature went through a separate neural network). Such feature corridors in the iLSTM layer have a specific size, defined by the internal element size of the corresponding sequential feature m, defining how much vector space is reserved for the sequential feature value computation, from the gates to the hidden state. Similar to a vanilla LSTM [44], the iLSTM uses a forget gate, an input gate, and an output gate, as well as a candidate state,resultinginthevectorsft,it,ot,and ˜ ct,respectively.The information of the sequence is then stored in a cell state ct, andahiddenstateht.Technically,thisrestriction,tonotallow uncontrolled interactions, is realized by multiplying weight matrices with masking matrices. A masking matrix includes only 0 or 1 values. If an element of a weight matrix should be considered, the corresponding element in the masking matrix is set to 1, else it has the value 0. Mathematically, the iLSTM can be formalized as ft=σxt⊗(Uf∗Um)+ht⊗(Vf∗Vm)+bf,(5) it=σ(xt⊗(Ui∗Um)+ht⊗(Vi∗Vm)+bi),(6) ot=σ(xt⊗(Uo∗Um)+ht⊗(Vo∗Vm)+bo),(7) ˜ ct=tanh (xt⊗(U˜c∗Um)+ht⊗(V˜c∗Vm)+b˜c),(8) ct+1=ft∗ct+it∗˜ ct,(9) ht+1=ot∗tanh (ct+1).(10) Here, σdenotes the sigmoid activation, ⊗is the matrix multiplication, and ∗denotes the element-wise multiplication. Umand Vmare masking matrices that ensure that the individual features are computed independently using values from their corridor and, therefore, omitting interactions between sequential features. By contrast, a traditional, noninterpretable LSTM [44] does not use such masking matrices and, therefore, allows any interaction between features for which values are to be computed. As output, the iLSTM layer returns for each sequential feature Xseq[j]the vector hj∈Rm, that is, the last hidden state of the iLSTM for the sequential feature j.Let fj iLSTM ∈R→Rmdenote this function, which maps the j-th sequential feature onto the corresponding hidden state, and let fiLSTM ∈Rp→Rp∗m denote the function that maps all sequential features to the complete hidden state vector. Beyond single sequential features, the iLSTM layer can encode values of a pairwise sequential feature interaction (j,k)in (1,...,p)×(1,...,p)into a vector hj,k∈Rm. Mathematically, the iLSTM computes such interactions as an additional sequential feature that does not interact with other features. The interactions to be used in PatWay-Net can be chosen manually or can be detected automatically using heuristics. We describe such a heuristic in Appendix A. 4.2.3 Connection module In the connection module, the information from the static module and the sequential module are then combined to compute the estimations ˆ yfor the target outcomes y. Mathematically, we use the hidden outcome values o1,...,oqfor static features, the hidden state values h1,...,hpfor sequential features, and potentially hj,kfor interacting sequential features j,k∈(1,...,p)×(1,...,p). These values are then concatenated and mapped onto the output neuron to provide the estimations ˆ y. The mapping is performed using a single feed-forward layer with sigmoid activation, as the prediction of patient activities (in our case ICU admission) is defined as a binary classification task. 4.3 Parameter optimization of the DNN model All parameters from the three modules are combined into one DNN model in which these are optimized simultaneously. Let fPatWay-Net denote this DNN model with parameters β. Depending on the task, the fit of fPatWay-Net to the target outcomes yis then measured by a loss function L. In our realworld data application, we use binary cross-entropy, as ICU admissionrepresentsabinarydecision.Overall,weminimize the empirical risk, that is β∗=argmin β s i=1 T t=1 LfPatWay-NetX(i) static,X(i) seq[: t]; β,yi,(11) where we iterate over the patient pathways sand over the prefixes for each patient pathway T. We address this optimization problem using an adaptive moment estimation (Adam) optimizer [58] with default hyperparameters. For every epoch, we perform a mini-batch gradient descent to optimize the internal parameters batchwise efficiently. 4.4 Interpretations of the DNN model Based on the architectural design of PatWay-Net’s DNN model, different interpretation plots can be created, allowing an interpretation of how the model input affects the model output. The interpretation plots are part of a comprehensive dashboard, that serves as a decision support tool for clinical decision-makers (cf. Section 5.4). Table 3provides an overview of the four interpretation plots that we propose in this paper, including plot names, the underlying equations, and short descriptions of the plots’ purposes. In the real-life data application that follows, we prefer the designation (medical) indicator over feature because it is more comprehensible for decision-makers in the medical domain. Accordingly, we name our four interpretation plots 123 A machine learning framework for interpretabl e…
150 prediction for ICU admission increases considerably. Thus, we can see a substantial alteration in the leukocyte count. Moreover, an elevated leukocyte count can be a typical indicator of an ongoing systemic inflammatory response to an infection,likesepsis [e.g.,65].However,a decreaseinLeukocytes canalso occur inseverecaseswherethe immune system is overwhelmed, indicating a worsening of the patient’s condition [e.g., 66]. In such an acute case, there exists a potential necessity for the patient to receive intensive care. Themedical indicator transition plot(lower right inFig. 6) shows how the prediction changes from the previous to the current Leukocytes value measurement. The figure illustrates thata decrease inthe Leukocytes value (from a previous value of 0.0 to a current value of 1.0) corresponds to an increased probability of the patient requiring ICU admission. This is consistentwith clinicalunderstanding,as adecreasein leukocytes often denotes a heightened vulnerability to developing an infection like sepsis, suggesting a more severe disease course that may require intensive care. Conversely, if there was a low Leukocytes value at the previous time step that subsequently increases by the current time step to a normal value, prediction indicates a lower likelihood of the patient being transferred to the ICU. This could suggest that the patient’s immune response is stabilizing, or the infection is being effectively controlled, thus reducing the necessity for intensive care. The medical indicator development plot (upper plot in Fig. 6) shows what effect the sequential medical indicator Leukocytes has on the model prediction over time. Up to time step three (2014-09-18 13:46 - 2014-09-18 13:56), the effect of Leukocytes is high since no measurement has been taken yet. From time step three to four (2014-09-18 13:56 - 2014-09-18 14:11), the effect on the prediction decreases, as a medium-high Leukocytes value of 0.51 has been measured in this time period. 6 Discussion and future work 6.1 Implications for healthcare management and practice Our research has multiple implications for healthcare management and practice. First, PatWay-Net supports a straightforward analysis of patient pathways using patients’ historicaleventdata.Inthisway,subjectivityisavoided,andmanual effort can be reduced to a minimum in decision-making. Likewise, our model provides high predictive performance in the context of patients with symptoms of sepsis without relying on explicit process knowledge. This allows flexibility for decision support applications in highly complex and dynamic healthcare environments. In our case, experiments have shown that the predictive performance is superior to traditional approaches by combining patients’ static features with sequential features in a DNN architecture that remains fully interpretable. This is a great advantage because prediction tasks in the healthcare sector are usually dominated by linear and logistic regression models with underlying static featurestoensureahighdegreeoftransparency[e.g., 12–15]. At the same time, PatWay-Net can improve decisionmaking in both patient-specific and administrative decision contexts. For example, in a patient-specific decision context, a model interpretation for admission to ICU prediction may indicate an increase in a patient’s probability of being transferred to the ICU after being treated with a certain medication.Basedonthis insight, medical expertshavethe chance to intervene and apply corrective treatments to prevent worse consequences. In an administrative context, model interpretation could reveal shortcomings in the hospital’s IT system. For instance, conflicting predictions between PatWay-Net and clinicians can be traced down to potentially missing patient information within an ERP system, allowing for optimization of hospital operations. Finally, PatWay-Net provides timely decision support. From a technical point of view, PatWay-Net’s inference time is similar to one of the shallow interpretable models as the underlying model of PatWay-Net represents a function mapping the data input to the prediction output. Compared to the inference time, the training time of PatWay-Net is considerably higher than the training time of the shallow interpretable models as PatWay-Net is a DNN with a recurrent iLSTM cell. Further, the training time increases with each sequential feature as each sequential feature is passed through a single corridor in the iLSTM cell. However, for our purpose, the inference time is far more important than the training time as the models are created and trained before they are applied in an online mode where the models are used for providing effective decision support. On the other hand, PatWay-Net’s interpretations can be immediately retrieved from the model itself. In doing so, it is considerably faster than applying a post-hoc explanation method such as SHAP for reconstructing explanations for non-interpretable DNN models. 6.2 Implications for research PatWay-Net combines two crucial streams of research. The first stream follows the idea that more complex models, such as DNNs, can naturally model specific structures of the underlying data and, thereby, increase predictive performance. Thus, PatWay-Net employs an iLSTM in its sequential module to model temporal structures of sequential data, and several MLPs in its static module to model nonlinear structures of static data. The second stream follows the idea that explanations of complex models can never provide the same understanding as that of intrinsically interpretable models [33,49]. Consequently, approximated explanations 123 S.Zilker et.al.
151 of complex models should be avoided or used carefully. As a remedy, PatWay-Net remains fully interpretable and prevents uncontrolled interactions of static and sequential features by incorporating the main principle of GAMs into its entire DNN architecture. As such, the model also provides an extension to traditional GAMs, which are unable to capture sequential data structures in their natural form [e.g., 51–55]. Within the realm of medical research, this work is aligned with emerging trends advocating for a shift from static, tabular data to multimodal data representation [67]. Traditional approaches often simplify complex health data such as images and vital signs into aggregated statistics or explicit features, thereby losing important information and only capturing a snapshot of the patient’s health. Our framework addresses this gap by accurately modeling health trajectories through both, sequential and static data. The architecture is not limited to merely processing patient pathway data but it can also be adapted to other temporal sequences commonly encountered in healthcare, such as data from wearable and ambient biosensors [68]. By facilitating a more rigorous representationofhumandata,weimprovenotonlythepredictive performance but also the clinical utility of ML models in healthcare settings. 6.3 Limitations and outlook Aswith anyresearch,ourworkis notfreeoflimitations.First, we focused in this paper on a use case of patients with symptoms of sepsis to demonstrate the benefits of PatWay-Net in caseswheretrustintheMLsystemiscrucialtoallowforpractical applications. However, the application of PatWay-Net is not limited to this use case but can also be used to predict process-related outcomes in other tasks or domains involving static and sequential features, as shown by the results of the additional use cases in Appendix D. Here, we find mixed results, highlighting that full generalizability in other contexts requires further work. Second, the event log sample from the real-life data application was relatively small, and using this sample for training PatWay-Net showed a performance decrease from validation to test scores. This difference could be an indicator of model selection criterion overfitting [69], which might affect PatWay-Net’s generalizability to unseen data. However, to mitigate the effect of this overfitting type, we followed the suggestion from Cawley and Talbot [69] and adopted solutions for the problem of overfitting to the training criterion. In particular, we tested model regularization, hyperparameter minimization, and early stopping [69–71]. Among these solutions, performing early stopping achieved the best predictive performance for our use case of patients withsymptomsof sepsis. In addition, results thatweobtained fromafurther use case on loanapplications(seeAppendixD) confirm that this type of overfitting is likely to be less present when the event log size is larger. However, despite these overfitting concerns, PatWay-Net achieved relatively high predictive performance, and being a neural network, it can be expected that its predictive performance will further improve when trained with more data [22]. Third, PatWay-Net’s sequential medical indicator plots provide interpretations that are tied to a patient’s individual pathway. This limitation is necessary because the predictive effect of a sequential medical indicator within our iLSTM cell is determined by the patient-specific trajectory over previous time steps. As a result, varying historical trajectories can lead to different outcomes, which may also affect the results of the interpretation plots. However, at this point, it is not practical for clinical decision support to include global interpretation plots for all conceivable trajectory variants across all patients in a single dashboard. Therefore, we decided to focus on developing a patient-specific dashboard with all relevant information to support clinicians in an easily accessible way. Nonetheless, future research should address this limitation to identify new ways of how feature effects of multiple sequences over several time steps can be visualized in a comprehensive, yet fully understandable manner. This may require new visualization techniques (e.g., interactive filter mechanisms) or additional abstraction layers (e.g., clustering of patient trajectories leading to similar outcomes and interpretation plots), which offer promising directions for future work. Fourth, PatWay-Net’s mechanism to automatically detect andintegrateinteractionscoverspairwiseinteractions among sequential features. For the use case addressed in this paper, we can show that the predictive performance of PatWay-Net with this mechanism is close to the predictive performance of an unrestricted LSTM cell (see Appendix A). Nevertheless, we assume that other types of interactions (e.g., more complex interactions between sequential features or interactions between sequential and static features) are more present in other use cases. The results obtained for a further use case on hospital billings (see Appendix D) give the first indication for this assumption and therefore provide an entry point for future research. Fifth, the proposed version of PatWay-Net does not currently consider a mechanism for selecting relevant features. This may become relevant when dealing with a large collection of features in other real-world applications, where the full set of features may lead to impractically large computational costs and a higher risk of overfitting. Future research could follow up on this point to investigate which feature selection methods are appropriate for combining static and sequential features. Nevertheless, PatWay-Net already providessomeguidanceforselectingthemostimportantfeatures 123 A machine learning framework for interpretabl e…
152 (or medical indicators) through its medical indicator importanceplot,thusfacilitatingcliniciansorhospitalmanagement when dealing with a large collection of medical indicators. Sixth, as with any ML model, PatWay-Net’s results are only as good as the data it consumes. That is, PatWay-Net is not only a reflection of possibly biased decisions made in the past but also of any data quality issues embedded in the data set. For example, in our use case of patients with symptoms of sepsis, some interpretation plots showed counter-intuitive relationships between medical indicators and the prediction target that may not be reflected in the medical literature. These findings underscore the need for rigorous data management in hospital operations to enable analytics tools like PatWay-Net to enhance decision-making substantially. Similarly, we want to emphasize that the learned feature effects should not be interpreted causally, as they are still based on correlations. Thus, it is not possible to say with certainty why some of the effects shown in the interpretation plots are present. This could be due to correlations with other(unmeasured)features,orotherunderlyingphenomena. However, despite these limitations, PatWay-Net still offers a fully transparent model that can be used to allow clinicians to compare the model results with their domain knowledge to iteratively debug and improve the model, identify underlying data quality issues, or initiate further investigations for the identification of causal relationships. Finally, our current approach pertains to the creation of a clinical dashboard that relies solely on the provided interpretation plots generated from the shape functions of PatWay-Net. While these plots offer exact insights into the model’s decision logic, they may still lack the level of context and intuitiveness required for effective clinical application. In the next steps, we intend to address this limitation by harnessingthe capabilities of largelanguage models [72,73]. By incorporating a large language model in an adaptive dialogue system,weaimtoprovidemoreintuitiveandcontextuallyrelevant explanations for clinical professionals when presenting theinterpretationplots.Thisenhancementwillnotonlymake the model’s outputs more accessible but also foster improved communication between the model and the healthcare practitioners, thereby enhancing the model’s utility in real-world clinical settings. Appendix A Further details and experiments on the main use case In this appendix, we provide further details on the use case we address to evaluate the clinical utility of PatWay-Net, our proposed ML framework for interpretable predictions in patient pathways. In what follows, we provide details on the preprocessing of the data set (Appendix A.1), before we present a heuristic for automatic interaction detection (Appendix A.2). After that, we provide details on model tuning, model evaluation, and model selection (Appendix A.3), and present further results on statistical tests (Appendix A.4), predictive performance (Appendix A.5,A.6, and A.7), interpretation quality (Appendix A.8), and runtime performance (Appendix A.9). A.1 Preprocessing of the data set We remove outliers in our data set by only considering completed patient pathways that are longer than two but shorter or equal to 50 patient events. We also remove patient pathways that do not start with activity ER registration because we assume this activity to be the central entry point into the patient pathway. As a result, the event log contains 724 patient pathways with 675 different variants over a period of 1.5 years. Our real-life data set comprises binary, categorical, or continuous medical indicators. The values of a binary medical indicator are mapped to 0 or 1, categorical values are onehot-encoded, and continuous values are scaled into the range [0,1]. Patient activities, such as Measure Leukocyte count, are either encoded by standard onehot encoding or by a custom encoding. In standard one-hot encoding, the patient activityais encoded as a vector containing only zeros, except for a single position that corresponds to a, which is set to 1. In our custom encoding, we explicitly model the relationship betweenactivitiesandtheir existingcontinuousmedicalindicators in the data. In detail, if an activity can be described by a continuous value, we set the corresponding position in the vector to this continuous value. Further, to provide more meaningful interpretations for sequential medical indicators, we also keep this continuous value for subsequent patient activities, as long the value does not change. We extract all prefixes from Xstatic ×Xseq; that is, we extract all subsequences of the sequential data, denote those as Xsub seq , and retain the static data as is, to predict at each time step. This step increases the number of training samples, which enables us to evaluate the ML models on how early they can already make accurate patient pathway predictions in the future. For each patient pathway prefix, the target labels yare created as follows: Given a prefix (X(i) static,X(i) seq[: t∗])and the patient activity of interest (e.g., Admission to ICU), we check the activities of patient pathway i. Then, if the activity of interest appears in the patient pathway’s activities, we set the target label to 1 and else to 0. We discard all prefixes and corresponding labels where the activity of interest is part of the sequential data. This is important to avoid data leakage problems in patient pathway predictions. 123 S.Zilker et.al.
153 A.2 Automatic search for interactions Given all subsequences of the sequential data Xsub seq , PatWay- Net’s interaction detection iterates 100 times to identify the most relevant pairwise feature interactions in these data (see Algorithm 1). Per iteration, a feature pair (j,k)is randomly determined from sequential features Dseq, and the sequential data for the features jand kare retrieved and reshaped. Then, the sequential data and label data are split into an 80% training set and 20% test set, an XGBoost [43] model fXGB with standard parameters is trained based on this data, and the trained model is applied to the test set to calculate an AUCROC value. Subsequently, the current interaction is added together with the respective AUCROC value to r, from which the best interactions are selected. In addition, the current interaction is added to Kso that the same interaction cannot be used again in future iterations of this procedure. After performing all iterations, the interactions with the highest AUCROC values are first selected based on rand k∗(number of best interactions) and then transferred to the iLSTM layer, in which they are considered as additional sequential features. Algorithm 1 PatWay-Net’s interaction detection. Given:Xsub seq ,y,k∗,Dseq,fXG Boost . 1for i←1to 100 do 2(j,k)←feature pair(Dseq). 3if j= k∧(j,k)/∈Kthen 4Xsub seq ←(Xsub seq [j],Xsub seq [k]). 5Xsub seq ←reshape(Xsub seq ). 6Xtrain seq ,ytrain ←split(Xsub seq ,y). 7Xtest seq ,ytest ←split(Xsub seq ,y). 8f∗ XGBoost ←train(fXG Boost ,Xtrain seq ,ytrain). 9per fi←test(f∗ XGBoost ,Xtest seq ,ytest). 10 r←r+(per fi,(j,k)). 11 K←K∪{(j,k)}. 12 end 13 end 14 return top-k-inter(r,k∗). A.3 Model tuning, model evaluation, and model selection Table 6reports the tuning parameters used in our grid search. We performed initial experiments for all models to find Table 6 Hyperparameters used in grid search ML approach Hyperparameter Hyperparameter range PatWay-Net Hidden size per sequential feature 4, 8 (with interaction) Hidden size per static feature 4, 8 Learning rate 0.001, 0.01 Batch size 32, 128 PatWay-Net Hidden size per sequential feature 4,8 (without interaction) Hidden size per static feature 4,8 Learning rate 0.001,0.01 Batch size 32, 128 LSTM network Hidden size sequential features 4, 32, 128 (with static module) Hidden size per static feature 4,8 Learning rate 0.001,0.01 Batch size 32, 128 Decision tree Max. depth 2, 3, 4 Logistic regression Regularization strength 10−3,10−2,...,100...,10+3 K-nearest neighbor Number of neighbors 3, 5, 10 Naïve Bayes Variance smoothing 10−9,10−8,...,10−4,...,1 XGBoost Max. depth 2,6,12 Learning rate 0.3, 0.1, 0.03 Random forest Max. depth 2, 6, 12 Number of estimators 100, 200, 400 Max. leaf nodes 2, 6, 12 Note: Best values over the five folds are marked in bold 123 A machine learning framework for interpretabl e…
154 appropriate value ranges for the hyperparameters. For the PatWay-Net models, the hyperparameters hidden size per sequential feature and hidden size per static feature define the used vector space for each sequential and static feature, respectively. For example, if the hidden size per static feature is set to 4, each model’s MLP has a hidden layer with 4 neurons. In contrast to the PatWay-Net models, the baseline LSTM model uses a traditional unrestricted LSTM cell instead of the proposed, restricted iLSTM cell, and therefore, the hidden size per sequential feature defines the used vector space for all sequential features. Further, we set the maximum number of neurons for the LSTM model with the unrestricted LSTM cell to 8 and for the PatWay-Net models with the restricted iLSTM cell to 128 as experiments in the use case showed that more vector space is required for the unrestricted LSTM to identify arbitrary dependencies between all sequential features and less vector space is sufficient for the restricted LSTM to compute each sequential feature effectively and efficiently. For the shallow ML models, we set the value ranges of the hyperparameters such that the resulting models are still interpretable. For instance, we bounded the depth of the decision tree or the number of neighbors in the K-nearest neighbor model. The procedure for model evaluation is summarized in Algorithm 2. Given the static data Xstatic, sequential data Xseq, and target outcome y, we start our evaluation by performing a five-fold stratified cross-validation with random shuffling on patient pathway level. That is, training and validation of each run are performed based on an entire pathway, without randomly shuffling sequentially ordered events, to avoid any temporal data leakage that could result from erroneously using future events as part of the evaluation of past events. After retrieving the training set and test set in each fold, we split the training set into a sub-training set (train*) and a validation set. Subsequently, we select the best model through a grid search using the sub-training and validation set and apply the best model to the test set to compute the test performance. We perform this procedure five times (k=5). In total, we repeat the entire model evaluation procedure five times, each with a different seed, and calculate the average performance and standard deviation over the performance values of these executions. To measure the performance of the ML models on the validation and the test set, we calculate the AUCROC [74](primary measure) and the weighted F1-score (secondary measure). AUCROC determines how well a classifier can distinguish between classes [75], and remains unbiased when dealing with highly imbalanced class distributions [76]. F1-score is the harmonic mean of precision and recall [74]. Moreover, we select the model with the highest test AUCROC from the model evaluation procedure to retrieve the interpretation. For model selection, we perform a grid search, as formalized in Algorithm 3. For each hyperparameter constellation Algorithm 2 Model evaluation. Given:Xstatic,Xseq,y,k,f 1(tr1,te1),...,(trk,tek)←split-k-stratified(Xstatic,y). 2for i←1to kdo 3Xtrain static,Xtrain seq ,ytrain ←retrieve(Xstatic,Xseq,y,tri). 4Xtest static,Xtest seq ,ytest ←retrieve(Xstatic,Xseq,y,tei). 5Xtrain∗ static ,Xtrain∗ seq ,ytrain∗←split(Xtrain static,Xtrain seq ,ytrain). 6Xval static,Xval seq,yval ←split(Xtrain static,Xtrain seq ,ytrain). 7fbest ← grid-search(f,Xtrain∗ static ,Xtrain∗ seq ,ytrain∗,Xval static,Xval seq,yval). 8per fi←test(fbest ,Xtest static,Xtest seq ,ytest). 9end 10 return 1 kk i=1per fi. (p1,...,pl)∈P, the model fis trained and validated. During training, we perform early stopping based on the validation set after 10 epochs to avoid overfitting. After training, we apply the model to the validation set and calculate the AUCROC.TheAUCROC is used as our selection criterion for the grid search, and the model with the highest AUCROC value on the validation set is selected for model testing. Algorithm 3 Model selection (grid-search). Given:Xtrain static,Xtrain seq ,ytrain,Xval static,Xval seq,yval,f,P. 1for (p1,...,pl)∈P1×,...,×Pldo 2f(p1,...,pl)←train(f,(p1,...,pl), Xtrain static,Xtrain seq ,ytrain). 3AUCROC,(p1,...,pl)← validate(f(p1,...,pl),Xval static,Xval seq,yval). 4if AUCROC,(p1,...,pl)>AUCROC,best then 5fbest ←f(p1,...,pl). 6AUCROC,best ←AUCROC,( p1,...,pl). 7end 8end 9return fbest A.4 Statistical tests Weperformstatistical testsforthe interpretable ML approaches used in our real-life data application. In particular, for our main metric (test AUCROC), we conduct a Friedman test showing significant differences among the results (statistic = 56.45143, p= 6.56037e-11). As a post-hoc test, we conduct a Wilcoxon signed-rank test with Holm p-value adjustment [60]. Table 7shows the pairwise p-values. PatWay-Net with one interaction outperforms the decision tree (p=.001503), logistic regression (p=.000593), and K-nearest neighbor (p=.000001) models significantly with α=1% and the setting with one interaction significantly outperforms the naïve Bayes (p=.044226) models with α=10%. PatWay-Net’s setting without interaction also outperforms the decision tree (p=.000637), logistic regression (p= 123 S.Zilker et.al.
155 Table 7 Overview of pairwise p-values for test AUCROC for Wilcoxon signed-rank test p-values for Test AUCROC PatWay-Net PatWay-Net Decision K-nearest Naïve Logistic (with int.) (without int.) tree neighbor Bayes regression PatWay-Net (with int.) .379944 .001503 .000001 .044226 .000593 PatWay-Net (without int.) .379944 .000637 .000001 .118250 .002009 Decision tree .001503 .000637 .001461 .340556 .915693 K-nearest neighbor .000001 .000001 .001461 .000160 .000383 Naïve Bayes .044226 .118250 .340556 .000160 .915693 Logistic regression .000593 .002009 .915693 .000383 .915693 .002009), and K-nearest neighbor (p=.000001) models significantly with α=1%. A.5 Predictive performance for controlled vs. uncontrolled interactions Table 8describes the results for PatWay-Net, in which the number of interactions varies. To provide a fair comparison, the experiments for PatWay-Net with two and three interactions are performed in the same way as described in Appendix A.3. Overall, we observe robust results with only marginal differences. We find that in our real-life use case, a single sequential interaction leads to the highest AUCROC and F1-score on the test set. Further, Fig. 7illustrates the most relevant pairwise sequential medical indicator interaction in terms of the predictive performance of PatWay-Net (with one interaction). The figure demonstrates the interaction between the indicators CRP (x-axis) and LacticAcid (y-axis). If both indicator values are low, the interaction effect is low (black-colored region in the lower left corner). However, the interaction effect becomes stronger with increasing values of both indicators. So, if the values for both indicators are close to 1, the interaction effect is high (yellow-colored region in the upper right corner). Fig. 7 Feature interaction between the indicators CRP and LacticAcid A.6 Predictive performance for different training sample sizes Figure 8shows the test AUCROC scores of different training sample sizes for PatWay-Net with one interaction, the best-performing variant. More specifically, we create training samples with 10%, 20%, 40%, 60%, 80%, and 100% of the instances of the complete training set. To avoid data leakage issues, we create the training set samples based on Table 8 Comparison of PatWay-Net with different interactions and the LSTM network (with static module) ML approach F1-score (weighted) AUCROC Validation Test Validation Test Controlled Interactions in Our Approach PatWay-Net (one interaction) 0.886 (±.016) 0.896 (±.016) 0.820 (±.028) 0.734 (±.058) PatWay-Net (two interactions) 0.886 (±.018) 0.894 (±.013) 0.821 (±.035) 0.720 (±.045) PatWay-Net (three interactions) 0.884 (±.017) 0.892 (±.015) 0.817 (±.035) 0.718 (±.042) Uncontrolled Interactions in Non- Interpretable Machine Learning LSTM network (with static module) 0.890 (±.018) 0.898 (±.014) 0.840 (±.028) 0.757 (±.049) 123 A machine learning framework for interpretabl e…
156 Fig. 8 Predictive performance for different training sample sizes complete patient pathways and split each of them into a train set and validation set before the prefixes are created from the complete patient pathways. For each sample, we perform a five-fold cross-validation, as described in Appendix A.3,but with default hyperparameters. The figure shows that the test AUCROC increases steadily with more training instances. Given that, we conclude that the event log size has an impact on the predictive performance and more training instances lead to a higher predictive performance. Further, as PatWay-Net outperforms all interpretable shallow ML baseline models in our comparison, we conclude that the use of the complete training set is appropriate for PatWay-Net to create predictions that outperform the predictions of interpretable ML baselines regarding AUCROC. Finally, as PatWay-Net belongs to the family of DNNs [22], we assume that it achieves an even higher predictiveperformance and maybe a greater differenceinpredictive performance compared to the interpretable ML baselines when more data are used for model training. A.7 Predictive performance over time Figure9showsthetest AUCROC scoresofthefirst12timesteps of patient pathways from the use case for PatWay-Net (without and with interaction) and the baselines. At each time step, prefixesof patient pathways of thecorresponding size are considered. For calculating the AUCROC scores per time step, we tune, evaluate,andselectthemodelsas described in Section A.3. The figure shows that PatWay-Net (with and without interaction) is already able to create predictions in the first steps of patient pathways, which are considerably more accurate in terms of AUCROC than the predictions of the interpretable ML baselines. Only the not-interpretable ML approach random forest outperforms PatWay-Net (with and without interaction)in nearly allof the consideredtime steps. By considering only longer sequences (longer than 8), a significant number of patient paths are filtered out from this experiment, and thus the performance varies for all ML models. Overall, this experiment empirically shows that the size of the given event log is appropriate for PatWay-Net to create timely predictions. Fig. 9 Predictive performance over time 123 S.Zilker et.al.
157 Fig. 10 Shape plot (left) and SHAP plot (right) for medical indicator Age A.8 Intrinsic interpretability vs. post-hoc explainability Figure 10 shows two interpretation plots for the same static medical indicator Age of our use case. While the medical indicator shape function retrieved from a PatWay-Net model with one interaction is illustrated on the left side, post-hoc generated SHAP values for a black-box XGBoost model are shown on the right side. While the trend of both plots is similar, the SHAP plot shows a high variance for the effect on the model prediction for a single value of the static medical indicator Age. In contrast, the shape plot for the same indicator shows a continuous function that is presumably easier to comprehend for non-technical users. A.9 Runtime performance Table 9compares the training and inference time of the static model of PatWay-Net’s DNN model, which is end-to-end trained, with a two-step version of the static module, which is not end-to-end trained. For the latter, we first trained an MLP model per static feature and then used the output of all MLP models as input for training and applying a subsequent logistic regression model. The experiments for this comparison were conducted on a workstation with 12 CPU cores and 128 GB RAM. Table 9 Runtime performance Training time (sec) Inference time (sec) PatWay-Net (static module) 51.569 (±30.790) 0.005 (±.000) MLPs + logistic regression 1646.935 (±125.027) 0.024 (±.001) The results show that the training time of PatWay-Net’s static module with end-to-end training is on average 51.569 seconds, whereas the training time for the variant without end-to-end training takes on average 1.646,935 seconds. In other words, PatWay-Net’s static module is about 33 times faster than training the individual models. Thus, the training of all MLP models in one single architecture is considerably more efficient. Concerning inference, the average time for both variants is very low. Appendix B Simulation study In this appendix, we verify and demonstrate PatWay-Net’s validity concerning the generated interpretation plots. More specifically, we follow the idea of other interpretable model proposals [e.g., 53,54], in which the authors perform simulation studies based on synthetic data with controllable feature effects.Indoingso,wepreferlikeinourreal-lifedataapplica- tion the designation (medical) indicator over feature because it is more comprehensible for decision-makers in the medical domain. In the following, we provide details on the simulation data creation (Appendix B.1), the used experimental setting (Appendix B.2), the obtained results ( Appendix B.3), and additional results (Appendix B.4). B.1 Simulation data creation For our simulation study, we create a synthetic event log containing static and sequential medical indicators.7The event log consists of 50,000 patient pathways with 12 events each. Every patient pathway begins with a start activity, called ER 7For the purpose of reproducibility, additional material can be found in the repository. 123 A machine learning framework for interpretabl e…
158 Registration, followed by three measurements of the Heart Rate and Blood Pressure each, and then the administration of medication (four times Aand one time Bin random order). Further, the data set contains four static and two sequential medical indicators. Two of the static medical indicators are numerical, namely Age and BMI (Body Mass Index), whereas the remaining two are binary, namely Gender and Foreigner. They are all set randomly. The sequential medical indicator Heart Rate occurs whenever the respective activity Heart Rate appears. Thus, it has different values that either solely increase or decrease over time. In our simulation study, the increase or decrease is exemplarily always set to 30%. The sequential medical indicator Blood Pressure occurs whenever the respective activity Blood Pressure appears. The values can increase and decrease randomly over time for one instance. We created a continuous label for solving a regressionlike prediction task. The label contains five different additive parts to introduce five different medical indicator effects concerning the static and sequential attributes in our event log: ygender ,yage,ypattern,yhr−nl, and yhr (see Eq. B1). Each part can take values between 0 and 0.2, thus, the label ycan take values between 0 and 1. The remaining medical indicators (e.g., Foreigner and BMI) are only included as noise terms and do not have any effect on the target variable. y=ygender +yage +ypattern +yhr−nl +yhr .(B1) Thefirst part ygender showsthe influence ofthe static medical indicator gender, where bgender static is 1 if the patient is female and 0 otherwise: ygender =0.2∗bgender static ,with bgender static =1,if xgender static =1, 0,otherwise.(B2) The second part, yage, demonstrates the effect of the static medical indicator age on the label, which we model as a downward open parabola: yage =−0.8∗(xage static −0.5)2+0.2.(B3) Besides the influence of static medical indicators on the label y, we also include an influence of sequential medical indicators. First, bpattern seq is 1 if an instance contains a certain pattern in its sequence of activities regarding the administrationofthemedication,namely “Medication A, Medication A, Medication A, Medication A, Medication B”: ypattern =0.2∗bpattern seq ,with bpattern seq =⎧ ⎪ ⎨ ⎪ ⎩ 1,if xmeda t,seq =1,∀t∈{8,...,11} ∧xmedb 12,seq =1, 0,otherwise. (B4) Further, yhr−nl shows the effect of the medical indicator Heart Rate at time step t2as a downward open parabola: yhr−nl =−0.8∗(xhr 2,seq −0.5)2+0.2.(B5) Lastly, yhr describes the effect of the behavior of the medical indicator Heart Rate over time. The values can increase or decrease over time for a single patient pathway. If the values increase, bhr seq will take the value 1, otherwise it will be 0: yhr =0.2∗bhr seq,with bhr seq =1,if xhr t,seq −xhr t−1,seq >0,for t∈{2,3,4}, 0,otherwise. (B6) B.2 Experimental setting We optimize PatWay-Net based on the complete simulation data set comprising 50,000 patient pathways and then generate interpretation plots based on 1,000 patient pathways of the same data set. Furthermore, we do not consider any sequential medical indicator interaction as we are interested in investigating how well the model can learn the five additive parts described in the previous section. Finally, we set the hidden size per sequential and static medical indicator to 16, and the batch size, number of epochs, and learning rate to 32, 1,000, and 0.001, respectively, as the loss converged well with these values. B.3 Results We present PatWay-Net’s interpretation plots for the created synthetic event log and, based on the interpretation plots, we validate how well PatWay-Net captures the effects we modeled in the synthetic event log data. B.3.1 Importance of medical indicators Figure11showsthemedicalindicatorimportanceplotattime step t12. As expected, the medical indicators Gender,Heart Rate,Age,Medication A, and Medication B show an effect on the model output. The remaining indicators correctly show no effect, as they were only included in the simulation to serve as irrelevant noise terms. 123 S.Zilker et.al.
159 Fig. 11 Importance for static and sequential medical indicators B.3.2 Static medical indicator shape Figure 12 compares the global effect of the model output for the static categorical medical indicators Gender (left) andForeigner (right).PatWay-Netcorrectlydetectstheeffect of the medical indicator Gender with a constant value of 0.2, which equals the magnitude of the simulated coefficient whenever the indicator value is 1 (= female). By contrast, the indicator Foreigner has no effect. Similarly, Fig. 13 shows the static medical indicator shape plot for indicator Age (left) compared to the indicator BMI (right).8PatWay-Net can correctly detect the parabolic effect on the label of the medical indicator Age, with the strongest effect of 0.2, being achieved at a value of 0.5, as modeled in Eq. B3. By contrast, no influence was correctly detected for the indicator BMI. 8For better interpretability, we applied post-processing to the effect values by adjusting the y-axis to 0. B.3.3 Sequential medical indicator shape For the sequential medical indicators Heart Rate and Blood Pressure, the indicator effect on the prediction can be evaluated for each time step. Figure 14 shows the sequential medical indicator shape for Heart Rate (left) and Blood Pressure (right) exemplarily for the time step t4and t7, respectively.9PatWay-Net can correctly detect the quadratic effect on the label for the indicator Heart Rate. Indicator Blood Pressure, on the other hand, has no effect, as correctly detected by PatWay-Net. B.3.4 Sequential medical indicator transition As sequential medical indicators may change over time, it is crucialtoinvestigatetheireffectnotonlyat acertaintimestep but also the transition between two consecutive time steps. Figure15showsexemplarily thesequentialmedical indicator transition of the indicator value as well as the change of effect on the prediction for the indicators Heart Rate and Blood Pressure from time step t3(previous) to time step t4 (current). For the medical indicator Heart Rate, the indicator value can either linearly increase or linearly decrease by 30% from time step t3(x-axis) to t4(y-axis). For increasing cases, we can observe the effect changes stronger negatively, that is, decreases from t3to t4, from any indicator value to a higher indicator value (blue region). If the increase concerns two indicator values with a lower value, the change in the indicator effect is lower. This is due to the parabolic shape of the medical indicator Heart Rate (see Fig. 14) as defined in Eq. B6. For the decreasing cases, we can observe the effect changes stronger positively, that is, increases from t3to t4, from a higher to a lower indicator value (red region). As for increasing cases, if the decrease concerns two indicator values with a lower value, the change in the indicator effect is lower. Again, this is due to the parabolic shape of the medical indicator Heart Rate (see Fig. 14). For the medical indicator Blood Pressure, the indicator values can either randomly increase or decrease from time step t6(x-axis) to t7(y-axis), as demonstrated by the plot (see Fig. 14). Further, the effect on the prediction of the medical indicator does not change over time, as the complete region is colored white, indicating a change in the indicator effect of 0. Thus, in summary, we can see that PatWay-Net correctly detects the effect of the transition of the Heart Rate asmodeledin Eq. B6,whereasBlood Pressure doesnotaffect the model output. 9We select these time steps because they are the last time steps at which aHeart Rate and Blood Pressure activity occur. However, the remaining figures can be found in the repository. 123 A machine learning framework for interpretabl e…
166 29. Rello J, Valenzuela-Sánchez F, Ruiz-Rodriguez M, Moyano S (2017) Sepsis: A review of advances in management. Adv Therapy 34:2393–2411. https://doi.org/10.1007/s12325-017-0622-8 30. Schuurman AR, Sloot PM, Wiersinga WJ, van der Poll T (2023) Embracing complexity in sepsis. Critical Care 27(1):102. https:// doi.org/10.1186/s13054-023-04374-0 31. Thorsen-Meyer H-C, Nielsen AB, Nielsen AP, Kaas-Hansen BS, Toft P, Schierbeck J, Strøm T, Chmura PJ, Heimann M, Dybdahl L et al (2020) Dynamic and explainable machine learning prediction of mortality in patients in the intensive care unit: A retrospective study of high-frequency data in electronic patient records. Lancet Digit Health 2(4):179–191. https://doi.org/10.1016/S2589- 7500(20)30018-2 32. Loh HW, Ooi CP, Seoni S, Barua PD, Molinari F, Acharya UR (2022) Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011–2022). Comput Methods Prog Biomed 226:107161. https://doi.org/10. 1016/j.cmpb.2022.107161 33. Rudin C (2019) Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell 1(5):206–215. https://doi.org/10.1038/s42256- 019-0048-x 34. Moreno RP, Metnitz PGH, Almeida E, Jordan B, Bauer P, Campos RA, Iapichino G, Edbrooke D, Capuzzo M, Le Gall J-R, SAPS 3 Investigators (2005) SAPS 3–From evaluation of the patient to evaluation of the intensive care unit. Part 2: Development of a prognostic model for hospital mortality at ICU admission. Intensive Care Med 31(10): 1345–1355. https://doi.org/10.1007/ s00134-005-2763-5 35. Galanti R, Coma-Puig B, d Leoni M, Carmona J, Navarin N (2020) Explainable predictive process monitoring. In: Proceedings of the 2nd international conference on process mining, pp 1–8 36. Weinzierl S, Zilker S, Brunk J, Revoredo K, Matzner M, Becker J (2020) XNAP: Making LSTM-based next activity predictions explainable by using LRP. In: Proceedings of the BPM 2020 international workshop, pp 129–141 37. Breuker D, Matzner M, Delfmann P, Becker J (2016) Comprehensible predictive models for business processes. MIS Quarterly 40(4):1009–1034 38. Lakshmanan GT, Shamsi D, Doganata YN, Unuvar M, Khalaf R (2015) A Markov prediction model for data-driven semi-structured business processes. Knowl Inf Syst 42(1):97–126. https://doi.org/ 10.1007/s10115-013-0697-8 39. Kaji DA, Zech JR, Kim JS, Cho SK, Dangayach NS, Costa AB, Oermann EK (2019) An attention based deep learning model of clinical events in the intensive care unit. PloS one 14(2):0211057. https://doi.org/10.1371/journal.pone.0211057 40. Zhang D, Yin C, Hunold KM, Jiang X, Caterino JM, Zhang P (2021)Aninterpretable deep-learning model for early prediction of sepsis in the emergency department. Patterns 2(2):100196. https:// doi.org/10.1016/j.patter.2020.100196 41. Esteban C, Staeck O, Baier S, Yang Y, Tresp V (2016) Predicting clinical events by combining static and dynamic information using recurrent neural networks. In: Proceedings of the 2016 IEEE international conference on healthcare informatics, pp 93–101 42. BreimanL(2001)Randomforests.MachLearn45(1):5–32.https:// doi.org/10.1023/A:1010933404324 43. Chen T, Guestrin C (2016) XGBoost: A scalable tree boosting system. In: Proceedings of the 22nd international conference on knowledge discovery and data mining, pp 785–794 44. Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735–1780. https://doi.org/10.1162/neco. 1997.9.8.1735 45. Rai A (2020) Explainable AI: From black box to glass box. J Acad Market Sci 48(1):137–141. https://doi.org/10.1007/s11747-019- 00710-5 46. Lundberg SM, Lee S-I (2017) A unified approach to interpreting model predictions. In: Proceedings of the 30th conference on advances in neural information processing systems, pp 4765–4774 47. Senoner J, Netland T, Feuerriegel S (2022) Using explainable artificial intelligence to improve process quality: Evidence from semiconductor manufacturing. Manage Sci 68(8):5704–5723. https:// doi.org/10.1287/mnsc.2021.4190 48. Babic B, Gerke S, Evgeniou T, Cohen IG (2021) Beware explanations from AI in health care. Science 373(6552):284–286. https:// doi.org/10.1126/science.abg1834 49. Zschech P, Weinzierl S, Hambauer N, Zilker S, Kraus M (2022) GAM(e) changer or not? An evaluation of interpretable machine learning models based on additive model constraints. In: Proceedings of the 30th European Conference on Information Systems, pp 1–18 50. Hastie T, Tibshirani R (1986) Generalized additive models. Stat Sci 1(3):297–310. https://doi.org/10.1214/ss/1177013604 51. Lou Y, Caruana R, Gehrke J, Hooker G (2013) Accurate intelligible models with pairwise interactions. In: Proceedings of the 19th ACMSIGKDDInternationalConferenceonKnowledgeDiscovery and Data Mining, pp 623–631 52. Lou Y, Caruana R, Gehrke J (2012) Intelligible models for classification and regression. In: Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp 150–158 53. Yang Z, Zhang A, Sudjianto A (2021) Gami-net: An explainable neural network based on generalized additive models with structured interactions. Pattern Recogn 120:108192. https://doi.org/10. 1016/j.patcog.2021.108192 54. AgarwalR,MelnickL,FrosstN,ZhangX,LengerichB,CaruanaR, Hinton GE (2021) Neural additive models: Interpretable machine learning with neural nets. In: Proceedings of the 34th Conference on Advances in Neural Information Processing Systems, pp 4699– 4711 55. Kraus M, Tschernutter D, Weinzierl S, Zschech P (2023) Interpretable generalized additive neural networks. Eur J Oper Res. https://doi.org/10.1016/j.ejor.2023.06.032 56. Marquez-Chamorro AE, Resinas M, Ruiz-Cortes A (2018) Predictive monitoring of business processes: A survey. IEEE Trans Serv Comput 11(6):962–977. https://doi.org/10.1109/TSC.2017. 2772256 57. Huang Z, Lu X, Duan H, Fan W (2013) Summarizing clinical pathways from event logs. J Biomed Inform 46(1):111–127. https://doi. org/10.1016/j.jbi.2012.10.001 58. Kingma D, Ba J (2015) Adam: A method for stochastic optimization. In: Proceedings of the International Conference on Learning Representations 59. Moreno R, Miranda DR (1998) Nursing staff in intensive care in europe: The mismatch between planning and practice. Chest 113(3):752–758. https://doi.org/10.1378/chest.113.3.752 60. Demšar J (2006) Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res 7:1–30. https://doi.org/10.5555/ 1248547.1248548 61. Klein SJ, Lehner GF, Forni LG, Joannidis M (2018) Oliguria in critically ill patients: A narrative review. J Nephrol 31:855–862. https://doi.org/10.1007/s40620-018-0539-6 62. Hotchkiss RS, Moldawer LL, Opal SM, Reinhart K, Turnbull IR, Vincent J-L (2016) Sepsis and septic shock. Nat Rev Dis Prim 2(1):1–21. https://doi.org/10.1038/nrdp.2016.45 63. Urrechaga E, Bóveda O, Aguirre U (2018) Role of leucocytes cell population data in the early detection of sepsis. J Clin Pathol 71(3):259–266. https://doi.org/10.1136/jclinpath-2017-204524 64. Pera A, Campos C, López N, Hassouneh F, Alonso C, Tarazona R, Solana R (2015) Immunosenescence: Implications for response to infection and vaccination in older people. Maturitas 82(1):50–55. https://doi.org/10.1016/j.maturitas.2015.05.004 123 S.Zilker et.al.
167 65. Riley LK, Rupert J (2015) Evaluation of patients with leukocytosis. Am Fam Phys 92(11):1004–1011 66. Belok SH, Bosch NA, Klings ES, Walkey AJ (2021) Evaluation of leukopenia during sepsis as a marker of sepsis-defining organ dysfunction. PLoS One 16(6):0252206. https://doi.org/10.1371/ journal.pone.0252206 67. Acosta JN, Falcone GJ, Rajpurkar P, Topol EJ (2022) Multimodal biomedical AI. Nat Med 28(9):1773–1784. https://doi.org/ 10.1038/s41591-022-01981-2 68. van Weenen E, Banholzer N, Föll S, Zueger T, Fontana FY, Skroce K, Hayes C, Kraus M, Feuerriegel S, Lehmann V et al (2023) Glycaemic patterns of male professional athletes with type 1 diabetes during exercise, recovery and sleep: Retrospective, observational study over an entire competitive season. Diabetes, Obesity and Metabolism. https://doi.org/10.1111/dom.15147 69. Cawley GC, Talbot NL (2010) On over-fitting in model selection and subsequent selection bias in performance evaluation. J Mach Learn Res 11:2079–2107. https://doi.org/10.5555/1756006. 1859921 70. CawleyGC, Talbot NL(2007)Preventingover-fittingduringmodel selection via bayesian regularisation of the hyper-parameters. J Mach Learn Res 8(4). https://doi.org/10.5555/1248659.1248690 71. Qi Y, Minka TP, Picard RW, Ghahramani Z (2004) Predictive automatic relevance determination by expectation propagation. In: Proceedings of the 21st International Conference on Machine Learning, p 85 72. Slack D, Krishna S, Lakkaraju H, Singh S (2023) Explaining machine learning models with interactive natural language conversations using TalkToModel. Nat Mach Intell 5(8):873–883. https:// doi.org/10.1038/s42256-023-00692-8 73. Feuerriegel S, Hartmann J, Janiesch C, Zschech P (2023) Generative AI. Business & Information Systems Engineering. https://doi. org/10.1007/s12599-023-00834-7 74. Davis J, Goadrich M (2006) The relationship between precisionrecall and ROC curves. In: Proceedings of the 23rd International Conference on Machine Learning, pp 233–240 75. McClish DK (1989) Analyzing a portion of the ROC curve. Med Dec Making 9(3):190–195. https://doi.org/10.1177/ 0272989X89009003 76. Bradley AP (1997) The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recogn 30(7):1145–1159. https://doi.org/10.1016/S0031- 3203(96)00142-2 77. Teinemaa I, Dumas M, Rosa ML, Maggi FM (2019) Outcomeoriented predictive process monitoring: Review and benchmark. ACM Trans Knowl Dis Data 13(2):1–57. https://doi.org/10.1145/ 3301300 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Authors and Affiliations Sandra Zilker1,2 ·Sven Weinzierl2·Mathias Kraus3·Patrick Zschech4· Martin Matzner2 Sven Weinzierl [email protected] Mathias Kraus [email protected]urg.de Patrick Zschech [email protected] Martin Matzner [email protected] 1Technische Hochschule Nürnberg Georg Simon Ohm, Professorship for Business Analytics, Hohfederstraße 40, 90489 Nuremberg, Germany 2Friedrich-Alexander-Universität Erlangen-Nürnberg, Chair of Digital Industrial Service Systems, Fürther Straße 248, 90429 Nuremberg, Germany 3University of Regensburg, Chair for Explainable AI in Business Value Creation, Bajuwarenstraße 4, 93053 Regensburg, Germany 4Leipzig University, Professorship for Intelligent Information Systems and Processes, Grimmaische Straße 12, 04109 Leipzig, Germany 123 A machine learning framework for interpretabl e…