scieee AI-readable full text Open interactive document viewer

Machine learning predictions of unplanned readmissions using electronic medical records: Predictor importance across medical and surgical patient populations

Havranek, Michael; Hwang, Aljoscha Benjamin; Funke, Ilona; Kuhlen, Dominique Emmanuelle; Liedtke, Daniel; Boes, Stefan

Abstract

Hospital readmissions prolong patient suffering and increase healthcare expenditures. While several studies have attempted to develop prediction models to reduce readmissions, most have demonstrated modest predictive accuracy. To improve upon prior approaches, we conducted an overview of systematic reviews to identify the most relevant predictor variables, then subsequently developed machine learning models in a retrospective, multisite study across eight hospitals. The patient sample comprised 200,799 inpatient stays from eligible hospitalizations, based on the Centers for Medicare and Medicaid Services (CMS) definition of unplanned readmissions within 30 days of discharge. We constructed random forest models and evaluated out-of-sample performance using the area under the receiver operating characteristic curve (AUC) across different train–test splits. The hospital-wide sample was divided into medical and surgical cohorts to investigate predictor importance across different patient populations. The average AUC score was 0.78 ± 0.01 (mean ± standard deviation [SD]). Patients’ diagnoses were the most important predictor variables (contributing 18.4% ± 0.15 to the model’s decision, mean ± standard error [SE]), followed by nursing assessments (11.2% ± 0.04, mean ± SE) and procedural information (10.8% ± 0.09, mean ± SE). Comparing medical and surgical patients, we found that medications and prior healthcare use (e.g., prior emergency encounters) were more important in the medical compared with the surgical cohort, whereas procedural information and healthcare provider information (e.g., physician caseload) were more relevant in the surgical relative to the medical cohort. In conclusion, we have established the feasibility of using Swiss electronic medical record (EMR) data to accurately predict unplanned readmissions. The reported variable importances may guide future research and inform development of clinical decision support systems aimed at reducing readmissions.

Full text

PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 1 / 17 OPEN ACCESS Citation: Havranek MM, Hwang AB, Funke I, Kuhlen D, Liedtke D, Boes S (2025) Machine learning predictions of unplanned readmissions using electronic medical records: Predictor importance across medical and surgical patient populations. PLoS One 20(9): e0331263. https://doi.org/10.1371/journal.pone.0331263 Editor: Jan Chrusciel, Centre Hospitalier de Troyes, FRANCE Received: January 7, 2025 Accepted: August 6, 2025 Published: September 4, 2025 Copyright: © 2025 Havranek et al . This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data availability statement: The administrative data that support the findings of this study are available from the Swiss Federal Office of Statistics (contact via [email protected]. ch) for researchers who meet the criteria for access to confidential data. The electronic RESEARCH ARTICLE Machine learning predictions of unplanned readmissions using electronic medical records: Predictor importance across medical and surgical patient populations Michael M. Havranek 1*, Aljoscha B. Hwang2,3, Ilona Funke2, Dominique Kuhlen4, Daniel Liedtke 4, Stefan Boes3 1 Competence Center for Health Data Science, Faculty of Health Science and Medicine, University of Lucerne, Lucerne, Switzerland, 2 Medical Department, Hirslanden Group, Zurich, Switzerland, 3 Faculty of Health Science and Medicine, University of Lucerne, Lucerne, Switzerland, 4 Group Executive Board, Hirslanden Group, Zurich, Switzerland * [email protected] Abstract Hospital readmissions prolong patient suffering and increase healthcare expenditures. While several studies have attempted to develop prediction models to reduce readmissions, most have demonstrated modest predictive accuracy. To improve upon prior approaches, we conducted an overview of systematic reviews to identify the most relevant predictor variables, then subsequently developed machine learning models in a retrospective, multisite study across eight hospitals. The patient sample comprised 200,799 inpatient stays from eligible hospitalizations, based on the Centers for Medicare and Medicaid Services (CMS) definition of unplanned readmissions within 30 days of discharge. We constructed random forest models and evaluated out-of-sample performance using the area under the receiver operating characteristic curve (AUC) across different train–test splits. The hospital-wide sample was divided into medical and surgical cohorts to investigate predictor importance across different patient populations. The average AUC score was 0.78 ± 0.01 (mean ± standard deviation [SD]). Patients’ diagnoses were the most important predictor variables (contributing 18.4% ± 0.15 to the model’s decision, mean ± standard error [SE]), followed by nursing assessments (11.2% ± 0.04, mean ± SE) and procedural information (10.8% ± 0.09, mean ± SE). Comparing medical and surgical patients, we found that medications and prior healthcare use (e.g., prior emergency encounters) were more important in the medical compared with the surgical cohort, whereas procedural information and healthcare provider information (e.g., physician caseload) were more relevant in the surgical relative to the medical cohort. In conclusion, we have established the feasibility of using Swiss electronic medical record (EMR) data to accurately predict unplanned readmissions. The reported variable importances may guide PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 2 / 17 future research and inform development of clinical decision support systems aimed at reducing readmissions. Introduction Hospital readmissions lead to patient suffering and increase healthcare expenditures. Many countries therefore use readmission rates as quality indicators, most prominently the unplanned readmission rates within 30 days of discharge as developed and popularized by the American Centers for Medicare and Medicaid Services (CMS) [1,2]. These unplanned readmission rates have been incorporated into Switzerland’s national quality monitoring since 2022 [3]. For this reason, Swiss hospitals aiming to improve quality of care are in need of effective tools and strategies to reduce unplanned readmissions. Multiple studies have developed prediction models for readmissions to improve quality of care and reduce healthcare expenditures, although most of the resulting models have shown modest prediction accuracy (see, e.g., [4]). The comparability of many previous studies is complicated by the fact that they use different definitions and operationalizations of readmissions (e.g., all-cause readmissions vs. condition-specific readmissions [5]). However, two studies should be mentioned in an exemplary manner in the context of the present study: The regression-based risk-adjustment models used by CMS for hospital-wide unplanned readmissions using the entire eligible US population enrolled in CMS achieved area under the receiver operating characteristic curve (AUC) values of 0.64–0.68 [6]. Another recent large-scale study from our neighboring country Germany using machine-learning models with over 4 million hospital stays attained AUC values of 0.68–0.69 [7]. The majority of previous approaches have relied on administrative data to predict readmissions. On the one hand, this has the advantage of including information that is available only after discharge (e.g., all coded diagnoses, procedures, and diagnosis-related groups [DRGs]). On the other hand, administrative data lacks other crucial information such as vital signs and medication. Electronic medical records (EMRs) contain a broader range of data sources but are limited to information available before discharge, and their use is often hindered by missing values and other challenges associated with real-world data [8,9]. In line with previous reports [10], we hypothesized that unplanned readmissions could be predicted more accurately using the EMR data available before discharge, provided that expert knowledge on predicting readmissions is combined with detailed feature engineering efforts to derive a comprehensive set of predictor variables. To achieve this, we adopted a two-step approach. Previously, we conducted an extensive overview of systematic reviews across 440 previous studies to identify the most relevant variables for predicting unplanned readmissions (see [11] and the discussion section for more details). Now, we report the development of tree-based machine learning models to predict readmissions across hospital-wide, medical, and surgical patient populations at eight hospitals in Switzerland. As part of this second step, we invested medical record data of the patients cannot be shared publicly. Funding: The author(s) received no specific funding for this work. Competing interests: The authors have declared that no competing interests exist. PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 3 / 17 considerable effort in feature engineering to maximize the utility of Swiss EMRs and address two central research questions: First, can unplanned readmissions within 30 days of discharge be accurately predicted using Swiss EMR data available before discharge? Second, does feature importance of predictor variables vary across different patient populations? Materials and methods Study design and data This retrospective multisite study was conducted across eight hospitals from a leading private hospital group in Switzerland. The participating hospitals provided administrative medical data to identify unplanned readmissions using a version of the CMS algorithm adapted for the Swiss medical coding system (see below), along with medical record data to build the prediction models for the years 2018–2024. The administrative dataset [12] contained all inpatient stays treated by the hospitals during the study period, including up to 50 diagnosis codes for each stay (from the International Statistical Classification of Diseases and Related Health Problems, 10th Revision, German Modification, ICD-10-GM [13]), up to 100 procedure codes (from the Swiss classification of surgical interventions, CHOP [13]), the DRG (from the SwissDRG system [14]), and other clinically relevant variables such as admission and discharge conditions, and patients’ demographic information. The EMR data contained all available information from patient documentation (e.g., discharge letters, surgery reports, nursing notes, vital signs, medications, and administrative information on patients and hospitals, as well as radiology, laboratory, and other clinical results). The study was approved through a jurisdictional inquiry by the Ethics Committee Northwest & Central Switzerland (May 19, 2022; ID: Req-2022–00616). The EMR data was accessed on October 12, 2024 by an employee from the hospital group. Informed consent was not required from the patients as the researchers only received fully anonymized data that did not contain any information to identify individual patients. Sample, outcome, and predictors The investigated sample consisted of 200,799 inpatient stays from eligible hospitalizations based on the CMS inclusion/ exclusion criteria for a hospital-wide population (i.e., adults older than 18 years, excluding psychiatric and rehabilitation stays, etc.; see [15,16]). This hospital-wide sample was further divided into a medical cohort (n = 45,073, including cardiovascular, cardiorespiratory, neurological, and other medical patients) and a surgical/gynecologic cohort (n = 155,726; hereafter, “surgical cohort”) based on CMS definitions to build separate prediction models for these distinct patient populations. The outcome of interest was “unplanned readmissions” within 30 days of discharge, defined as readmissions resulting from acute clinical events requiring urgent rehospitalization (i.e., not planned or foreseen during the index hospitalization [15]). Unplanned readmissions were flagged according to the CMS definitions (version 2020 [15,16]), using the CMS method as previously translated into the Swiss medical coding system and slightly modified for the Swiss healthcare setting (conceptually proposed in [17], described in [3], and validated in [18]). See S1 Appendix for a comparison between the original CMS method and our adapted version. The timeframe of 30 days was chosen because it is the most frequently employed timeframe in quality monitoring as well as prediction of unplanned readmissions [4], which has been explained by the fact that patients are particularly vulnerable to readmissions during that timeframe [19]. Only internal readmissions within the hospital group could be identified in the data from our collaborating hospitals; external readmissions outside the hospital group could not be identified within the data. Table 1 provides a description of all predictor variables, grouped into main categories and subcategories. Selection of these variables was based on our initial overview of systematic reviews [11] and included all structured and unstructured information available in the EMRs of the collaborating hospital group. The variables include patient demographics, diagnoses, procedures, medications, prior healthcare use, admission and procedural information, vital signs, laboratory and radiology results, nursing assessments, health behaviors and living conditions, and available information on healthcare providers. PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 4 / 17 Table 1. Details of variables, groupings, and missing data. Variable Definition Type N cataMissing values Demographics Age Age in years at admission Continuous – – Sex Sex Categorical 2 – Insurance Insurance category during the hospital stay Categorical 3 – Foreigner Foreign nationality/language Binary 2 – Admission information Admission day Day of week of index admission Categorical 7 – Admission month Month of index admission Categorical 12 – Admission reason Reason for hospitalization Categorical 4 – Residence Patient’s residence before index admission Categorical 9 – Diagnosis-related groups DRG Working DRG at index admission [14] Categorical 711 0.5% Prior healthcare use Prior hospitalizations Yes/No variable, plus number of hospitalizations within 90/180/365 days prior to index admission date Binary/ continuous – – Prior hospitalization, LOS ≥ 10 days Yes/No variable, plus count variable of hospitalizations with LOS ≥ 10 days within 90/180/365 days prior to index admission date Binary/ continuous 2 – Prior ED encounters Yes/No variable, plus number of emergency encounters within 90/180/365 days prior to index admission date Binary/ continuous – – Prior outpatient visits Yes/No variable, plus number of outpatient visits within 90/180/365 days prior to index admission date Binary/ continuous – – Diagnoses and comorbidities All diagnoses Known diagnoses (3-digit ICD-10 GM codes) [13] Categorical 1,152 1.2%c Condition categories Condition category diagnosis groups that are used in risk adjustment of the unplanned readmission rates by CMS [20] Categorical 239 – Specific diagnosis groups Specific diagnosis groups related to unplanned readmissions mentioned in the scientific literature (anemia, fluid and electrolyte disorders, renal failure, drug abuse, etc.) [11] Each binary 2 – Comorbidity count Count of all/selected diagnoses at admission (e.g., Elixhauser groups) Continuous – – Comorbidity scores Comorbidity scores such as the Royal College of Surgeons Charlson Comorbidity Score and Elixhauser Comorbidity Index [6,21] Categorical/ continuous 4 – Procedures All procedures Performed procedures (3-digit CHOP codes) [13] Categorical 670 3.5%c Planned procedures count Count of planned procedures at admission Continuous – – Procedure risk Surgical procedure risk classification [22] Categorical 6 – Procedural information LOS Length of stay at the day of model computation Continuous – – LEP Patient-specific total acute care nursing services provided, in minutes [23] Continuous – – Operating room Utilization of the operating room as a Yes/No variable, plus LOS in minutes Binary/ continuous 2 – Incision suture Incisions suture time in minutes Continuous – – Anesthesia Application of anesthesia as a Yes/No variable, plus duration in minutes Binary/ continuous 2 – Ventilation Artificial ventilation provided as a Yes/No variable, plus duration in hours Binary/ continuous 2 – Recovery room Utilization of the recovery room as a Yes/No variable, plus LOS in minutes Binary/ continuous 2 – (Continued) PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 5 / 17 Variable Definition Type N cataMissing values Intermediate care unit Utilization of intermediate care as a Yes/No variable, plus LOS in minutes Binary/ continuous 2 – Intensive care unit Utilization of intensive care as a Yes/No variable, plus LOS in minutes Binary/ continuous 2 – ASA ASA classification [24] Categorical 8 – NEMS A score evaluating nine specific nursing services (effort) commonly provided in intensive care units [25] Continuous – – SAPS Simplified Acute Physiology Score II [26] Continuous – – Vital signs Blood pressure Systolic and diastolic blood pressure (most recent, mean, maximum, minimum, and median values), in mm Hg Continuous – 30.5% Heart rate Heart rate (most recent, mean, maximum, minimum, and median values), in bpm Continuous – 30.5% Oxygen saturation Oxygen saturation (most recent, mean, maximum, minimum, and median values), in % Continuous – 31.2% Temperature Temperature (most recent, mean, maximum, minimum, and median values), in degrees Celsius Continuous – 31.2% Laboratory results Labs norm Flags any laboratory results outside the normal range, including hematocrit, creatinine, urea, phosphate, potassium, C-reactive protein, calcium, lactate, alanine aminotransferase, alkaline phosphatase, aspartate aminotransferase, chloride, cystatin C, gamma-glutamyltransferase, HDL cholesterol, LDL cholesterol, total cholesterol, triglycerides, international normalized ratio, potassium, creatinase, lactate dehydrogenase, leukocytes, lithium, magnesium, methotrexate, natriuretic peptide, thrombocytes, troponin T Each binary 2 – Labs high Flags any laboratory results higher than the upper normal value, for the same analyses as Labs norm Each binary 2 – Labs low Flags any laboratory results below the lower normal value, for the same analyses as Labs norm Each binary 2 – Labs Yes/No Flags the presence of laboratory results of any kind, for the same analyses as Labs norm Each binary 2 – POCT blood sugar Point-of-care testing of blood sugar as a Yes/No variable, plus recorded values (most recent, minimum, maximum, and mean), in mmol/L Binary/ continuous 2 9.7%d Medication Medi all Three-digit ATC classification codes of all administered medication during hospitalization, including reserve Categorical 2 – Medi specific Specific administered medication (including reserve) of certain anatomical groups or alimentary system and metabolism, related to unplanned readmissions (e.g., beta-blockers, corticosteroids, drugs used in diabetes), on entry and during hospitalization Categorical 2 – Medi count Count of distinct ATC codes present on entry or administered during hospitalization, including reserve Continuous – – Medi poly Polymedication status on entry or during hospitalization. Polymedication refers to the number of administered medications >6, including reserve Categorical 2 – Radiology Rad measure count Count of radiology measures (grouped by device and/or anatomical region) Continuous – – Rad urgency Max urgency of radiology measures (grouped by device and/or anatomical region) Continuous – – Nursing assessment EPA mobility Most recent values of movement, exhaustion, balance disorder, etc. [27] Categorical 2–5 18.2–18.5% EPA hygiene Most recent values of personal hygiene and getting dressed Categorical 4 18.2–18.5% EPA diet Most recent values of drinking, eating, nausea, artificial nutrition, etc. Categorical 2–5 18.2–18.4% EPA excretion Most recent values of urine excretion, control of urine excretion, stool excretion, etc. Categorical 2–4 18.2–18.4% EPA cognition Most recent values of vigilance, orientation, attention, fall-risk-increasing medication intake, etc. Categorical 2–5 18.2–18.4% Table 1. (Continued) (Continued) PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 6 / 17 Variable Definition Type N cataMissing values EPA communication Most recent values of hearing, vision, communication, challenging behavior, etc. Categorical 3–5 18.2–18.4% EPA sleep Most recent values of falling asleep, sleep–wake rhythm Categorical 3 18.2% EPA respiration Most recent values of tracheostomy, ventilation >24 h, chronic disorder of respiratory system, etc. Categorical 2–3 18.4–18.6% EPA pain Most recent values of chronic pain, sadness, anxiety, etc. Categorical 3–7 18.2% EPA medical aids Most recent values of medical aids with regard to hearing, vision, mobility, and nutrition. Categorical 2–3 16.5% EPA other Includes pneumonia risk, fall risk, and malnutrition risk (minimum, maximum, most recent, and mode) Categorical 2 50.9–70.6% DOS Yes/No variable for delirium observation screening, plus values for most recent, maximum, minimum, and mode [28] Binary/ continuous 2 12.0%d GCS Yes/No variable for Glasgow Coma Scale score, plus values for most recent, maximum, minimum, and mode [29] Binary/ continuous 2 3.9%d NRS-M/R Numeric Pain Rating Scale scores when moving/resting (most recent, mean, maximum, minimum, and median) [30] Continuous – 58.4–64.9%d NRS Yes/No variable for nutritional risk screening, plus values for most recent, maximum, minimum, and mode [31] Binary/ continuous 2 4.4%d SPI Yes/No variable for self-care index, plus values for most recent, maximum, minimum, mean, and median [27] Binary/ continuous 2 18.4% PainbPain (severe) within 36 h before discharge Binary 2 4.2–40.6%d BleedingbBleeding (severe) within 36 h before discharge Binary 2 5.4–16.7%d EmotionsbEmotional distress Binary 2 1.7%d Health behaviors and living conditions BMI Height (in cm), weight (in kg), and BMI Continuous – 39.4–44.8% NoxaebDescribes regular consumption of alcoholic beverages, use of addictive substances, and smoking status with pack years Binary/ continuous 2 41.1%c Discharge destination Planned discharge destination Categorical 2–8 – Living conditionsbExisting social network and available support at home Categorical 2 3.2–11.4%d Healthcare provider information Hospital name Hospital name Categorical 8 – Hospital type Federal Bureau of Statistics hospital typology [32] Categorical 3 – Patient mix Hospital-specific proportion of foreign/private and semi-private/out-of-canton inpatients (2022) Continuous – – Beds & occupancy Hospital-specific average number of hospital beds in service and bed occupancy rate (2022) Continuous – – Staff Hospital-specific full-time nursing FTEs and FTEs per patient bed (2022) Continuous – – Inpatient volume Hospital-specific number of hospitalizations for 90/180/365 days prior to index admission date Continuous – – Outpatient volume Hospital-specific number of outpatients treated (2022) Continuous – – Physician caseload Physician caseload and procedure volume performed 90/180/365 days prior to index admission date Continuous – – ASA, American Society of Anesthesiologists; ATC, anatomical therapeutic chemical; BMI, body mass index; CMS, Centers for Medicare and Medicaid Services; DRG, diagnosis-related group; DOS, delirium observation screening; ED, emergency department; EPA, outcome-oriented nursing assessment; FTE, full-time equivalent; GCS, Glasgow Coma Scale; ICD-10 GM, International Classification of Diseases, 10th Revision, German Modification; LEP, nurse activity record; LOS, length of stay; NEMS, Nine Equivalents of Nursing Manpower; NRS, nutritional risk screening; NRS-M/R, Numeric Pain Rating Scale – Moving/Resting; POCT, point-of-care testing; SAPS, Simplified Acute Physiology Score II; SPI, self-care index. aIn each case, the number of categories (N cat) was counted before exclusion of rare characteristics. bInformation was extracted from nursing notes using an LLM. cPercentage of subjects without at least one diagnosis or procedure. dPercentage of subjects with at least one data point. https://doi.org/10.1371/journal.pone.0331263.t001 Table 1. (Continued) (Continued) PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 7 / 17 Only information available before 9:00 AM on the day of discharge was used. Unstructured information (e.g., nursing notes) was processed using an internally employed large language model (LLM). Categorical variables were converted into binary dummy variables, while ICD-10 GM, CHOP, and anatomical therapeutic chemical (ATC) codes were restricted to three digits before transforming them into dummy variables. Binary variables with rare occurrences (<200 instances) were excluded from modeling. Continuous variables were used both in their original form and as binary variables indicating whether a certain event or measurement occurred (e.g., number of prior hospitalizations, alongside yes/no information on whether a previous hospitalization occurred within a specific look-back period; see Table 1). For continuous variables with repeated measurements during stays (e.g., vital signs), different aggregation methods (including mean, median, maximum, minimum, and most recent values) were evaluated in a preselection step on the training data prior to modeling, as outlined below. Similarly, continuous variables assessing different time frames (e.g., prior hospitalizations within 3, 6, or 12 months) were preselected before modeling. Missing values were imputed using the median, separately for the training and test data (see below). As shown in Table 1, certain variable groups had large quantities of missing data. Nevertheless, these variables were included during modeling to allow them to be used for patient subsets where they were available and relevant. Model building, evaluation, and interpretability Models were built in Python (version 3.9.13, [33]) using random forests (i.e., tree-based machine learning algorithms) from the scikit-learn library [34]. Random forests were used because tree-based algorithms have been found to be well-suited for structured data and have shown an ability to handle missing data well (see, e.g., [35]). The data was split into training and test datasets at an 80:20 ratio. Preselection of comparable variables (e.g., variables with different aggregation methods, see above) was performed with SelectKBest from the scikit-learn library using analysis of variance (ANOVA) F-values [34]. To address class imbalance of the outcome, overand under-sampling were evaluated for the training data but were not used during final modeling because they did not provide performance benefits on the training data. GridSearch was employed to select the number of estimators (n = 500), maximum depth (15), maximum features (“sqrt”), minimum number of samples for splitting (10), minimum number of samples for leaves (5), and the selection criterion (“Gini”) across three cross-validation folds of the training data. Modeling was repeated with 20 different train–test splits to ensure reproducibility. During each run, two models were built on the training data: a first model using all available variables, and a second (final) model selecting only the 75% of variables with the highest importance from the first run. Model performance was evaluated on the test data using AUC values. Sensitivity and specificity were calculated using the optimal threshold identified via Youden’s J statistic [36]. Model interpretability was derived from analysis of Gini feature importances (i.e., the normalized total reduction of the evaluation criterion due to a particular feature). To demonstrate the results’ robustness independent of the employed machine learning algorithm, additional results are provided in S5 Appendix comparing the model performance of XGBoost (eXtreme Gradient Boosting) with those of the random forests. To build the XGBoost models, the same pipeline was used (see first paragraph) and GridSearch was employed again to select the number of estimators (n = 500), maximum depth (15), learning rate (0.01), subsample ratio (0.8), column sample by tree (0.8), L1 regularization (1), L2 regularization (5), and scale of positive weight (1). The code used in this study is available on GitHub (https://github.com/MMH-999/Pred_Read). Results Population characteristics Table 2 summarizes the main characteristics of the three patient populations. As is typical for private hospitals in Switzerland, the sample includes a large number of surgical relative to medical patients (155,726 vs. 45,073), along with a majority of elective hospitalizations (79.3%) and, most importantly, a substantial proportion of privately insured patients (49.7%). PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 8 / 17 Out of 200,799 eligible hospitalizations, 7,755 (3.9%) were followed by an unplanned readmission within the hospital group. The percentage of unplanned readmissions was higher in the medical cohort than in the surgical cohort (6.5% vs. 3.1%). The average age was also higher for the medical compared with the surgical cohort (mean ± standard deviation [SD]: 69 ± 17 vs. 59 ± 17 years, respectively). In contrast, the proportion of elective admissions was significantly lower in the medical cohort (42.7% vs. 88.9%). Model performance Evaluation of out-of-sample model performance from the random forests was performed using test data across the different train–test splits, yielding average AUC scores of 0.78 ± 0.01 (mean ± standard deviation [SD]:) for the hospital-wide sample, 0.72 ± 0.01 for the medical cohort, and 0.78 ± 0.01 for the surgical cohort. The models yielded an average sensitivity and specificity of 0.75 ± 0.04 and 0.68 ± 0.04 for the hospital-wide sample, 0.70 ± 0.06 and 0.62 ± 0.06 for the medical cohort, 0.73 ± 0.02 and 0.71 ± 0.02 for the surgical cohort. The average model performances are visualized for the three patient populations in Fig 1. Individual performances across the various train– test splits are provided in S2–S4 Figs. In addition, the performance of the random forest models is compared with that of the XGBoost models in S5 Fig. The two machine learning algorithms yield nearly identical average AUC scores across the 20 train-test splits. Model interpretability Tables 3 and 4 show the average relative feature importance for all predictor groups and subgroups across the three patient populations, based on the full set of prediction models for the 20 different train–test splits. As shown in Table 3, patients’ diagnoses are the most important predictors for unplanned readmissions (mean ± standard error Table 2. Main characteristics of the included patient populations. Patient cohorts Characteristics HWR MED SURG Subjects, n 200,799 45,073 155,726 Unplanned readmissions, n (%) 7,755 (3.9%) 2,924 (6.5%) 4,831 (3.1%) Age (years), mean ± SD 61.7 ± 17.6 69.4 ± 16.5 59.4 ± 17.3 Sex, % Female 51.1% 47.9% 52.1% Male 48.9% 52.1% 47.9% Insurance, % General 50.3% 47.3% 51.2% Semi-private/private 49.7% 52.7% 48.8% Admission type, % Elective 79.3% 42.7% 89.9% Emergency 19.8% 55.6% 9.5% Other 0.9% 1.7% 0.6% ALOS (days), mean ± SD 4.3 ± 4.9 4.5 ± 4.5 4.2 ± 5.0 Most frequent working major diagnostic category (%) Diseases and disorders of the musculoskeletal system and connective tissue (28.4%) Diseases and disorders of the circulatory system (35.6%) Diseases and disorders of the musculoskeletal system and connective tissue (33.3%) ALOS, average length of stay; HWR, hospital-wide readmissions; MED, medical cohort; SURG, surgical cohort. https://doi.org/10.1371/journal.pone.0331263.t002 PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 9 / 17 [SE]: making 18.4% ± 0.15, 20.2% ± 0.14, and 17.2% ± 0.07 contribution to the model’s decision for the hospital-wide, medical, and surgical cohorts, respectively), followed by nursing assessments (11.2% ± 0.04, 12.0% ± 0.11, and 11.8% ± 0.04) and procedural information (10.8% ± 0.09, 7.7% ± 0.03, and 12.1% ± 0.03). In contrast, working DRGs (1.4% ± 0.02, 1.3% ± 0.02, 1.4% ± 0.01), patient demographics (2.5% ± 0.02, 2.6% ± 0.05, 2.5% ± 0.02), and admission information (2.6% ± 0.04, 2.6% ± 0.03, 2.3% ± 0.02) contributed least to the model’s decision. Direct comparison of the medical and surgical patient samples showed that prior healthcare use and medications were more important in the medical than the surgical cohort (9.4% ± 0.37 vs. 6.6% ± 0.06 and 9.2% ± 0.04 vs. 8.7% ± 0.02, respectively). Conversely, procedural information and healthcare provider information were more relevant in the surgical compared with the medical cohort (12.1% ± 0.03 vs. 7.7% ± 0.03 and 8.9% ± 0.04 vs. 8.3% ± 0.13, respectively). Looking more closely at the predictor subgroups in Table 4, the most important variables for prediction were all diagnoses, all procedures, and all medications (i.e., Medi all) used as dummy variables from ICD-10 GM (8.0% ± 0.12, 8.6% ± 0.07, and 7.6% ± 0.08 for the hospital-wide, medical, and surgical cohorts, respectively), CHOP (6.0% ± 0.01, 4.9% ± 0.03, and 6.6% ± 0.03), and ATC codes (5.3% ± 0.02, 5.4% ± 0.02, and 5.3% ± 0.03). These were followed Fig 1. Average receiver operating characteristic (ROC) curves for the three patient populations. https://doi.org/10.1371/journal.pone.0331263.g001 PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 16 / 17 Writing – original draft: Michael M. Havranek. Writing – review & editing: Michael M. Havranek, Aljoscha B. Hwang, Ilona Funke, Dominique Kuhlen, Daniel Liedtke, Stefan Boes. References 1. Desai NR, Ross JS, Kwon JY, Herrin J, Dharmarajan K, Bernheim SM, et al. Association Between Hospital Penalty Status Under the Hospital Readmission Reduction Program and Readmission Rates for Target and Nontarget Conditions. JAMA. 2016;316(24):2647–56. https://doi. org/10.1001/jama.2016.18533 PMID: 28027367 2. Ibrahim AM, Nathan H, Thumma JR, Dimick JB. Impact of the Hospital Readmission Reduction Program on Surgical Readmissions Among Medicare Beneficiaries. Ann Surg. 2017;266(4):617–24. https://doi.org/10.1097/SLA.0000000000002368 PMID: 28657948 3. Havranek MM. Nationaler Vergleichsbericht «Ungeplante Rehospitalisationen». ANQ. 2023. Available from: https://www.anq.ch/de/fachbereiche/ akutsomatik/messinformation-akutsomatik/ungeplante-rehospitalisationen/ 4. Kansagara D, Englander H, Salanitro A, Kagen D, Theobald C, Freeman M, et al. Risk prediction models for hospital readmission: a systematic review. JAMA. 2011;306(15):1688–98. https://doi.org/10.1001/jama.2011.1515 PMID: 22009101 5. Zhou H, Della PR, Roberts P, Goh L, Dhaliwal SS. Utility of models to predict 28-day or 30-day unplanned hospital readmissions: an updated systematic review. BMJ Open. 2016;6(6):e011060. https://doi.org/10.1136/bmjopen-2016-011060 PMID: 27354072 6. Readmission Measures Methodology. CMS. 2024. Available from: https://qualitynet.cms.gov/inpatient/measures/readmission/methodology 7. KI-Projekt zu Vorhersagen beim Entlassmanagement zeigt durchwachsene Ergebnisse. Deutsches Ärzteblatt. Available from: https://www.aerzteblatt.de/nachrichten/155852/KI-Projekt-zu-Vorhersagen-beim-Entlassmanagement-zeigt-durchwachsene-Ergebnisse. 2024. 8. Nijman S, Leeuwenberg AM, Beekers I, Verkouter I, Jacobs J, Bots ML, et al. Missing data is poorly handled and reported in prediction model studies using machine learning: a literature review. J Clin Epidemiol. 2022;142:218–29. https://doi.org/10.1016/j.jclinepi.2021.11.023 PMID: 34798287 9. Edmondson ME, Reimer AP. Challenges frequently encountered in the secondary use of electronic medical record data for research. CIN: Computers, Informatics, Nursing. 2020;38(7). 10. Mahmoudi E, Kamdar N, Kim N, Gonzales G, Singh K, Waljee AK. Use of electronic medical records in development and validation of risk prediction models of hospital readmission: systematic review. BMJ. 2020;369:m958. https://doi.org/10.1136/bmj.m958 PMID: 32269037 11. Koch JJ, Beeler PE, Marak MC, Hug B, Havranek MM. An overview of reviews and synthesis across 440 studies examines the importance of hospital readmission predictors across various patient populations. J Clin Epidemiol. 2024;167:111245. https://doi.org/10.1016/j.jclinepi.2023.111245 PMID: 38161047 12. Medical Statistic of Hospitals. SFSO. 2020. Available from: https://www.bfs.admin.ch/bfs/de/home/statistiken/gesundheit/erhebungen/ms.html 13. Instruments of Medical Coding. SFSO. 2023. Available from: https://www.bfs.admin.ch/bfs/de/home/statistiken/gesundheit/nomenklaturen/medkk/ instrumente-medizinische-kodierung.html 14. SwissDRG. Swiss DRG classification 2019. SwissDRG. 2019. Available from: https://www.swissdrg.org/de/akutsomatik/archiv-swissdrg-system/ swissdrg-system-802019 15. Horwitz LI, Grady JN, Cohen DB, Lin Z, Volpe M, Ngo CK, et al. Development and Validation of an Algorithm to Identify Planned Readmissions From Claims Data. J Hosp Med. 2015;10(10):670–7. https://doi.org/10.1002/jhm.2416 PMID: 26149225 16. Horwitz LI, Partovian C, Lin Z, Grady JN, Herrin J, Conover M, et al. Development and use of an administrative claims measure for profiling hospital-wide performance on 30-day unplanned readmission. Ann Intern Med. 2014;161(10 Suppl):S66–75. https://doi.org/10.7326/M13-3000 PMID: 25402406 17. Ellimoottil C, Khouri RK, Dhir A, Hou H, Miller DC, Dupree JM. An Opportunity to Improve Medicare’s Planned Readmissions Measure. J Hosp Med. 2017;12(10):840–2. https://doi.org/10.12788/jhm.2833 PMID: 28991951 18. Havranek MM, Dahlem Y, Bilger S, Rüter F, Ehbrecht D, Oliveira L, et al. Validity of different algorithmic methods to identify hospital readmissions from routinely coded medical data. J Hosp Med. 2024;19(12):1147–54. https://doi.org/10.1002/jhm.13468 PMID: 39051630 19. Dharmarajan K, Hsieh AF, Kulkarni VT, Lin Z, Ross JS, Horwitz LI, et al. Trajectories of risk after hospitalization for heart failure, acute myocardial infarction, or pneumonia: retrospective cohort study. BMJ. 2015;350:h411. https://doi.org/10.1136/bmj.h411 PMID: 25656852 20. Pope GC, Kautter J, Ellis RP, Ash AS, Ayanian JZ, Lezzoni LI, et al. Risk adjustment of Medicare capitation payments using the CMS-HCC model. Health Care Financ Rev. 2004;25(4):119–41. PMID: 15493448 21. Armitage JN, van der Meulen JH, Royal College of Surgeons Co-morbidity Consensus Group. Identifying co-morbidity in surgical patients using administrative data with the Royal College of Surgeons Charlson Score. Br J Surg. 2010;97(5):772–81. https://doi.org/10.1002/bjs.6930 PMID: 20306528 22. Hirslanden. Präoperative Abklärung durch den Hausarzt: Hirslanden; 2018. Available from: https://www.hirslanden.ch/content/dam/salem-spital/ downloads/Pr%C3%A4operative%20Abkl%C3%A4rung.pdf 23. LEPAG LEP. Lep: Lep ag. Available from: https://www.lep.ch/de/lep-produkte. 2024. PLOS One | https://doi.org/10.1371/journal.pone.0331263 September 4, 2025 17 / 17 24. ASA. Statement on ASA Physical Status Classification System: American Society of Anesthesiologists; 2024. Available from: https://www.asahq. org/standards-and-practice-parameters/statement-on-asa-physical-status-classification-system 25. Reis Miranda D, Moreno R, Iapichino G. Nine equivalents of nursing manpower use score (NEMS). Intensive Care Med. 1997;23(7):760–5. https:// doi.org/10.1007/s001340050406 PMID: 9290990 26. Le Gall JR, Lemeshow S, Saulnier F. A new Simplified Acute Physiology Score (SAPS II) based on a European/North American multicenter study. JAMA. 1993;270(24):2957–63. https://doi.org/10.1001/jama.270.24.2957 PMID: 8254858 27. epaCC GmbH. epaCC. Available from: https://www.epa-cc.de/bewertungssysteme/epaac/. 2021. 28. Universitätsspital B. Delirium Observation Screening Scale (DOS). Universitätsspital Basel. 2007. Available from: https://www.unispital-basel.ch/ dam/jcr:63bb6ad1-64aa-4d59-b382-7759af67c7ae/medizinische-direktion_praxisent 29. Teasdale G, Jennett B. Assessment of coma and impaired consciousness. A practical scale. Lancet. 1974;2(7872):81–4. https://doi.org/10.1016/ s0140-6736(74)91639-0 PMID: 4136544 30. McCaffery M. Pain: clinical manual for nursing practice. Mosby. 1989. 31. Kondrup J, Rasmussen HH, Hamberg O, Stanga Z, Ad Hoc ESPEN Working Group. Nutritional risk screening (NRS 2002): a new method based on an analysis of controlled clinical trials. Clin Nutr. 2003;22(3):321–36. https://doi.org/10.1016/s0261-5614(02)00214-5 PMID: 12765673 32. SFSO. Krankenhaustypologie. 2006. Available from: https://dam-api.bfs.admin.ch/hub/api/dam/assets/23546402/master 33. Rossum GV, Drake FL. Python 3 Reference Manual. CreateSpace. 2009. 34. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O. Scikit-learn: Machine Learning in Python. J Mach Learn Res. 2011;12:2825–30. 35. Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, California, USA, 2016. 785–94. 36. Youden WJ. Index for rating diagnostic tests. Cancer. 1950;3(1):32–5. https://doi.org/10.1002/1097-0142(1950)3:1<32::aid-cncr2820030106>3.0.co;2-3 PMID: 15405679 37. Talwar A, Lopez-Olivo MA, Huang Y, Ying L, Aparasu RR. Performance of advanced machine learning algorithms overlogistic regression in predicting hospital readmissions: A meta-analysis. Explor Res Clin Soc Pharm. 2023;11:100317. https://doi.org/10.1016/j.rcsop.2023.100317 PMID: 37662697 38. Huang Y, Talwar A, Chatterjee S, Aparasu RR. Application of machine learning in predicting hospital readmissions: a scoping review of the literature. BMC Med Res Methodol. 2021;21(1):96. https://doi.org/10.1186/s12874-021-01284-z PMID: 33952192 39. Dafrallah S, Akhloufi MA. Factors Associated with Unplanned Hospital Readmission after Discharge: A Descriptive and Predictive Study Using Electronic Health Record Data. BioMedInformatics. 2024;4(1):219–35. https://doi.org/10.3390/biomedinformatics4010014 40. Li Q, Yao X, Échevin D. How Good Is Machine Learning in Predicting All-Cause 30-Day Hospital Readmission? Evidence From Administrative Data. Value Health. 2020;23(10):1307–15. https://doi.org/10.1016/j.jval.2020.06.009 PMID: 33032774 41. Berchtold P, Sims EA, Horton ES, Berger M. Obesity and hypertension: epidemiology, mechanisms, treatment. Biomed Pharmacother. 1983;37(6):251–8. PMID: 6367845 42. Bonanni S, Chang KC, Scuderi GR. Should Body Mass Index Be Considered a Hard Stop for Total Joint Replacement?: An Ethical Dilemma. Orthop Clin North Am. 2025;56(1):13–20. https://doi.org/10.1016/j.ocl.2024.05.004 PMID: 39581641 43. Rasmussen LF, Grode L, Barat I, Gregersen M. Prevalence of factors contributing to unplanned hospital readmission of older medical patients when assessed by patients, their significant others and healthcare professionals: a cross-sectional survey. Eur Geriatr Med. 2023;14(4):823–35. https://doi.org/10.1007/s41999-023-00799-6 PMID: 37222865 44. Tsai TC, Joynt KE, Orav EJ, Gawande AA, Jha AK. Variation in surgical-readmission rates and quality of hospital care. N Engl J Med. 2013;369(12):1134–42. https://doi.org/10.1056/NEJMsa1303118 PMID: 24047062 45. Schnipper JL, Kirwin JL, Cotugno MC, Wahlstrom SA, Brown BA, Tarvin E, et al. Role of pharmacist counseling in preventing adverse drug events after hospitalization. Arch Intern Med. 2006;166(5):565–71. https://doi.org/10.1001/archinte.166.5.565 PMID: 16534045