Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [250] AUGMENTED INTELLIGENCE IN HEALTHCARE: BRIDGING MACHINE LEARNING, HUMAN EXPERTISE, AND BUSINESS PROCESS AUTOMATION Md. Kamruzzaman MBA in Data Analytics, University of New Haven, CT, USA Email:
[email protected] ORCID: 0009-0005-0671-6397 Sujoy Saha MSc. in Business Analytics, University of New Haven, CT, USA Email:
[email protected] ORCID: 0009-0000-9358-7813 Md Kamrul Islam MBA in Banking and Insurance, University of Dhaka, Dhaka, Bangladesh Primary Email:
[email protected] ORCID:0009-0001-8906-630X Corresponding Author: Md. Kamruzzaman,
[email protected] ABSTRACT A recent study has shown advanced intelligence dramatically transforms modern healthcare by incorporating machine learning technologies with human clinical knowledge, and software for business processes automation. This study focuses on how advanced intelligence strategies can help improve diagnostic accuracy, simplify operational process and enhance evidence-based decision-making over diverse clinical contexts. It addresses a recent advancement of predictive analytics and natural language processing, and the real-time decision support systems through a review of recent research work on how collaborative interaction between algorithmic models and expert supervision results in medical errors, timely clinical intervention, and improves outcomes for patients. But the study further explores the impact of business process automation in increasing administrative burdens, lessening operating costs, and making it easier for clinicians to allocate more time to complex, high-value patient care activities. It also suggests that augmenting intelligence does not objectify human professionals but more often builds a stronger capacity by creating adaptive, efficient and ethically grounded health care systems. Overall, this study puts the augmented intelligence paradigm in place as a transformative paradigm that fosters innovation, operating excellence, and sustainable progress in next-generation digital health ecosystems. Keywords: Augmented Intelligence; Machine Learning; Human-in-the-Loop Systems; Clinical Decision Support; Business Process Automation; Predictive Analytics 1. INTRODUCTION Healthcare systems are facing increased demand for patients and complex diagnostics but most of the workflow still relies on manual judgment and fragmented processes. In the absence of clinical context or ethical awareness, machine learning can detect risk and analyze large datasets at scale, but algorithms alone cannot be applied to the scale of the large amounts of data. This has led to a shift away from automation only practices to augmented intelligence, which ML improves, not replaces, human expertise. In this paradigm, clinicians respond to contextual reasoning, ML carries out pattern recognition and real-time analysis. The evidence has shown that unsupervised AI has an increased bias increase in unsupervised models, and human-in-the-loop models improve safety in practice (Parikh et al., 2019; Topol, 2019). Another dimension is workflow automation. Tools such as robotic process automation can route triage and manage routine tasks but can only be used to lead a patient on implementing results via clinically validated
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [251] outputs. In practice, automation reduces delays and maintains professional control (Kellermann & Jones, 2013; Verghese et al., 2018). Most research focus on algorithms or workflows separately. This study examines the full ecosystem of machine learning, human supervision and BPA through a real-life clinical dataset. We show how clinician review improves accuracy and mitigates false positives; and how automation transforms the validated information into faster, more efficient care. Augmented intelligence is a practical model of digital medicine, based on a partnership of computation and judgments. 2. RELATED WORK The science of computational systems in healthcare can be structured in four streams: 1) integration between automation, AI, and hybrid (augmented) approaches 2) formal models of human-AI collaboration and human-inthe-loop systems and human-in-the-loop (HITL) systems (3) domain-specific evidence of machine learning in diagnostic support (radiology, pathology and intensive care) (4) business process automation and BPA. Here, there is a consistent observation of many studies reporting excellent algorithmic performance in isolation, but not as many in the case of an integrated landscape where ML models, clinician supervision, and automated workflow execution interact in practice. AI, Automation, and Hybrid (Augmented) Approaches Early and recent reviews look into a fundamental concept of conceptual distinction. This is often an analogy to “automatation,” a rule-based system that performs certain tasks without human intervention; but modern AI (particularly deep learning) construct probabilistic, data driven responses, which can enable advanced diagnostic or predictive capabilities, but are usually lacking contextual grounding or interpretation. Augmented intelligence or hybrid approaches implicitly view computational systems as partners that enhance clinician’s capacities while preserving human judgment and accountability (Topol; High-performance medicine). This reframing is important because performance metrics must not only be predictive accuracy, but must be interpretable, humanover-human oversight and operational fit. Dimension Traditional AI Systems Automation (BPA / RPA) Augmented Intelligence (Hybrid) Primary Goal Predict or classify outcomes Execute repetitive or rules-based tasks Enhance human decision-making with automation support Human Role Limited or none Task oversight Central supervisory layer; final decision authority Typical Tools ML models, NLP, CV, deep learning RPA bots, workflow rules, orchestration engines ML + clinician review + automated routing Clinical Use Cases Diagnosis prediction, imaging classification Scheduling, referrals, claims, triage queues ICU risk alerts, diagnostic decision support, optimized triage Strengths Scalability, pattern detection, speed Operational efficiency, consistency, nonclinical load reduction Safety, interpretability, fewer errors, actionable outputs Failure Mode Incorrect predictions → patient harm Stalled workflows → inefficiency Rare errors, mitigated by clinician veto Table 1. Comparison of AI, Automation, and Augmented Intelligence in Healthcare Human–AI Collaboration Models and HITL Architectures This is the evolving work of defining how humans and AI should interact. Human-in-the-loop architectures place clinicians in multiple phases of life cycle data curation, model selection, validation, and post-deployment auditing to improve automation bias, error correction, and for continuous learning. Recent reviews elaborate on assessment frameworks and domain-specific techniques for feedback and decisions overrides, and retraining triggers. These studies also explore methods of toolmaking such as standardized measures of human–AI
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [252] evaluation, experimental design testing for behavior change, and governance that utilizes accountability and responsibility. ML in Diagnostic Support: Radiology, Pathology, and ICU Applications The best evidence for ML in healthcare is from images-intensive domains and prediction tasks in the ICU. Convolutional neural networks and radiomics pipelines have achieved high sensitivity for tasks in radiology such as detecting pulmonary nodules or screening for hemorrhage screening, but external validation and clinical workflow integration remain key barriers to translation to practice. Hosny et al. summarize developments in technical terms, but caution that models generalizability and prospects for clinical assessment are urgent needs. But, in digital pathology the “clinical-grade” slide screening can be accomplished through weakly supervised and multiple-instance learning, and Campanella et al. demonstrated that large scale, weakly supervised models provide pathologist-level performance on a large scale, thereby decreasing the time of re-indexing if triage used as triage tools. But, the deployment of such systems requires human verification and effective workflow design to avoid error and accidental failure. The intensive care setting has been the area that enables predictive analysis, particularly when using time series and derived features, particularly for early detection of deterioration or sepsis in the intensive care setting. Systemsatic reviews offer promise but also illustrate heterogeneity in methodology, low reproducibility across sites, and clinical utility problems if models are evaluated in only retrospective terms. Moor et al. and Desautels et al. emphasize the need for prospective, clinician-involved tests to determine safety and impact. Business Process Automation and Workflow Optimization Robotic Process Automation and intelligent orchestration are increasingly being deployed in hospitals to take care of scheduling, claims processing, EHR reconciliation, and automated notification processes. Repay can cut time and operational error, and when combined with AI (aka intelligent automation) it can route the output of ML into operational tasks (e.g. triage alerts specific to teams). But, domain studies also advocate that automation must be integrated into the medical management process to avoid routinization of low quality predictions and retain clinician control. Case studies and reviews also point to organizational factors, governance, interoperaability, workforce impact, as key to success or failure. 3. MATERIALS AND DATASET This analysis utilizes the Medical Information Mart for Intensive Care IV (MIMIC-IV) database in the context of evaluating the performance of an augmented intelligence framework in a practice setting. MIMIC-IV is a publicly available, de-identified clinical data set developed by the Beth Israel Deaconess Medical Center in Boston, Massachusetts, USA and based on MIT Laboratory for Computational Physiology together with MIT Laboratory for Computational Physiology. The database is a complete record of critical care admissions in a large U.S. hospital system, providing information on fluid and physiological details of over 200,000 hospitalized
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [253] patients between 2008 and 2019. Its scale, clinical realism, and diversity have made MIMIC the most used model in AI-enabled healthcare research, especially in risk prediction, ICU triage, sepsis detection, and deterioration forecasting (Johnson et al., 2023). 3.1 Dataset Composition and Structure MIMIC-IV is composed of three components: hospital-level data, ICU level data, and log-based clinical events all modeled to reflect the actual care experience of critically ill patients in the United States. ⚫ Hospital Module: demographic characteristics, emergency admissions, discharge diagnoses, billing codes (ICD-9 / ICD-10), and encountered metadata. ⚫ ICU Module: time-varying vital signs, heart rate, arterial pressure, respiratory rate, laboratory tests - lactate, creatinine, hematocrit, ventilators, and medication administration. ⚫ Event logs: radiology reports, medication orders, lab request timestamps, procedures records, and clinical notes. The dataset is structured and unstructured, allowing for a variety of tabular ML models, deep learning architectures, and natural language processing (NLP) for contextual risk assessment. 3.2 Target Population and Inclusion Criteria In order to assess enhanced intelligence in high risk decisions, our research is based on adults, requiring sufficient physical symptoms to adequately perform the supervision tasks of machine learning with an initial focus on adult ICU admissions. By excluding patients who were not required for vitals in the first 24 hours of admission, models were trained and biased to avoid imputation bias. This sampling approach corresponds to practice in intensive care units in U.S. intensive care hospitals, where early physiological instability is strongly predictive of mortality, organ failure, and emergency intervention requirements. Inclusion parameters: ⚫ Adults (>18 years) ⚫ First ICU visit in single hospitalization. ⚫ Core vitals (HR, SBP/DBP, SpO2, RR, temperature) ⚫ LOS > 4 hours to exclude ambulatory transfers or post-operative holding. Exclusion parameters: A. Neonatal ICU and pediatric cohorts. B. severe missingness (>40% of clinical features) C. Patients whose procedures were not considered due to procedural observation were admitted to undergo only procedural observation. This model corresponds to common predictive analytics standards used in ICU risk research and mitigates the confounding that occurs in datasets, a common problem encountered in retrospective EHR modeling (Shickel et al., 2018). 3.3 Clinical Variables Used These characteristics are clinically recognizable and frequently evaluated in early stratification for early patients in U.S. ICUs with the assistance of interpretation of their clinical significance. ⚫ Demographics: age, sex, ethnicity, admission type (ED, OR, inpatient). ⚫ Vital signs: heart rate, respiratory rate, systolic/diastolic pressure, oxygen saturation. ⚫ Laboratory markers: white blood cell count, lactate, hematocrit, hemoglobin, creatinine. ⚫ Clinical outcomes: ICU mortality, discharge disposition, length of stay. These variables enable static and time-series modeling. The early-window statistics measures mean, trend, variance, first and last value in response to established ICU prediction methods were included as temporal information. 3.4 Ground Truth Medical Outcomes This baseline endpoint is in-hospital or ICU mortality, a clinically relevant measure commonly used in AI-based risk stratification research. Mortality remains constant in MIMIC-IV using electronic medical records and appropriate billing information resulting in simplified classification. Other results, such as 48-hour deterioration or prolonged ICU stay of > 72 hours, may be available to support multi-objective modeling.
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [254] Because mortality prediction is a choice. ⚫ It is the true risk bias decisions that are based in U.S. care systems. ⚫ It has established literature-based benchmarks that are in order to be compared. ⚫ It is intuitively accessible to clinicians and hospital administrators. 3.5 Ethical Access and Regulatory Compliance MIMIC-IV is fully de-identified according to HIPAA standards and is acceptable under a Data Use Agreement (DUA) that requires ethical handling of, secure storage and training on human subject research. Because all name, address and admission timestamps from the dataset are deleted and randomly selected, a dataset that does not encompass human-subject research is not human-subject research as defined by the United States Office for Human Research Protections. In digital health research, this ethical model is commonly accepted and has been supported for algorithmic evaluation tests that do not reveal sensitive personal information (Johnson et al., 2023). 3.6 Justification for USA Context MIMIC-IV, from a tertiary U.S. medical school serving an urban population with heterogeneous populations, is well supported by the client’s need to look to the U.S. This dataset contains the specific characteristics of American health systems. ⚫ ICD-9/ICD-10 diagnostic billing requirements ⚫ Emergency-based ICU transfers common in U.S. hospitals ⚫ High incidence of multi-comorbid chronic patients These contextual factors in these activities greatly influence workflow pressures, care coordination, automation possibilities and clinician-AI interaction. The results from this dataset reflect actual problems inherent to U.S. critical care, not simulated or abstract healthcare conditions. 4. METHODOLOGY The research methodology for this study is constructed around three core components, 1) developing predictive models using real clinical data; 2) human-in-the-loop oversight through which algorithmic outputs are evaluated and adjusted; and 2) intelligent workflow automation that implements validated insights about the hospital processes. This approach is intended to capture the fundamental principles of augmented intelligence: machine capability, clinical judgment, and execution at scale rather than evaluate each component individually. 4.1 Data Acquisition and Preprocessing As discussed in Section 3, patient records were retrieved from the MIMIC-IV database based on adult ICU encounters. Preprocessing methods adapted the best practices from previous EHR-based ML studies (Shickel et al., 2018; Purushotham et al., 2016). Initial observations of early patient instability in the first 24 hours after admission to ICU were combined with data from raw physiological and laboratory measurements into observation windows for the first 24 hours before admission, since early patient instability is linked strongly with mortality risk. For missing data, the data was processed by a hierarchical methodology: forward imputation for time-series gaps in the same admission; median imputation for population-level missingness; and feature removal when not more than 40% of missing data was available. In addition to z-score normalization, continuous features such as heart rate, lactate, systolic blood pressure were standardized by z-score normalization. Among categorical variables - admission type and ethnicity - there was one-hot encoded and demographic variables were not implicitly encoded ordinal bias. By sparing data leakage, scaling parameters were only learned from the training partition. 4.2 Feature Engineering and Temporal Representation This is a temporal phenomenon of clinical degradation. Rather than extracting linear time-scale data into fixed points, the study gathered statistical descriptors of physiological change: mean, slope, variance, minimum, maximum, and last measured value per variable. This handcrafted descriptors perform better than static averages in early ICU risk prediction as published research has demonstrated in Ghassemi et al. and Desautels et al. (2015). For those with high-frequency measurement such as HR, the time differences were addressed with fixed interval resampling for high frequency measures, such as HR every 5 minutes. This minimizes artificial interpolation in
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [255] the form of each time window showing a time-window character, and allows each of the time-window features to be representative of real world measurements. While clinical notes did not were modeled directly but are used as a human feedback channel as the evidence for human assessment, augmented intelligence design principles include the treatment of the natural language through clinicians rather than self-taught models. 4.3 Model Development Using interpretable and high-capacity learning models to avoid overloading a single algorithm type, we introduced models of interpretable and high capacity learning. ⚫ Logistic Regression (baseline model) Selected for interpretability and clinical transparency. ⚫ Random Forests Suitable for nonlinear interactions between vitals and labs, historically effective with ICU tabular data (Desautels et al., 2016). ⚫ Gradient Boosting Models (XGBoost) Widely used in EHR prediction tasks due to robustness to missing patterns and sparse distributions. nested cross-validation resulted in optimistic hyperparameters. The training partition was split into 80/20 with five-fold validation in the training segment and no tuning of the test set. All experiments were done in Python using standard machine learning libraries on a workstation with controlled access ensuring compliance with MIMIC-IV’s DUA requirements. Model Key Hyperparameters Advantages (Technical + Clinical) Clinical Interpretability Level Logistic Regression (Baseline) Regularization: L2; C=1.0; Solver: liblinear Highly interpretable; stable with linear relationships; provides coefficient-level clinical relevance; often preferred by clinicians for transparency High — coefficients map directly to physiological variables Random Forest (100– 300 trees) Trees: 200; Max depth: 8; Min samples split: 4; Bootstrapping: Yes Captures nonlinear interactions; robust to missingness; stable for tabular ICU data; consistent performance in small-to-medium datasets Moderate — feature importance available, but decision paths opaque Gradient Boosting (XGBoost) Estimators: 300; Learning rate: 0.05; Max depth: 6; Subsample: 0.8; Colsample_bytree: 0.7 Handles sparse and imbalanced patterns; superior accuracy for ICU risk prediction; minimizes overfitting via shrinkage and column sampling Low–Moderate — requires post-hoc interpretability (SHAP) to explain Neural Network (Feedforward MLP) Layers: 3–4 dense; Hidden units: 64–256; Dropout: 0.2–0.5; Optimizer: Adam Learns complex multivariate interactions; scalable to multimodal features; high performance in ICU mortality tasks Low — black-box behavior without explainable AI overlays Table 2. Model Architectures, Hyperparameters, Advantages, and Clinical Interpretability
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [256] 4.4 Clinician Oversight: Human-in-the-Loop Mechanism Machine predictions were not accepted as final outputs. Instead, they consulted clinicians. I. High-risk prediction (top decile probability), II. inconsistent classifications between baseline and ensemble models III. false positives with interpreted features IV. clinical settings in which automated systems are infrequently used. The human review protocol was based on two principles: ⚫ Historical analysis of statistical predictions Clinicians studied physiological progression, medication history and admission conditions. ⚫ Corrective feedback Model errors, including false positives, were punctuated and reinserted into additional training messages. In the example of radiology triage systems and computational pathology, expertise validation reduces automation bias and optimizes downstream model performance (Campanella et al., 2019; Sendak et al., 2020). 4.5 Intelligent Automation Workflow Human-approved model outputs were routed into a prototype Business Process Automation (BPA) layer as it was designed with U.S. clinical operations. I. High-risk cases → routed to ICU triage alerts II. Moderate-risk cases → flagged for enhanced monitoring III. Low-risk cases → regular scheduling queue It did not create medical orders or prescribe treatment, but expedited administrative work and remained clinically authoritative. This is consistent with current U.S. health policy practice, where automation complements human monitoring, not competes with it (Kellermann & Jones, 2013). 4.6 Model Evaluation Metrics Both predictive accuracy and operational benefit were assessed in both performance. As reported in clinical safety studies, error asymmetry was emphasised. ⚫ Area Under ROC Curve (AUC) — discrimination capabilit ⚫ Precision & Recall — risk of missed deterioration vs unnecessary alarms ⚫ F1-Score — balanced harm analysis ⚫ Brier Score — calibration of probabilistic risk ⚫ False Positive Burden — directly impacts clinician workload During BPA simulation, operational metrics were recorded. ⚫ Notification response latency ⚫ Queue congestion reduction ⚫ Alert fatigue reduction ⚫ Task routing success rate These measures reflect actual hospital trade-offs, not abstract model performance. 5. EXPERIMENTS AND RESULTS The experiment was aimed at testing the performance of an augmented intelligence framework against the sole use of algorithmic or automation-only approaches. We evaluated model performance, clinical oversight, and operational applications of routing validations vetted outputs via an automated layer. Research using a stratified cohort of adult ICU admissions from the MIMIC-IV database was conducted, with mortality prediction as the primary endpoint, and workflow efficiency as the secondary endpoint. The analyses were conducted in a controlled manner, meeting specified reproducibility criteria in healthcare informatics (Johnson et al. 2023; Shickel et al. 2018). 5.1 Experimental Design To verify accuracy, a three-layer experiment structure was used: ⚫ Baseline Machine Learning (ML-only): Human-to-human prediction and analysis. As explained above, this is the common AI assessment paradigm in digital health research.
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [257] ⚫ HITL: Clinicians used model outputs to identify implausible suggestions, and adapted high-risk cases. These structures are similar to augmented clinical decision-making structures (Topol, 2019; Sendak et al., 2020) . ⚫ AI+BPA: Tested predictions were routed through an automation layer for triage priority scheduling, alert scheduling, or monitoring escalation. This methodology allows us to distinguish performance improvement from performance by algorithm alone, clinician control, or operating orchestration. 5.2 Model Performance on ICU Mortality Prediction AUC, F1-score, recall, and calibration error were safety measures applied to the held-out test cohort to compare models in the held-out test cohort. These measures reflect not only discriminative capacity but also asymmetrical damage to misclassification that is correlated with distortions in misclassification, such as false assurance or alarm fatigue. Model AUC F1-Score Recall (Sensitivity) Calibration Error (Brier) Logistic Regression 0.78 0.63 0.61 0.19 Random Forest 0.84 0.67 0.65 0.15 XGBoost 0.87 0.71 0.69 0.12 MLP Neural Network 0.86 0.70 0.66 0.13 Table 3. Model Performance Across Architectures (ML-only) In addition, these results follow evidence of the superiority of tree-based ensemble models over linear and neural models on tabular EHR data in the presence of sparsity and mixed scale variables with the addition of tree-based ensemble models. 5.3 Impact of Human-in-the-Loop Supervision Raw performance does not equal clinical safety. The most significant improvement took place when the clinician looked at predictions. Physicians regularly identified inflammatory markers and mild hypoxemia on a regular basis, where EHR models often underestimate risk (Obermeyer et al., 2019).
Volume-07 Issue 06, June-2023 ISSN: 2456-9348 Impact Factor: 6.736 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [258] Metric Before Review After Review AUC 0.87 0.89 Recall 0.69 0.75 False Positive Rate 0.18 0.11 Error Concentration (Top 10% Risk) 31% 19% Table 4. Improvement After Clinician Review (XGBoost) 1. Two of them are particularly noteworthy: 2. The sensitivity increased with proportional inflation in false alarms. 3. Error concentration diminished, and serious algorithmic errors were rarer. This is consistent with the previous HITL findings on computational pathology and radiology in which expert review corrects systematic bias and improves downstream stability (Campanella et al., 2019; Kiani et al., 2020). 5.4 Automation Layer: Workflow and Operational Benefits In other words, the machine was not used to make medical decisions but performed credible risk prioritization. For test purposes, we conducted a triage experiment where: Mortality probability > 0.80 immediate clinician attention 0.60-0.80 Monitoring escalation queue 0.60 Standard workflow Automation actually made much more performance gains in tasks allocation efficiency.