scieee AI-readable full text Open interactive document viewer

Enhancing Clinical Trial Insights with Advanced Tools for Treatment Effect Heterogeneity

Guilatco, Ruffy; Dalam, Alexis Bernard; Ramos, Paula Angelica; Rubrico, Roxanne Jean; Torres, Keith Russel; Cadatal, Mary Jane; Reyes, Christian Russel; Zhao, Yuxi; Kudela, Maria; Gamalo, Margaret

Abstract

Clinical trials evaluate efficacy and safety in defined populations, yet treatment responses often vary due to biological, clinical, and demographic factors. Understanding this heterogeneity is critical for characterizing therapeutic effects across subgroups and guiding regulatory decision-making (e.g., FDA guidance, May 2023). We present a flexible analytical framework that integrates traditional and modern machine learning (Random Forests, Elastic Net, Neural Networks) with Bayesian extensions and causal inference methods—including Virtual Twins (VT) for individualized treatment effects, Targeted Maximum Likelihood Estimation (TMLE) and Debiased Machine Learning (DBL) for robust causal estimation, and conformal prediction for principled uncertainty quantification. Applied to synthetic ulcerative colitis trial data, the framework (i) identifies influential covariates, distinguishing prognostic from predictive effects; (ii) uncovers clinically meaningful subgroups; and (iii) produces uncertainty – calibrated estimates under rigorous statistical inference. By addressing model misspecification, overfitting, and interpretability, these synergistic approaches enhance the robustness of inference and support evidence-based clinical development and regulatory assessments. Key Words: Bayesian, machine learning, Virtual Twins, causal inference, conformal prediction, efficacy, ulcerative colitis, clinical trials

Full text

Enhancing Clinical Trial Insights with Advanced Tools for Treatment Effect Heterogeneity Ruffy Guilatco1, Alexis Bernard Dalam1, Paula Angelica Ramos1 Roxanne Jean Rubrico1, Keith Russel Torres1, Mary Jane Cadatal1, Christian Russel Reyes1, Yuxi Zhao2, Maria Kudela2 Margaret Gamalo2 1Pfizer Inc, Philippines 2Pfizer Inc, USA Abstract Clinical trials evaluate efficacy and safety in defined populations, yet treatment responses often vary due to biological, clinical, and demographic factors. Understanding this heterogeneity is critical for characterizing therapeutic effects across subgroups and guiding regulatory decision-making (e.g., FDA guidance, May 2023). We present a flexible analytical framework that integrates traditional and modern machine learning (Random Forests, Elastic Net, Neural Networks) with Bayesian extensions and causal inference methods—including Virtual Twins (VT) for individualized treatment effects, Targeted Maximum Likelihood Estimation (TMLE) and Debiased Machine Learning (DBL) for robust causal estimation, and conformal prediction for principled uncertainty quantification. Applied to synthetic ulcerative colitis trial data, the framework (i) identifies influential covariates, distinguishing prognostic from predictive effects; (ii) uncovers clinically meaningful subgroups; and (iii) produces uncertainty – calibrated estimates under rigorous statistical inference. By addressing model misspecification, overfitting, and interpretability, these synergistic approaches enhance the robustness of inference and support evidence-based clinical development and regulatory assessments. Key Words: Bayesian, machine learning, Virtual Twins, causal inference, conformal prediction, efficacy, ulcerative colitis, clinical trials 1. Introduction Drug development is a complex, multiphase process critical to advance medical innovation and enhance patient care. It begins with the discovery and development of potential compounds, followed by preclinical research involving in vitro and in vivo testing. As the process unfolds, clinical research becomes central, generating safety and efficacy data that inform regulatory decisions. The FDA review process plays a pivotal role, determining a drug's intended use and approvals and further post-marketing requirements to ensures that its benefits continue to outweigh risks. At the heart of this process lies the discipline of statistics, which provides essential tools to derive insights and evidence of safe and effective use from diverse data sources—including clinical trials, biomarker profiles, real-world evidence, and published literature. Clinical trials are designed to evaluate the efficacy and safety of new interventions within defined patient populations; however, treatment responses may vary due to biological, clinical, and demographic factors such as genetics, biomarkers, age, and comorbidities. Careful assessment whether heterogeneity exists is essential to ensure that therapeutic benefits are appropriately characterized across all clinically relevant subpopulations intended for labeling. From the sponsor’s perspective, effective clinical development depends on understanding patient heterogeneity and identifying covariates that influence treatment efficacy (FDA guidance, May 2023). By identifying impactful covariates, quantifying uncertainty, assessing bias, and leveraging available data, statisticians can support evidence-based decisions and enhance the probability of program success. This manuscript proposes a flexible statistical framework that integrates traditional methods with modern approaches—such as causal inference, Bayesian modeling, machine learning, and conformal prediction—to predict treatment efficacy, to identify impactful covariates, and to quantify uncertainty. While established techniques were implemented as an example, the framework can be extended, allowing inclusion of other methods of interest. The proposed framework was applied to a case study using synthetic ulcerative colitis trial data to demonstrate its ability to address questions from multiple perspectives, highlighting why this approach is particularly effective for drug development—where understanding treatment-effect heterogeneity, identifying predictive covariates, and quantifying uncertainty are critical for evidence-based decisions. The paper is organized as follow: Section 2 provides a concise overview of methods within the proposed framework. This section highlights advantages and disadvantages of each method and its Bayesian extension. Section 3 details the real-world application through a case study, demonstrating practical utility and effectiveness. Section 4 presents the results of the case study. Finally, Section 5 concludes the paper by providing summary and discussion. 2. Methods In clinical trials, key questions often arise regarding population heterogeneity: • Are there covariates that explain treatment effects? • Are these covariates prognostic or predictive? • Do subpopulations exist with notably higher or lower treatment effects? In the following section, we provide a high-level overview of methods designed to address these questions, all of which have been implemented within our proposed framework. Figure 1 provides a schematic overview of the framework, illustrating how these components interact to support our analytic objectives. As an overview, this framework adopts a hybrid approach that leverages the traditional ML algorithms and its Bayesian extensions with the complementary strengths of VT, DML, TMLE, and Conformal Prediction to robustly estimate treatment effect heterogeneity in clinical research. • Exploratory Phase (ML, Conformal Prediction, and VT): The exploratory phase begins by applying machine learning models to identify features that may be prognostic or predictive of treatment effects. This initial step can be coupled with variable selection to enhance model stability, interpretability, and predictive accuracy, while also reducing computational burden. To further assess the reliability of the predictive models, conformal prediction is applied post-modeling to evaluate prediction accuracy and provide valid, distribution-free uncertainty quantification. These insights lay a stronger foundation for the subsequent application of VT, which estimates individual treatment effects (ITEs) and identifies covariate-defined subgroups with heterogeneous responses. While VT offers intuitive insights, it does not provide formal inferential guarantees, making it best suited for hypothesis generation. • Confirmatory Phase (DML and TMLE): Building on the exploratory findings, DML and TMLE are used to formally estimate causal effects and validate identified subgroups. These methods incorporate flexible machine learning for nuisance parameter estimation while maintaining statistical rigor through orthogonalization (DML) and targeted updating (TMLE). Both approaches offer double robustness, bias reduction, and asymptotic efficiency, ensuring reliable and interpretable causal inference even in high-dimensional settings. By integrating these methods, the framework balances exploratory flexibility with confirmatory rigor, enabling a comprehensive understanding of treatment effect heterogeneity and supporting robust, data-driven decision-making in clinical development. In the subsequent sections, we briefly describe the use of Bayesian Machine Learning including Random Forests, Bayesian Elastic Net, and Bayesian Neural Networks, along with Virtual Twins, Debiased Machine Learning, Targeted Maximum Likelihood Estimation (TMLE), and Conformal Prediction. Figure 1 Schematic of the framework 2.1 Machine leaning techniques and uncertainty quantification To identify key covariates or underlying subgroups that predicts the outcome of interest, it necessitates the construction of a well-performing predictive model. Traditional and widely used machine learning techniques, such as random forest, elastic net, and neural networks, are powerful tools for identifying complex and embedded patterns in data. Each of these approaches leverages different mathematical principles to model relationships and make predictions: random forests utilize ensemble decision trees for robustness, elastic net applies a combination of regularization techniques to manage multicollinearity and facilitate variable selection, and neural networks capture intricate nonlinear dependencies. Notably, each of these methods also has a counterpart under the Bayesian framework, where model parameters are treated probabilistically. The Bayesian versions not only preserve the strengths of their classic forms but also enhance them by enabling principled uncertainty quantification, regularization based on prior knowledge, and improved robustness to overfitting, ultimately providing more stable inference and greater reliability in prediction (Table 1). Table 1: ML Techniques (RF, EN, NN): Review of traditional versions and Bayesian versions ML Technique Methodology Advantages of Bayesian Version vs Traditional Version Traditional Version Bayesian Version Random Forest (RF) RF constructs multiple decision trees using Bootstrap Aggregating (Bagging), where each tree is trained on a randomly sampled subset of the data. For classification, the final prediction is determined by majority vote, while for regression, it is the average of all tree outputs (Breiman, 2001). Scornet and Hooker (2025) provide a comprehensive review of RF variants, including Mondrian and Median Forests, and discuss their statistical properties such as excess risk and central limit theorems. Bayesian Random Forest (BRF) replaces traditional bootstrap sampling with probabilistic techniques like binomial and hypergeometric sampling and uses Metropolis-Hastings algorithms for posterior estimation. • Split points, where weights may be determined using deviance-based calculations. • Terminal nodes, where weights are derived from prior predictive densities (Raynal et al., 2019; He et al., 2025). Bayesian approach allows for probabilistic reasoning and prior knowledge integration, making BRF particularly effective for sparse, high-dimensional data (Olaniran & Abdullah, 2019). The potential advantages include: • Quantification of predictive uncertainty, enabling more informed decision-making (McAlexander & Mentch, 2020). • Improved control over overfitting, especially in noisy or complex datasets. • Enhanced performance in high-dimensional spaces, where traditional RF may struggle (Sulik et al., 2023). Elastic Net (EN) EN is a regularized regression technique that integrates both L1 (lasso) and L2 (ridge) penalties. This hybrid approach enables Elastic Net to effectively address scenarios involving multiple correlated predictors, striking a balance between the variable selection capability of lasso and the coefficient shrinkage of ridge regression (Zou & Hastie, 2005). Bayesian Elastic Net (BEN) extends the classical EN by introducing a probabilistic framework for parameter estimation (Li & Lin, 2010). Unlike the traditional approach, which relies on sequential cross-validation to select penalty parameters that often leads to excessive shrinkage of coefficients, the Bayesian method estimates both penalty parameters simultaneously using either full or empirical Bayes procedures (Huang, 2025). The Bayesian approach translates the regularization penalties into specific prior distributions on the model parameters. Instead of merely penalizing coefficients, BEN places a hierarchical prior on β that effectively induces both L1 and L2 shrinkage. While both methods exhibit comparable prediction accuracy, BEN provides enhanced variable selection capabilities due to its hierarchical Bayesian framework, which allows for more nuanced regularization and better handling of multicollinearity (Hussein & Mohammed, 2022; Hans & Liu, 2024). This is especially beneficial in scenarios involving structured sparsity or correlated predictors, where classical EN may struggle to distinguish relevant features. Kimura (2025) and Lu et al. (2025) also demonstrated that BEN variants outperform EN in robustness and variable selection, especially under heteroscedasticity and structured sparsity. Moreover, BEN models exhibit greater interpretability and robustness, especially when extended with spike-and-slab priors or heteroscedastic error structures (Lu et al., 2025; Kimura, 2025). In this study, BEN for binary response was implemented using the EBglmnet R package. Neural Network (NN) NN is a popular machine learning model designed to mimic the function and structure of the human brain. It has layers of nodes – an input layer, one or more hidden layers, and an output layer. The input layer consists of the predictors in the model, the hidden layers are where the estimation of the weights and biases happen, and then the output layer is the computed point estimate (Rumelhart, 1986; Miikkulainen, 2021). Bayesian Neural Network (BNN) extends traditional NN by integrating principles from Bayesian statistics. While NNs aim to find a single set of optimal weights that minimize a loss function, BNNs treat these weights as random variables, inferring a probability distribution over them rather than a deterministic point estimate. This probabilistic approach enables BNNs to quantify the uncertainty associated with their predictions, a critical feature often lacking in conventional deep learning models that tend to produce overconfident outputs (Neal, 2012; Gawlikowski et al., 2023). Moreover, by leveraging Bayesian methods, BNNs are more robust when working with small datasets, in contrast to traditional neural networks that typically require relatively large amounts of data for effective training (Jospin et al., 2022; Arbel et al., 2023). Potential benefits of BNN over NN include: • Robust against overfitting, especially with limited data: BNN can capture a wider range of possible models consistent with the data, leading to more robust predictions compared to point-estimate NNs that might be overfit to noise, especially in small datasets (Arbel et al., 2023; Dabiran, 2023). • Better decision making: With predictive distribution, decision makers are better informed on how confident a model is in its output and may act with caution when uncertainty is high. This is crucial for safety-critical applications where the "blackbox" nature of traditional deep learning is a concern. Examples include personalized treatment plans in healthcare, where BNNs dynamically update predictions and provide uncertainty for cases needing monitoring (Ngartera, 2024). Despite their compelling advantages, BNNs come with their own set of significant disadvantages as below: • Computational complexity: In training, BNNs involves approximating complex posterior distributions. Exact Bayesian inference methods like Markov Chain Monte Carlo (MCMC) are often computationally infeasible for modern deep neural networks (DNN) with millions or billions of parameters (Chandra, 2023; Wiese et al., 2023). • Challenges in prior formulation: In the high-dimensional weight spaces of DNN, formulating sensible and informative priors is extremely difficult, and a poor choice of prior can significantly impact results. Seemingly sensible priors (such as independent Gaussian priors, Uniform priors, etc.) can lead to unintended artifacts in the output function distribution, and subjective priors are often absent (Fortuin, 2022). • Approximation inference introducing errors: Most practical BNNs rely on approximate inference methods (like Variational Inference) because exact inference is intractable. These approximations inherently introduce errors, meaning the estimated uncertainty might not perfectly reflect the true uncertainty, potentially leading to underor over-confident predictions (Bakhouya, 2023). While machine learning models are often highly effective at generating predictions and identification of impactful covariates, they frequently lack mechanisms for conveying the reliability or confidence of those outputs. Conformal prediction overcomes this limitation by delivering statistically sound and distribution-free techniques for quantifying uncertainty at the individual prediction level (Angelopoulos & Bates, 2021; Huang et al., 2024). Applied to machine learning outputs, conformal prediction creates prediction intervals or sets that correspond to a user-defined confidence threshold, offering greater clarity and utility for real-world decision-making. The following section details core aspects of this approach. Conformal Prediction Conformal prediction (CP) is a statistical framework that provides finite-sample coverage guarantees for machine learning models. In classification tasks, CP constructs prediction sets that contain the true label with a user-specified probability, regardless of the underlying data distribution. Recent research has focused on improving the efficiency and robustness of these prediction sets through different families of methods, including Score-based, Rank-based, and Adaptive Prediction Set (APS) approaches. Each of these methods differs in how they define the nonconformity measure and how they balance model calibration, set size efficiency, and robustness to uncertainty (Angelopoulos & Bates, 2021; Huang et al., 2024). The details of these methods are listed in Table 2. Table 2: Overview of Conformal Prediction Approaches Method Score-based Rank-based APS Methodology Score-based method relies on the probability outputs of a trained classifier, typically using softmax scores as nonconformity measures. A common formulation computes the nonconformity score as (1 − 𝑝𝑦), where 𝑝𝑦 is the predicted probability of the true class. As compared to score-based method, the rank-based method replaces the reliance on absolute probability values with the relative ordering of class predictions and defines the nonconformity score using the rank position of the true label in the model’s sorted list of predictions. Developed by integrating both score-based and rank-based concepts to harness the strengths of robustness and probabilistic interpretability (Angelopoulos et al., 2021), APS sequentially includes the top-ranked labels until the cumulative probability mass exceeds a defined threshold, ensuring that each prediction set achieves the desired marginal coverage. As a further refinement, RAPS (Regularized APS) adds a regularization term that penalizes the inclusion of very low-probability labels, effectively controlling the tail behavior and reducing average set size. Advantage Intuitive, simple to implement, and directly applicable to any probabilistic classifier; produce compact prediction sets with reliable coverage (Angelopoulos et al., 2021) Robust to miscalibration and model scaling (Huang et al., 2024); Often produce smaller and more stable prediction sets when the classifier ranks classes accurately regardless if the probability magnitudes are unreliable (Luo, 2024), which is useful in large-class or imbalanced settings, where probability estimates tend to be noisy; Recent advancements in rankcalibrated conformal prediction and class-conditional coverage extensions (RC3P) that enables complex classification tasks (Wang et al., 2024). These methods balance flexibility and efficiency by combining probability-based calibration with rank-aware selection. Empirical studies on large-scale vision tasks, such as ImageNet classification, demonstrate that RAPS achieves significantly smaller prediction sets compared to classical Score-based CP while maintaining formal coverage guarantees (Angelopoulos & Bates, 2023). Disadvantage Highly sensitive to poor calibration (often observed in DNN), likely leading to overly conservative or inefficient prediction sets (Angelopoulos & Bates, 2023); Often suffer from the "tinytail" effect, where many lowprobability classes inflate the size of the prediction set; Their real-world performance depends heavily on effective probability calibration or posthoc adjustment mechanisms (Luo & Zhou, 2024) Ignore valuable information contained in well-calibrated probability magnitudes and in most cases underperform when the model’s probability outputs are meaningful; Unstable ranks for neartied class scores, leading to nonsmooth (Hanselle, 2024). Notably, APS and RAPS still rely on the quality of probability magnitudes; in cases of poor calibration, performance may degrade. Consequently, newer variants have integrated rank-only adjustments or hybrid calibration strategies to mitigate this dependence (Huang et al., 2024). In summary, Score-based methods are straightforward and effective when probabilities are trustworthy, Rank-based methods are robust under miscalibration, and APS/RAPS methods offer a practical balance by exploiting both probability and rank information. In terms of calibration dependence, Score methods are the most sensitive, followed by APS (medium sensitivity), while Rank methods are the most robust (Angelopoulos & Bates, 2023; Huang et al., 2024). With respect to prediction set size efficiency, RAPS and hybrid Rank-APS models generally outperform both Rank and Score methods, particularly in high-dimensional or multi-class settings. Although all methods guarantee marginal coverage, recent studies emphasize the importance of extending toward conditional and class-wise coverage. This direction is critical for promoting fairness and reliability in modern AI systems (Wang et al., 2024). 2.2 Causal Inference Causal inference serves as the bedrock in clinical and epidemiological research, aiming to uncover the true effects of interventions by disentangling correlation from causation. At its core lies Rubin’s Causal Model, grounded in the recognition that for any individual, we can observe only one potential outcome—a principle often referred to as Rubin’s Law. This fundamental limitation necessitates methodological strategies to approximate the unobserved counterfactual, with randomization standing as the gold standard to ensure treatment assignment is independent of potential outcomes. Traditionally, statistical approaches have focused on estimating the Average Treatment Effect (ATE), offering population-level insights. However, modern demands for precision and personalization have shifted attention toward the Individual Treatment Effect (ITE), which seeks to capture heterogeneity in treatment response. Building on this, tools like virtual twins enable subgroup identification by pairing observed outcomes with predicted counterfactuals. Further extending the ITE framework, recent advancements such as TMLE and debiased machine learning integrate flexible machine learning models with rigorous statistical theory to produce robust, asymptotically unbiased estimates, even in complex, high-dimensional settings. Together, these developments represent a continuum from classical to contemporary causal inference, driven by the pursuit of more granular and reliable decision-making tools. Virtual Twins (VT) Results from machine learning techniques discussed in the previous section provide a stronger foundation for VT method, which estimates individual treatment effects (ITEs) and identifies covariate-defined subgroups with heterogeneous responses. The VT method provides a data-driven framework to address the key question in clinical research: Which subgroups of patients benefit more (or less) from a treatment, and how can we identify them using observed data? By estimating ITEs and identifying covariate-defined subgroups with heterogenous treatment effects (HTEs). Originally proposed by Foster, Taylor, and Ruberg (2011), VT operates in two main steps as below: a. Estimation of Individual Treatment Effects: Predictive models (typically random forests) are used to estimate the probability of a positive outcome for each individual under both treatment and control conditions. The difference between these predictions represents the estimated treatment effect for each subject. b. Subgroup Discovery via Tree-Based Models: The estimated treatment effects are used as the response variable in a regression or classification tree, with baseline covariates as predictors. This tree partitions the population into subgroups based on covariate patterns associated with higher or lower treatment effects. To ensure clinical relevance, identified subgroups are evaluated using criteria such as effect size, biological plausibility, consistency across modeling approaches, and expert review by clinicians. This process supports the identification of meaningful patient populations that inform precision medicine strategies. The main advantages of the method as discussed by Foster et al. (2011) are as follows: • Model Flexibility: VT leverages machine learning (e.g., random forests) to capture complex, nonlinear relationships without requiring explicit interaction terms. • Interpretability: The tree-based subgroup identification yields simple, rule-based definitions of subgroups, enhancing clinical interpretability. • Improved Sensitivity: Compared to logistic regression with forward selection, VT methods demonstrate better sensitivity and predictive value in simulation studies. • Bias Correction Options: The method includes strategies (e.g., bootstrap bias correction) to estimate subgroup treatment effects more reliably. • Scalability: Additionally, VT can handle moderate to high-dimensional covariate spaces effectively (Xie & Loh, 2017). Since its introduction, VT has been extended and benchmarked against other subgroup identification techniques in precision medicine research (Xie & Loh, 2017; Wang, 2015). Recent studies have reinforced VT’s relevance and adaptability. Deng et al. (2023) demonstrated that VT’s performance is highly sensitive to modeling choices in its first step, and that ensemble methods like SuperLearner significantly improve accuracy and robustness. Furthermore, VT’s intuitive structure and compatibility with modern machine learning make it a preferred choice over more rigid parametric approaches. Building on VT’s foundation, hybrid approaches such as VG (Virtual Twins + GUIDE) has been developed to enhance its variable selection capabilities (Jia et al., 2020). Additionally, VT shares conceptual goals with emerging frameworks in causal inference and machine learning, such as causal forests (Wager & Athey, 2018) and generalized random forests (Athey et al., 2019), which offer formal inference tools and extend tree-based methods to broader estimation tasks. Moreover, recursive partitioning methods for causal inference (Athey & Imbens, 2016) emphasize the importance of separating model fitting from effect estimation—a principle that VT operationalizes effectively through its two-step structure. In summary, VT stands out as a robust, interpretable, and adaptable method for subgroup identification in clinical trials. Within the proposed framework, VT complements Bayesian machine learning algorithms and causal inference techniques by uncovering clinically meaningful subgroups with differential treatment responses, contributing to more informed, data-driven decisions throughout the clinical development lifecycle. Debiased Machine Learning (DML) is a modern framework designed to address key challenges in estimating causal effects from observational data, especially in the presence of high-dimensional covariates. Traditional machine learning methods, while flexible, often suffer from bias in treatment effect estimation due to regularization and overfitting in nuisance parameter models. DML, developed by Chernozhukov et al. in 2018, directly tackles these limitations by combining machine learning with rigorous statistical principles. The core innovation of DML lies in its use of orthogonalized estimating equations and cross-fitting, which together mitigate bias introduced by imperfect nuisance parameter estimation. By leveraging Neyman orthogonality, DML ensures that small errors in estimating nuisance functions have minimal impact on the final treatment effect estimates. This results in root-n consistent and asymptotically normal estimators, even in complex, high-dimensional settings. Key advantages include: • Robustness to moderate model misspecification. • Flexibility in using modern machine learning techniques such as random forests, lasso, and neural networks. • Valid inference with reliable confidence intervals for causal parameters. In this study, the methodology was selected for its ability to combine predictive power with rigorous causal inference, enabling accurate and interpretable treatment effect estimates without strong parametric assumptions. DML was implemented using the DoubleML R package, which integrates regression and classification learners via the mlr3 framework. Recent extension (Chernozhukov et al., 2021) further enhanced the method and developed a general framework in the presence of unmeasured confounding, to quantify and bound omitted variable bias (OVB) in causal machine learning models. Targeted Maximum Likelihood Estimation (TMLE) is a robust and efficient method designed to address the limitations of traditional estimation methods, which often suffer from bias and inefficiency due to reliance on parametric assumptions or the use of flexible machine learning algorithms for nuisance parameter estimation. Developed by van der Laan and colleagues in 2006 year, TMLE integrates machine learning with a targeted updating step that enhances both bias Table 8. Average Treatment Effect (90% CI) by Method and Drug Method ATE (%) 90% CI Lower Upper Drug 1 Debiased ML 19.37 10.34 28.31 TMLE 16.20 6.14 26.25 Bayesian Random Forest 19.30 10.75 27.85 Traditional method 19.45 3.24 32.19 Drug 2 Debiased ML 13.49 9.57 17.41 TMLE 13.01 9.85 16.17 Bayesian Random Forest 18.00 12.57 23.43 Traditional method 14.32 7.75 20.79 The comparative evaluation of ATE estimation methods revealed notable differences in estimated treatment effects and associated uncertainty across methods examined for Drug 1 and Drug 2. The traditional method was a conventional approach estimating the absolute risk difference, defined as the difference in event proportions between treatment and control groups, with corresponding 95% confidence intervals. Among the alternative estimators, DML demonstrated the closest agreement with the traditional estimates for both Drug 1 (19.37 vs. 19.45) and Drug 2 (13.49 vs. 14.32). This consistency suggests that DML effectively reduces bias while preserving alignment with observed trial outcomes. With respect to interval estimation, DML also provided narrower confidence intervals (CIs) compared with the traditional method, indicating improved precision. For Drug 2, TMLE achieved the smallest CI (9.85–16.17), reflecting superior efficiency in variance reduction, although its point estimate (13.01) was slightly lower than the traditional method. In contrast, Bayesian Random Forest produced estimates comparable to DML for Drug 1 but substantially overestimated the ATE for Drug 2 (18.00 vs. 14.32), accompanied by a wider CI (12.57–23.43). This pattern may indicate sensitivity to treatment effect heterogeneity or prior specification, highlighting the need for careful model calibration when applying Bayesian approaches in similar contexts. Overall, these results emphasize the value of modern causal inference techniques in complementing traditional analyses. DML appears to offer the most balanced performance in terms of accuracy and precision, while TMLE provides the greatest efficiency in uncertainty reduction (van der Laan et al., 2006; Díaz et al., 2024). However, Bayesian methods may require additional refinement to avoid systematic overestimation (DiTraglia & Liu, 2025). In practice, these can be implemented using open-source tools like Python's econml library for DML/TMLE or R's bartCause for Bayesian approaches, facilitating integration with clinical datasets (European Journal of Epidemiology, 2024). 5. Discussion In this paper, we propose a novel framework to enhance the understanding of treatment effect heterogeneity and support evidence-based decision-making in clinical research. This hybrid framework integrates traditional machine learning algorithms and their Bayesian extensions with Conformal Prediction, VT, DML, and TMLE to robustly estimate and validate treatment effect heterogeneity. In the exploratory phase, machine learning models are used to identify prognostic and predictive covariates, with variable selection improving model stability and interpretability. Conformal prediction is then applied to assess prediction accuracy and quantify uncertainty. These insights inform the application of VT to generate hypotheses by estimating individual treatment effects and identifying subgroups with heterogeneous responses. Building on these explorations, DML and TMLE are used in the confirmatory phase to formally estimate causal effects, offering double robustness, bias reduction, and statistical efficiency in high-dimensional settings. Together, this integrated framework balances exploratory flexibility with inferential rigor, providing a comprehensive approach to causal inference in clinical research. We demonstrated the utility of the proposed framework through a case study using synthetic Ulcerative Colitis clinical trial data. In the exploratory phase, we have demonstrated the potential merit of Bayesian Machine Learning that adopts probabilistic formulation in the framework of additional ML techniques, yielding more stable estimation and robust prediction, effectively mitigating overfitting. For example, BRF demonstrated more consistent and reliable performance across training and test sets, whereas traditional RF showed greater variability and limited ability to capture the underlying signal. Following the exploratory ML modeling, conformal prediction was applied to assess prediction accuracy and model reliability across different algorithms—Scorebased, Rank-based, and APS. Among these, the Score-based method offered a balanced trade-off between accuracy and efficiency, supporting reliable uncertainty quantification. For subgroup identification using Virtual Twins, we identified consistent subgroups aligned with clinically meaningful covariates related to disease severity, progression, and biological pathways. In the confirmatory phase, causal inference was assessed by comparing ATE estimates from DML, TMLE, and BRF against the benchmark using observed population from the randomized clinical trial. Covariate adjustment using DML, TMLE, and BRF produced narrower confidence intervals than unadjusted analyses, highlighting efficiency gains from reducing residual variability and improving precision in clinical trials. DML demonstrated the least bias relative to the benchmark, though with slightly lower efficiency compared to TMLE. Both methods confirmed treatment efficacy with consistent trends and effect size magnitudes. These findings underscore the strength of the integrated framework in uncovering treatment effect heterogeneity and generating reliable, interpretable causal insights. When implementing this framework, several practical and methodological considerations should be acknowledged. In the exploratory phase, predictive modeling using machine learning must be carefully selected based on study design and sample size. Unbalanced designs and small sample sizes can increase the risk of misleading findings, and certain data-driven algorithms—such as deep neural networks—may be unsuitable in such contexts. Additionally, model tuning and design parameters require thoughtful calibration to ensure stability and interpretability. For subgroup identification using VT, it is important to note that the method does not provide formal inference and should be used strictly for hypothesis generation. In the confirmatory phase, both DML and TMLE offer robust causal inference, each with distinct strengths—DML excels in bias reduction, while TMLE provides greater efficiency. However, both methods assume that all confounders are observed, which may not hold in real-world settings. To further strengthen the framework, future work should explore methods such as recent extension in DML (Chernozhukov et al., 2021) for addressing unmeasured confounding, such as sensitivity analysis or negative control strategies, to assess the robustness of causal findings. Moreover, while this case study focused on clinical trial data, the framework could be extended to incorporate real-world data (RWD) as an additional source for validating and triangulating insights. References 1. Athey, S., & Imbens, G. W. (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27), 7353–7360. https://doi.org/10.1073/pnas.1510489113 2. Angelopoulos, A. N., & Bates, S. (2023). Conformal prediction: A gentle introduction. arXiv preprint arXiv:2306.13726. 3. Angelopoulos, A. N., Bates, S., Malik, J., & Jordan, M. I. (2021). Uncertainty sets for image classifiers using conformal prediction. International Conference on Learning Representations (ICLR). 4. Angelopoulos, A. N., & Bates, S. (2023). Practical uncertainty quantification with conformal prediction. Journal of Machine Learning Research, 24(83), 1–52. 5. Arbel, J., Pitas, K., Vladimirova, M., & Fortuin, V. (2023). A primer on Bayesian neural networks: Review and debates. arXiv:2309.16314. 6. Athey, S., Tibshirani, J., & Wager, S. (2019). Generalized random forests. The Annals of Statistics, 47(2), 1148–1178. https://doi.org/10.1214/18-AOS1709 7. Bach, P., Chernozhukov, V., Kurz, M. S., & Spindler, M. (2024). DoubleML: An objectoriented implementation of double machine learning in R. Journal of Statistical Software, 108(3), 1–56. https://doi.org/10.18637/jss.v108.i03 8. Bakhouya, M., Ramchoun, H., Hadda, M., & Tawfik, M. (2023). A review of variational inference for Bayesian neural network. In Artificial Intelligence and Industrial Applications (Lecture Notes in Networks and Systems, Vol. 772, pp. 231–243). Springer. 9. Bonnet, D., Hirtzlin, T., Majumdar, A., Dalgaty, T., Esmanhotto, E., Meli, V., Castellani, N., Martin, S., Nodin, J.-F., Bourgeois, G., Portal, J.-M., Querlioz, D., & Vianello, E. (2023). Bringing uncertainty quantification to the extreme-edge with memristor-based Bayesian neural networks. Nature Communications, 14(7530). 10. Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. 11. Chandra, R., & Simmons, J. (2023). Bayesian neural networks via MCMC: A Python-based tutorial. arXiv. 12. Chen, J., Dunson, D. B., & Carin, L. (2009). Bayesian elastic net for multi-task learning with application to gene expression classification. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence (UAI). https://www.cs.ubc.ca/~murphyk/Teaching/CS540-Spring10/projects/carin09BayesianEnet.pdf 13. Chen, M., Carlson, D., Zaas, A., Woods, C., Ginsburg, G. S., Hero III, A., Lucas, J., & Carin, L. (2009). The Bayesian Elastic Net: Classifying Multi-Task Gene-Expression Data. Duke University Technical Report. 14. Chernozhukov, V., et al. (2017). Orthogonal Machine Learning for Causal Inference. arXiv preprint arXiv:1608.00060 15. Chernozhukov, V., Newey, W. K., & Singh, R. (2022). Automatic debiased machine learning of causal and structural effects. Econometrica, 90(3), 967–1023. https://doi.org/10.3982/ECTA17414 16. Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., & Robins, J. (2018). Double/debiased machine learning for treatment and causal parameters. The Econometrics Journal, 21(1), C1–C68. https://doi.org/10.1111/ectj.12097 17. Chernozhukov, V., Cinelli, C., Newey, W., Sharma, A., & Syrgkanis, V. (2022). Long story short: Omitted variable bias in causal machine learning. Journal of Machine Learning Research, 23(1), 1–48. https://www.jmlr.org/papers/v23/21-0862.html 18. Dabiran, N., Robinson, B., Sandhu, R., Khalil, M., Poirel, D., & Sarkar, A. (2023). Sparse Bayesian neural networks for regression: Tackling overfitting and computational challenges in uncertainty quantification. arXiv:2310.15614. 19. Deng, C., Wolf, J. M., Vock, D. M., Carroll, D. M., Boatman, J. A., Hatsukami, D. K., Leng, N., & Koopmeiners, J. S. (2023). Practical guidance on modeling choices for the virtual twins method. Journal of Biopharmaceutical Statistics, 33(5), 653–676. https://doi.org/10.1080/10543406.2023.2170404 20. Dijkhuis, T. B., & Blaauw, F. J. (2022). Transferring targeted maximum likelihood estimation for causal inference to survival analysis. Entropy, 24(8), Article 1060. https://doi.org/10.3390/e24081060 21. Fontana, M., Zeni, G., & Vantini, S. (2023). Conformal prediction: A unified review of theory and new challenges. Bernoulli, 29(1), 1–23. 22. Fortuin, V. (2022). Priors in Bayesian deep learning: A review. International Statistical Review, 90(3), 563–591. 23. Foster, J. C., Taylor, J. M. G., & Ruberg, S. J. (2011). Subgroup identification from randomized clinical trial data. Statistics in Medicine, 30(24), 2867–2880. https://deepblue.lib.umich.edu/bitstream/handle/2027.42/87004/sim4322.pdf 24. Gawlikowski, J., Tassi, C. R. N., et. al. (2023). A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56, 1513–1589. 25. Gruber, S., & van der Laan, M. J. (2012). tmle: An R package for targeted maximum likelihood estimation. Journal of Statistical Software, 51(13), 1–35. https://doi.org/10.18637/jss.v051.i13 26. Guha, E., Natarajan, S., Möllenhoff, T., Khan, M. E., & Ndiaye, E. (2024). Conformal prediction via regression-as-classification. International Conference on Learning Representations (ICLR). 27. Hahn, P. R., Murray, J. S., & Carvalho, C. M. (2020). Bayesian regression tree models for causal inference: Regularization, confounding, and heterogeneous effects. Bayesian Analysis, 15(3), 965–1056. https://doi.org/10.1214/19-BA1195 28. Hans, C. M., & Liu, N. (2024). Sampling the Bayesian Elastic Net. arXiv preprint arXiv:2501.00594. https://arxiv.org/abs/2501.00594 DOI: 10.48550/arXiv.2501.00594 29. Hanselle, J. (2024). Conformal prediction without nonconformity scores. arXiv preprint arXiv:2405.10231 30. He, J., Li, Z., & Yin, L. (2025). An efficient and accurate random forest node-splitting algorithm based on dynamic Bayesian methods. Machine Learning and Knowledge Extraction, 7(3), 70. 31. Huang, J., Li, R., & Zhao, X. (2024). Conformal prediction for deep classifiers via label ranking. Proceedings of the 41st International Conference on Machine Learning (ICML) 32. Huang, Z., Rossi, S., Yuan, R., & Hannagan, T. (2025). From predictions to confidence intervals: An empirical study of conformal prediction methods for in-context learning. Proceedings of the 7th Symposium on Advances in Approximate Bayesian Inference, 289, 67–90. 33. Hussein, A. A., & Mohammed, B. K. (2022). Comparison between Bayesian and classical elastic net regression models: A simulation study. Al-Qadisiyah Journal for Administrative and Economic Sciences, 24(2), 1–15. https://iasj.rdd.edu.iq/journals/uploads/2024/12/19/6f5b7ebaed9190fc316bbc79b9abc81e.pdf 34. Jia, J., Tang, Q., & Xie, W. (2020). A novel method of subgroup identification by combining Virtual Twins with GUIDE (VG) for development of precision medicines. In Design and Analysis of Subgroups with Biopharmaceutical Applications (pp. 167–180). Springer. https://doi.org/10.1007/978-3-030-40105-4_7 35. Jospin, L. V., Laga, H., Boussaid, F., Buntine, W., & Bennamoun, M. (2022). Hands-on Bayesian neural networks – A tutorial for deep learning users. IEEE Computational Intelligence Magazine, 17(2), 29–48. 36. Kato, Y., Tax, D. M. J., & Loog, M. (2023). A review of nonconformity measures for conformal prediction in regression. Proceedings of the Twelfth Symposium on Conformal and Probabilistic Prediction with Applications, 204, 369–383. 37. Kimura, M. (2025). Heteroscedastic double Bayesian elastic net for high-dimensional regression. arXiv preprint arXiv:2502.02032. https://arxiv.org/pdf/2502.02032 38. Luo, R. (2024). Trustworthy classification through rank-based conformal prediction. arXiv preprint arXiv:2403.08765 39. Lu, X., Zhang, Y., & Wang, J. (2025, August). Robust Bayesian elastic net with spike-andslab priors for high-dimensional data. Paper presented at the Joint Statistical Meetings (JSM) 2025. https://ww3.aievolution.com/JSMAnnual2025/Events/viewEv?ev=3360 40. McAlexander, R. J., & Mentch, L. (2020). Predictive inference with random forests: A new perspective on classical analyses. Research & Politics, 7(1), 1–7. 41. Miikkulainen, R. (2021). Topology of a Neural Network. In C. Sammut & G. I. Webb (Eds.), Encyclopedia of Machine Learning (pp. 988–989). Springer. 42. Neal, R. M. (2012). Bayesian learning for neural networks (Vol. 118). Springer Science & Business Media. 43. Ngartera, L., Issaka, M. A., & Nadarajah, S. (2024). Application of Bayesian neural networks in healthcare: Three case studies. Machine Learning and Knowledge Extraction, 6(4), 2639– 2658. 44. Olaniran, O. R., & Abdullah, M. A. A. (2019). BayesRandomForest: An R implementation of Bayesian random forest for regression analysis of high-dimensional data. In M. A. A. Abdullah & M. A. M. Ali (Eds.), Proceedings of iCMS2017 (pp. 379–388). Springer. 45. Olaniran, O. R., & Abdullah, M. A. A. (2023). Bayesian weighted random forest for classification of high-dimensional genomics data. Kuwait Journal of Science, 50(4), 477–484. https://doi.org/10.1016/j.kjs.2023.06.008 46. Polley, E. C., & van der Laan, M. J. (2010). Super Learner in Prediction. U.C. Berkeley Division of Biostatistics Working Paper Series 47. Raynal, L., Marin, J.-M., Pudlo, P., Ribatet, M., Robert, C. P., & Estoup, A. (2019). ABC random forests for Bayesian parameter inference. Bioinformatics, 35(10), 1720–1728. 48. Romano, Y., Sesia, M., & Candès, E. J. (2020). Classification with Valid and Adaptive Coverage. Advances in Neural Information Processing Systems, 33 (NeurIPS 2020). 49. Rose, S., & van der Laan, M. J. (2011). Targeted maximum likelihood estimation of effect modification parameters in survival analysis. The International Journal of Biostatistics, 7(1), Article 13. https://doi.org/10.2202/1557-4679.1304 50. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. 51. Scornet, E., & Hooker, G. (2025). Theory of Random Forests: A Review. HAL Archives. Retrieved from https://hal.science/hal-05006431v1/file/main.pdf 52. Sulik, J., Banger, K., Janovicek, K., Nasielski, J., & Deen, B. (2023). Comparing random forest to Bayesian networks as nitrogen management decision support systems. Agronomy Journal, 115(6), 1–12. 53. U.S. Food and Drug Administration. Ulcerative Colitis: Clinical Trial Endpoints Guidance for Industry. Draft Guidance. Silver Spring, MD: Center for Drug Evaluation and Research; August 2016. Available at: https://www.fda.gov/files/drugs/published/Ulcerative-Colitis-- Clinical-Trial-Endpoints-Guidance-for-Industry.pdf 54. Wager, S., & Athey, S. (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113(523), 1228–1242. https://doi.org/10.1080/01621459.2017.1319839 55. Wang, Y. (2015). Subgroup identification in clinical trials: A review and comparison of methods. Contemporary Clinical Trials, 42, 113–119. https://doi.org/10.1016/j.cct.2015.03.005 56. Wang, Y. (2015). A comparison of multiple approaches to subgroup analysis in clinical trials. University of Texas Repository. https://repositories.lib.utexas.edu/items/72bf192e-34b4-4d5cbc2d-5b8dd65ddd91 57. Wang, Y., Chen, T., & Gao, W. (2024). Rank-calibrated conformal prediction with class-wise coverage guarantees (RC3P). Advances in Neural Information Processing Systems 58. Wiese, J. G., Wimmer, L., Papamarkou, T., Bischl, B., Günnemann, S., & Rügamer, D. (2023). Towards efficient MCMC sampling in Bayesian neural networks by exploiting symmetry. In Machine Learning and Knowledge Discovery in Databases: Research Track (ECML PKDD 2023) (pp. 459–474). Springer. 59. Xie, N., & Loh, W.-Y. (2017). Combining Virtual Twins with GUIDE for Precision Medicine. arXiv preprint arXiv:1708.04741. https://arxiv.org/pdf/1708.04741 60. Yao, J., Pan, W., Ghosh, S., & Doshi-Velez, F. (2019). Quality of uncertainty quantification for Bayesian neural network inference. In Proceedings of the ICML Workshop on Uncertainty and Robustness in Deep Learning. arXiv:1906.09686. 61. Zou, H., & Hastie, T. (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2), 301–320. 62. van der Laan, M. J., & Rubin, D. (2006). Targeted maximum likelihood learning. The International Journal of Biostatistics, 2(1), Article 11. https://doi.org/10.2202/1557-4679.1043 63. van der Laan, M. J., & Gruber, S. (2023). Developing a targeted learning-based statistical analysis plan. Journal of Biopharmaceutical Statistics, 33(3), 1–20. https://doi.org/10.1080/19466315.2022.2116104 64. van der Laan, M. J., & Rose, S. (2011). Targeted learning: Optimal statistical data-adaptive methods [Monograph]. Springer. (For broader TMLE theory; see also PMC articles like PMC3083138 for applications.) 65. van der Laan, M. J., & Gruber, S. (2009). Targeted Maximum Likelihood Estimation: A Gentle Introduction. U.C. Berkeley Division of Biostatistics Working Paper Series