scieee AI-readable full text Open interactive document viewer

Why Meta-analysis Cannot Establish Generalisability

Reidpath, Daniel

Abstract

This paper presents two independent critiques demonstrating that meta-analysis cannot establish generalisability of research findings. The first critique, based on the causal uncertainty principle, shows that restriction operations in primary studies contract evidential breadth in ways that pooling cannot reverse. The second critique argues that randomisation, while ensuring internal validity within trials, does not harmonise latent distributions of unobserved effect modifiers across different study contexts, making estimands potentially incommensurable. Claims to generalisability must therefore rest on judgements about population similarity rather than on properties of the meta-analytic estimator itself.

Full text

Why Meta-analysis Cannot Establish Generalisability Daniel D. Reidpath∗ December 8, 2025 Meta-analysis is commonly taken to increase the precision of an estimand and, at the same time, to enhance the generalisability of study findings [6, 4, 9]. The classical view is that effects observed across multiple studies warrant broader substantive causal claims. This note examines two independent critiques of that position on the generalisability of meta-analyses. The first critique considers the problem through the evidential-state framework of the causal uncertainty principle [7]. The second critique is specifically of meta-analyses based on randomised control trials (RCTs) and relies on the incommensurability of the latent mixture of the unobserved variables. Background A meta-analysis is a weighted average of effects ˆ θma estimated from a series of nempirical studies. ˆ θma = n ∑ i=1 wiˆ θi, where wiis the weight (inverse-variance, random-effects, etc.) and ˆ θithe estimand of the i-th study. Causal uncertainty principle Each primary study begins with a broad evidential state E, which is a representation of the information available for causal inference, and applies restrictions Rthat contract this state to some state, Ei. Restrictions may include eligibility criteria, geographic delimitation, feasibility constraints, and post-baseline exclusions. These restrictions necessarily remove heterogeneity and thereby limit the evidential breadth of the resulting data. The contraction is essential ∗Institute of Global Health and Development, Queen Margaret University, Edinburgh, Scotland, UK. email: [email protected] 1 Meta-analysis and generalisability for securing causal identification, but it also eliminates information about the larger Ethat cannot be recovered through further analytic procedures on Ei. Meta-analysis pools estimates ˆ θ1,ˆ θ2, . . . , ˆ θnderived from the evidential states E1, E2, . . . , Enof the nincluded studies. The aggregation of estimands from each study does not act on the original evidential state E. Instead, it combines outputs from already-restricted evidential states. As a result, information removed by restrictions in the primary studies remains absent from the meta-analytic result. For example, if certain subpopulations are consistently excluded across studies (e.g., women in clinical trials [3]), the meta-analysis provides no evidential basis for inferences about those groups. By extension, meta-regression may interpolate across observed covariate patterns, but it cannot supply variation that is structurally absent [10, 5]. Meta-regression can characterise how effects vary across the observed studylevel covariates, but it cannot support inferences for covariate patterns that do not occur in any primary study. When certain subpopulations or covariate combinations are absent from all evidential states {Ei}, estimates for those regions are extrapolations driven by model form rather than empirical variation. They therefore do not recover evidential breadth lost through restriction. Rather, the models extend beyond the data using assumptions external to the process. Within the evidential-state framework, this limitation reflects the general principle that operations which contract evidential breadth cannot be undone. Meta-analysis increases precision within the contracted evidential domains it receives, but it does not expand those domains. A further development to try and overcome these restrictions uses transportability estimators to assign a causal interpretation to pooled trial results [1, 2]. The idea is to obtain an estimate for a specified unstudied target population by introducing its covariate distribution into the modelling. The approach requires that the recorded covariates Xcapture all effect modification that distinguishes each trial from the target population, so that unmeasured modifiers Udo not vary in ways relevant to the contrast. Because the restriction operations that define the trials have already removed the variation needed to examine this requirement, the procedures cannot detect when the target population contains effect-modifying structure that was absent from the trials. The resulting estimand remains an extrapolation from restricted evidential states, with the shift from precision-weights to X-weights leaving the underlying structural uncertainty not only unchanged but unobservable. Latent unobserved mixture A second limitation, independent of the causal uncertainty principle, concerns the role of randomisation in the presence of contextual heterogeneity. Randomisation ensures internal validity within a given trial by balancing both observed and unobserved variables between the treatment and control arms. It does not, however, harmonise the distributions of unobserved variables across DOI: 10.5281/zenodo.17854326 2 Meta-analysis and generalisability different trials. Each trial is conducted in a particular context and necessarily draws its participants from that context’s underlying mixture of unmeasured characteristics, including potential effect modifiers. Even when trials employ the same eligibility criteria, the unobserved configurations of those who satisfy the criteria can differ markedly across contexts. Two trials may therefore estimate average treatment effects that are internally valid for their respective study populations but defined over different latent mixtures of unobserved factors. The resulting estimands are not guaranteed to be commensurable, even if the observed covariate structures appear similar. To illustrate the problem, consider the two estimates θAand θBarising from trials that are “identical” with respect to their observable design features. Let UAand UBdenote the unobserved contextual mixtures of the participants enrolled in each study. The two effects may be represented as: θA=E[Y(1) −Y(0) |UA], θB=E[Y(1) −Y(0) |UB]. where UA=UBbecause the contexts from which the participants are drawn differ in ways that are neither measured nor controlled. Randomisation balances unobserved variables within each trial, but it cannot harmonise the latent distributions across trials. Random-effects models do not resolve this difficulty. Their use relies on the idea that study-level effects are drawn from a common distribution, θi∼ N(µ, τ2), where µrepresents a population average effect. However, this formulation presupposes what it must demonstrate, that a common causal universe exists from which studies randomly sample. When UA=UB, the estimands are defined on different latent structures and are not variations around a shared parameter. The random-effects model performs a mechanical computation regardless of whether its ontological premise holds. It will produce ˆµand ˆτ2 whether studies estimate a common effect with random variation or fundamentally distinct causal quantities. The model form cannot adjudicate between these possibilities. It merely assumes commensurability. No statistical adjustment, weighting scheme, or modelling strategy can correct for context-specific differences in unobserved effect modifiers that are determined by the populations available to each trial. This limitation arises from heterogeneous realities, is independent of the evidential-state framework and applies even when all constituent studies are rigorously conducted RCTs. Conclusion Taken together, the two critiques show that meta-analysis lacks the information required for deductive claims about generalisability. The first critique, based on the causal uncertainty principle, demonstrates that the restriction operations required for causal identification contract evidential breadth in ways that subsequent pooling, weighting, or modelling cannot reverse. Meta-analysis DOI: 10.5281/zenodo.17854326 3 Meta-analysis and generalisability therefore increases precision only within the narrowed domains defined by the primary studies. The second critique, independent of the evidential-state framework, shows that even within those narrowed domains the pooled estimands need not correspond to a single underlying causal quantity. Randomisation balances unobserved variables within each trial, but it does not harmonise the latent distributions of unobserved effect modifiers across trials. Identical eligibility criteria can therefore yield causally distinct study populations, making the trial-level effects incommensurable. Meta-analysis in such circumstances averages treatment effects defined on different latent mixtures, rather than refining a common effect. Because meta-analysis neither restores evidential breadth nor guarantees estimand coherence across studies, its claims to generality cannot rest on the deductive logic of causal identification. Generalisation beyond the study contexts depends on inductive judgement about the similarity of populations and conditions, not on properties of the meta-analytic estimator itself. Meta-analysis provides a more precise estimate of what has been studied, but it does not, by itself, establish what will hold more broadly. This captures Rothwell’s point that external validity necessarily rests on clinical and contextual judgement [8]. Whether effects identified in restricted and context-specific studies extend to other populations is an ontological question about the likeness of worlds. The statistical machinery of meta-analysis offers no evidence for, nor adjudication of, that assumption. Funding Not applicable Acknowledgement Three artificial intelligence models (ChatGPT 5.1 (OpenAI), Claude Sonnet 4.1 (Anthropic), and Kimi 2 (Moonshot AI)) were used extensively as editorial aids in the drafting and redrafting of this paper. The tools were also used to check the logic, flow and coherence of the argument, tighten expressions, and manage the word count. Interchanging their use provided different “editorial voices”, and independent checks on the logic and flow of arguments. They were also used for code correcting the L A T EX layout. References [1] Dahabreh, I.J., Petito, L.C., Robertson, S.E., Hernán, M.A., Steingrimsson, J.A., 2020. Toward Causally Interpretable Meta-analysis: Transporting Inferences from Multiple Randomized Trials to a New Target Population. Epidemiology 31, 334–344. doi:10.1097/EDE.0000000000001177. DOI: 10.5281/zenodo.17854326 4 Meta-analysis and generalisability [2] Dahabreh, I.J., Robertson, S.E., Petito, L.C., Hernán, M.A., Steingrimsson, J.A., 2023. Efficient and Robust Methods for Causally Interpretable Meta-Analysis: Transporting Inferences from Multiple Randomized Trials to a Target Population. Biometrics 79, 1057–1072. URL: https://doi.org/10.1111/biom.13716, doi:10.1111/biom.13716. [3] Daitch, V., Turjeman, A., Poran, I., Tau, N., Ayalon-Dangur, I., Nashashibi, J., Yahav, D., Paul, M., Leibovici, L., 2022. Underrepresentation of women in randomized controlled trials: a systematic review and meta-analysis. Trials 23, 1038. doi:10.1186/s13063-022-07004-2. [4] Hedges, L.V., 2019. Statistical Considerations, in: Cooper, H.M., Hedges, L.V., Valentine, J.C. (Eds.), The handbook of research synthesis and metaanalysis. 3rd edition ed.. Russell Sage Foundation, New York, pp. 37–48. [5] Higgins, J.P.T., López-López, J.A., Aloe, A.M., 2021. Meta-Regression, in: Schmid, C.H., Stijnen, T., White, I.R. (Eds.), Handbook of meta-analysis. first edition ed.. CRC Press, Taylor & Francis Group, Boca Raton London New York. Chapman & Hall/CRC handbooks of modern statistical methods, pp. 129–149. [6] Matt, G.E., Cook, T.E., 2019. Threats to the Validity of Generalized Inferences from Research Syntheses, in: Cooper, H.M., Hedges, L.V., Valentine, J.C. (Eds.), The handbook of research synthesis and meta-analysis. 3rd edition ed.. Russell Sage Foundation, New York, pp. 489–516. [7] Reidpath, D.D., 2025. The Causal Uncertainty Principle. URL: http://arxiv.org/abs/2511.22649, doi:10.48550/arXiv.2511.22649. arXiv:2511.22649 [stat]. [8] Rothwell, P.M., 2005. External validity of randomised controlled trials: ”to whom do the results of this trial apply?”. Lancet (London, England) 365, 82–93. doi:10.1016/S0140-6736(04)17670-8. [9] Shadish, W.R., Cook, T.D., Campbell, D.T., 2002. Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin Company, Boston, MA. [10] Thompson, S.G., Higgins, J.P.T., 2002. How should meta-regression analyses be undertaken and interpreted? Statistics in Medicine 21, 1559–1573. doi:10.1002/sim.1187. DOI: 10.5281/zenodo.17854326 5