scieee AI-readable full text Open interactive document viewer

The Modified-Arm Fragility Quotient: An Improved Metric for Assessing Robustness in Clinical Trials

Heston, Thomas F

Abstract

Update v5: MAJOR REVISION - Formula modified for 0-1 normalization and NBF integration Critical Formula Change: NEW: MFQ = (fragility count) / n_mod, where fragility count equals the number of outcome toggles to reverse statistical significance by Fisher's exact p-value. The Fragility Index (FI) may be used to determine the fragility count for legacy purposes; however, the MFQ also applies to FI variants. OLD: MFQ = FI / (2 × n_mod) (range: 0-0.5) Relationship to FQ: Balanced trials using the FI for the fragility count: MFQ / 2 = FQ Unbalanced trials: MFQ maintains allocation-invariant interpretation The FQ = FI / N, where N = total sample size MFQ = (fragility count) / n_mod, where Fragility count is the number of toggles to reverse statistical significance; this can be the FI or a modified version of the FI n_mod is the sample size of the arm undergoing toggles Key Revisions: Updated MFQ formula to achieve 0-1 normalization consistent with Neutrality Boundary Framework (NBF) Formula will work regardless of the method utilized to calculate the number of toggles required to flip significance; historically this has been the FI but modifications can also be used to determine the fragility count Updated mathematical analysis: replaced strict MSE dominance claims with rigorous bias-variance tradeoff analysis Revised Theorem 3: Proves MFQ is unbiased for allocation-invariant fragility, while FQ introduces systematic bias under unequal allocation Updated simulation results: All MFQ values scaled by a factor of 2 (new range: 0.065-0.163 vs 0.033-0.082 in v4) Empirical validation: MFQ demonstrates 2.5-fold variation vs FQ's 5.7-fold variation across allocation ratios Enhanced NBF integration discussion with cross-domain robustness comparison Corrected mathematical claims to match proven theorems (removed unsupported strict MSE dominance) Mathematical Framework: MFQ = (fragility count) / n_mod where: fragility count = number of outcome toggles required to reverse significance Legacy fragility count is the Fragility Index (FI), although alternatives can also be used Framework accommodates current or future modifications to FI (such as label-invariant variants) n_mod = sample size of arm receiving outcome toggles 0-1 range enables integration with Neutrality Boundary Framework metrics (RQ, MeCI, DTI, ANOVA_NBF) Theoretical Results: MFQ is unbiased for allocation-invariant fragility θ* = (fragility count)/n_mod When FI is used as the fragility count: FQ is biased with Bias(FQ) = -θ*/(r+1) for allocation ratio r ≠ 1 FQ has lower variance than MFQ Empirical evidence across 96 trial configurations demonstrates MFQ's practical superiority despite the trade-off in variance. Abstract: The Fragility Index (FI) quantifies the robustness of statistically significant results in randomized controlled trials by identifying the minimum number of outcome reversals required to change significance. The Fragility Quotient (FQ) normalizes FI by total sample size to enable cross-study comparisons; however, neither metric accounts for unequal randomization ratios, which are increasingly common in clinical trials. We introduce the Modified-arm Fragility Quotient (MFQ), defined as the fragility count divided by the sample size of the arm receiving outcome toggles during fragility calculation. Through mathematical analysis and simulation studies across 96 trial configurations (10,000 replications each), we demonstrate that MFQ maintains stability across allocation ratios (range 0.065-0.163 across ρ=0.5 to 3.0, 2.5-fold variation) while FQ varies more dramatically (range 0.019-0.109, 5.7-fold variation). We prove that MFQ is unbiased for allocation-invariant fragility, whereas FQ introduces systematic bias under unequal allocation, though it achieves lower variance. Simulations demonstrate MFQ's practical superiority across realistic trial parameters. MFQ normalizes to a 0-1 scale consistent with the Neutrality Boundary Framework. For balanced trials using the Fragility Index as the fragility count, MFQ/2 = FQ, enabling straightforward conversion while maintaining standardized interpretation across allocation ratios. MFQ is a more accurate metric of fragility in trials employing unequal randomization. Citation: Heston TF. The Modified-Arm Fragility Quotient: An Improved Metric for Assessing Robustness in Clinical Trials. Version 5. Zenodo. 2025. DOI: 10.5281/zenodo.17014261 Heston, Thomas F, The Modified-Arm Fragility Quotient: An Improved Metric for Assessing Robustness in Clinical Trials (November 2025). Available at SSRN: https://ssrn.com/abstract=5425334 or http://dx.doi.org/10.2139/ssrn.5425334 Previous Versions: Version 4 (typo fix only) Version 3 (MSE dominance proofs, MFQ = FI/(2×n_mod), range 0-0.5) Version 2 (nIFQ → MFQ terminology) Version 1 (IFQ terminology, obsolete) License: CC-BY-4.0 Keywords: fragility index, modified-arm fragility quotient, MFQ, robustness, clinical trials, randomized controlled trials, unequal randomization, neutrality boundary framework, NBF, allocation-invariant fragility, bias-variance tradeoff Note on Version Numbering: Version 5 represents a methodological revision, not a minor update. Users of v3/v4 (old formula) should migrate to v5 (new formula) for NBF-consistent 0-1 scaling. The factor of 2 difference is mathematically significant.

Full text

The Modified-Arm Fragility Quotient: An Improved Metric for Assessing Robustness in Clinical Trials Version 2.0 Thomas F. Heston1,2 1Univ. of Washington School of Medicine, Seattle, WA, USA 2Elson S. Floyd College of Medicine, Washington State Univ., Spokane, WA, USA ORCID: 0000-0002-5655-2512 November 2025 Version History Version 2.0 (November 2025): Modified MFQ formula from FI/(2nmod) to (fragility count)/nmod for 0–1 normalization (Neutrality Boundary Framework). For balanced trials using FI, MFQ/2 = FQ. All results updated accordingly. Version 1.0 (October 2025): Initial publication. Abstract The Fragility Index (FI) quantifies the robustness of statistically significant results in randomized controlled trials by identifying the minimum number of outcome reversals required to change significance. The Fragility Quotient (FQ) normalizes FI by total sample size to enable cross-study comparisons. However, neither metric accounts for unequal randomization ratios, which are increasingly common in clinical trials. We introduce the Modified-arm Fragility Quotient (MFQ), defined as the fragility count divided by the sample size of the arm receiving outcome toggles during fragility calculation. Through mathematical analysis and simulation studies across 96 trial configurations (10,000 replications each), we demonstrate that MFQ maintains stability across allocation ratios (range 0.065–0.163 across ρ=0.5 to 3.0, 2.5-fold variation) while FQ varies more dramatically (range 0.019–0.109, 5.7-fold variation). We prove that MFQ is unbiased for allocation-invariant fragility while FQ introduces systematic bias under unequal allocation, though FQ achieves lower variance. Simulations demonstrate MFQ’s practical superiority across realistic trial parameters. MFQ normalizes to a 0–1 scale consistent with the Neutrality Boundary Framework. For balanced trials using the Fragility Index as the fragility count, MFQ/2 = FQ, enabling straightforward conversion while maintaining standardized interpretation across allocation ratios. MFQ is a more accurate metric of fragility in trials employing unequal randomization. 1 Keywords: Fragility index, Randomized controlled trials, Unequal randomization, Clinical trial design, Statistical robustness 1 Introduction Statistical significance testing via p-values remains the primary framework for evaluating randomized controlled trial (RCT) results1, yet p-values alone provide limited information about result robustness2. The Fragility Index (FI), introduced by Walsh et al.3, quantifies robustness by identifying the minimum number of outcome changes required to reverse statistical significance. This intuitive metric has been widely applied across medical specialties4,5. To enable comparisons across trials of different sizes, Ahmed et al.6proposed the Fragility Quotient (FQ), defined as FI divided by total sample size. While FQ successfully normalizes for sample size, it does not account for allocation ratios. Modern RCTs increasingly employ unequal randomization (e.g., 2:1, 3:1) to maximize patient access to novel therapies, reduce costs, or improve recruitment7–9. Under unequal allocation, FQ may provide misleading interpretations because it dilutes the modified-arm fragility signal with the unmodified arm size. We address this limitation by introducing the Modified-arm Fragility Quotient (MFQ), defined as the fragility count divided by the sample size of the arm receiving outcome toggles during fragility calculation. The MFQ framework uses the Fragility Index as the standard fragility count, though it accommodates current or future modifications to the FI (such as label-invariant variants). MFQ normalizes to a 0–1 scale, enabling integration with the Neutrality Boundary Framework (NBF)—a unified system for measuring geometric distance from therapeutic neutrality (e.g., RR=1, ∆=0) across different data types10—for crossdomain robustness comparison. Our contributions are threefold. First, we establish the mathematical relationship between FI, FQ, and MFQ, demonstrating their quantitative relationship under balanced allocation and analyzing their bias-variance tradeoff under unequal allocation (Section 2). Second, we demonstrate through simulations that MFQ maintains stable interpretation across allocation ratios while FQ varies systematically with randomization imbalance (Section 4). Third, we provide practical guidance for metric selection based on trial design characteristics (Section 5). 2 Mathematical Framework 2.1 Notation and Setup Consider a two-arm RCT with binary outcome: Define the allocation ratio: ρ=nI/nC. Standard designs: •Balanced: ρ= 1 (1:1 randomization) •Unbalanced: ρ= 1 (e.g., ρ= 2 for 2:1) 2 Event No Event Total Intervention a b nI=a+b Control c d nC=c+d Total a+c b +d N =nI+nC 2.2 Fragility Metrics Definition 1 (Fragility Index).The Fragility Index (FI) is the minimum number of outcome reversals (event ↔no event) in one arm required to change statistical significance, determined by Fisher’s exact test with α= 0.05 (two-sided). By convention, reversals occur in the arm with fewer events; if tied then in the arm with smaller sample size3. Row totals remain fixed while column totals change during FI calculation. Definition 2 (Fragility Quotient). FQ = FI N=FI nI+nC (1) Definition 3 (Modified-arm Fragility Quotient).Let nmod denote the sample size of the arm receiving outcome toggles during fragility calculation (the arm with fewer events, or if tied, the smaller arm), and let “fragility count” denote the number of outcome toggles required to reverse significance. Then: MFQ = fragility count nmod (2) The standard fragility count is the Fragility Index (FI), though the MFQ framework accommodates current or future modifications to the FI (such as label-invariant variants). The 0–1 normalization enables integration with the Neutrality Boundary Framework. 2.3 Relationship Between Metrics Theorem 1 (Relationship to FQ).Under balanced randomization (ρ= 1,nI=nC=n, N= 2n), when FI is used as the fragility count: MFQ =FI n=2·FI 2n=2·FI N= 2 ·FQ (3) Equivalently: MFQ/2=FQ Proof. When nI=nC=n, we have N= 2nand nmod =n. Therefore: MFQ = FI nmod =FI n=FI N/2= 2 ·FI N= 2 ·FQ (4) 3 Theorem 2 (Behavior Under Unequal Allocation).For any allocation ratio ρ, when FI is used as the fragility count: FQ =nmod N·MFQ (5) Specifically: •If the intervention arm is modified (nmod =nI): FQ =nI nI+nC ·MFQ =ρ 1+ρ·MFQ (6) •If the control arm is modified (nmod =nC): FQ =nC nI+nC ·MFQ =1 1+ρ·MFQ (7) Which arm is modified depends on event frequencies, not allocation ratio. Under unequal allocation, FQ systematically underestimates or overestimates modified-arm fragility depending on which arm has fewer events. Proof. By definition: FQ = FI N(8) MFQ = FI nmod (9) Therefore: FQ = FI N=nmod N·FI nmod =nmod N·MFQ (10) When nmod =nI:nI N=nI nI+nC=nI nI+nI/ρ =ρ 1+ρ When nmod =nC:nC N=nC nI+nC=nI/ρ nI+nI/ρ =1 1+ρ 2.4 Bias-Variance Properties of MFQ and FQ We now analyze the statistical properties of MFQ and FQ as estimators of allocationinvariant fragility. Theorem 3 (Bias and Variance of FQ and MFQ).Let the target of inference be the allocationinvariant fragility θ⋆:= FI nmod ∈[0,1],(11) where FI is the fragility index and nmod is the size of the arm to which toggles are applied. Define r:= nmod nother ≥1, n =nmod +nother,(12) where nother is the size of the other arm. Then: 4 1. MFQ is unbiased: Bias(MFQ) = 0 2. FQ is biased for r= 1:Bias(FQ) = −θ⋆/(r+ 1) 3. FQ has lower variance than MFQ: Var(FQ) <Var(MFQ) for r > 1 Proof. Link both quotients to θ⋆: MFQ = θ⋆(13) FQ = FI n=FI nmod ·nmod n=θ⋆·1 1+nother/nmod =θ⋆·1 1+1/r (14) Bias: Bias(MFQ|θ⋆) = 0 (15) Bias(FQ |θ⋆, r) = E[FQ −θ⋆] = θ⋆1 1+1/r −1=−θ⋆1 r+ 1,(16) which equals 0 only at r= 1 and is strictly negative for r > 1. Variance: Treat FI as a random integer under the data-generating process. Using scaling, Var(MFQ) = Var(FI) n2 mod (17) Var(FQ) = Var(FI) n2=Var(FI) n2 mod(1 + 1/r)2(18) Therefore Var(FQ) Var(MFQ) =1 (1 + 1/r)2=r2 (r+ 1)2<1 (19) for all r > 1, with limit 1 as r→ ∞. Thus FQ has lower variance than MFQ, but is biased, while MFQ has higher variance but is unbiased. Corollary 4 (Bias-Variance Tradeoff).Under equal allocation (r= 1),MFQ = 2 ·FQ = θ⋆ when FI is the fragility count, and both estimators are unbiased with proportional variances. For any unequal allocation (r= 1), MFQ remains unbiased while FQ becomes biased, but FQ achieves lower variance. The practical superiority of MFQ depends on whether bias reduction outweighs the variance increase in realistic trial settings. Remark 1. Theorem 3 establishes that MFQ and FQ represent different points on the bias-variance frontier. The choice of θ⋆=FI/nmod as the estimand reflects our position that allocation-invariant fragility—normalized to the modified arm—is the appropriate target for cross-trial comparison. Given this target, MFQ is unbiased while FQ introduces systematic bias proportional to allocation imbalance. However, FQ achieves lower variance. Across 96 simulated trial configurations (Section 4), MFQ demonstrates superior stability: 2.5-fold variation (0.065–0.163) versus 5.7-fold for FQ (0.019–0.109). This empirical evidence supports MFQ as the preferred metric for comparing trials with different allocation ratios, as the bias reduction consistently outweighs the variance increase in realistic settings. 5 Corollary 5 (FQ Divergence Under Imbalance).For any unbalanced trial (ρ= 1), when FI is the fragility count, FQ differs from MFQ by: •Factor of ρ 1+ρif intervention arm modified •Factor of 1 1+ρif control arm modified The divergence increases as ρmoves further from 1 in either direction. 3 Interpretation and Clinical Meaning 3.1 MFQ Interpretation MFQ measures fragility on a 0–1 scale consistent with the Neutrality Boundary Framework, enabling cross-domain comparison with other robustness metrics. For balanced trials (ρ= 1) using FI as the fragility count: MFQ = 2 ·FQ, providing straightforward conversion to traditional fragility quotient values. For unbalanced trials: MFQ focuses on the modified arm (where toggles occur) while maintaining the 0–1 normalization that enables comparison across different allocation ratios and integration with other NBF metrics. Example: MFQ = 0.10 indicates that 10% of the modified arm switching outcomes would eliminate significance. For a balanced trial using FI, this corresponds to FQ = 0.05 (5% of total sample). 3.2 Comparison Across Metrics Consider a trial with nI= 100, nC= 50, experimental arm has 30 events, control arm has 10 events, FI = 5: Metric Value Interpretation FI 5 5 patients must switch FQ 5/150 = 0.033 3.3% of total sample MFQ 5/50 = 0.10 10% of modified arm (0–1 NBF scale) Key insight: MFQ provides a normalized 0–1 measure that accounts for which arm receives toggles. For balanced trials, MFQ/2 enables direct comparison to traditional FQ values. Theorem 3 proves this normalization is statistically optimal. 4 Simulation Study 4.1 Design We conducted Monte Carlo simulations to evaluate fragility metric behavior across trial designs. Factors: 6 •Allocation ratios:ρ∈ {0.5,1.0,2.0,3.0}(1:2, 1:1, 2:1, 3:1 randomization) •Intervention arm sizes:nI∈ {50,100,200} •Baseline event rates:pC∈ {0.2,0.3,0.4} •Effect sizes: Odds ratio ∈ {1.5,2.0,3.0} •Replications: 10,000 per condition Note: Scenarios with nI= 50 and OR = 1.5 were not run, yielding 96 conditions instead of the full 4 ×3×3×3 = 108 grid. Total conditions: 96 Software Simulations were run in R 4.5.1 (R Foundation for Statistical Computing) using RStudio. Fisher’s exact tests used stats::fisher.test. Procedure: 1. Generate a 2 ×2 table from binomial distributions 2. Apply two-sided Fisher’s exact test at α= 0.05 (R stats::fisher.test) 3. If p≤0.05, compute FI, FQ, MFQ using canonical toggle rules (fixed row totals) 4. Analyze metric distributions Terminology mapping In the simulation output, MFQ appears as nIFQ; thus MFQ = nIFQ = FI/nmod. A legacy column IFQ corresponds to the conventional FQ. 4.2 Results 4.2.1 Metric Stability Across Allocation Ratios Table 1 shows mean metric values across allocation ratios for nI= 100, pC= 0.3, OR = 2.0. Table 1: Fragility metrics by allocation ratio Allocation FI FQ MFQ 1:2 (ρ= 0.5) 32.6 0.109 0.163 1:1 (ρ= 1.0) 6.5 0.033 0.065 2:1 (ρ= 2.0) 3.5 0.023 0.070 3:1 (ρ= 3.0) 2.5 0.019 0.075 Key finding: Across allocation ratios, FQ varies 5.7-fold (0.019 to 0.109), while MFQ varies only 2.5-fold (0.065 to 0.163). For the canonical 1:1 allocation, MFQ = 2·FQ (0.065 = 2 ×0.033), confirming Theorem 1. As allocation becomes increasingly unbalanced, FQ varies systematically due to dilution by total sample size. For example, at ρ= 0.5 (1:2 randomization with more control patients), 7 FQ=0.109 reflects the large total sample, while MFQ=0.163 provides a normalized measure focused on the modified arm. At ρ= 3.0 (3:1 randomization), FQ=0.019 is depressed by the large total sample, while MFQ=0.075 maintains consistent 0–1 scale interpretation. The stability of MFQ across allocation ratios (2.5-fold range vs 5.7-fold for FQ) demonstrates its superiority for comparing trials with different randomization schemes, empirically validating the bias-variance analysis in Theorem 3. 5 Discussion 5.1 When to Use Each Metric Use MFQ when: •Unequal randomization (ρ= 1) •Comparing trials with different allocation ratios •Standardized fragility reporting across diverse trial designs •Meta-analysis of fragility across studies •Unbiased estimation of allocation-invariant fragility is prioritized (Theorem 3) •Integration with Neutrality Boundary Framework metrics Use FQ when: •Historical convention matters for comparison to prior literature using traditional FQ scale •Total sample fragility perspective is specifically desired Note: For balanced trials using FI as the fragility count, MFQ/2 = FQ enables direct conversion between metrics. 5.2 Practical Recommendations For trial reporting, we recommend: 1. Always report FI (primary metric) 2. Report MFQ for all trials (balanced and unbalanced) 3. For balanced trials, note that MFQ/2 equals traditional FQ for historical comparison 4. Use MFQ as the standard for meta-analyses across different allocation ratios The unbiasedness for allocation-invariant fragility established in Theorem 3, combined with empirical evidence of superior stability, provides strong justification for preferring MFQ in contemporary trial designs. 8 5.3 Integration with the Neutrality Boundary Framework The 0–1 normalization range enables MFQ to integrate with the Neutrality Boundary Framework (NBF), a unified system for measuring geometric distance from therapeutic neutrality across different data types10. Within this framework, MFQ joins other 0–1 bounded metrics including Risk Quotient (RQ), Meaningful Change Index (MeCI), Distance to Independence (DTI), and ANOVA Neutrality Distance (ANOVANBF). This standardization facilitates comparison of evidence robustness across binary, continuous, and correlation-based outcomes. The NBF emphasizes that robustness metrics should measure geometric distance from neutrality (e.g., RR=1, ∆=0) independently of study design choices like sample size or allocation ratio. MFQ achieves this by normalizing to the modified arm—the arm where outcome changes would move results toward or away from neutrality. This design-invariant approach enables cross-study comparison on a common 0–1 scale. For balanced trials where FI is used as the fragility count, the relationship MFQ/2 = FQ provides a conversion factor for comparing to the traditional Fragility Quotient literature while maintaining NBF-consistent 0–1 scaling. 5.4 Limitations MFQ measures fragility relative to the arm receiving toggles during fragility calculation. The identity of this arm depends on event frequencies, not study design. Trials where different arms would be modified under different outcome scenarios may require careful interpretation. The simulation was limited to binary outcomes; extension to survival and continuous outcomes requires future work. Additionally, our analysis assumes fixed row totals per canonical FI methodology; alternative toggle approaches may yield different metric behaviors. 6 Conclusion The Modified-arm Fragility Quotient provides superior interpretability for unequally randomized trials while maintaining a quantitative relationship to FQ under balanced allocation (MFQ/2 = FQ when using FI). By normalizing to the modified arm on a 0–1 scale, MFQ enables direct comparisons across trials regardless of allocation ratio and integrates seamlessly with the Neutrality Boundary Framework. We have proven that MFQ is unbiased for allocation-invariant fragility while FQ introduces systematic bias under unequal allocation. Across 96 simulated trial configurations, MFQ demonstrates superior stability (2.5-fold variation) compared to FQ (5.7-fold variation), empirically validating its practical superiority. We recommend MFQ as the standard fragility measure for modern RCT designs employing any allocation ratio. Data and Code Availability R simulation code and complete results (CSV format) are available at Zenodo11:https: //doi.org/10.5281/zenodo.17386376 9