scieee AI-readable full text Open interactive document viewer

HERALD: High-resolution Early Recognition of Antigenic Landscape Divergence

Davis, Bee Rosa

Abstract

HERALD: High-resolution Early Recognition of Antigenic Landscape Divergence A theoretical framework for geometry-based viral surveillance that enables early detection of immune-escape variants before widespread transmission. HERALD constructs a Riemannian pullback manifold where Euclidean distances in latent space approximate antigenic relationships within bounded-distortion regimes. Key Contributions: Manifold Construction: Defines a contrastive learning objective with Jacobian/Laplacian regularization that induces smooth pullback metrics, enabling Euclidean computations to approximate geodesic distances reflecting antigenic divergence Real-Time Drift Detection: Specifies a probability-integral transform (PIT) fusion scheme combining sequence, antigenic, and structural signals into a scalar drift statistic with O(log n) amortized complexity Formal Guarantees: Derives margin-to-separation results for InfoNCE objectives and establishes Cantelli-based probability bounds requiring only finite variance (no sub-Gaussian assumptions), composing into conditional end-to-end dominance bounds with explicit error budgets Evaluation Protocols: Provides falsifiable protocols for retrospective time-slice replay, prospective streaming emulation, distortion audits, and equity/parity analysis across pathogens (SARS-CoV-2, influenza, HIV) Ethics Framework: Includes comprehensive governance templates addressing dual-use risks, information hazards, data sovereignty, abstention policies, and oversight structures Scope: This manuscript is entirely theoretical—it presents definitions, assumptions, theorems, and evaluation protocols but reports no empirical results, performance metrics, or case studies. Applications extend beyond viral surveillance to bacterial pathogen monitoring (STEC, antibiotic resistance) where antigenic/functional divergence precedes clinical detection. Author: Bee Rosa Davis (NASA Mission Systems Engineer, IBM X-Force Red Principal Adversarial Intelligence Engineer) Keywords: viral surveillance, Riemannian geometry, contrastive learning, immune escape, early warning systems, algorithmic epidemiology, information geometry, public health AI License and Patent Disclaimer Copyright License: This manuscript is licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0). You are free to share and adapt this work for non-commercial purposes with appropriate attribution, provided derivative works use the same license. Patent Notice: The methods, systems, and algorithms described herein are subject to U.S. Provisional Patent Application No. 63/919,595 (filed November 18, 2025). This copyright license does NOT grant any rights under patent law. Implementation, commercial use, or deployment of the HERALD framework may require separate patent licensing arrangements. Clarification: The CC BY-NC-SA 4.0 license governs only the manuscript text and documentation—the right to read, cite, and build upon these ideas academically. The patent covers the technical implementation of the framework. For research and educational use, no patent license is required. For commercial deployment or production systems, contact the inventor regarding patent licensing. Contact for Patent Licensing: [email protected]

Full text

HERALD: High-resolution Early Recognition of Antigenic Landscape Divergence Learned Riemannian Geometry for Real-Time Viral Surveillance Bee Rosa Davis NASA Mission Systems Engineer [email protected] Abstract Public-health surveillance remains fundamentally reactive: phylogenetic trees and string-edit distances detect concerning variants only after substantial spread. We present HERALD as a theoretical framework for constructing an antigenically meaningful latent geometry from sparse neutralization and sequence-derived signals, and for mapping geometric drift to dominance risk under explicit assumptions. Formally, we (i) define a pullback-metric construction with Jacobian/Laplacian regularization under which Euclidean latent displacements approximate functional geodesics within regimes of bounded geodesic–Euclidean distortion; (ii) specify a probability-integral transform (PIT) fusion scheme that yields a drift statistic D ( t )with O ( log n )amortized per-update complexity; (iii) derive a margin-to-separation result for contrastive objectives (InfoNCE) and obtain a distance gap ∆; (iv) establish Cantelli-based probability bounds mapping ∆to immune-escape risk, requiring only finite second moments (no sub-Gaussian tails); and (v) compose these elements into a conditional end-to-end dominance lower bound with an explicit error budget covering drift detection, escape linkage, calibration, and coverage via abstention. We also formalize data-sufficiency requirements and provide component-wise complexity bounds that enable streaming deployment. Scope: This manuscript is conceptual—it states definitions, assumptions, theorems, and proposed evaluation protocols for future empirical work. It does not report retrospective or prospective performance, system latency, or calibration metrics. Within this scope, HERALD provides a principled mathematical foundation for geometry-based viral surveillance with compositional error bounds and OOD abstention. Contents 1 Introduction 4 2 HERALD Overview 4 3 Novelty Framing: Construction vs Traversal 5 4 Methods 6 4.1 Background&Notation .................................. 6 1 4.2 Data Model & Signals (assumptions) . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 4.3 Stage 1 — Manifold Construction (Pullback + Smoothness) . . . . . . . . . . . . . . 7 4.4 Stage 2 — Multi-space Drift (PIT-normalized fusion) . . . . . . . . . . . . . . . . . . 7 4.5 Stage 3 — Risk mapping, interpretability, and abstention . . . . . . . . . . . . . . . 8 4.6 Theory I — Distance ⇒Escape (finite-variance bounds) . . . . . . . . . . . . . . . . 8 4.7 Theory II — Escape ⇒Growth (monotone link and coverage) . . . . . . . . . . . . . 9 4.8 Theory III — Compositional (conditional) dominance bound . . . . . . . . . . . . . 9 4.9 Systems & Complexity (asymptotic analysis) . . . . . . . . . . . . . . . . . . . . . . 10 5 Evaluation Protocols (No Results Reported) 10 5.1 Evaluation Principles & Pre-Registration . . . . . . . . . . . . . . . . . . . . . . . . . 10 5.2 Retrospective Time-Slice Replay (Protocol) . . . . . . . . . . . . . . . . . . . . . . . 10 5.3 Prospective Streaming Emulation (Protocol) . . . . . . . . . . . . . . . . . . . . . . . 11 5.4 Metrics(DefinitionsOnly)................................. 11 5.5 Pathogen-Specific Protocol Templates (No Results) . . . . . . . . . . . . . . . . . . . 11 5.5.1 SARS-CoV-2 (Template) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 5.5.2 Influenza (H3N2) (Template) . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 5.5.3 HIV Resistance (Template) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 5.6 Robustness, Baselines, and Confounders (Protocol) . . . . . . . . . . . . . . . . . . . 12 5.7 Statistical Power, Label Noise, and αVerification (Protocol) . . . . . . . . . . . . . . 12 5.8 Transfer Learning (Protocol) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 5.9 OOD Detection & Abstention (Protocol) . . . . . . . . . . . . . . . . . . . . . . . . . 13 5.10 Geographic Parity (Protocol) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 5.11 Recombination Robustness (Protocol) . . . . . . . . . . . . . . . . . . . . . . . . . . 13 5.12 Lead-Time vs Alert-Volume Frontier (Design Study) . . . . . . . . . . . . . . . . . . 13 5.13 Model Updates & Maintenance (Protocol) . . . . . . . . . . . . . . . . . . . . . . . . 13 5.14 Manifold Flattening Validation (Distortion Audits; Protocol) . . . . . . . . . . . . . 13 6 Discussion 13 6.1 Why Geometry Anticipates Epidemiology . . . . . . . . . . . . . . . . . . . . . . . . 13 6.2 Cost-Benefit Thresholding and the Error Budget . . . . . . . . . . . . . . . . . . . . 14 6.3 Practitioner Playbook: Tiers and Lifecycle (Policy Template) . . . . . . . . . . . . . 15 6.4 Threat Model, Negative Controls, and Failure Modes . . . . . . . . . . . . . . . . . . 15 6.5 LimitationsandRedFlags................................. 16 6.6 Why Construction Matters (Not Just Traversal) . . . . . . . . . . . . . . . . . . . . 16 7 Ethics & Governance 16 7.1 Scope of Use and Ethical Principles . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 7.2 Dual-Use Risk and Information-Hazard Mitigations . . . . . . . . . . . . . . . . . . . 16 7.3 Data Stewardship, Privacy, and Sovereignty . . . . . . . . . . . . . . . . . . . . . . . 17 7.4 Decision Governance and Pre-Registered Policies . . . . . . . . . . . . . . . . . . . . 18 7.5 Negative-Control Catalog (Operations) . . . . . . . . . . . . . . . . . . . . . . . . . . 18 7.6 Abstention Monitoring and Fallback Surveillance . . . . . . . . . . . . . . . . . . . . 18 7.7 Oversight, Accountability, and Auditability . . . . . . . . . . . . . . . . . . . . . . . 18 7.8 Transparency, Risk Communication, and Release Policy . . . . . . . . . . . . . . . . 19 2 7.9 Equity and Parity Commitments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 7.10 Incident Response, Rollback, and Model Update Governance . . . . . . . . . . . . . 20 7.11 Operational Overhead and Resourcing . . . . . . . . . . . . . . . . . . . . . . . . . . 20 7.12 Limitations, Compliance, and Community Engagement . . . . . . . . . . . . . . . . . 21 7.13 Ethics Checklist (Operational Template) . . . . . . . . . . . . . . . . . . . . . . . . . 21 8 Conclusion 21 A Proofs and Technical Details 22 A.1 Proof of Lemma 1 (InfoNCE margin ⇒separation) . . . . . . . . . . . . . . . . . 22 A.2 Explicit constants in the corollary . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 A.3 Cantelli bound: odds-form (algebraically equivalent) . . . . . . . . . . . . . . . . . . 25 A.4 Independence conditions, union bound, and vacuity . . . . . . . . . . . . . . . . . . . 26 3 1 Introduction Phylogenetic pipelines and flat sequence distances conflate genotype similarity with functional similarity, obscuring early immune-escape signals until after substantial spread. The key obstacle is that the genotype → phenotype map is curved and nonlinear: Hamming or tree distance is a poor proxy for antigenic effect, especially when immune-relevant mutations are sparse, epistatic, or context-dependent. Key idea. Learn a Riemannian pullback manifold whose geodesic distances capture antigenic relationships, then monitor movement in that space. We apply smoothness regularization to control geodesic–Euclidean distortion, enabling Euclidean computations in latent space to approximate geodesic relationships within quantifiable bounds. We formalize this approach through the following contributions: • C1 (Construction). We define a contrastive+proxy objective with Jacobian/Laplacian regularization that induces a smooth pullback metric. We state the regime of validity for the Euclidean≈geodesic approximation via a geodesic–Euclidean distortion bound (see §4.3). • C2 (Drift statistic and complexity). We specify a probability-integral transform (PIT) fusion of sequence-, antigenic-, and structural-drift signals into a scalar D ( t )with O ( log n )amortized per-update complexity, enabling streaming deployment in principle while keeping the analysis theoretical here. • C3 (Risk and abstention). We formalize a risk aggregator R that combines drift, structural features, and convergence signals, and we link R to growth via a hierarchical monotone calibration; we include an abstention mechanism for out-of-distribution or high-uncertainty regimes, emphasizing coverage control rather than unconditional prediction. • C4a (Margin-to-separation). We derive a margin-to-separation result for InfoNCE contrastive objectives, obtaining a distance gap ∆between “escape” and “similar” pairs. • C4b (Compositional guarantee). We establish Cantelli-based probability bounds that map ∆to immune-escape risk, requiring only finite second moments (no sub-Gaussian tails), and we compose these pieces into a conditional end-to-end dominance lower bound with an explicit error budget (for drift detection, escape linkage, and calibration), together with a conservative union-bound fallback when conditional independence cannot be verified. • C5 (Evaluation protocols). We specify falsifiable evaluation protocols for future empirical work: metrics and thresholds, required baseline comparisons, data-sufficiency checks, and distortion audits to quantify when the Euclidean≈geodesic approximation holds. 2 HERALD Overview Training (construction). We ingest sequences, sparse neutralization assays, and structural/clinical context, and learn an embedding fϕ with Jacobian Jϕ = ∂fϕ/∂x , whose pullback metric gϕ = J⊤ ϕJϕ is regularized for smoothness. The objective is to control geodesic–Euclidean distortion so that Euclidean computations in latent space approximate antigenic geodesics within a bounded–distortion regime (Section 4.3 formalizes the distortion bound). 4 Runtime (traversal). At runtime, given sequences at time t , we compute latent displacements and derive drift statistics for each signal type (sequence, antigenic, and structural; see Section 4.4). Each signal is normalized via a probability-integral transform (PIT) and fused into a scalar drift score D ( t ). We map D ( t )to decision - relevant risk through a hierarchical monotone link with time - asymmetric calibration (i.e., using only historical data; Section 4.7), and abstain under out - of - distribution (OOD) or high - uncertainty conditions. We derive per - update complexity bounds for the algorithmic components (summarized in Section 4.9); this paper provides theoretical foundations and complexity analysis but does not report implementation benchmarks. Figure 1: HERALD dataflow. Construction: learn a smooth pullback manifold with geodesic– Euclidean distortion control. Traversal: compute Euclidean distances in latent space, fuse signals via PIT to obtain D(t), and map to risk with OOD abstention. For Microbiologists: What This Framework Does Traditional surveillance waits for a pathogenic variant (e.g., a new STEC serotype with novel virulence factors) to spread before detection. HERALD proposes detecting such variants earlier by: 1. Learning antigenic relationships from sparse neutralization data (analogous to serotyping, but geometric) 2. Monitoring genetic drift in real-time as sequences arrive (like tracking stx gene variants, but multi-dimensional) 3. Predicting dominance risk before widespread clinical presentation (proactive rather than reactive) The mathematics formalizes how to construct a "map" where distance between variants reflects their functional difference (immune escape, virulence), not just genetic similarity. This could apply to tracking antibiotic resistance emergence in enteric pathogens or predicting which STEC strains might cause outbreaks. 3 Novelty Framing: Construction vs Traversal What is new (construction vs. traversal). The novelty lies in how the space is constructed, not in how distances are computed at runtime. We optimize a composite objective L=αLbase +βLantigenic +γLproxy +λLsmooth, to learn an embedding whose pullback metric is explicitly regularized (Jacobian/Laplacian smoothness). This construction controls geodesic–Euclidean distortion so that, within a boundeddistortion regime formalized in Section 4.3, Euclidean computations in latent space approximate geodesic distances that reflect antigenic divergence. Runtime traversal is Euclidean by design; Euclidean distances acquire antigenic meaning because the geometry is constructed to ensure this correspondence. Claim (theoretical statements). Under the regularization framework and data conditions formalized in Section 4, we (i) derive a margin-to-separation result for InfoNCE that yields a latent distance gap ∆; (ii) establish Cantelli-based probability bounds mapping ∆to immuneescape risk, requiring only finite second moments (no sub-Gaussian tails); and (iii) compose 5 these into a conditional end-to-end dominance lower bound with an explicit error budget (drift detection, escape linkage, and calibration). Status (testable implications). We provide definitions and theorems in Sections 4.6–4.8 and specify distortion audits in Section 5.14 to test when the geodesic–Euclidean approximation holds. When distortion exceeds the stated bounds, the theoretical guarantees may not apply; Section 4.5 details OOD detection and abstention mechanisms for such cases. This manuscript is theoretical and reports no empirical results. 4 Methods 4.1 Background & Notation Let X denote sequence space and fϕ : X → Rd a learned embedding. The Jacobian is Jϕ ( x ) = ∂fϕ(x)/∂x and the pullback metric is gϕ(x)=Jϕ(x)⊤Jϕ(x).(1) Let dg ( ·,· )be the geodesic distance induced by gϕ , and δ ( x, x′ ) = ∥fϕ ( x ) −fϕ ( x′ ) ∥2 the latent Euclidean distance. Section roadmap. Sections 4.2–4.3 define the construction phase; Sections 4.4–4.5 specify runtime procedures; Sections 4.6–4.8 state theoretical guarantees; Section 4.9 analyzes computational complexity. Bounded-distortion regime (assumption). There exist ( ε, L )and a class of mutation paths Pepi ( L )consisting of at most L single-amino-acid substitutions restricted to predefined epitope positions such that, for any x, x′connected by a path π∈ Pepi(L), (1 −ε)dg(x, x′)≤δ(x, x′)≤(1 + ε)dg(x, x′).(2) All algorithmic uses of δ are interpreted within this bounded-distortion regime. Section 4.3 formalizes distortion control; Section 5 specifies prospective distortion audits (this manuscript is theoretical and reports no empirical results in Section 5). 4.2 Data Model & Signals (assumptions) We consider three signal types defined on sequences or sequence pairs: sequence,antigenic, and structural. When available, sparse neutralization assays provide binary/ordinal pair labels Y∈ { + ,−} (“+”: similar; “ − ”: escape) used only in the construction objective (Section 4.3) and theory (Sections 4.6–4.8). We assume: •A1 (Finite variance). Var(δ|Y=±)<∞. • A2 (Leakage control). Any time-indexed estimates use strictly historical data (no look - ahead). • A3 (Label noise). Symmetric mislabeling at rate η < ηmax <1 2 ; separation bounds include an O(η)slack (cf. Lemma 1). 6 • A4 (Signal informativeness). Each signal has non - zero association with antigenic divergence under the true metric (otherwise guarantees are vacuous). Concrete feature choices (e.g., protein language-model embeddings, structural proxies) are illustrative; the theory depends on A1–A4 and the construction/regularization below. 4.3 Stage 1 — Manifold Construction (Pullback + Smoothness) We learn fϕby minimizing the composite objective L=αLbase +βLantigenic +γLproxy +λLsmooth,(3) Lantigenic =E"−log exp(s(x, x+)/τc) exp(s(x, x+)/τc) + PK k=1 exp(s(x, x− k)/τc)#,(4) with score s ( x, x′ ) = −∥fϕ ( x ) −fϕ ( x′ ) ∥2 2/ (2 τ2 c ), temperature τc , and K negatives. Proxy terms softly encode biophysical heuristics where available. Smoothness regularization (curvature proxies). To encourage smooth geometry we penalize Jacobian energy and/or graph Laplacian smoothness: Lsmooth =Ex∇xfϕ(x)2 For Lsmooth = Tr(F⊤LF),(5) where F = [ fϕ ( x1 ) , . . . , fϕ ( xn )] ⊤ and L is a k NN Laplacian in sequence space. We do not claim certified global Lipschitz bounds; statements are restricted to the bounded-distortion regime. Biological Analogy: Why Geometry Matters Consider two E. coli isolates that differ by: - Scenario A: 10 mutations in housekeeping genes - Scenario B: 2 mutations in stx2 epitopes Traditional phylogenetic distance treats these equally (Hamming distance ≈ 10 vs 2). But functionally, Scenario B may cause greater immune escape or altered toxin activity—hence greater public health risk. HERALD’s construction learns to place Scenario B farther in latent space than Scenario A, despite fewer mutations, because the geometry is trained to reflect antigenic/functional divergence. The "pullback metric" formalizes this: we’re not just counting mutations, we’re learning which mutations matter for phenotype. Clinical parallel: This is like distinguishing STEC O157:H7 from non-O157 STEC—serotype alone (genotype proxy) doesn’t predict virulence; you need stx profiles, eae presence, and clinical context (phenotype). 4.4 Stage 2 — Multi-space Drift (PIT-normalized fusion) For each time t we compute raw statistics per signal type: SΣ (sequence/latent), SA (antigenic), SS (structural). Each is mapped through a historical null CDF Fi (estimated on a leakage-free window) and Gaussianized via the probability-integral transform (PIT): Zi(t) = Φ−1 Fi(Si(t)), i ∈ {Σ, A, S}.(6) 7 We fuse by a linear or logistic stacker with a rank-fusion fallback: D(t) = wΣZΣ(t)+wAZA(t)+wSZS(t),(7) and treat the alert threshold θ as a policy parameter to be set by decision - makers (no empirical operating points are reported here). 4.5 Stage 3 — Risk mapping, interpretability, and abstention We define a scalar risk R that aggregates distal evidence (e.g., geodesic distance to vaccine anchors), epitope overlap, structural fitness proxies, and convergence indicators, and link R to a growth signal via a hierarchical monotone calibration map (Section 4.7). Interpretability (e.g., coarse epitope - level attributions) follows dual - use constraints (Section 7). Abstention is triggered under out - of - distribution (OOD) or high - uncertainty regimes; abstention affects coverage in the error budget (Sections 4.8 and 6.2). 4.6 Theory I — Distance ⇒Escape (finite-variance bounds) Let Y∈ { + ,−} denote neutralization labels, δ = ∥fϕ ( x ) −fϕ ( x′ ) ∥2 , class means µ± = E [ δ|Y = ± ], and gap ∆ = µ−−µ+>0. Lemma 1 (InfoNCE margin ⇒ separation).With score s ( δ ) = −δ2/ (2 τ2 c )and K negatives, the population InfoNCE objective yields an effective score margin meff ≥c1 ( τc ) log K−c2 ( H ) pC/n−c3η , where C is a capacity/complexity measure of the function class and η is the label - noise rate from A3. If sis Leff-Lipschitz in δon the bounded domain of interest, then ∆≥meff Leff . Corollary (simplified). There exist constants a, b > 0(depending on τc , Leff , and domain diameter) such that ∆≥alog K−bqC/n −O(η). Full definitions and proof are provided in Appendix A. Proposition 1 (Cantelli bound for escape probability).Assume finite second moments for δ| Y = ± . Let d⋆ = ( µ+ + µ− ) / 2, variances σ2 ± = Var ( δ|Y = ± ), and class priors π± . Then the misclassification tail α≡P Y= + |δ > d⋆ admits the explicit upper bound α≤ π+ σ2 + σ2 ++ (∆/2)2 π− (∆/2)2 σ2 −+ (∆/2)2+π+ σ2 + σ2 ++ (∆/2)2 +rn, where rn = O ( pC/n )captures estimation error. A sub - Gaussian corollary yields exponential decay in ∆ 2 , but the finite - variance bound does not require sub - Gaussian tails. (Derivation in Appendix A.) 8 What the Probability Bounds Mean in Practice Lemma 1 (InfoNCE →separation): If we train on neutralization assays that correctly label "similar" vs "escape" pairs, the algorithm learns to separate them in latent space by distance ∆. Translation: Like using antisera to distinguish serotypes—if your assays are good (low noise η), you get clean separation. Proposition 1 (Cantelli bound): Given separation ∆, we can upper-bound the probability of misclassifying an "escape" variant as "similar." Translation: If a new isolate is far from vaccine strains in learned space, we can quantify confidence that it truly escapes immunity—even with noisy, finite assay data. The bound holds under "finite variance" (heavy-tailed assay noise), not requiring Gaussian assumptions. Clinical relevance: This is like saying "if this STEC isolate’s toxin profile differs substantially from reference strains, here’s our statistical confidence it will cause severe disease"—but with explicit uncertainty quantification. 4.7 Theory II — Escape ⇒Growth (monotone link and coverage) We posit a region-specific monotone link sr=gr(R)+ϵr, gr(R) = gglobal(R)+br(R),(8) with hierarchical pooling for data - sparse regions. Time asymmetry means fitting gr on historical data and freezing it prospectively (no look - ahead). Conformal methods provide coverage - calibrated prediction sets; e.g., under exchangeability, P p(t+T)∈[L, U]≥1−α, but they do not in general calibrate point probabilities without additional structure. 4.8 Theory III — Compositional (conditional) dominance bound Let p ( t )denote prevalence at time t , and let P∗∈ (0 , 1) be a policy - relevant threshold (e.g., 10%). Define ε1=P δ≤d⋆D(t)> θ, alert issued, ξ =P s<s∗R. Under conditional independence of the error sources ( ε1, α, ξ )and abstention residual ζ , we obtain the conditional lower bound P p(t+T)≥P∗D(t)> θ, alert issued≥(1 −ε1)(1 −α)(1 −ξ).(9) Without independence, a conservative union-bound fallback gives P p(t+T)≥P∗alert issued≥1−(ε1+α+ξ). An unconditional version that includes coverage multiplies the right - hand side by (1 −A ( θ ; τOOD, wmax )) (see Section 6.2). 9 6.5 Limitations and Red Flags • Assay noise and fat tails. Neutralization assays are noisy and heavy - tailed. The Cantelli bound trades exponential for polynomial decay; Section 5.7 defines how to compare empirical α to the finite-variance and sub-Gaussian bounds. No empirical values are reported here. • Calibration drift. NPIs, prior immunity, and vaccination coverage shift the R→s mapping. We enforce strict time asymmetry, leakage control, and region - wise pooling policies (Sections 4.7 and 5). • Geometry stability. Biology can reconfigure local geometry (e.g., large antigenic jumps). Distortion audits (Section 5.14) specify tests for when retraining is required. • Compute trade - offs. Runtime computations are Euclidean in latent space; training includes Jacobian/Laplacian penalties. Section 4.9 provides asymptotic complexity; throughput budgets are systems work beyond this manuscript’s scope. 6.6 Why Construction Matters (Not Just Traversal) A natural concern is that if runtime computations are Euclidean, a pullback metric may be unnecessary. Our position is structural: Euclidean distances acquire antigenic meaning only because the geometry is constructed to ensure correspondence in a bounded - distortion regime. Section 5.14 specifies distortion audits to test this hypothesis (e.g., geodesic–Euclidean error vs. path length, association with calibration error). This manuscript proposes those audits but reports no ablation results. 7 Ethics & Governance 7.1 Scope of Use and Ethical Principles HERALD is a decision-support framework for public-health surveillance; it does not replace clinical judgment or regulatory authority. We adopt four principles: beneficence (maximize early-warning benefit), non-maleficence (minimize harm from false alerts and information hazards), justice (equitable performance across regions and populations), and accountability (transparent, auditable decisions with human oversight). All uses must comply with applicable laws and data-use agreements. 7.2 Dual-Use Risk and Information-Hazard Mitigations Predictive surveillance can create information hazards if it reveals actionable mutation “recipes” or prioritizes potentially dangerous variants. We therefore propose: 1. Redaction of sensitive detail. Report attributions at coarse granularity (epitope/region level); omit unpublished combinatorial mutation sets and stepwise “paths.” 2. Operational rule for unobserved combinations. Redact attributions that involve mutation sets not present in the reference surveillance database at alert time or flagged as OOD; display only epitope-category summaries in such cases. 3. Aggregation and delay. Public summaries aggregate across sequences/regions with optional short disclosure lags; per-sequence reports are restricted to authorized officials. 16 4. Abstention-first policy. Where support is weak (OOD) or uncertainty is wide, abstain and recommend laboratory characterization rather than issuing quantitative risk. 5. No optimization over sequences. Forbid sequence search/optimization for increased risk; counterfactuals are limited to observed mutations and reported at aggregate levels. 6. Publication policy. Do not release ranked mutation combinations absent prior observation in surveillance data; do not publicly release geolocated per-sample alerts. Dual-Use Considerations for Clinical Microbiologists The information-hazard mitigations in Section 7.2 parallel biosafety concerns in clinical labs: Don’t publish "recipes": Just as we don’t publish step-by-step protocols for culturing Francisella or synthesizing botulinum toxin, HERALD avoids revealing combinatorial mutation paths that could guide gain-of-function experiments. Coarse-grained reporting: Report "epitope region X shows escape" not "mutation A123T + B456G confers immune evasion"—analogous to reporting "carbapenem-resistant Klebsiella" without detailing exact bla gene configurations in public databases. Abstention for novel combinations: If the system detects a variant with mutation combinations never seen in surveillance (analogous to a novel toxin fusion gene), it abstains from public risk scoring and routes to BSL-3 characterization. Laboratory parallel: This is like Select Agent regulations—some information requires restricted access and oversight, not public dissemination. 7.3 Data Stewardship, Privacy, and Sovereignty Data minimization. Inputs are limited to fields necessary for surveillance; personal identifiers are excluded. Location metadata are coarsened to predefined administrative levels before processing. Sovereignty and benefit sharing. Origin countries retain control over dissemination of alerts derived from their sequences. Data-sharing agreements (DSAs) specify permissible uses, redistribution constraints, retention windows, and attribution. Privacy by design. Access follows least-privilege role-based access control (RBAC). Logs and derived analytics are pseudonymized and encrypted in transit/at rest. Aggregated outputs are de-identified and coarsened (geographic, temporal) to mitigate re-identification risk; where genomic sequences may be inherently identifying, additional safeguards—secure enclaves and/or differential privacy—are applied per DSA. The phrase “ k -anonymity or equivalent threshold” denotes a policy target enforced via these mechanisms. Table 2: Data retention and access policy (template). Artifact Access Purpose Retention Raw sequences & metadata Restricted (RBAC) Inference, audit 12 months (renewable by DSA) Derived embeddings Restricted Reproducibility, drift audits 24 months Alert logs (AuditCard) Restricted Oversight, incident response 36 months Public summaries Public Transparency Indefinite 17 7.4 Decision Governance and Pre-Registered Policies We separate model estimation from decision policy. Thresholds on drift ( θ ), OOD gates ( τOOD ), and conformal-width limits ( wmax ) are pre-registered per evaluation window and updated only at scheduled reviews. The cost matrix in Table 1 is set by the competent authority; operating thresholds are then selected by maximizing per-sample expected utility in Equation (??) . Regional threshold pooling uses the shrinkage rule in Section 6.2 to avoid overfitting in data-sparse regions. Cross-border coordination under regional thresholds. Because ˜ θr varies by region, a variant may trigger in one jurisdiction but not a neighbor. To avoid blind spots, any event with D ( t ) exceeding the global θ⋆ (regardless of ˜ θr ) is mirrored to a cross-border watchlist and routed to designated contacts under multilateral DSAs for coordinated assessment. 7.5 Negative-Control Catalog (Operations) Maintain a Negative-Control Catalog of highD variants that later failed to dominate (from protocolized analyses in Section 5.6), including pattern descriptors (e.g., epitope clusters, structural penalties, sampling artifacts) and their adjudications. Governance: • Ownership and updates. Curated by the ML Lead and Policy Lead; updated quarterly and after major incidents; reviewed by the ORB and IESB. • Usage. Consulted during alert adjudication (Algorithm 2) to flag recurrent false-positive motifs. It is not a blacklist; entries inform human judgment and may be overruled with justification. • Traceability. Each entry includes provenance (time-slice, datasets), confounder assessment, and links to laboratory follow-up where available. Specification is in Supplementary Materials S7. 7.6 Abstention Monitoring and Fallback Surveillance Abstentions reduce overconfident extrapolation but can produce “silent” periods with few alerts. We therefore monitor the joint-policy abstention rate A(θ;τOOD, wmax): •Trigger. If A(θ;τOOD, wmax)exceeds a control threshold TA(t) = max0.40, µA(t)+2σA(t) for two consecutive weekly windows—where µA, σA are the rolling 8-week mean and standard deviation, respectively—default to intensified conventional surveillance (increase sentinel sequencing, targeted neutralization assays, wastewater monitoring) until uncertainty resolves. Thresholds are reviewed and may be adjusted by the ORB based on retrospective analysis. • Public reporting. Weekly transparency reports include coverage 1 −A ( θ ; τOOD, wmax )to make abstention visible to stakeholders. 7.7 Oversight, Accountability, and Auditability Two-layer oversight (templates): • Operational Review Board (ORB)—approves alert issuance, threshold updates, and incident responses on a scheduled or ad hoc basis. 18 • Independent Ethics & Safety Board (IESB)—periodic review of dual-use mitigations, equity metrics, and red-team findings; authority to pause deployment. Constitution: members nominated by public-health authorities, WHO-collaborating centers, and affected regions; conflictof-interest disclosures and term limits required. Funding: independent multi-party consortium budget (not controlled by the operating entity) with public annual reports. Dual authorization (two-person integrity). Tier 2/3 alerts require dual authorization: the Duty Officer initiates and the ORB Chair (or designated deputy) countersigns prior to issuance. Enforce this by a cryptographic countersignature recorded in the AlertCard and audit log; the RACI matrix (Table 3) reflects this requirement. Audit artifacts (system-wide vs variant-specific). Each alert produces a cryptographically signed AlertCard containing: timestamp; model/version IDs; input digests; D ( t );variant-specific risk R with confidence intervals; system-wide error budget from the latest validation window ((1 −ε1 ) , (1 −α ) , (1 −ξ )) and coverage 1 −A ( θ ; τOOD, wmax ); abstention rationale (if any); explanation artifacts (redacted per Section 7.2); and the final human decision. Detailed specifications appear in Supplementary Materials S4–S6 (Model Card, Calibration Dossier, Security/Privacy Datasheet). Table 3: Governance roles and responsibilities (RACI template). Activity Responsible Accountable Consulted Informed Data ingestion & QC Data Lead ORB Chair Legal, MOH IESB Model training & evaluation ML Lead ORB Chair Biosafety IESB Threshold policy update Policy Lead ORB Chair ML Lead, MOH IESB Tier 1 alert issuance Duty Officer ORB Chair ML Lead MOH Two-person sign-off (Tier 2/3) Duty Officer + ORB Chair ORB Chair IESB (ad hoc) MOH Negative-control catalog maintenance ML Lead + Policy Lead ORB Chair IESB Public (summary) Model drift monitoring ML Lead ORB Chair IESB, Policy Lead All Abstention override1ORB Chair IESB Chair ML Lead, Lab Lead MOH Incident response & rollback Incident Commander ORB Chair IESB, Vendors Public 7.8 Transparency, Risk Communication, and Release Policy Public communications prioritize clarity and harm minimization: 1. Context. State that HERALD is decision-support; decisions rest with public-health authorities. 2. Quantified uncertainty. Report the end-to-end conditional confidence bound in Equation (??) and current coverage 1−A(θ;τOOD, wmax). 1 Break-glass policy. An abstention override is an exceptional procedure permitted only to correct a confirmed false-positive abstention (e.g., a lab-confirmed high-risk variant flagged OOD due to a trivial artifact). It requires unanimous consent from the ORB and IESB chairs and records a cryptographic override annotation in the AlertCard. Overrides do not change model thresholds or gates and are subject to post-incident review. 19 3. Explanations. Provide coarse-grained attributions (epitope/region level) and convergence evidence; avoid novel combinatorial details (Section 7.2). 4. Action relevance. Map alerts to tiered actions (Section 6.3), linked to local guidance. 5. Versioning. Include model/card versions, calibration date, and a contact for queries. Specifications in Supplementary Materials S4–S6. 7.9 Equity and Parity Commitments Continuously monitor parity in discrimination, calibration (ECE), lead-time, and abstention across geographies and resource settings (Section 5.10). When gaps are detected, (i) increase hierarchical pooling, (ii) raise abstention (erring on caution) until support improves, and (iii) shift resources (assays, sequencing) toward under-performing regions. DSAs ensure in-country stakeholders control dissemination of alerts derived from their data. 7.10 Incident Response, Rollback, and Model Update Governance Maintain a playbook with triggers and a rollback path to a safe baseline. Algorithm 2 Alert Governance Protocol (tiered with safety checks) Require: Drift threshold θ, OOD gate τOOD, width limit wmax 1: if SOOD > τOOD or width > wmax then 2: Abstain; route to lab; log AlertCard; return 3: end if 4: Compute D ( t ), risk R , explanations, convergence, and record system-wide error budget & coverage 5: if D(t)≤θthen 6: No alert;return 7: end if 8: Apply confounder checklist (sampling spikes, NPIs); check recombination consistency (Section 5.11); consult Negative-Control Catalog (Section 7.5) 9: Map to tier (Watch/Act/Mobilize) per Section 6.3 10: if Tier ∈ {2,3}then 11: ORB convenes; two-person integrity sign-off; issue alert; mirror to cross-border watchlist if D(t)> θ⋆ 12: else 13: Duty Officer issues Tier 1 alert 14: end if 15: Post-issue monitoring: if calibration drift or distortion control limits exceeded (Section 5.14), schedule retrain; if high-cost false positives accumulate under stable costs, reassess θ via Equation (??) 7.11 Operational Overhead and Resourcing Governance introduces overhead that must be budgeted. Cryptographic signing and immutable audit logging add latency; monitoring dashboards add compute. As a design target, overhead should 20 remain in the single-digit percentage range at approximately 10 6 sequences/day, supported by a small operations team (e.g., 3–6 FTEs at that scale). These are templates, not measured outcomes; specific figures belong to systems work outside this manuscript’s scope (see protocols in Section 5). 7.12 Limitations, Compliance, and Community Engagement HERALD relies on observational data and cannot guarantee causal fitness advantage absent laboratory confirmation; outputs are advisory. Compliance with biosafety, privacy, and cross-border data-transfer regulations is mandatory under DSAs. Encourage community feedback via a public issue tracker (for non-sensitive items) and routine stakeholder workshops, especially with low-resource sites. 7.13 Ethics Checklist (Operational Template) Before issuing an alert: 1. Support check: SOOD ≤τOOD, width ≤wmax. 2. Error budget: report system-wide (1 −ε1 ),(1 −α ),(1 −ξ )and coverage 1 −A ( θ ; τOOD, wmax ). 3. Confounders addressed: sampling/NPIs reviewed; recombination consistency verified; NegativeControl Catalog consulted. 4. Equity review: parity metrics reported; if gaps, adjust pooling/abstention. 5. Dual-use screen: coarse attribution only; redact unobserved combinations or OOD details. 6. Governance: correct RACI sign-offs (including two-person integrity for Tier 2/3); AlertCard finalized; communication pack prepared. 8 Conclusion This manuscript presents HERALD as a theoretical framework for geometry-based viral surveillance. The central idea is construction over traversal: learn a smooth pullback metric so that Euclidean computations in latent space approximate geodesic distances that reflect antigenic divergence within a bounded-distortion regime. Within this scope, the paper contributes: (i) a construction objective with Jacobian/Laplacian regularization that targets geodesic–Euclidean distortion control; (ii) a probability-integral transform (PIT) fusion of sequence, antigenic, and structural signals into a drift statistic D ( t ); (iii) a risk mapping with abstention for out-of-distribution regimes; (iv) theory linking margin-to-separation (InfoNCE) to finite-variance escape bounds (Cantelli) and a conditional end-to-end dominance lower bound; and (v) pre-registered evaluation protocols that make these statements falsifiable in future work. What is—and is not—claimed. All results herein are formal statements under explicit assumptions (finite variance, leakage control, bounded label noise, and signal informativeness). We do not report empirical performance, latency, or calibration metrics. Appendix A collects proofs, constants, and algebraic variants of the probability bounds; Section 5 specifies evaluation protocols without numerical outcomes. 21 Assumptions and their audits. Key assumptions are made explicit (A1–A4) and paired with prospective checks: bounded-distortion is audited via geodesic–Euclidean error profiles; leakage is prevented by time-asymmetric estimation; label noise is stress-tested by pre-registered flips; and signal informativeness is assessed through baseline comparisons and negative controls. If distortion exceeds bounds or support is weak, the framework abstains and routes to laboratory follow-up (Sections 4.7, 4.8, and 5). Implications and next steps. The theory provides a path from constructed antigenic geometry to decision-relevant predictions with explicit error budgets and safe abstention. The next phase is prospective validation under the protocols of Section 5: time-slice replay, streaming emulation, parity audits, recombination gating, and error-budget reporting. Systems questions (throughput, latency, operational overhead) are outside this manuscript’s scope and are to be addressed in future engineering work guided by the governance templates in Section 7. Limitations. Finite-variance bounds trade exponential for polynomial tails; dominance guarantees compose conditional errors and may weaken when independence fails (a conservative union bound is provided and may be vacuous in some regimes). Geometry can shift abruptly under biology; retraining policies and distortion audits specify when the approximation ceases to hold. Closing. HERALD is offered as a rigorous blueprint: definitions, assumptions, and guarantees that can be put on the record now, and tested later. By emphasizing construction to endow Euclidean traversal with antigenic meaning—and by insisting on abstention when support is insufficient—the framework aims to make early-warning surveillance principled, interpretable, and auditable. Acknowledgments The author completed this work independently. No external funding, collaborations, or institutional resources beyond standard employment facilities were used. The author thanks the open scientific community for maintaining publicly available literature, preprint archives, and open-source toolchains that made this theoretical development possible. The views and conclusions expressed are solely those of the author and do not necessarily represent the views of any employer or affiliated organization. Author Contributions and Conflicts of Interest All conceptual, mathematical, and written components of HERALD were developed solely by the author. The author declares no competing financial interests or conflicts of interest. A Proofs and Technical Details A.1 Proof of Lemma 1 (InfoNCE margin ⇒separation) Population and empirical risks (notation). Let Lpop NCE ( ϕ )denote the population InfoNCE risk and b LNCE ( ϕ )its empirical counterpart on n i.i.d. tuples. We use γ∈ (0 , 1) for confidence; all “with high probability” statements mean probability at least 1−γ. 22 Setup. For a positive pair ( x, x+ )and K negatives {x− k}K k=1 (assumed i.i.d.; sampling without replacement yields analogous bounds up to 1/K corrections), define sϕ(x, x′) = −∥fϕ(x)−fϕ(x′)∥2 2 2τ2 c , δ(x, x′) = ∥fϕ(x)−fϕ(x′)∥2, and ∆k=sϕ(x, x− k)−sϕ(x, x+). Then Lpop NCE(ϕ) = E"log1 + K X k=1 e∆k#. Definitions used in the bound. •Effective score margin: meff(ϕ) := Ehsϕ(x, x+)−1 K K X k=1 sϕ(x, x− k)i=E[s+−¯s−],¯s−=1 K K X k=1 s− k. • Lipschitz constant on a bounded domain: If δ∈ [0 , D ](bounded-distortion regime) and s′(δ) = −δ/τ2 c, then Leff := sup δ∈[0,D] |s′(δ)|=D/τ2 c. • Capacity term and generalization: Assume there exist c > 0and a complexity measure C for the loss class such that, with probability ≥1−γ, Lpop NCE(ϕ)≤b LNCE(ϕ)+csC+ log(1/γ) n. • Symmetric label noise (A3): Observed labels are corrupted at rate η < 1 2 : the observed positive-pair law is (1 −η)P++ηP−and the observed negative-pair law is (1 −η)P−+ηP+. Step 1: InfoNCE risk ⇒ score margin (AM ≥ GM/log-sum-exp chain). Since 1+ Pke∆k≥ Pke∆kand log Pkeuk≥log K+1 KPkuk(AM≥GM), log1 + K X k=1 e∆k≥logK X k=1 e∆k≥log K+1 K K X k=1 ∆k. Taking expectations gives Lpop NCE(ϕ)≥log K−meff(ϕ)⇒meff(ϕ)≥log K− Lpop NCE(ϕ).(15) (Tighter variant.) One can also show Lpop NCE ( ϕ ) ≥log (1 + Ke−meff (ϕ) ), which implies the monotone relation meff ≥log K−log(eL−1); the linear form (15) suffices here. 23 Step 2: Generalization (population vs. empirical). With probability ≥1−γ, Lpop NCE(ϕ)≤b LNCE(ϕ)+csC+ log(1/γ) n. Combining with (15) for an empirical minimizer ˆ ϕyields meff(ˆ ϕ)≥log K−b LNCE(ˆ ϕ)−csC+ log(1/γ) n.(16) Small-loss clarification. When both the empirical risk b LNCE ( ˆ ϕ )and the capacity term cp(C+ log(1/γ))/n are ≪1(well-separated regime), the bound is non-vacuous and tightens linearly in log K. Step 3: Score margin ⇒ distance separation (Lipschitz step). Since s ( δ ) = −δ2/ (2 τ2 c )is strictly decreasing with derivative s′(δ) = −δ/τ2 c, for any u, v ∈[0, D]: s(u)−s(v) = v2−u2 2τ2 c =(v−u)(v+u) 2τ2 c ≤D τ2 c (v−u) = Leff (v−u). Apply with u=δ(x, x+)and v=1 KPkδ(x, x− k)and take expectations to get ∆:=E[δ|Y=−]−E[δ|Y= +] ≥meff(ϕ) Leff .(17) Step 4: Symmetric label noise η .Under A3, the observed margin contracts by (1 − 2 η ): mnoisy eff = (1 −2η)m(0) eff . Combining with (17) gives ∆(η)≥(1 −2η) Leff m(0) eff . Substituting (16) yields the finite-sample form with an optional O(η)slack. Connection to Lemma 1. Equation (17) is the population statement ∆ ≥meff/Leff claimed in Lemma 1. Step 2 supplies the O ( pC/n )term via uniform convergence, and Step 4 provides the label-noise correction. Writing a = 1 /Leff = τ2 c/D and bundling constants gives the corollary form used in Section 4.6: ∆≥alog K−bqC/n −O(η). □ Remarks. (i) Well-separated regime. The bound becomes informative when b LNCE ( ˆ ϕ )and the capacity term are small; when b LNCE ≈log(K+ 1) (random scores), the bound is vacuous. (ii) Negatives. The i.i.d. negative sampling assumption simplifies exposition; batch sampling without replacement (e.g., SimCLR/MoCo) yields similar bounds with O(1/K)corrections. (iii) Explicit constants. Appendix A.2 provides explicit expressions and typical numerical ranges for a(τc, D)=τ2 c/D and b(·)that appear in the simplified corollary. 24 A.2 Explicit constants in the corollary Recall from Appendix A.1 that, on the operational domain δ∈ [0 , D ], the score s ( δ ) = −δ2/ (2 τ2 c )is Leff-Lipschitz with Leff =D τ2 c ⇒a(τc, D) = 1 Leff =τ2 c D. Let the uniform-convergence bound for the InfoNCE risk be Lpop NCE(ϕ)≤b LNCE(ϕ)+csC+ log(1/γ) n, (with the confidence parameter absorbed into the constant), where C is a capacity proxy for the score class. Combining Appendix A.1 (Steps 1–3) yields the finite-sample separation ∆≥1 Leff log K−b LNCE(ϕ)−c Leff |{z} =: b(F) sC+ log(1/γ) n−Oη Leff . Thus, a(τc, D) = τ2 c D, b(F) = c τ2 c D. Example scales (illustrative). For τc∈[0.05,0.2],D∈[1,5], and K= 64 negatives, alog K=τ2 c Dlog 64 ∈2×10−3,1.7×10−1. These ranges are illustrative only; no empirical calibration is implied in this manuscript. A.3 Cantelli bound: odds-form (algebraically equivalent) Let α = P ( Y = + |δ > d⋆ )be the misclassification tail, d⋆ = ( µ+ + µ− ) / 2, class priors π± , and variances σ2 ± = Var ( δ|Y = ± ), with ∆ = µ−−µ+> 0. The population (finite-variance) Cantelli argument (Proposition 1) is equivalent to the following odds-form bound: α 1−α≤π+ π− ·σ2 + (∆/2)2·σ2 −+ (∆/2)2 σ2 ++ (∆/2)2. In the finite-sample setting, add the estimation term rn=OpC/nfrom uniform convergence: α 1−α≤π+ π− ·σ2 + (∆/2)2·σ2 −+ (∆/2)2 σ2 ++ (∆/2)2+rn, which is algebraically equivalent to the nested-fraction form stated in Proposition 1 and is often numerically friendlier. When to use odds-form (optional guidance). The odds formulation is numerically stable near the midrange ( α≈ 0 . 5), where α/ (1 −α ) ≈ 1; the direct probability form is sometimes easier 25