scieee AI-readable full text Open interactive document viewer

The Davis Manifold: Geometry-First Detection with Compositional Error Budgets

Davis, Bee Rosa

Abstract

The Davis Manifold: Geometry-First Detection with Compositional Error Budgets Conceptual Framework Paper This manuscript introduces Davis manifolds and Davis systems as a unified mathematical framework for safety-critical detection in domains with identity-preserving temporal structure. The framework is designed for applications where an underlying configuration (e.g., viral antigenic state, human identity, physical pose) evolves along continuous paths while only indirect, noisy observations are available. Core Contributions Geometric Foundation: Defines Davis manifolds as Riemannian state spaces with bounded geodesic–Euclidean distortion profiles ε(L) along benign path families, soft/hard configuration margins (κ_hard, κ_soft), and explicit ambiguity bands triggering abstention. Existence Theory: Proves that contrastive training (InfoNCE) with smoothness regularization yields Davis manifolds with explicit distortion bounds ε(L) ≤ K(λ)L, establishing a tunable trade-off between geometric fidelity and representation flexibility. Detection Guarantees: Derives finite-variance Cantelli bounds mapping separation in a scalar detection statistic to misclassification risk, with a compositional error budget decomposing failure into geometry error (E_geom), feature linkage error (E_link), calibration error (ξ), and abstention failure (ζ). Path-Horizon Optimization: Formalizes the trade-off between path length (detection coverage) and distortion (geometric stability), providing operational guidelines for selecting the horizon L*. Operational Protocols: Supplies complete validation workflows including distortion audits, margin estimation, error-budget estimation, calibration methods, and deployment monitoring procedures. Scope and Instantiations This is a conceptual and theoretical manuscript. It presents definitions, assumptions, theorems, and empirical protocols but does not report performance metrics. The framework is instantiated through two worked examples: HERALD: Viral antigenic drift surveillance using pullback Riemannian geometry on sequence space VIDAR: Deepfake detection via identity trajectories on the hypersphere Keywords Riemannian geometry, metric learning, contrastive learning, temporal detection, safety-critical AI, abstention, error budgets, calibration, viral surveillance, deepfake detection, explainable AI Document Type Theoretical framework paper with operational protocols Author Bee Rosa Davis, NASA Mission Systems Engineer & IBM X-Force Red Principal Adversarial Intelligence Engineer Note: This work provides a reusable theoretical foundation for geometry-first detection systems across domains including biosurveillance, media forensics, robotics, and medical monitoring. Related Work HERALD (one of the two instantiating systems discussed in this framework) is patent-pending: Davis, B. R. (2025). HERALD: High-resolution Early Recognition of Antigenic Landscape Divergence. Zenodo. https://doi.org/10.5281/zenodo.17640400 U.S. Provisional Patent Application No. 63/919,595, filed November 18, 2025.

Full text

The Davis Manifold Geometry-First Detection with Compositional Error Budgets Bee Rosa Davis NASA Mission Systems Engineer [email protected] Abstract Many safety-critical detection problems have a common structure: an underlying configuration (e.g., antigenic profile, real human identity, physical pose) evolves along identity-preserving paths, while only indirect, noisy observations are available. Contemporary machine-learning systems typically treat these problems as static classification or black-box sequence modeling, conflating Euclidean similarity in a learned embedding with semantic stability, and providing few explicit regimes of validity, abstention rules, or decomposed error budgets. We introduce Davis manifolds and Davis systems as a general framework for geometry-first detection in identity-preserving temporal domains. A Davis manifold is a Riemannian state space equipped with (i) path classes P ( L )describing “benign” identity-preserving trajectories of bounded length L , (ii) a bounded geodesic–Euclidean distortion profile ε ( L )along these paths, and (iii) soft/hard configuration margins ( κhard, κsoft )that carve out an explicit ambiguity band where the system must abstain. A Davis system couples such a manifold to (iv) monotone path-based features and a scalar statistic S , (v) a finite-variance tail bound (Cantelli inequality) mapping separation in S to misclassification risk, and (vi) a compositional error budget over geometry, feature linkage, calibration, and abstention failures. Formally, we (1) define Davis manifolds and Davis systems and state core assumptions (A1– A6) together with non-vacuity conditions that ensure configuration margins are not destroyed by distortion; (2) prove an existence theorem showing that contrastive training (InfoNCE) plus smoothness regularization yields Davis manifolds with explicit ε ( L ) ≤K ( λ ) L dependence on regularization strength and path horizon; (3) derive finite-variance detection bounds in which posterior correctness decomposes into a Cantelli term and a multiplicative error budget (1 −Egeom )(1 −Elink )(1 −ξ )(1 −ζ )corrected by an independence slack δindep ; (4) analyze the trade-off between path length L and distortion ε ( L )and describe how to choose an operational horizon L⋆ that balances coverage against geometric stability; and (5) instantiate the framework in two domains—HERALD for viral antigenic drift and VIDAR for deepfake detection via identity trajectories—as worked Davis systems. Scope: This manuscript is conceptual. It states definitions, assumptions, theorems, and proposed empirical protocols (including distortion audits, error-budget estimation, and abstention monitoring) but does not report retrospective or prospective performance metrics. Within this scope, Davis manifolds provide a unifying geometric foundation for identity-preserving temporal detection with explicit regimes of validity, abstention as a first-class behavior, and compositional error budgets that make failures legible and falsifiable. 1 1 Introduction Many high-stakes learning problems involve identity-preserving temporal processes: viruses drifting through antigenic space while remaining the “same” lineage; faces moving through pose and illumination while remaining the same person; robots traversing a workspace while preserving object identity. In all of these cases, what matters is not just what an observation looks like in isolation, but how it moves in a latent space where distance has semantic meaning. Standard pipelines blur three distinct questions: 1. Construction: How do we construct a representation in which distances and paths correspond to meaningful functional changes? 2. Traversal: Given such a representation, how do we monitor trajectories and build detection statistics? 3. Guarantees: Under what explicit assumptions can we bound risk, expose failure modes, and abstain when the regime of validity is violated? Most deep learning work optimizes traversal—architectures and losses—on top of whatever representation emerges, and then reports empirical metrics. Geometry, if present, is implicit and un-audited; uncertainty, abstention, and error budgets are usually bolted on post hoc. Key idea. We propose to reverse this order. A Davis manifold is a learned (or inherited) Riemannian manifold on which: •geodesic distance dgis constructed to reflect semantic change along identity-preserving paths; • Euclidean computations in an ambient embedding are guaranteed to approximate dg within a bounded-distortion regime parameterized by ε ( L ), the distortion radius for paths of length at most L; • configuration regions (“similar”, “changed”, “ambiguous”) are separated by soft/hard margins (κhard, κsoft), with ambiguity explicitly mapped to abstention; • downstream detection statistics admit finite-variance risk bounds with decomposed error budgets— geometry, linkage, calibration, and abstention each have their own terms. ADavis system couples such a manifold with path families, features, aggregation, and abstention policies to form an end-to-end detector whose guarantees are conditional, falsifiable, and operational. HERALD—a framework for geometry-based viral surveillance—and VIDAR—a framework for Riemannian identity dynamics in deepfake detection—are concrete instantiations of this blueprint in two very different domains. In this manuscript we abstract their common structure into a domain-agnostic theory. 1.1 Construction vs. Traversal We distinguish sharply between: Construction (training time). Learn a representation fϕand induced metric gϕsuch that: 1. InfoNCE-style or margin-based objectives separate configurations in distance (Theory I); 2 2. smoothness regularization on fϕ controls curvature and yields explicit distortion bounds ε ( L )on a path family P(L); 3. soft/hard configuration margins ( κhard, κsoft )and an ambiguity band are identified and validated. Traversal (runtime). Given observations (xt), we: 1. embed them via fϕ and follow paths on M along P ( L⋆ ), where L⋆ is an operational path horizon; 2. extract windowed features that are monotone functions of distance or curvature along these paths; 3. fuse them via probability-integral transforms (PIT) into a scalar statistic S , map S through a calibrated link to risk, and trigger abstention when distortion, separation, or support fall outside validated regimes. In a Davis system, Euclidean traversal acquires semantic meaning only because the geometry was constructed and regularized to make it so, and only within a bounded-distortion regime made explicit and auditable. 1.2 Running Examples: HERALD and VIDAR We ground the theory in two previously proposed systems that follow this construction-first philosophy. HERALD (viral surveillance). HERALD learns a pullback Riemannian manifold on viral sequences such that geodesic distances capture antigenic relationships under neutralization data, with Jacobian/Laplacian regularization controlling geodesic–Euclidean distortion. A PIT-fused drift statistic D ( t )monitors movement relative to vaccine anchors; Cantelli bounds map separation in D ( t )to immune-escape and dominance risk with an explicit error budget over drift detection, escape linkage, calibration, and coverage. VIDAR (deepfake detection). VIDAR treats face-recognition embeddings as points on the hypersphere Sd−1 , views short clips as identity trajectories on a Riemannian identity manifold, and constructs Riemannian identity manifold (RIM) features (velocity, acceleration, principal-geodesic residuals, subspace leakage) alongside auxiliary detectors. PIT-normalized features are fused into a statistic T ( Z ), calibrated to a synthesis probability ˆp , and combined with bounded-distortion and separation assumptions into a compositional error budget for high-risk alerts. Both HERALD and VIDAR obey the same high-level pattern: contrastive + smoothness ⇒Davis manifold with ε(L) ⇒path features and separation ∆S ⇒finite-variance risk bounds + error budget. The present work abstracts this pattern, yielding a reusable theory for any domain with identitypreserving temporal structure. 3 1.3 Contributions Within this unified perspective, the paper makes the following contributions. • C1 (Davis manifolds and systems). We define Davis manifolds as Riemannian manifolds with path-class-specific distortion bounds ε ( L )and soft/hard configuration margins ( κhard, κsoft ), and Davis systems as detectors that operate on such manifolds with explicit ambiguity and abstention behavior. • C2 (Existence via InfoNCE + smoothness). We prove a two-part existence result: (i) InfoNCE-style contrastive training yields an identity distance gap ∆ emb ; (ii) smoothness regularization on the embedding’s Jacobian/metric yields curvature bounds that translate into explicit distortion functions ε ( L ) ≤C′K ( λ ) L along path families P ( L ). Together these yield Davis manifolds with tunable trade-offs between expressivity and distortion. • C3 (Finite-variance detection guarantees with error budgets). We derive finite-variance Cantelli bounds that map separation ∆ S in a scalar detection statistic to misclassification risk, decomposing failure into four components: geometry error Egeom , linkage error Elink , calibration error ξ , and abstention failure ζ , with a correlation slack δindep capturing departures from independence. • C4 (Path-horizon optimization). We formalize the trade-off between path length L (coverage and detection power) and distortion ε ( L )(geometric stability), and define an operational path horizon L⋆ as the maximizer of a lower bound on detection performance subject to a distortion constraint. A worked toy example illustrates how L⋆emerges from this trade-off. • C5 (Operational protocols and instantiations). We provide protocols for estimating each error-budget component, monitoring distortion and separation, and enforcing abstention; we show how HERALD and VIDAR instantiate all abstract objects via explicit tables mapping symbols to concrete components. 4 Table 1: Contributions of Davis systems relative to representative prior work. Aspect Typical prior work Davis manifolds / systems Geometry Embeddings with implicit geometry; Euclidean metrics used for convenience Explicit Riemannian geometry with path-class distortion bounds ε ( L )and validated regimes of validity Temporal structure Sequence models (RNNs, transformers) without geometric guarantees Manifold-valued paths with soft/hard configuration margins and identity-preserving path families P(L) Separation Empirical margins or AUC/F1; no analytic link to training objective Margin-to-separation results: InfoNCE ⇒ distance gap ∆ emb ⇒ statistic separation ∆S Risk bounds Concentration inequalities or asymptotics detached from system design Finite-variance Cantelli bounds integrated into the detector, with explicit non-vacuity conditions Error budgets Single aggregate error metric; failure modes implicit Decomposed ( Egeom, Elink, ξ, ζ, δindep )with estimation protocols and abstention thresholds Abstention Often omitted or heuristic confidence thresholds Abstention as a designed outcome tied to distortion, feature range, separation, and error-budget vacuity Cross-domain framing Separate methods per domain (e.g., surveillance, forensics) Single theoretical framework instantiated by HERALD (antigenic drift) and VIDAR (identity dynamics) Throughout, we emphasize that all guarantees are conditional: they hold only for clips/variants/trajectories that lie within empirically validated distortion, separation, and calibration regimes, and are paired with explicit procedures for auditing those regimes. 1.4 Related Work and Historical Context Metric learning and contrastive objectives. Classical metric learning and contrastive methods learn distances so that similar pairs are close and dissimilar pairs are far, often in Euclidean space with Mahalanobis metrics or angular margins. Davis systems build on this tradition but differ in what is guaranteed: instead of stopping at embedding quality, we use contrastive margins as one step in a chain that leads to path-wise distortion bounds, separation in a scalar statistic, and explicit probability guarantees. 5 Riemannian geometry in statistics and machine learning. Riemannian geometry has long been used to model curved parameter spaces and structured data: from Fisher-information manifolds and information geometry, to learning on SPD matrices, Lie groups, and hyperbolic spaces. Existing Riemannian ML typically focuses on optimization or representation advantages; Davis manifolds instead emphasize regimes of validity (via ε ( L )), configuration margins, and path-based guarantees for temporal processes. Temporal ML, sequence modeling, and manifold-valued time series. Recurrent networks, transformers, and dynamical systems models provide powerful sequence models, and there is growing work on manifold-valued time series (e.g., pose trajectories, SPD flows). However, these methods rarely expose explicit distortion regimes, abstention policies, or decomposed error budgets. Davis systems treat temporal evolution as paths on a purpose-built manifold and make the mapping from geometry to detection guarantees explicit. Uncertainty, selective prediction, and conformal methods. Calibration, selective prediction, and conformal prediction give tools for uncertainty-aware decisions and coverage guarantees. Davis systems integrate these ideas but tie abstention and calibration back to geometric assumptions: when distortion or separation estimates fail, the system is required to abstain, and this behavior appears as a term (ζ) in the error budget. Domain-specific precursors: HERALD and VIDAR. HERALD and VIDAR are domainspecific blueprints for geometry-first surveillance and forensics, respectively, each combining learned or inherited geometry, PIT-normalized fusion, finite-variance Cantelli bounds, and explicit error budgets. This paper extracts the common theory that underlies both, generalizes it, and provides a framework that can be applied to other domains such as robotics, medical time series, or audio-visual monitoring. Historical context. Conceptually, Davis manifolds sit at the intersection of several intellectual lineages: • Riemannian geometry, originating with Riemann and Cartan, which formalized curved spaces and geodesics as models for physical and statistical structure; • information geometry, which interprets statistical models as Riemannian manifolds endowed with Fisher metrics; •metric and representation learning, which learn distances and embeddings tailored to tasks; • safety and reliability in ML, including conformal prediction, selective classification, and robustness analysis. Davis manifolds contribute a missing piece: a unified, path-based framework that links representation construction, distortion control, temporal dynamics, and compositional error budgets, with abstention and vacuity conditions treated as first-class objects. 6 1.5 When to Use Davis Systems vs. Standard ML Davis systems are not a universal replacement for standard ML pipelines. They are most useful when geometry, temporal structure, and explicit guarantees matter more than raw accuracy on a static benchmark. Table 2: When standard ML suffices vs. when a Davis system is appropriate. Scenario Standard ML is usually enough when. . . Davis systems are appropriate when. . . Temporal structure Inputs are i.i.d. or orderless; no notion of identity-preserving trajectories Identity or functional state persists over time and trajectories carry semantic information Stakes and guarantees Errors are low-stakes; empirical test accuracy is sufficient Decisions are safety-critical and require auditable assumptions, abstention, and lower bounds on correctness Geometry Distances are proxies for similarity only empirically You need distances and paths that mean something (e.g., antigenic drift, identity change) with a regime where Euclidean approximations are certified Multi-signal fusion A single dominant signal drives performance; simple ensembling works Multiple heterogeneous detectors must be fused with principled normalization (PIT), linkage guarantees, and interpretable error budgets Uncertainty and abstention It is acceptable to always output a label, even when uncertain The system must abstain when assumptions fail (distortion high, features OOD, separation collapses), and coverage must be tracked explicitly Cross-domain transfer Model is deployed in a narrow, fixed context You anticipate re-using the geometric construction across domains (e.g., between pathogens, media types, sensors) with revalidated assumptions Concrete examples: • Image classification on static benchmarks, generic recommendation systems, and short-text sentiment analysis typically fall on the “standard ML” side. • Viral surveillance, deepfake forensics, medical monitoring where identity and physiology evolve smoothly, and robotics with identity-preserving world models are natural candidates for Davis systems. 7 1.6 Roadmap and Reading Guide The remainder of the paper is structured in three acts. • Act I (Framework). Section 2 defines Davis manifolds and Davis systems, introduces path families P ( L ), distortion functions ε ( L ), and configuration margins ( κhard, κsoft ), and provides a parameter table and notation index together with HERALD/VIDAR instantiation tables. • Act II (Theory). Section 3 proves existence and construction results (Theorem 1) from InfoNCE + smoothness; Section 4 derives finite-variance detection bounds (Theorem 2); Section 5 analyzes path-horizon optimization and the trade-off between coverage and distortion. • Act III (Synthesis). Section 7 discusses operational protocols for error-budget estimation, distortion audits, and abstention monitoring; examines limitations, failure modes, and the Davis universality conjecture; and outlines open problems and future directions. Reading guide. • Theorists may focus on Sections 2–5 (framework, existence, detection bounds, and path optimization) and the proofs in the appendices. • Practitioners can skim this introduction, then read the definitions and parameter tables in Section 2, the operational guidelines in Section 7 (error-budget estimation, distortion audits, and abstention policies), and the HERALD/VIDAR instantiation tables in Section 2. • Domain experts (e.g., virology, forensics, robotics) may start with the running examples in Section 1.2 and the instantiation tables (Tables 4 and 5), then consult the theory sections as needed. • Reviewers and skeptics may wish to jump to the limitations and open problems in Section 7 first, particularly the failure modes and Davis universality conjecture discussions. 2 Davis Manifolds and Systems This section introduces the geometric objects that underlie Davis manifolds and Davis systems. We start with a visual, coordinate-free picture of what the framework is trying to capture, then formalize the observation space, manifold, distances, configuration structure, and path families. We conclude with (i) a compact assumptions box, (ii) parameter ranges grounded by two running examples (HERALD and VIDAR), (iii) a practitioner checklist, and (iv) a notation index. 2.1 Geometric Intuition and Running Examples At a high level, a Davis manifold is a Riemannian space in which: • points z∈ M encode functional states of a system (e.g., antigenic state of a viral variant; identity state of a face in a video); • geodesic distance dg ( z, z′ )reflects a task-relevant notion of change (e.g., immune escape, identity change); • observed data x∈ X (sequences, frames) are mapped into M by an encoder ϕ whose geometry is constructed at training time; 8 Figure 1: Geometric intuition for Davis manifolds. Left: Curved manifold with geodesic vs. chord distances and bounded-distortion band. Middle: Configuration regions with hard and soft margins and an ambiguous band. Right: Short benign paths P ( L ), some of which cross configuration boundaries and correspond to high-risk events. • at runtime, we traverse short paths on M and compute Euclidean or tangent-space statistics that inherit semantic meaning because the underlying geometry was built to make geodesic–Euclidean approximations valid in an audited regime. For intuition, we imagine a three-panel schematic (Figure 1): • Panel A (Geometry). A curved 2D surface embedded in R3 represents M . A geodesic arc connecting z and z′ traces the shortest path on the surface, with length dg ( z, z′ ). A straight ambient-space chord between the same points has length δ ( z, z′ ). In a bounded-distortion regime, these agree up to a small multiplicative error: (1 −ε)dg(z, z′)≤δ(z, z′)≤(1+ε)dg(z, z′). • Panel B (Configuration structure). Regions of M are colored by a configuration map c : M→C (e.g., antigenic cluster; identity class). A coarser map h : C → { 0 , 1 ,amb} partitions states into “non-event” (0), “event” (1), and “ambiguous” ( amb ), with inner and outer radii κhard, κsoft defining a fuzzy margin: inside κhard we are confident in label 0or 1; between κhard and κsoft the theory recommends abstention. • Panel C (Path families). Short, smooth paths γ of length at most L represent benign dynamics (e.g., realistic antigenic drift; natural identity motion). Most paths stay within a single configuration region; a subset cross a margin and correspond to high-risk events. The Davis framework focuses on such path families P ( L )and on how reliably Euclidean/tangent computations along them reflect geodesic behavior. HERALD and VIDAR will serve as running examples throughout: HERALD operates on a pullback manifold induced by a sequence encoder for viral spikes, with geodesic distance approximating antigenic distance; VIDAR operates on the unit hypersphere Sd−1 induced by a face-recognition encoder, with geodesic distance approximating identity difference. 2.2 Basic Objects: Observation Space, Manifold, and Distances We now formalize the core objects. Definition 1 (Observation space and manifold).Let X denote the observation space (e.g., sequences, frames, clips), and let M be a d -dimensional Riemannian manifold representing functional states. An encoder ϕ : X → M maps observations to states; a coordinate map ψ : M → Rd provides an ambient representation. We write z=ϕ(x)∈ M, u =ψ(z)∈Rd. Definition 2 (Metric, geodesic distance, and ambient distance).Let g denote the Riemannian metric on M with associated geodesic distance dg : M × M → [0 ,∞ ), and define the ambient 9 Symbol Meaning XObservation space (sequences, frames, clips) MRiemannian manifold of functional states ϕ:X → M Encoder from observations to manifold ψ:M→RdCoordinate map / ambient representation gRiemannian metric on M dg(z, z′)Geodesic distance induced by g δ(z, z′)Euclidean distance in ambient space Rd P(L)Benign path family with geodesic length ≤L ε(L)Distortion radius at path horizon L CConfiguration set (fine-grained states) c:M→C Configuration map h:C → {0,1,amb}Coarse label map (non-event, event, ambiguous) κhard, κsoft Hard and soft configuration margins YCoarse label Y=h(c(z)) ∈ {0,1,amb} SScalar detection statistic derived from features and auxiliaries ∆SClass separation in S(difference in class means) Egeom Probability of geometric failure affecting S Elink Probability of feature-linkage failure ξCalibration error in mapping Sto probabilities ζAbstention failure probability δindep Independence slack between error sources L⋆Operational path horizon ε⋆ε(L⋆), distortion bound at horizon L⋆ τvac Vacuity threshold for total error budget Table 6: Notation index for Davis manifolds and systems. with parameter vector θ∈ Θ. The image Mθ = fθ ( X )is equipped with the pullback metric gθ induced by the ambient Euclidean metric, and geodesic distance dgθon Mθ. We write δθ(x, x′) := fθ(x)−fθ(x′)2 for the induced Euclidean distance between embedded observations. In-distribution region. The theoretical guarantees in this section are intended to hold only on an in-distribution region Ω in ⊂ X , defined operationally by quality filters and out-of-distribution (OOD) detectors (Section 6). Concretely, Ω in is the subset of X on which: (i) we have sufficient training coverage to estimate null distributions and calibration maps; and (ii) runtime distortion audits and quality checks are expected to pass with high probability. All statements below are restricted to x, x′∈Ωin unless otherwise noted. Configuration structure and path families. We assume a configuration map c : Mθ→ C together with a coarse map h : C → { 0 , 1 ,amb} as in Definition 4, and a family of paths P ( L )as in Definition 5, consisting of curves γ : [0 , 1] → Mθ of geodesic length at most L . In applications, P ( L )encodes “benign” transformations (e.g., short mutation paths, smooth identity trajectories) 16 that preserve configuration except at explicit transition points. Our goal in this section is to show that for suitable training hyperparameters ( λ, L⋆ )there exist: • a distortion profile ε ( L )with ε ( L⋆ )small enough that Euclidean distances approximate geodesic distances along P(L⋆), and •hard/soft configuration margins (κhard, κsoft)that are not eaten by distortion, so that (Mθ, gθ,P(L⋆), ε, c, h)satisfies Definition 8 and the non-vacuity condition κsoft −2R ε(L⋆)>0,(1) where R is a radius controlling the local neighborhood in which we operate (e.g., a bound on dgθ between reference configurations and points of interest). Training objective. We assume θis obtained by minimizing a composite loss L(θ) = αLbase(θ)+βLNCE(θ)+λLsmooth(θ),(2) where: •Lbase is a task-specific loss (e.g., cross-entropy on labels derived from cor h), •LNCE is an InfoNCE-style contrastive loss over configuration-consistent pairs ( x, x+ )and configuration-changing negatives (x, x−), defined below, and •Lsmooth is a smoothness regularizer that penalizes rapid variation of fθ and the induced metric gθon Ωin. We now show how LNCE yields a configuration-dependent distance gap (Part A), and how Lsmooth bounds distortion along P ( L )(Part B), before combining them into an existence theorem for Davis manifolds. 3.2 Part A: Contrastive Training and Configuration Distance Gaps This subsection proves the first component of Theorem 1: a small InfoNCE loss on configurationrespecting pairs induces a margin in the embedding space that separates “same-configuration” from “different-configuration” points. Let Dcfg be a distribution over tuples (x, x+, x− 1, . . . , x− K)satisfying: (A1) c(fθ(x)), c(fθ(x+)) lie in the same coarse configuration under h(e.g., both mapped to 0), (A2) each c(fθ(x− k)) lies in a configuration that hmaps to a different label (e.g., 1), and (A3) all points fall in the in-distribution region: x, x+, x− k∈Ωin. Define a similarity score sθ(x, x′) = −1 2τ2fθ(x)−fθ(x′)2 2, with temperature τ > 0, and the InfoNCE objective LNCE(θ) = E(x,x+,x− 1,...,x− K)∼Dcfg "−log exp{sθ(x, x+)} exp{sθ(x, x+)}+PK k=1 exp{sθ(x, x− k)}#.(3) 17 Lemma 1 (Contrastive margin ⇒configuration separation).Assume: (B1) sθ ( x, x′ )is Leff -Lipschitz in δθ ( x, x′ ) = ∥fθ(x)−fθ(x′)∥2 on the range of distances attained on Ωin, (B2) the empirical InfoNCE loss b LNCE ( θ )is within η of the population loss LNCE ( θ )with high probability, and (B3) the negative sampling in Dcfg respects the configuration structure, in the sense that each x− k is drawn from a distribution whose ( c, h )-labels differ from those of ( x, x+ )with probability at least 1−ρfor some ρ < 1 2. Then there exist constants c, C > 0such that, with high probability, meff ≥log K−b LNCE(θ)−csC n−η−O(ρ), where meff is an effective margin between expected positive and negative scores under Dcfg . Moreover, the corresponding expected configuration distance gap ∆cfg := Eδθ(x, x−)h◦c(fθ(x)) =h◦c(fθ(x+))−Eδθ(x, x+)h◦c(fθ(x))=h◦c(fθ(x+)) satisfies ∆cfg ≥meff Leff . Proof sketch. The proof follows the standard InfoNCE margin-to-separation argument, adapted to the configuration-aware setting and restricted to Ω in . Using log 1+ Pke∆k≥log K + 1 KPk ∆ k and ∆ k = sθ ( x, x− k ) −sθ ( x, x+ ), one shows that the population InfoNCE loss lower bounds log K−meff , so that meff ≥log K−LNCE ( θ ). Uniform convergence then gives a deviation term cpC/n + η . Lipschitz continuity of sθ in δθ ( ·,· )converts a score margin into a distance margin, yielding ∆ cfg ≥meff/Leff up to the contamination factor ρ from occasional mis-sampled negatives. A full proof with explicit constants appears in Appendix A. Intuitively, Lemma 1 says: if contrastive training can reliably distinguish same-configuration from different-configuration pairs, then the embedding must allocate a non-trivial distance gap between them. This will later serve as the raw material for configuration margins κhard, κsoft in Definition 4. 3.3 Part B: Smoothness Regularization and Bounded Distortion We now connect the smoothness term Lsmooth in (2) to a distortion profile ε ( L )on the path family P(L). Smoothness proxy. On the in-distribution region Ωin, we consider a regularizer of the form Lsmooth(θ) = Ex∼µin ∥∇xfθ(x)∥2 For Lsmooth(θ) = Tr(F⊤LF),(4) where F = [ fθ ( x1 ) , . . . , fθ ( xn )] ⊤ and L is a graph Laplacian over Ω in . In both cases, the effect of increasing λ in (2) is to penalize rapid variation of the Jacobian Jθ ( x ) = ∂fθ ( x ) /∂x and thus to control the curvature of the pullback metric gθ=J⊤ θJθon Ωin. 18 Proposition 1 (Smoothness ⇒ bounded distortion along P ( L )).Fix λ > 0in (2) . Suppose that on Ωin: (C1) the Jacobian is uniformly bounded: ∥Jθ(x)∥F≤B(λ)for all x∈Ωin, and (C2) the covariant derivative of gθ along any path in P ( L )is bounded: for all γ∈ P ( L )with image in Ωin,∇˙γgθ(γ(t))≤K(λ)for all t∈[0,1], where K(λ)is non-increasing in λ. Then there exists a constant C′> 0(depending on the geometry of Mθ on Ω in ) such that for every path γ∈ P(L)with γ([0,1]) ⊂Ωin, ℓgθ(γ)−ℓδθ(γ)≤C′K(λ)L2, and hence, for any endpoints z, z′∈γ([0,1]), (1 −ε(L)) dgθ(z, z′)≤δθ(z, z′)≤(1 + ε(L)) dgθ(z, z′), ε(L):=C′′K(λ)L, (5) for some constant C′′ > 0. In particular, increasing λ decreases K ( λ )and hence ε ( L )at fixed path length L, while increasing Lworsens ε(L)linearly. Proof sketch. The assumptions bound both the norm of the Jacobian (controlling local scaling) and the rate at which the metric gθ can change along paths in P ( L ). Standard Riemannian geometry arguments (e.g., comparison with a constant-curvature model space) then imply that the geodesic length of γ deviates from its chordal length by at most a term proportional to K ( λ ) L2 , yielding the first inequality. Since dgθ ( z, z′ )is the infimum of ℓgθ ( γ )over all connecting paths, the relative distortion between dgθ ( z, z′ )and δθ ( z, z′ )can be bounded by a constant multiple of K ( λ ) L , giving (5) . The dependence of K ( λ )on λ follows from the role of Lsmooth as a curvature proxy; a detailed argument and references are provided in Appendix A. Intuition. Smoothness regularization makes the learned geometry “flatter” on Ω in by penalizing sharp changes in fθ and hence in gθ . When curvature is small, geodesics and Euclidean chords agree up to a controlled error: short paths on a nearly-flat manifold are well-approximated by straight lines in the embedding. Proposition 1 captures this intuition quantitatively: stronger regularization (larger λ ) shrinks K ( λ )and therefore the distortion radius ε ( L ), while longer path horizons L inevitably pay a price in distortion. 3.4 Combined Existence Theorem for Davis Manifolds We now combine Lemma 1 and Proposition 1 to obtain an existence theorem: under suitable hyperparameters, training with (2) produces a Davis manifold with a non-vacuous configuration margin. Theorem 1 (Existence of Davis manifolds under contrastive + smoothness training).Fix a target path horizon L⋆>0and radius R > 0. Assume: (D1) The contrastive conditions of Lemma 1 hold on Ω in , and the resulting configuration distance gap ∆cfg is strictly positive. 19 (D2) The smoothness conditions of Proposition 1 hold on Ω in for some λ > 0, yielding a distortion profile ε(L) = C′′K(λ)Lon P(L). (D3) The path family P ( L⋆ )consists of paths whose images lie in {z∈ Mθ : dgθ ( z, z0 ) ≤R} for some center(s) z0 associated with reference configurations, so that dgθ between relevant points is at most 2R. Define hard and soft configuration margins κhard := 1 4∆cfg, κsoft := 1 2∆cfg. Then there exists λ⋆ (and hence an associated distortion bound ε⋆ := ε ( L⋆ )) such that for all λ≥λ⋆ : (E1) The tuple (Mθ, gθ,P(L⋆), ε, c, h)is a Davis manifold in the sense of Definition 8 and Definition 5. (E2) The non-vacuity condition κsoft −2R ε(L⋆)>0 holds, so that configuration changes cannot be entirely eaten by distortion within the operational horizon L⋆. Proof sketch. By Lemma 1, a small InfoNCE loss yields ∆ cfg > 0, so that embeddings of different configurations are, on average, at least ∆ cfg farther apart than same-configuration pairs. Choosing κhard = 1 4 ∆ cfg and κsoft = 1 2 ∆ cfg ensures that: (i) points within κhard of a reference configuration z0 are unlikely to belong to a different coarse configuration, and (ii) points at distance at least κsoft from all same-configuration references are strongly suggestive of a change under h , matching Definition 4. By Proposition 1, for any fixed L⋆ and radius R we can increase λ until K ( λ ), and hence ε ( L⋆ ), is small enough that κsoft −2R ε(L⋆) = 1 2∆cfg −2R ε(L⋆)>0. This guarantees that even worst-case distortion along paths of length at most L⋆ within the radiusR region cannot erase the soft margin. The bounded-distortion inequality in (5) and the path-family properties in Definition 5 then verify the remaining conditions of Definition 8. A complete proof with explicit parameter dependencies appears in Appendix A. Theorem 1 formalizes the informal picture from Section 2: contrastive training carves out configuration separation in the embedding, smoothness regularization keeps the geometry flat enough on P ( L⋆ ), and a non-vacuity condition ensures that configuration margins survive bounded distortion. 3.5 Remarks and Instantiations Trade-offs between λ and L⋆ .The distortion profile ε ( L ) = C′′K ( λ ) L makes explicit the core trade-off: for a fixed path horizon L⋆ , increasing λ decreases K ( λ )and hence ε ( L⋆ ), tightening the Davis-manifold regime but potentially reducing representation flexibility; for fixed λ , increasing L⋆ improves coverage of long-range transformations but worsens distortion. Section 5 returns to this trade-off and formulates an optimization problem for choosing L⋆ to maximize downstream detection guarantees (Section 4) subject to a target distortion budget. 20 Learned vs. inherited manifolds. Theorem 1 applies most directly when both separation and smoothness are learned (e.g., via InfoNCE and a Jacobian/Laplacian regularizer). In some applications, however, part of the geometry is inherited from a frozen backbone (e.g., a face-recognition or protein-language model). In that case, Lemma 1 may be replaced by an empirical assumption about ∆ cfg (as in the inherited-margin regime of HERALD and VIDAR), and Proposition 1 is applied only to fine-tuning layers or to the restricted region Ω in where distortion audits empirically validate a small ε(L⋆). Connection to downstream detection guarantees. The Davis manifold produced by Theorem 1 is a geometric object: it says nothing yet about detection statistics or risk bounds. Section 4 will introduce features extracted along paths in P ( L⋆ ), show how a monotone feature mapping propagates configuration separation to separation in a scalar statistic S , and apply finite-variance inequalities to obtain misclassification-risk bounds. The existence theorem here ensures that those features are computed in a regime where Euclidean operations have a meaningful Riemannian interpretation. HERALD and VIDAR as concrete instances. HERALD and VIDAR provide two domainspecific instantiations of Theorem 1: • In HERALD, fθ embeds viral sequences into a latent space with a pullback metric regularized via Jacobian and Laplacian terms; P ( L )consists of epitope-restricted mutation paths of length at most L ; and contrastive pairs ( x, x+ ),( x, x− )are formed from neutralization assays distinguishing “similar” vs. “escape” variants. Empirical distortion audits on such paths support a small ε ( L⋆ ) and justify using Euclidean displacements as proxies for antigenic geodesics. • In VIDAR, fθ is an ArcFace-style encoder whose outputs lie on the hypersphere Sd−1 ; P ( L ) consists of short identity trajectories (arcs) with bounded geodesic length; and contrastive structure comes from samevs. different-identity pairs. Small-angle regimes on Sd−1 provide a natural bounded-distortion region, and temporal-smoothness regularization further reduces K(λ). In both cases, the existence theorem explains why the Euclidean computations used at runtime are meaningful: they operate inside a Davis manifold whose geometry has been constructed to reflect the relevant notion of “identity” or “configuration” within a bounded-distortion regime. 4 Detection Guarantees for Davis Systems Section 3 showed how contrastive training with smoothness regularization can construct Davis manifolds with a bounded-distortion regime and nontrivial configuration margins. We now turn to traversal: how statistic-level separation on such a system can be translated into finite-variance risk bounds and an end-to-end correctness guarantee for high-risk alerts, together with an explicit error budget. Throughout this section we fix a Davis system D=X,M, g, φ, ψ, c, h, P(L⋆), S 21 satisfying Assumptions A1–A6 from Section 2 on an in-distribution region Ω in ⊂ X , and we restrict attention to non-abstaining inputs (i.e., clips/samples for which all quality and OOD gates pass). We write Y∈ { 0 , 1 } for the coarsened label, Y = 1 for the high-risk class (e.g., escape, synthetic, anomalous), with class priors πy=P(Y=y), and let S=S(X)∈R denote the scalar detection statistic produced by the system for an input X. 4.1 From Geometric Margins to Statistic Separation The existence result (Theorem 1) ensures that, on Ω in , Davis manifolds admit (i) a boundeddistortion path regime with radius ε⋆ = ε ( L⋆ )and (ii) hard/soft configuration margins ( κhard, κsoft ) with an explicit non-vacuity gap. We now show how these geometric quantities induce separation at the level of the scalar statistic Sunder the monotonicity assumptions on features. Recall that in a Davis system the detection statistic is constructed as S=w⊤F+b, Fk=ϕk(D), k = 1,...,K, where D is some scalar path-level or distance-like quantity on M (e.g., an aggregated geodesic or Euclidean distance along paths in P ( L⋆ )), the feature maps ϕk : R+→R are nondecreasing with derivative bounded below on an operational range, and the stacker weights wk≥0. Proposition 2 (Geometric margin ⇒ statistic separation).Let Dbe a Davis system satisfying A1–A5. Assume: (a) (Soft configuration margin) For non-ambiguous configurations h(c(z)) ∈ {0,1}we have h(c(z)) = 0 ⇒dg(z, z0)≤κhard, h(c(z)) = 1 ⇒dg(z, z1)≥κsoft for some reference points or sets z0, z1∈ M and κsoft > κhard ≥0. (b) (Bounded distortion on paths) For all paths γ∈ P(L⋆)and points z, z′∈γ([0,1]), (1 −ε⋆)dg(z, z′)≤δψ(z), ψ(z′)≤(1+ε⋆)dg(z, z′). (c) (Feature monotonicity) There exist constants mk> 0and an operational range [ dmin, dmax ]such that for all d∈[dmin, dmax], ϕ′ k(d)≥mk, k = 1, . . . , K, and D(X)∈[dmin, dmax]for all non-abstaining X. (d) (Nonnegative fusion) The stacker weights satisfy wk≥0. Then there exists α > 0, depending only on ( wk, mk ), such that the class-conditional means µy := E[S|Y=y]satisfy ∆S:= µ1−µ0≥α(1 −ε⋆)κsoft −κhard+, 22 where ( · ) + = max{·, 0 } . In particular, whenever the non-vacuity condition κsoft > κhard holds and ε⋆<1, we have ∆S>0. Proof sketch. By the configuration margin (a), any two non-ambiguous points z0, z1with different coarse labels satisfy dg ( z0, z1 ) ≥κsoft −κhard along some path in P ( L⋆ )(or along a concatenation of such paths). Bounded distortion (b) then implies that for the associated distance-like quantity D we have D1−D0≳(1 −ε⋆)κsoft −κhard+, up to constants that can be absorbed into the definition of α. Feature monotonicity (c) and nonnegative fusion (d) imply that for any D1> D0 in the operational range, S(D1)−S(D0) = K X k=1 wkϕk(D1)−ϕk(D0)≥K X k=1 wkmk(D1−D0). Taking expectations over Y = 1 and Y = 0 and combining with the geometric gap above yields the stated lower bound with α := Pkwkmk . A full derivation, including explicitly tracking constants and conditioning on the ambiguous band [κhard, κsoft], appears in Appendix B. Intuition. The Davis construction ensures that different configurations are geodesically separated; bounded distortion says Euclidean distances see (almost) the same gap; feature monotonicity and nonnegative fusion ensure that larger distances monotonically increase the scalar statistic. Proposition 2 formalizes the chain configuration margin ⇒geodesic gap ⇒Euclidean gap ⇒statistic gap. In practice, we will take ∆ S as an empirically estimated quantity, but Proposition 2 explains why geometry plus monotone features cannot yield ∆ S≈ 0unless either the margin collapses or the distortion bound is violated. 4.2 Finite-Variance Risk Bounds for a Scalar Statistic Given a Davis system with statistic-level separation ∆ S> 0, we next derive a finite-variance bound on misclassification probabilities using a Cantelli inequality. This mirrors the HERALD and VIDAR analyses, but we keep notation abstract. Let µy=E[S|Y=y], σ2 y= Var(S|Y=y), y ∈ {0,1}, and ∆S=µ1−µ0>0. Consider the midpoint threshold s⋆:= µ0+µ1 2=µ0+∆S 2. Proposition 3 (Cantelli bound for Davis statistics).Assume A1 (finite variance) and ∆ S> 0. Define one-sided Cantelli tails c0:= σ2 0 σ2 0+ (∆S/2)2, c1:= σ2 1 σ2 1+ (∆S/2)2, 23 and let πy=P(Y=y). Then: P(S > s⋆|Y= 0) ≤c0, P(S≤s⋆|Y= 1) ≤c1, and the posterior correctness of high-score points satisfies PY= 1 |S > s⋆≥π1(1 −c1) π0c0+π1(1 −c1). When (∆S/2)2≫σ2 0, σ2 1, both c0and c1are small and the lower bound is close to one. Proof. Cantelli’s inequality states that for any real-valued random variable X with mean µ and variance σ2<∞and any a>0, P(X−µ≥a)≤σ2 σ2+a2. Apply this with X = S|Y = 0, µ = µ0 and a = ∆ S/ 2to obtain the bound on P ( S > s⋆|Y = 0), and with X=−S|Y= 1 (so that X−µ=µ1−S) to bound P(S≤s⋆|Y= 1). For the posterior, Bayes’ rule gives P(Y= 1 |S > s⋆) = π1P(S > s⋆|Y= 1) π0P(S > s⋆|Y= 0) + π1P(S > s⋆|Y= 1). Using P(S > s⋆|Y= 0) ≤c0and P(S > s⋆|Y= 1) ≥1−c1yields the stated fraction. Remark (Finite-variance vs. sub-Gaussian). Proposition 3 requires only finite second moments (A1); it does not assume sub-Gaussian tails. Sub-Gaussian assumptions would yield exponentially decaying bounds in ∆ 2 S , but are often unrealistic for heavy-tailed statistics. Cantelli instead offers a robust, polynomial-decay bound that matches the HERALD and VIDAR finite-variance guarantees. For later use, we define the Cantelli term BCantelli(∆S, σ0, σ1, π0, π1) := π1(1 −c1) π0c0+π1(1 −c1). In practice, ( µy, σ2 y, πy )are estimated on a held-out validation set, and an additional estimation slack εstat is subtracted from the right-hand side; we return to this in Theorem 2. 4.3 Error Decomposition for Davis Systems The Cantelli bound quantifies how separation in S controls misclassification conditional on using the correct statistic. To obtain a system-level guarantee for high-risk alerts, we must account for several distinct sources of error. Let H denote the event that the system issues a high_risk decision on an input X , under some operational threshold τ for S and auxiliary gates (Section 6 in the outline). Let C denote the event that this decision is correct, i.e. Y= 1. We decompose error into four components: 24 • Geometry error Egeom .Probability that the Davis manifold geometry behaves outside its validated regime on a non-abstaining input—e.g., bounded-distortion or smoothness assumptions fail, or the path-class choice P ( L⋆ )does not capture the actual transformation, in a way that materially corrupts S. • Linkage error Elink .Probability that the detection pipeline linking geometry to S fails even when geometry is good—e.g., feature extraction breaks, the PIT nulls Fk are misspecified, or the learned fusion fails to preserve the monotone relationship between distance and event indicator. • Calibration error ξ .Discrepancy between the calibrated output and the true conditional probability P ( Y = 1 |S )in the high-risk region, e.g. measured by an expected calibration error (ECE) or Brier decomposition restricted to high scores. • Abstention failure ζ .Probability that the system does not abstain when it should—e.g., when distortion audits, feature-range checks, or support conditions indicate that the Davis assumptions are violated, but a non-abstaining prediction is nonetheless issued. Let Ggeom , Glink , Gcal , and Gabst denote the complements of these error events (geometry behaves, linkage is faithful, calibration is adequate, abstention policy behaves), and write G:= Ggeom ∩Glink ∩Gcal ∩Gabst. By definition, P(Gc geom)=Egeom,P(Gc link)=Elink,P(Gc cal)=ξ, P(Gc abst)=ζ. We will also use the shorthand Esum := Egeom +Elink +ξ+ζ. Assumption A6 asserts that these errors are approximately independent after conditioning on being in-regime and non-abstaining, or at least that their joint failure probability is close to the product of their marginals. To capture deviations from independence, we introduce an independence slack term: Definition 10 (Independence slack).The independence slack δindep ∈[0,1] is defined as δindep := P(Gc)−1−(1 −Egeom)(1 −Elink)(1 −ξ)(1 −ζ)+. Equivalently, δindep measures how much more often some component fails than would be expected under independence. When the four error sources are genuinely independent on the evaluation distribution, δindep ≈ 0; when they correlate (e.g., geometry and linkage both fail when curvature is high), δindep is positive and must be estimated empirically. 25 The squared separation factor is 0.852 0.852+1 ≈0.72 1.72 ≈ 0 . 42. Thus g LB (3) ∝ 0 . 39 × 0 . 42 ≈ 0 . 16 (arbitrary units). •L = 6: ε (6) = 0 . 30,so1 −ε (6) = 0 . 70. Detection power is Pdetect (6) ≈ 1 −e−1≈ 0 . 63. The squared separation factor is 0.702 0.702+1 ≈0.49 1.49 ≈0.33. Thus g LB(6) ∝0.63 ×0.33 ≈0.21. •L = 10: ε (10) = 0 . 50, so 1 −ε (10) = 0 . 50. Detection power is Pdetect (10) ≈ 1 −e−10/6≈ 0 . 81. The squared separation factor is 0.502 0.502+1 =0.25 1.25 = 0.20. Thus g LB(10) ∝0.81 ×0.20 ≈0.16. •L = 15: ε (15) = 0 . 75, so 1 −ε (15) = 0 . 25. Detection power is Pdetect (15) ≈ 1 −e−15/6≈ 0 . 92. The squared separation factor is 0.252 0.252+1 = 0.0625 1.0625 ≈ 0 . 059. Thus g LB (15) ∝ 0 . 92 × 0 . 059 ≈ 0.054. In this cartoon, g LB ( L )increases from L = 3 to a peak near L≈ 6and then degrades as distortion overwhelms the benefit of additional detection power. The precise numerical values are not important; the qualitative shape (a single interior maximum) is robust across reasonable choices of ( a, L0, σ2 )with aL0 neither too small nor too large. Figure 2 plots a typical instance of (11). 0 2 4 6 8 10 12 14 16 18 20 0 0.1 0.2 L⋆≈6 Path horizon L Normalized bound g LB(L) Toy bound Key horizons Figure 2: Trade-off between path horizon and detection bound (toy model). The normalized lower bound g LB ( L )from Equation (11) exhibits a single interior maximum at L⋆≈ 6, balancing detection power Pdetect ( L )(which increases with L ) against distortion ε ( L )(which also increases with L). Parameters: a= 0.05,L0= 6,σ2= 1. This example illustrates why the path horizon L⋆ should be treated as a tunable parameter rather than fixed a priori: too small and the system misses many events (low Pdetect ); too large and the geometric approximation degrades (large ε ( L )), shrinking the effective separation and pushing the Davis bound toward vacuity. 5.4 Guidelines for choosing P(L⋆)in practice We now translate the abstract trade-offs above into concrete guidelines for two representative domains, using the parameter ranges in Section 2.7 as reference. 32 HERALD-style antigenic drift. In HERALD-like settings, a natural path family is the class of mutation paths restricted to epitope positions: PHER(L) = nγinduced by at most Lsingle-amino-acid substitutions in a pre-specified epitope seto. Guidelines: • Start with the biological horizon. Survey historical antigenic drift to identify the typical number of substitutions associated with meaningful immune escape (e.g., 3–10 changes in key epitopes over a season). This suggests a plausible range for L⋆. • Audit distortion vs. L .For candidate horizons (e.g., L∈ { 3 , 5 , 8 , 10 } ), simulate or replay mutational paths in PHER ( L )and measure empirical distortion ratios along each path. Choose L⋆such that ε(L⋆)≤εtarget, where εtarget is set by a non-vacuity condition of the form κsoft − 2 Rε ( L⋆ ) > 0and by the domain’s tolerance for geometric error. • Estimate detection power. Using retrospective data or simulations, estimate Pdetect ( L )as the fraction of historically important trajectories whose relevant configuration changes can be realized within PHER(L). This empirically populates the toy trade-off in (11). • Select L⋆ via validation. For each candidate L , instantiate the Davis system, estimate the components of (10) on held-out data (separation ∆ S ( L ), error terms, coverage), and pick L⋆ that maximizes a surrogate of LBDavis(L)subject to ε(L⋆)and error-budget constraints. Empirically, horizons in the range L⋆≈ 5–10 mutations are plausible for seasonal respiratory viruses, but the framework does not assume any specific value. VIDAR-style identity dynamics. In VIDAR-like settings on Sd−1 , the path family consists of short arcs traced by identity trajectories: PVID(L) = nγwith arc length ℓg(γ)≤Lalong an identity trajectory on Sd−1o. Guidelines: • Anchor L in typical motion. Measure frame-to-frame geodesic distances along real identity trajectories to estimate typical per-frame angles and the distribution of cumulative arc lengths over windows. This yields a natural scale for L(e.g., 0.1–0.5 radians for short clips). • Check the small-angle regime. Distortion and the validity of tangent-space approximations depend on remaining in a small-angle regime. Distortion audits analogous to Section 6.2 should verify that ε(L)remains below a target threshold on real trajectories up to L⋆. • Tie to feature behavior. Evaluate how RIM features (F1–F4) behave as functions of L ; for example, whether subspace leakage (F4) starts to saturate beyond a certain arc length, indicating diminishing returns from larger L. • Tune L⋆ on validation clips. As in the HERALD-style case, instantiate the Davis system for several candidate L values, estimate the Davis bound components, and pick L⋆ that balances 33 detection power (ability to see identity anomalies) with geometric fidelity and error-budget constraints. In both domains, the key point is that L⋆ is not a free, purely heuristic parameter: it is constrained by distortion audits, margin non-vacuity, and empirical detection power. 5.5 Limitations of the path-optimization view The toy analysis above is intentionally simplified. In real systems: • The distortion function ε ( L )may deviate from linearity (e.g., saturating for very short paths or accelerating beyond a curvature threshold). • The geodesic separation ∆ g ( L )is not strictly independent of L ; enlarging P ( L )can change which configurations are reachable and how often they occur. •Detection power Pdetect(L)is shaped by domain-specific constraints (e.g., epitope masks, occlusions, motion patterns) that may not follow a simple exponential saturation. • Error terms Egeom ( L )and ζ ( L )may increase sharply once L crosses a regime boundary, making the bound in Theorem 2 vacuous even if separation remains non-zero. • Slices of the data (by demographic, capture condition, or generator family) may exhibit different optimal horizons L⋆, suggesting that a single global L⋆is suboptimal. For these reasons, Equation (11) should be viewed as a conceptual guide rather than a prescriptive formula. The practical recommendation is to: (i) parameterize P ( L )in a domain-appropriate way; (ii) perform distortion audits and feature-behavior checks as L varies; (iii) estimate the components of the Davis bound on held-out data; and (iv) treat L⋆ as a hyperparameter chosen to maximize a validated lower bound, not merely an empirical score. In summary, path-class design is where geometric theory, domain knowledge, and operational constraints meet. The Davis framework provides the language and structure for this choice; the actual selection of P ( L⋆ )is an empirical design decision that must be documented, audited, and revisited as data and deployment conditions evolve. 6 Operational Protocols and Validation Sections 3–5 provide the theoretical foundation for Davis systems: existence guarantees from contrastive training and smoothness (Theorem 1), finite-variance detection bounds (Theorem 2), and path-horizon optimization. This section translates those results into concrete protocols for building, validating, and monitoring Davis systems in deployment. The central operational challenge is that all Davis guarantees are conditional: they hold only when geometry, features, calibration, and abstention behave as assumed on an in-distribution region Ω in . In practice, we cannot verify these assumptions exhaustively at training time, nor can we guarantee they persist in deployment. Instead, we must: (i) Audit the bounded-distortion regime and feature behavior on held-out validation data; (ii) Estimate each component of the error budget ( Egeom, Elink, ξ, ζ, δindep )with confidence intervals; 34 (iii) Monitor these quantities during deployment and trigger abstention or retraining when they drift; and (iv) Document all assumptions, thresholds, and procedures in a version-controlled dossier (Section 7). This section provides step-by-step protocols for each of these tasks, organized around the errorbudget decomposition from Theorem 2. Algorithm boxes summarize key procedures; tables provide decision rules; and worked examples show how HERALD and VIDAR instantiate these protocols. 6.1 Overview: The Validation Workflow Table 7 (conceptual) summarizes the end-to-end validation workflow for a Davis system. The process begins with a trained encoder fθand proceeds through five phases: Table 7: Davis validation workflow (overview). Each phase produces estimates or gates that feed into the error budget and abstention policies. Phase Goal Output 1. Distortion audit Validate ε ( L⋆ ) < εtarget on held-out paths Empirical ˆε ( L )curve, distortion gate thresholds 2. Margin estimation Verify κsoft −2Rˆε(L⋆)>0Estimates (ˆκhard,ˆκsoft)with CIs 3. Feature validation Check monotonicity and linkage on operational range Feature slopes {mk} , linkage failure rate ˆ Elink 4. Calibration Fit calibrator and measure ECE/Brier in high-risk region Calibration map, ˆ ξestimate 5. Error budget Estimate all terms and independence slack ( ˆ Egeom,ˆ Elink,ˆ ξ, ˆ ζ, ˆ δindep )and vacuity check Data splits. All validation protocols assume access to three disjoint held-out sets: •Dval : primary validation set for distortion audits, margin estimation, and feature checks ( ≈ 20% of data); •Dcal: calibration set for fitting the calibrator and estimating ξ(≈10%); •Dtest: final error-budget estimation and integration test (≈10%). These should be stratified by any known slices (demographics, capture conditions, generator families) and temporally separated from training data when time-series structure is present. 6.2 Phase 1: Distortion Audits The bounded-distortion assumption—that Euclidean distances approximate geodesic distances along paths in P ( L⋆ )within a factor (1 ±ε⋆ )—is the geometric foundation of Davis systems. Distortion audits empirically validate this assumption and identify regimes where it fails. 35 6.2.1 Protocol: Measuring Distortion Ratios Algorithm 1 Distortion Audit on Validation Paths Require: Encoder fθ, path family P(L), validation set Dval Ensure: Empirical distortion profile ˆε(L), per-path distortion ratios {ρj} 1: Sample Npath paths {γj}Npath j=1 from P(L⋆)using Dval 2: for each path γjwith waypoints (xj,1, . . . , xj,Tj)do 3: Embed: zj,t =fθ(xj,t)for t= 1, . . . , Tj 4: Compute geodesic length: ℓg(γj) = PTj−1 t=1 dg(zj,t, zj,t+1)▷Using metric g 5: Compute Euclidean length: ℓδ(γj) = PTj−1 t=1 ∥zj,t −zj,t+1∥2 6: Distortion ratio: ρj=ℓδ(γj)/ℓg(γj) 7: end for 8: Compute summary statistics: ¯ρ=mean(ρj),σρ=std(ρj), percentiles 9: Estimate distortion radius: ˆε(L⋆) = max{|ρj−1|} or 95th percentile of |ρj−1| 10: return ˆε(L⋆),{ρj}, slice-wise breakdowns Implementation notes. • Geodesic approximation. For general pullback manifolds, dg ( z, z′ )is not available in closed form. Use: –Riemannian metric integration: Numerically integrate Rpg( ˙γ, ˙γ)dt along the path; – Geodesic solver: Run a Riemannian optimizer (e.g., using geomstats or pymanopt ) to find the shortest path, then measure its length; or –Inherited geometry: For hypersphere Sd−1(VIDAR), use dg(z, z′) = arccos⟨z, z′⟩directly. •Path sampling. Sample paths that reflect operational usage: –HERALD: mutation trajectories with ≤L⋆substitutions in validated epitope positions; –VIDAR: contiguous frame windows of length Twith cumulative arc length ≤L⋆. • Slice-wise audits. Compute ˆε ( L⋆ )separately for each slice (e.g., by age group, lighting condition, or generator). Flag slices where ˆε>εtarget for elevated Egeom or mandatory abstention. 36 6.2.2 Decision Rules and Gates Table 8: Distortion audit decision rules. These thresholds are illustrative; actual values depend on domain and risk tolerance. Condition Interpretation Action ˆε(L⋆)<0.10 Low distortion, geometry reliable Proceed; use geometry-based features (F1–F4 in VIDAR) 0.10 ≤ˆε(L⋆)<0.20 Moderate distortion Proceed with caution; down-weight geometric features or widen uncertainty ˆε(L⋆)≥0.20 High distortion Abstain or suppress geometry; rely on auxiliary detectors only |ρj− 1 |> 0 . 30 for > 10% of paths Frequent outliers Flag slice for investigation; may indicate OOD inputs or tracking failures Example (VIDAR on deepfake validation data). Suppose Algorithm 1 yields: ¯ρ= 1.03, σρ= 0.08, ˆε(L⋆) = max j|ρj−1|= 0.18 (or 95th percentile: 0.14). This suggests moderate distortion. If using the 95 th percentile definition, ˆε = 0 . 14 < 0 . 20 and we proceed; if using the max, ˆε = 0 . 18 is borderline and we down-weight RIM features or increase abstention thresholds. 6.3 Phase 2: Margin Estimation and Non-Vacuity The configuration margins ( κhard, κsoft )from Definition 4 must be empirically validated to ensure that the non-vacuity condition κsoft −2R ε(L⋆)>0 holds on Dval, where Rbounds the geodesic radius of the operational region. 37 6.3.1 Protocol: Measuring Configuration Separation Algorithm 2 Margin Estimation on Validation Pairs Require: Encoder fθ, validation set Dval with coarse labels Y∈ {0,1} Ensure: Estimates (ˆκhard,ˆκsoft)with confidence intervals 1: Sample Nsame same-configuration pairs (x, x+)with h(c(fθ(x))) = h(c(fθ(x+))) 2: Sample Ndiff different-configuration pairs (x, x−)with h(c(fθ(x))) =h(c(fθ(x−))) 3: Compute distances: 4: Within-configuration: {dg(0) i }Nsame i=1 for same-config pairs 5: Across-configuration: {d(1) g,i }Ndiff i=1 for different-config pairs 6: Estimate margins: 7: ˆµ0=mean(d(0) g,i ),ˆµ1=mean(d(1) g,i ) 8: ˆ ∆g= ˆµ1−ˆµ0 9: ˆκhard = ˆµ0+ 2ˆσ0▷Conservative: mean + 2 std of within-config 10: ˆκsoft = ˆµ1−2ˆσ1▷Conservative: mean - 2 std of across-config 11: Compute bootstrap CIs for (ˆκhard,ˆκsoft)via resampling 12: Non-vacuity check: Verify ˆκsoft −2Rˆε(L⋆)>0using ˆεfrom Phase 1 13: return (ˆκhard,ˆκsoft), CIs, non-vacuity flag Interpretation. If ˆκhard <ˆκsoft and the non-vacuity condition holds, the system has a validated ambiguity band [ ˆκhard,ˆκsoft ]where abstention should trigger. If ˆκhard ≥ˆκsoft or non-vacuity fails, the configuration structure is not sufficiently separated for reliable detection at horizon L⋆ ; either reduce L⋆, increase λ(smoothness), or collect more contrastive training data. 6.4 Phase 3: Error Budget Estimation The compositional guarantee in Theorem 2 decomposes correctness into four error terms plus an independence slack. This subsection provides estimation protocols for each. 6.4.1 Geometry Error (Egeom) Definition. Egeom is the probability that the Davis manifold geometry behaves outside its validated regime in a way that materially corrupts the detection statistic S. Operational proxy. On Dtest, flag samples as geometry failures if: (a) Any path γused to compute features has distortion |ργ−1|> εmax (e.g., εmax = 0.25); (b) Smoothness checks fail (e.g., large second-order finite differences in embeddings suggest high local curvature); or (c) Paths exit the validated region (e.g., geodesic distance to training support exceeds a Mahalanobis threshold). Estimator. ˆ Egeom =1 Ntest Ntest X i=1 ⊮{sample iflagged as geometry failure}. 38 Compute separately on slices and worst-case across slices for a conservative bound. 6.4.2 Linkage Error (Elink) Definition. Elink is the probability that features fail to track geodesic distance even when geometry is good—e.g., monotonicity breaks, PIT nulls are misspecified, or feature fusion produces a statistic Suncorrelated with true configuration change. Operational proxy. On Dtest , restrict to samples passing geometry checks (not flagged in Phase 1). For each, compute: (a) Ground-truth configuration change ∆c(binary: same vs. different); (b) Predicted statistic Sand its sign relative to threshold s⋆. Flag as linkage failures any samples where S and ∆ c are anti-correlated (e.g., S < s⋆ but ∆ c = 1, or vice versa) despite passing geometry gates. Estimator. ˆ Elink =#{samples with anti-correlated (S, ∆c)} #{samples passing geometry checks}. 6.4.3 Calibration Error (ξ) Definition. ξ quantifies miscalibration in the mapping from the raw statistic S to the calibrated probability ˆp=calib(S)in the high-risk region (S > s⋆). Protocol: Expected Calibration Error (ECE). 1. Fit a calibrator (e.g., isotonic regression, Platt scaling, or temperature scaling) on Dcal to map S→ˆp. 2. On Dtest, bin samples by ˆp(e.g., 10 bins from 0 to 1). 3. For each bin Bkin the high-risk region (ˆp>0.5or S > s⋆), compute: ECEk=avg. predicted ˆpin bin −fraction of true positives in bin. 4. Set ˆ ξ= maxkECEkover high-risk bins, or use a weighted average. Alternative: Brier score decomposition. Decompose the Brier score restricted to S > s⋆ into calibration and refinement components; use the calibration component as ˆ ξ. 6.4.4 Abstention Failure (ζ) Definition. ζ is the probability that the system fails to abstain when it should—i.e., issues a non-abstaining prediction (‘info’, ‘warn’, or ‘high_risk’) despite being out-of-regime. 39 Operational proxy. Define an oracle abstention set Aoracle on Dtest as samples that fail any of: •Distortion gate (|ρ|−1> εmax); •Margin check (dg∈[κhard, κsoft], i.e., in ambiguous band); •Feature range check (Sor individual features outside [Fmin, Fmax]from training); •Support check (e.g., Mahalanobis distance to training manifold exceeds threshold). Let Apred be the set of samples for which the Davis system actually abstains. Then: ˆ ζ=#Aoracle \ Apred #Aoracle =should-abstain but didn’t total should-abstain . 6.4.5 Independence Slack (δindep) Definition. δindep captures the excess probability that some error occurs beyond what would be predicted by assuming (Egeom, Elink, ξ, ζ)are independent. Estimator. On Dtest, flag each sample for each error type. Compute: ˆεjoint =#{samples with ≥1error} Ntest , ˆεmult = 1 −(1 −ˆ Egeom)(1 −ˆ Elink)(1 −ˆ ξ)(1 −ˆ ζ). Then: ˆ δindep = max0,ˆεjoint −ˆεmult. If ˆ δindep >0.1, the multiplicative bound is optimistic and a union-bound fallback should be used. 6.5 Phase 4: Abstention and Quality Gates Abstention in Davis systems is not a post-hoc confidence threshold; it is a first-class outcome triggered when assumptions fail. This subsection formalizes abstention policies. 40 6.5.1 Multi-Stage Abstention Architecture Table 9: Abstention gates and decision logic. Gates are applied sequentially; failure at any stage triggers ‘insufficient_signal’. Gate Condition Interpretation G1: Quality Basic quality checks (resolution, blur, completeness) Filters obviously broken inputs G2: OOD Mahalanobis distance, reconstruction error, or novelty score Detects inputs far from training support G3: Distortion |ργ−1|≤εmax for all paths γEnsures bounded-distortion regime G4: Margin dg/∈[κhard, κsoft]Avoids ambiguous band G5: Feature range All features Fk∈[F(k) min, F (k) max]Checks features are in-range G6: Vacuity ˆ Egeom +ˆ Elink +ˆ ξ+ˆ ζ < τvac Global error budget below vacuity threshold Decision tree. 1. If any of G1–G5 fail for a sample, output ‘insufficient_signal’. 2. If G6 fails (error budget exceeds τvac ≈0.7), downgrade all alerts or abstain globally. 3. Otherwise, compute S and proceed to threshold-based decisions: ‘ info ’, ‘ warn ’, or ‘ high_risk ’. 6.6 Phase 5: Calibration Methods Calibration maps the raw detection statistic S to a probability ˆp≈P ( Y = 1 |S ). This subsection describes three common approaches. 6.6.1 Isotonic Regression (Non-Parametric) Fit a piecewise-constant, monotone map S7→ ˆp via isotonic regression on Dcal . This preserves the ranking of Swhile adjusting probabilities to match empirical frequencies. Advantages: Flexible, no parametric assumptions. Disadvantages: Can overfit on small calibration sets; requires careful binning. 6.6.2 Platt Scaling (Logistic) Fit a logistic model: ˆp(S) = 1 1 + exp(−aS −b), where (a, b)are learned on Dcal via maximum likelihood. Advantages: Simple, interpretable, works well when Sis approximately linear in log-odds. 41 Red flags. Empirically, the following are warning signs that a Davis instantiation is out of its validated regime: (i) distortion audits show large ε ( L⋆ )or heavy tails, but coverage remains high; (ii) measured separation ∆ S collapses on deployment data; (iii) the estimated error budget exceeds τvac; or (iv) slice analyses show severe disparities that cannot be explained by data scarcity alone. 7.4 Relation to Existing Frameworks Davis manifolds touch several established areas: metric learning, Riemannian machine learning, conformal prediction and selective classification, and domain-specific geometric systems such as HERALD and VIDAR. Table 13 summarizes the relationship at a high level. Framework Primary focus What Davis manifolds add Metric / contrastive learning Embeddings where similar pairs are close, dissimilar pairs are far Explicit Riemannian structure, boundeddistortion path families, and soft/hard configuration margins with ambiguity bands. Riemannian ML (SPD nets, hyperbolic embeddings, etc.) Optimization and representation on fixed manifolds Data-driven construction of the manifold via pullback metrics, plus deployment-time distortion audits and abstention policies tied to the geometry. Conformal / selective prediction, OOD detection Coverage guarantees, abstention, uncertainty quantification A geometric substrate that specifies when abstention should trigger (e.g., distortion or margin violations), and a compositional error budget that includes geometry and feature linkage. HERALD, VIDAR (domain-specific systems) Viral surveillance and deepfake forensics with geometry-based detectors A unifying abstraction: both are instances of Davis systems with problem-specific manifolds, path families, and statistics; the present work generalizes their theory and operational protocols. Table 13: Relation between Davis manifolds and selected prior frameworks. Davis manifolds can be seen as an interface layer: metric learning and Riemannian ML supply tools for constructing ( M, g ); conformal prediction and selective classification supply tools for abstention and coverage; and domain knowledge supplies task-specific path families, configuration maps, and features. 7.5 The Davis Universality Conjecture A central motivating question is whether Davis manifolds capture a broad class of safety-relevant temporal problems. Conjecture (Davis universality). Any supervised learning problem with identitypreserving temporal structure can be recast as path classification on a Davis manifold, provided that: 48 (U1) there exists an appropriate notion of “identity” or functional state that evolves continuously along benign trajectories; (U2) a contrastive signal (same/different or similar/dissimilar pairs) is available to learn or validate a separation margin; (U3) smoothness can be incentivized architecturally or via regularization so that boundeddistortion path families exist; and (U4) configuration changes of interest exhibit a positive geodesic jump (a margin κ > 0) in the learned geometry. Under these conditions, there exists a Davis manifold and path family for which the detection task can be formulated as classifying paths that cross configuration margins. This conjecture may not hold in full generality, but it provides a concrete research program: identify domains that satisfy (or fail) the conditions, characterize what identity and smoothness mean in each, and study how far the Davis construction can be pushed. Examples likely to satisfy (U1)–(U4). • Protein or viral evolution (identity = functional antigenic state; paths = mutation trajectories). •Video-based identity or pose tracking (identity = person or object; paths = short clips). • Robot or autonomous-vehicle telemetry (identity = system state; paths = control trajectories). •Speaker or instrument recognition over time (identity = source; paths = audio segments). Examples that may violate the conjecture. • Domains with rapidly mixing Markov dynamics and no meaningful notion of identity (violating (U1)). •Settings where contrastive labels are unavailable or untrustworthy (violating (U2)). • Adversarial environments that routinely induce discontinuous jumps in the relevant state, or where the generator explicitly targets the distortion and margin regimes (violating (U3) and (U4)). 7.6 Open Problems Several technical questions are left intentionally open. Sharper smoothness–distortion bounds. Section 3 uses a curvature-style bound ε ( L ) ≤ C′K ( λ ) L derived from smoothness penalties. Tightening this relationship—for example, via more precise control of sectional curvature or metric derivatives under Jacobian regularization—would sharpen the trade-off between distortion, path horizon, and regularization strength. Beyond finite-variance tails. Our main detection bound relies on Cantelli and finite second moments. In some domains, heavier tails may warrant sub-exponential or robust alternatives that still admit an interpretable decomposition and integrate cleanly into the error budget. 49 Learning the path family. We treat P ( L⋆ )as a designed object (e.g., epitope-restricted mutation paths, identity trajectories with bounded arc length). Learning path families from data—subject to smoothness and curvature constraints—could reduce reliance on hand-crafted structure while maintaining interpretability. Joint Davis geometry across modalities. HERALD and VIDAR each operate in a single embedding space. Many applications are naturally multi-modal (e.g., audio–visual, sequence– structure–clinical). Constructing joint manifolds and path families, with cross-modal distortion and margin guarantees, is an open challenge. Adaptive error budgets. Section 6 treats Egeom , Elink , ξ , ζ , and δindep as periodically reestimated but essentially static over a deployment window. Designing online estimators and control charts that adapt these quantities in real time without introducing look-ahead bias is an important systems problem. Benchmarks and community standards. The Davis framework suggests a style of evaluation centered on distortion audits, error budgets, and abstention behavior. Curating benchmark suites and reporting standards that exercise these aspects across domains—for example, Davis-style tasks for sequences, video, robotics, and audio—is a prerequisite for systematic comparison and progress. 7.7 Reproducibility, Dossiers, and Deployment Finally, translating Davis theory into practice requires structured documentation. Building on the operational protocols in Section 6, we recommend that any deployed Davis system maintain a Davis dossier containing at least: (D1) a description of the manifold construction: encoder architecture, training objectives, regularization, and the resulting distortion audits (empirical ε(L)curves); (D2) the configuration map and margins: how c and h are defined, values of ( κhard, κsoft ), and evidence for non-vacuity; (D3) the path family and features: definition of P ( L⋆ ), feature design, and empirical separation ∆ S with confidence intervals; (D4) error-budget estimates (Egeom, Elink, ξ, ζ, δindep)with methodology and slice-wise breakdowns; (D5) the resulting lower bound from Theorem 2 for the intended deployment context, together with a vacuity threshold and abstention policy; and (D6) version identifiers (code, weights, data snapshots) and a change log so that bounds and failures can be traced to specific system states. The Davis framework is not a panacea for all temporal detection problems, but it offers a common pattern: construct a geometry in which distances and paths encode semantics; traverse that geometry using bounded-distortion surrogates; and attach explicit, compositional error budgets and abstention rules. HERALD and VIDAR show that this pattern can span domains as distant as viral evolution and deepfake forensics. We hope that future work will both stress-test and extend 50 this template, clarifying where Davis manifolds truly help, where they fail, and how they might form part of a broader toolkit for safe, interpretable, and auditable AI systems. References [1] Shun-Ichi Amari. Information Geometry and Its Applications. Springer, 2016. [2] Anastasios N Angelopoulos and Stephen Bates. A gentle introduction to conformal prediction and distribution-free uncertainty quantification. In arXiv preprint arXiv:2107.07511, 2021. [3] Sanjeev Arora, Mikhail Khodak, Nikunj Saunshi, et al. A theoretical analysis of contrastive unsupervised representation learning. arXiv preprint arXiv:1902.09229, 2019. [4] Peter L Bartlett and Shahar Mendelson. Rademacher and Gaussian complexities: Risk bounds and structural results. Journal of Machine Learning Research, 3:463–482, 2002. [5] Stéphane Boucheron, Gábor Lugosi, and Pascal Massart. Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press, 2013. [6] Francesco Paolo Cantelli. Sui confini della probabilità. Atti del Congresso Internazionale dei Matematici, 1928. [7] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. [8] François Chollet. Xception: Deep learning with depthwise separable convolutions. In CVPR, 2017. [9] C Chow. On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1):41–46, 1970. [10] Bee Rosa Davis. Herald: High-resolution early recognition of antigenic landscape divergence, 2025. [11] Bee Rosa Davis. VIDAR: Video identity dynamics and riemannian forensics for deepfake detection. Manuscript in preparation. Instantiates the Davis manifold framework for deepfake forensics via identity trajectories on Sd−1., 2025. [12] Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. ArcFace: Additive angular margin loss for deep face recognition. In CVPR, 2019. [13] Manfredo Perdigão do Carmo. Riemannian Geometry. Birkhäuser, 1992. [14] Alexei J Drummond, Marc A Suchard, Dong Xie, and Andrew Rambaut. Bayesian phylogenetics with BEAUti and the BEAST 1.7. Molecular Biology and Evolution, 29(8):1969–1973, 2012. [15] Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. Calibrating noise to sensitivity in private data analysis. Theory of Cryptography Conference, pages 265–284, 2006. [16] Ran El-Yaniv and Yair Wiener. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11:1605–1641, 2010. 51 [17] William Feller. An Introduction to Probability Theory and Its Applications, volume 1. Wiley, 1968. [18] Octavian-Eugen Ganea, Gary Bécigneul, and Thomas Hofmann. Hyperbolic neural networks. In NeurIPS, 2018. [19] Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In NeurIPS, 2017. [20] Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In ICML, 2017. [21] James Hadfield, Colin Megill, Sidney M Bell, et al. Nextstrain: Real-time tracking of pathogen evolution. Bioinformatics, 34(23):4121–4123, 2018. [22] Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020. [23] Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In ICLR, 2017. [24] Zhiwu Huang and Luc Van Gool. Riemannian manifold learning. In IEEE TPAMI, volume 30, 2017. [25] Pang Wei Koh, Shiori Sagawa, Henrik Marklund, et al. WILDS: A benchmark of in-the-wild distribution shifts. ICML, 2021. [26] Brian Kulis et al. Metric learning: A survey. Foundations and Trends in Machine Learning, 5(4):287–364, 2013. [27] Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-DF: A large-scale challenging dataset for DeepFake forensics. In CVPR, 2020. [28] Shiyu Liang, Yixuan Li, and R Srikant. Enhancing the reliability of out-of-distribution image detection in neural networks. In ICLR, 2018. [29] Weiyang Liu, Yandong Wen, Zhiding Yu, et al. SphereFace: Deep hypersphere embedding for face recognition. In CVPR, 2017. [30] Mahdi Pakdaman Naeini, Gregory F Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In AAAI, 2015. [31] Maximilian Nickel and Douwe Kiela. Poincaré embeddings for learning hierarchical representations. In NeurIPS, 2017. [32] Xavier Pennec, Pierre Fillard, and Nicholas Ayache. A Riemannian framework for tensor computing. International Journal of Computer Vision, 66(1):41–66, 2006. [33] John Platt. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers, pages 61–74, 1999. 52 [34] Andrew Rambaut, Tommy T Lam, Luiz Max Carvalho, and Oliver G Pybus. Exploring the temporal structure of heterochronous sequences using TempEst. Virus Evolution, 2(1), 2016. [35] Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers generalize to ImageNet? In ICML, 2019. [36] Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, et al. FaceForensics++: Learning to detect manipulated facial images. In ICCV, 2019. [37] Sam T Roweis and Lawrence K Saul. Nonlinear dimensionality reduction by locally linear embedding. Science, 290(5500):2323–2326, 2000. [38] Florian Schroff, Dmitry Kalenichenko, and James Philbin. FaceNet: A unified embedding for face recognition and clustering. In CVPR, 2015. [39] Glenn Shafer and Vladimir Vovk. A tutorial on conformal prediction. Journal of Machine Learning Research, 9:371–421, 2008. [40] Shai Shalev-Shwartz and Shai Ben-David. Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014. [41] Joshua B Tenenbaum, Vin De Silva, and John C Langford. A global geometric framework for nonlinear dimensionality reduction. Science, 290(5500):2319–2323, 2000. [42] Ruben Tolosana, Ruben Vera-Rodriguez, Julian Fierrez, et al. Deepfakes and beyond: A survey of face manipulation and fake detection. Information Fusion, 64:131–148, 2020. [43] Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu. Contrastive learning, multi-view redundancy, and linear models. In ALT, 2021. [44] Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. In arXiv preprint arXiv:1807.03748, 2018. [45] Vladimir N Vapnik. Statistical Learning Theory. Wiley, 1998. [46] Vladimir Vovk, Alex Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer, 2005. [47] Hao Wang, Yitong Wang, Zheng Zhou, et al. CosFace: Large margin cosine loss for deep face recognition. In CVPR, 2018. [48] Kilian Q Weinberger and Lawrence K Saul. Distance metric learning for large margin nearest neighbor classification. In NeurIPS, 2006. Appendix A. Proofs and Technical Details A.1. Proof of Lemma 1 (InfoNCE margin ⇒embedding separation) For completeness we restate the setting. Let ( x, x+, x− 1, . . . , x− K )be drawn from a contrastive sampling distribution DNCE in which ( x, x+ )share the same fine-grained identity/configuration and 53 each x− khas a different identity/configuration. Let sθ(x, x′) = −∥fθ(x)−fθ(x′)∥2 2 2τ2 denote the similarity score induced by the embedding fθwith temperature τ > 0, and let δθ(x, x′) = ∥fθ(x)−fθ(x′)∥2 be the corresponding Euclidean distance. The InfoNCE population risk is LNCE(θ) = E(x,x+,x− 1,...,x− K)∼DNCE "−log exp{sθ(x, x+)} exp{sθ(x, x+)}+PK k=1 exp{sθ(x, x− k)}#. Define the effective score margin meff(θ) := Ehsθ(x, x+)−1 K K X k=1 sθ(x, x− k)i and the distance gap ∆emb := Eδθ(x, x−)Ydiff−Eδθ(x, x+)Ysame. Step 1: InfoNCE risk controls the score margin. For a single tuple, write ∆k:= sθ(x, x− k)−sθ(x, x+), k = 1, . . . , K. Since 1 + Pke∆k≥Pke∆k and logPkeuk≥log K + 1 KPkuk (AM–GM plus monotonicity of log), we obtain log 1 + K X k=1 e∆k≥log K+1 K K X k=1 ∆k. Taking expectations over DNCE and using the definition of meff(θ)gives LNCE(θ) = Ehlog 1 + Pke∆ki≥log K−meff(θ), so meff(θ)≥log K−LNCE(θ).(12) Step 2: Population vs. empirical risk. Let b LNCE ( θ )denote the empirical InfoNCE risk on n i.i.d. samples. Assume a standard uniform-convergence bound for the loss class: there exists a complexity term Cand universal constant c>0such that, with probability at least 1−γ, LNCE(θ)≤b LNCE(θ)+csC+ log(1/γ) n. 54 Combining with (12) gives, for any empirical minimiser ˆ θ, meff(ˆ θ)≥log K−b LNCE(ˆ θ)−csC+ log(1/γ) n.(13) Step 3: Lipschitz transfer from scores to distances. On the bounded operational domain δ∈ [0 , D ](the in-distribution region for identity/configuration pairs), the score s ( δ ) = −δ2/ (2 τ2 )is Leff-Lipschitz in δwith Leff = sup δ∈[0,D] |s′(δ)|=D τ2. Thus, for any two distances u, v ∈[0, D], |s(u)−s(v)| ≤ Leff |u−v|. Apply this with u = δθ ( x, x+ )and v = 1 KPkδθ ( x, x− k ), take expectations, and rearrange to obtain ∆emb ≥meff(θ) Leff . Combining with (13) yields the claimed margin-to-separation relation from Lemma 1, up to the usual OpC/ngeneralization slack. Step 4: Symmetric label noise. Under symmetric label noise with rate η < 1 2 (assumption A4), the effective margin contracts by a factor (1 − 2 η ); this simply rescales the right-hand side of (13) and yields the O(η)term in the lemma statement. A.2. Explicit constants in the margin corollary On the bounded domain δ∈[0, D]we have Leff =D τ2=⇒1 Leff =τ2 D. Writing a := τ2/D and absorbing the uniform-convergence constant into b := c/Leff = cτ2/D , the inequality in Lemma 1 can be expressed in the simplified form used in Section 3: ∆emb ≥alog K−bsC+ log(1/γ) n−O(η), with a, b depending only on the score temperature τ , the operational diameter D , and the capacity term C. 55 A.3. Cantelli bound for Davis systems (details for Proposition 3) We recall the quantities from Section 4. Let S be the scalar detection statistic, Y∈ { 0 , 1 } the class label (0 = benign, 1 = event), and µy=E[S|Y=y], σ2 y= Var(S|Y=y),∆S=µ1−µ0>0. Define the midpoint threshold s⋆=µ0+µ1 2=µ0+∆S 2, and the class priors πy=P(Y=y). Cantelli inequality. For a real-valued random variable X with mean µ and variance σ2<∞ , Cantelli’s (one-sided Chebyshev) inequality states that for any a > 0, P(X−µ≥a)≤σ2 σ2+a2. Apply this to S|Y= 0 with a= ∆S/2to obtain PS > s⋆|Y= 0=PS−µ0≥∆S/2|Y= 0≤c0:= σ2 0 σ2 0+ (∆S/2)2. Similarly for S|Y= 1, PS≤s⋆|Y= 1=Pµ1−S≥∆S/2|Y= 1≤c1:= σ2 1 σ2 1+ (∆S/2)2, so that PS > s⋆|Y= 1≥1−c1. Posterior correctness above the midpoint. The posterior probability of Y = 1 given that S exceeds s⋆is P(Y=1|S > s⋆) = π1P(S > s⋆|Y= 1) π0P(S > s⋆|Y= 0) + π1P(S > s⋆|Y= 1). Using P(S > s⋆|Y= 0) ≤c0and P(S > s⋆|Y= 1) ≥1−c1gives the population bound P(Y=1|S > s⋆)≥π1(1 −c1) π0c0+π1(1 −c1).(14) Finite-sample estimation. In practice, µy , σ2 y and πy are replaced by empirical estimates ˆµy , ˆσ2 y , ˆπy computed on a validation set. Standard concentration arguments (e.g., for sub-Exponential or finite-variance variables) imply that with high probability the plug-in version of (14) deviates from its population value by at most an Op1/n term. In Proposition 3 we bundle these effects into the single estimation slack εest, yielding the stated lower bound. 56 An equivalent “odds form” sometimes useful numerically is P(Y=1|S > s⋆) P(Y=0|S > s⋆)≥π1 π0 1−c1 c0 , from which (14) follows by algebra. A.4. Independence slack, union bound, and vacuity (supporting material for Theorem 2) Recall the four component error probabilities from Section 6: Egeom, Elink, ξ, ζ, corresponding to geometry failure, linkage/feature failure, calibration error and abstention failure, respectively. Write Ggeom ={geometry behaves as intended}, and similarly define Glink , Gcal , Gabst as the complements of the corresponding error events, so that P(Gc geom)=Egeom,P(Gc link)=Elink,P(Gc cal)=ξ, P(Gc abst)=ζ. Let G=Ggeom ∩Glink ∩Gcal ∩Gabst denote the global “good” event on which all components behave. Approximate independence and slack. If the four error events were independent, we would have P(G) = Y i∈{geom,link,cal,abst} (1 −Ei) =: Pindep. In practice, errors may be positively correlated (e.g., heavy OOD shifts that simultaneously break geometry and calibration). We therefore introduce a non-negative independence slack δindep ≥ 0 defined implicitly by P(G)≥Pindep −δindep.(15) When independence (or approximate independence) is empirically plausible, δindep is small; when errors co-occur frequently, δindep can be large and the product Pindep is overly optimistic. In Section 7.6 we estimate δindep from validation data via ˆεjoint =Pat least one error occurs,ˆεmult = 1 −Y i (1 −ˆ Ei), and set ˆ δindep = max{ 0 ,ˆεjoint −ˆεmult} ; this ensures (15) holds with high probability up to sampling noise. 57