COMPASS: A Web-Based COMPosite Activity Scoring System to Navigate Health and Disease Through Deterministic Digital Biomarkers
Abstract
Quantifying pathway activation in absolute, reproducible terms is central to systems biology and precision medicine. COMPASS (COMPosite Activity Scoring System) provides a deterministic, ontology-free framework that defines gene-specific thresholds and aggregates standardized deviations into composite digital biomarkers of pathway activation. By transforming raw expression data into interpretable, data-anchored metrics, COMPASS delivers a transparent, reproducible, and scalable means to integrate, align, and benchmark ‘humanness’ across digital and experimental model systems, bringing computational rigor, regulatory reproducibility, and experimental relevance at the fingertips of every biologist.
Full text
1 | Page Title 1 COMPASS: A Web-Based COMPosite Activity Scoring System to 2 Navigate Health and Disease Through Deterministic Digital Biomarkers 3 Authors 4 Saptarshi Sinha1, 3* and Pradipta Ghosh1-3* 5 Departments of 1Cellular and Molecular Medicine and 2Medicine, and 3Institute for Network Medicine, 6 University of California, San Diego, CA, 92093, USA. 7 *Correspondence to [email protected] (SS) or [email protected] (P.G). 8 Keywords (two to six) 9 Gene expression analysis · Data-driven thresholding · Composite activity scoring · Pathway 10 quantification · Systems biology · Precision medicine · Multi-omics integration · Reproducibility 11 12 13
2 | Page Abstract (71 words) 14 Quantifying pathway activation in absolute, reproducible terms is central to systems biology 15 and precision medicine. COMPASS (COMPosite Activity Scoring System) provides a 16 deterministic, ontology-free framework that defines gene-specific thresholds and aggregates 17 standardized deviations into composite digital biomarkers of pathway activation. By 18 transforming raw expression data into interpretable, data-anchored metrics, COMPASS delivers 19 a transparent, reproducible, and scalable means to integrate, align, and benchmark 20 ‘humanness’ across digital and experimental model systems, bringing computational rigor, 21 regulatory reproducibility, and experimental relevance at the fingertips of every biologist. 22
3 | Page Main Text (998 words) 23 Understanding how genes act collectively to define cellular states remains a central challenge in systems 24 biology. Yet, enrichment-based frameworks, such as GSEA1, GSVA2, ssGSEA3, PLAGE4, AUCell5, and VAM6 25 measure relative enrichment, not absolute activity. These approaches rely on evolving ontologies and 26 permutation statistics, which introduce blind spots and reproducibility failures (Fig. 1a-left). As pathway 27 definitions shift, scores fluctuate across datasets, platforms, and time, raising persistent concerns about the 28 reliability of ontology-dependent analyses in translational research7-9. 29 To overcome these limitations, we developed COMPASS (COMPosite Activity Scoring System), a 30 deterministic, transparent, and web-based framework that quantifies gene-set activity directly from expression 31 data (Fig. 1a-middle). COMPASS enables researchers, including those without coding expertise, to convert 32 gene-expression matrices into standardized, reproducible indices of pathway activation without random 33 permutations or curated hierarchies. It transforms biological variability into measurable logic, yielding stable, 34 interpretable, and transferable digital biomarkers of biological activity. 35 COMPASS operationalizes the Boolean principle that biological systems compute through thresholds 36 and feedback, capturing the continuum of cellular decisions from health to disease [expanded here10]. While 37 Boolean frameworks like BoNE11 have demonstrated reproducibility across species and diseases [reviewed 38 here10,12], enabling objective benchmarking of digital biomarkers, assessing the “humanness” of animal models 39 and NAMs, and tracking disease trajectories or therapeutic reversals10, their computational complexity has 40 limited accessibility. By embedding this logic in an intuitive web platform, COMPASS democratizes Boolean 41 logic, enabling users to convert raw expression data into deterministic digital biomarkers with a few clicks, 42 bridging foundational biological reasoning, computational rigor, regulatory reproducibility, and clinical and/or 43 experimental relevance (Fig. 1a-right). 44 Mathematical transparency, the defining feature of COMPASS, is achieved through a deterministic 45 workflow of three sequential steps: thresholding, standardization, and aggregation (Fig. 1b; see Online 46 Methods). Each step corresponds to a concrete, interpretable biological operation rather than an abstract 47 statistical transformation. In the thresholding step, COMPASS identifies for every gene its intrinsic expression 48 boundary—the inflection where its state transitions from “off” to “on.” This data-driven threshold, determined 49 automatically by adaptive segmentation algorithms, pinpoints where biological probability collapses into 50 commitment. The standardization step aligns genes of varying dynamic range onto a unified scale, 51 transforming noisy expression profiles into interpretable distance measures that indicate how far each gene 52 lies from its activation boundary. Finally, the aggregation step integrates these standardized deviations across 53 genes within a defined set, applying direction-specific weights to capture the balance of opposing influences. 54 The resulting composite score represents the net activation of a pathway in each sample. 55
4 | Page In practice, these mathematically transparent operations occur entirely in the background (Fig. 1c). 56 Users simply upload a tab-delimited expression matrix (txt, tsv, csv), define custom gene signatures and 57 sample groups, and initiate analysis through an intuitive point-and-click interface. No pathway databases, 58 metadata, or programming are required. The backend executes all computations automatically, generating 59 downloadable boxplots and ROC–AUC curves that benchmark pathway activation against clinical or biological 60 groups. Every run from identical input yields identical output, ensuring the reproducibility required for 61 regulatory-grade science. 62 The most important output is the framework’s ability to transform qualitative gene signatures into 63 quantitative digital biomarkers. By linking each computational step to a measurable biological principle—64 expression, threshold, deviation, or directionality—COMPASS replaces randomization with determinism. This 65 direct mapping of data to decision logic defines its mathematical transparency: users can trace each result to 66 its underlying threshold and deviation, eliminating both statistical opacity and annotation bias. Because 67 thresholds are derived directly from the data rather than from evolving ontologies, COMPASS is platform-68 agnostic and transferable across tissues, cohorts, and species. It thus aligns biological logic independent of 69 normalization or batch effects, converting pathway analysis from a statistical comparison into a mechanistic 70 measurement system. It harmonizes results across studies and captures subtle transitions such as health-to-71 disease progression, differentiation dynamics, or treatment-induced reversibility with reproducible precision. 72 Table 1 highlights the distinctive features of COMPASS relative to enrichment-based and latent-73 variable methods. Unlike GSEA, which compares only two discrete conditions, COMPASS models graded, 74 multi-class trajectories. Datasets spanning disease evolution, therapy response, or circadian cycles can be 75 analyzed along continuous activation axes, allowing longitudinal tracking of biological programs such as 76 epithelial–mesenchymal transition or immune exhaustion. By computing absolute rather than relative activity, 77 COMPASS provides quantitative grounding for digital biomarkers that trace biological and therapeutic 78 trajectories. But perhaps the most distinctive feature of COMPASS is its conceptual grounding in cellular 79 decision-making, which in turn improves reasoning and explainability of the conclusions drawn. COMPASS 80 extends Boolean logic into analog quantification. Whereas Boolean frameworks such as BoNE11 identifies 81 digital “if–then” rules that govern cellular decisions, COMPASS quantifies the transitions as populations 82 traverse those rules, thus integrating discrete and continuous representations of biology. Together, these 83 approaches operationalize the Five Rules of Biological Logic10—that life computes through thresholds, 84 maintains fidelity through feedback, and progresses along reversible, measurable continua, thereby providing a 85 quantitative bridge between binary logic and analog biology (see Logic of Life10). Unlike enrichment-based 86 methods (GSEA1, GSVA2, ssGSEA3, PLAGE4, AUCell5, and VAM6) that depend on ranked gene lists and 87 random permutations, COMPASS deterministically measures biological activity without reference to predefined 88 ontologies. Each variable—expression, threshold, deviation, or directionality—corresponds to an observable 89 biological quantity, not a statistical abstraction. In contrast to matrix-factorization or deep-learning models such 90
5 | Page as PCA, ICA, NMF, MOFA⁺, or variational autoencoders13,14), COMPASS avoids latent dimensions, retaining 91 interpretability while eliminating stochastic variation. 92 Because thresholds are derived directly from data, COMPASS is platform-agnostic and transferable 93 across tissues, cohorts, and species, aligning biological logic independent of normalization or batch effects. 94 This transparency turns COMPASS from a statistical model into a mechanistic measurement system, that 95 produces deterministic, reproducible metrics that can serve as digital biomarkers for disease state, drug 96 response, or biological fidelity. It also enables cross-study harmonization, generating quantitative measures 97 that capture subtle biological transitions with reproducible precision. 98 Unlike GSEA, which can only compare two discrete conditions, COMPASS models graded, multi-99 class trajectories. If datasets include samples representing the evolution from health to disease, normal to 1 00 carcinoma, or response to resistance, or even circadian rhythms, COMPASS can capture these transitions, 1 01 enabling longitudinal tracking of phenotypes. This makes COMPASS uniquely capable of quantifying 1 02 therapeutic trajectories (e.g., during biologic therapy) or disease progression (e.g., rising stemness along the 1 03 adenoma–carcinoma continuum). Thus, COMPASS transforms pathway activity into a universal digital 1 04 biomarker, enabling objective, reproducible alignment of human, organoid, and animal models within a single 1 05 logical framework for systems medicine. 1 06 Beyond biomarker discovery, COMPASS supports diverse use cases (Fig. 2): computing composite 1 07 activity scores predictive of outcomes15; tracking drug response or toxicity16; mapping continuum states such 1 08 as immune activation or differentiation11,15,17,18; and benchmarking organoids, animal models, or New Approach 1 09 Methodologies (NAMs) against human cohorts to quantify humanness and biological relevance15,16,19-23. 1 10 Through these capabilities, COMPASS bridges computation and bench biology, offering a scalable and user1 11 friendly platform for digital biomarker discovery and systems-level disease modeling10. Its accessibility allows 1 12 any biologist to navigate data using deterministic metrics rather than probabilistic inference, bringing the 1 13 mathematical logic of life to practical use in translational and regulatory settings (Fig. 2). 1 14 1 15 COMPASS has a few limitations. Although deterministic and interpretable, COMPASS depends on well1 16 normalized transcriptomic data and assumes unimodal gene distributions, potentially limiting performance in 1 17 sparse or noisy single-cell datasets. Its current weighting scheme treats all genes equally, without accounting for 1 18 causal topology or network hierarchy. Future versions will integrate context-dependent weighting based on 1 19 Boolean network connectivity to strengthen causal inference. At present, COMPASS models transcriptome-level 1 20 activity; extending its logic to proteomic, metabolomic, and epigenomic data will test its generalizability. 1 21 Integration with longitudinal and spatial datasets may enable trajectory mapping of therapeutic responses and 1 22 disease progression, while embedding COMPASS in interactive dashboards could facilitate real-time digital 1 23 biomarker tracking for translational and regulatory use. 1 24
6 | Page Online Methods: 1 25 Overview 1 26 The COMPASS (COMPosite Activity Scoring System) workflow transforms a raw gene-expression matrix into 1 27 deterministic, quantitative measures of pathway activity through a three-step process—thresholding, 1 28 standardization, and aggregation (Fig. 1b–c). Unlike enrichment-based methods that rely on predefined 1 29 ontologies or random permutations, COMPASS derives gene-specific thresholds directly from expression data, 1 30 converts deviations from these thresholds into standardized metrics, and integrates them to compute 1 31 composite activity scores. 1 32 All computations are closed form and deterministic: identical input produces identical output, ensuring 1 33 reproducibility, auditability, and mathematical transparency. 1 34 1 35 1. Threshold determination (biological decision boundaries) 1 36 At the core of COMPASS lies the StepMiner24 algorithm, an adaptive algorithm originally developed to identify 1 37 sharp transcriptional transitions from noisy expression data. For each gene g, the expression values across 1 38 samples are sorted, and a step function is fit to this ordered distribution. The transition index (Tg) represents 1 39 the inflection point that minimizes within-group variance—corresponding to the biological boundary between 1 40 low (“off”) and high (“on”) expression states. 1 41 For each gene 𝑔 and sample 𝑖, COMPASS computes a normalized deviation from the threshold as: 1 42 𝐸,=𝜇,,𝑖<𝑇 𝜇,,𝑖≥𝑇 1 43 Where 𝑇 denotes the transition index minimizing within-group variance. This threshold defines the binary 1 44 states “low (0)” and “high (1)” expression. This adaptive segmentation identifies the intrinsic activation 1 45 threshold for each gene directly from data, without requiring any external annotation or training. Conceptually, 1 46 this threshold identifies where biological probability collapses into commitment (“off” → “on”)10. In doing so, 1 47 thresholding reveals the cellular decision boundary. 1 48 To minimize stochastic fluctuations around this boundary, COMPASS applies a small confidence offset (σ) to 1 49 the threshold, defining the final decision boundary of each gene 𝑔 as: 1 50 𝑇∗ = 𝑇 + 𝜎 1 51 where, 𝜎=0.5 log 𝑈𝑛𝑖𝑡𝑠. This conservative shift in threshold ensures that only values confidently above or 1 52 below the step contribute to the score. 1 53
7 | Page 2. Standardization (normalization across genes): 1 54 Once thresholds are established, each gene’s deviation from its intrinsic boundary is standardized relative to 1 55 its variance, yielding a dimensionless measure of distance from the decision point: For every gene 𝑔 and 1 56 sample 𝑖, the deviation from this intrinsic threshold is standardized as: 1 57 𝑍,= 𝐸,− 𝑇 𝜎 1 58 where 𝐸, is the observed expression and 𝜎 the gene-specific standard deviation. 𝑍, thus measures the 1 59 distance from the decision boundary11, producing a hybrid metric: discrete when categorical precision is 1 60 needed (i.e., digital) and continuous when population gradients matter (i.e., analog). The standardized 1 61 deviation used in COMPASS is then calculated as: 1 62 𝑍,= 𝐸,− 𝑇∗ 3𝜎 1 63 Dividing by 3𝜎 (instead of 1×𝜎) compresses extreme outliers and brings all genes to a comparable dynamic 1 64 scale. This preserves biological directionality while reducing threshold-adjacent noise. 1 65 This standardization step converts heterogeneous expression profiles into a common coordinate system in 1 66 which each Z ₍ g,i ₎ representing how strongly a gene deviates from its activation threshold. 1 67 1 68 3. Composite aggregation (pathway-level integration) 1 69 For any gene-set, 𝐺=𝑔,𝑔,…,𝑔, COMPASS computes the composite activity score for each sample i 1 70 as the weighted mean of standardized deviations: 1 71 𝐶=1 𝑚 𝑤 𝑍, 1 72 where weights wⱼ encode gene directionality (+1 = activation, −1 = repression). The resulting composite score 1 73 (Cᵢ) reflects the net balance of activation and repression, providing an interpretable, quantitative index of 1 74 pathway activity. A higher positive value indicates dominance of activation, while a negative value indicates 1 75 repression. Published use cases include proversus anti-inflammatory macrophage states18, differentiation 1 76 versus stemness15, or metabolic activation versus inhibition, within a single continuous framework. 1 77 This closed-form calculation eliminates the randomness inherent to permutation-based enrichment 1 78 methods. It also ensures that each component—expression, threshold, deviation, and directionality— 1 79 corresponds to an observable biological quantity rather than a statistical abstraction. 1 80 1 81
8 | Page 4. Benchmarking and performance metrics 1 82 To assess biological and clinical relevance, COMPASS automatically benchmarks the resulting composite 1 83 activity scores using ROC–AUC analysis, comparing the separability of defined sample groups (for example, 1 84 disease vs control or responder vs non-responder). 1 85 The Area Under the Curve (AUC) is computed as: 1 86 𝐴𝑈𝐶 =𝑃 (𝐶> 𝐶) 1 87 where 𝐶 and 𝐶 denote the composite scores of state 𝐴 (e.g., treated or disease) and state 𝐵 (e.g., control or 1 88 healthy) samples, respectively. This provides a universal, model-agnostic measure of discriminatory power that 1 89 reflects the internal biological coherence of each gene-set’s activity pattern and serves as a statistical 1 90 summary the same. 1 91 All calculations are performed automatically within the web interface. Users can download publication1 92 ready boxplots and ROC curves along with associated t, F, and p statistics. The absence of stochastic 1 93 resampling ensures identical results across runs, datasets, and computational environments. 1 94 1 95 5. Implementation and accessibility 1 96 COMPASS is implemented as a web-based analytical engine with an intuitive graphical interface (Fig. 1c). All 1 97 analyses are performed locally within the web interface to preserve data privacy. The user interface will be 1 98 made available through the Institute for Network Medicine, UC San Diego (Link will be made public upon 1 99 acceptance). Users upload a tabor comma-delimited matrix (txt, tsv, csv), where genes occupy rows and 2 00 samples occupy columns, along with optional metadata to define comparison groups. Gene-set definitions are 2 01 fully customizable, with flexibility to assign directionality weights (+/−). Once data are uploaded, all backend 2 02 computations—including thresholding, normalization, and statistical validation—are executed automatically. 2 03 Outputs include: 2 04 • Boxplots summarizing pathway activity across user-defined groups, accompanied by pairwise t-test 2 05 statistics. 2 06 • ROC–AUC plots displaying discrimination between states with associated F and p values. 2 07 • Composite activity tables containing standardized Z-scores and composite scores for each pathway 2 08 and sample. 2 09 All results are downloadable in CSV or PDF format. The deterministic nature of COMPASS ensures that the 2 10 same input always yields identical output, enabling full traceability and reproducibility for regulatory 2 11 submissions or multi-site validation. 2 12
9 | Page In terms of accessibility, unlike conventional computational frameworks, COMPASS is designed to be intuitive 2 13 and accessible to both computational and non-computational users: 2 14 • Minimal Input Requirement: Only a gene expression matrix is required; no need for pathway 2 15 databases, sample identifier, or clinical annotations. 2 16 • Flexible Group Comparisons: Users can define and compare any number of sample groups directly 2 17 from their uploaded data (e.g., WT vs KO, responders vs non-responders, multiple treatment conditions). 2 18 • No Coding or Statistical Expertise Required: All statistical analyses, such as, composite score 2 19 calculation, group-wise comparisons, ROC–AUC, p-values are performed automatically. 2 20 • Interactive and User-Friendly: The interface enables point-and-click selection of genes, pathways, and 2 21 groups, generating publication-ready figures without manual scripting (Fig. 1c). 2 22 • Scalable and Reproducible: Every step is deterministic, ensuring identical results across runs which is 2 23 critical for regulatory submissions and clinical validation. 2 24 Reproducibility is guaranteed by the deterministic nature of the algorithm: identical inputs produce identical 2 25 numerical and graphical outputs, ensuring that results can be independently verified across laboratories or 2 26 regulatory reviews. 2 27 2 28 Summary 2 29 Through these three core operations, i.e., thresholding to define biological boundaries, standardization to 2 30 normalize deviations, and aggregation to integrate weighted gene effects, COMPASS converts raw gene2 31 expression matrices into quantitative, deterministic indices of pathway activity. 2 32 2 33 By linking mathematical transparency to biological interpretability, COMPASS enables objective, reproducible 2 34 quantification of cellular logic across model systems and clinical cohorts. 2 35 2 36
16 | Page 18 Ghosh, P. et al. Machine learning identifies signatures of macrophage reactivity and tolerance that predict disease outcomes. EBioMedicine 94, 104719 (2023). https://doi.org/10.1016/j.ebiom.2023.104719 19 Tindle, C. et al. A living organoid biobank of patients with Crohn's disease reveals molecular subtypes for personalized therapeutics. Cell Rep Med 5, 101748 (2024). https://doi.org/10.1016/j.xcrm.2024.101748 20 Sinha, S. et al. COVID-19 lung disease shares driver AT2 cytopathic features with Idiopathic pulmonary fibrosis. EBioMedicine 82, 104185 (2022). https://doi.org/10.1016/j.ebiom.2022.104185 21 Vo, D. T. et al. SPT6 loss permits the transdifferentiation of keratinocytes into an intestinal fate that resembles Barrett's metaplasia. iScience 24, 103121 (2021). https://doi.org/10.1016/j.isci.2021.103121 22 Penrose, H. M. et al. A Living Organoid Biobank of Crohn's Disease Patients Reveals Distinct Clinical Correlates of Molecular Subtypes of Disease. medRxiv (2025). https://doi.org/10.1101/2025.04.01.25325058 23 Tindle, C. et al. Adult stem cell-derived complete lung organoid models emulate lung disease in COVID19. Elife 10 (2021). https://doi.org/10.7554/eLife.66417 24 Sahoo, D., Dill, D. L., Tibshirani, R. & Plevritis, S. K. Extracting binary signals from microarray timecourse data. Nucleic Acids Res 35, 3705-3712 (2007). https://doi.org/10.1093/nar/gkm284