1 | Page PICASSO™: A Human-First Framework for Federally Mandated New Approach Methodologies Authors: Pradipta Ghosh1-4, Saptarshi Sinha1, 3, Courtney Tindle1, 2, Mark M. Garner5, Hans Clevers6-9 Affiliations: 1Department of Cellular and Molecular Medicine, University of California San Diego, University of California San Diego, La Jolla, CA, USA. 2UC San Diego HUMANOID™ Center, University of California San Diego, La Jolla, CA, USA. 3Institute for Network Medicine, University of California San Diego, University of California San Diego, La Jolla, CA, USA. 4Department of Medicine, University of California San Diego, University of California San Diego, La Jolla, CA, USA. 5Agilent Technologies, Mississauga, ON Canada. 6Hubrecht Institute, Royal Netherlands Academy of Arts and Sciences (KNAW) and UMC Utrecht, the Netherlands. 7Oncode Institute, the Netherlands. 8Pharma, Research and Early Development (pRED) of F. Hoffmann-La Roche Ltd., Basel, Switzerland. 9The Princess Máxima Center for Pediatric Oncology, Utrecht, the Netherlands. * Correspondence to:
[email protected] (P.G) ABSTRACT The bridge to humans is where drug discovery collapses. New Approach Methodologies (NAMs), powered by organoids, microphysiological systems, and AI/ML, are emerging as the foundation of a human-first drug discovery paradigm. Yet without standards, most NAMs remain fragmented, irreproducible, and detached from Phase 3–level rigor: diversity, reproducibility, and clinically anchored endpoints. PICASSO™ (Phenotype-Informed Clinical Abstraction for Systematic Simulation and Outcomes) closes this gap by algorithmically anchoring NAMs to large, diverse patient cohorts. By abstracting only the essential disease-driving features that align with clinical outcomes, PICASSO™ strips away irrelevant complexity while enforcing reproducibility and clinical fidelity. Rather than replacing animal models, it redeploys them strategically when they reflect human disease mechanisms, sharpening their translational value. Compact in scale yet anchored to Phase 3–sized populations, NAMs operating within PICASSO™ become standardized, scalable, and capable of outcome-level predictions with regulatory-grade confidence. PICASSO™ is inclusive but discerning, transforming NAMs—whether algorithmic, human, or animal — from boutique pilot tools into engines for discovery, trial design, and precision therapeutics, ushering in a truly human-centered biomedical future that passes the “humanness” test.
2 | Page GRAPHIC ABSTRACT From an era of reproducibility crises rooted in animal model–based research (top; adapted from Declan Butler, Nature 20081) to a future that is end-to-end human (bottom). Through deliberate abstraction, PICASSO™ integrates computational and systems biology approaches to transform NAMs from fragmented pilot models into standardized, simple, scalable, and human-first frameworks. Anchored to large, diverse patient cohorts, and to foundational models of cellular intelligence, NAMs operating within PICASSO™ are capable of predicting clinical outcomes with regulatory-grade confidence.
3 | Page In Brief: • Regulatory momentum: FDA Modernization Act and NIH initiatives are accelerating adoption of NAMs as credible alternatives to animal testing. • Ecosystem reality: Drug discovery is now a connected continuum; NAMs must function seamlessly across academia, biotech, pharma, and regulators. • Scientific promise vs. validation gap: Organoids, MPS, and AI models deliver human relevance but demand rigorous integration, standardization and regulatory alignment, particularly on SOPs, measurement parameters and evaluation criteria. • Academic imperative: Reproducibility and translational benchmarking are essential for credibility. • Biotech imperative: Platform builders must design with regulatory end-use in mind, not just innovation, to bridge discovery and approval. Highlights: • FDA and NIH are no longer asking for NAMs—they’re requiring them. • Organoids, MPS, and AI are proposed as replacements for animal models; not all scientists agree. • The biggest barrier(s) are not science; it’s the lack of validation consensus and regulatory clarity. • Academia must move beyond one-off “model papers” and deliver reproducibility at scale. • Biotech that does not design NAM platforms with regulatory end-use in mind will be left behind. • Real progress hinges on breaking silos; academia, industry, and regulators must co-develop the playbook. • Only NAMs (algorithmic, human and animal models) that may rely on smaller n for humanness and diversity than Phase 3 trials yet remain anchored in Phase 3–sized populations through abstraction—and that can recapitulate universal, measurable, scalable, and regulator-ready patterns—will define the future of translational science.
4 | Page Translational modeling in biomedicine has long struggled to bridge experimental predictions with clinical outcomes, fueling costly failures and reproducibility crises. The federally mandated remedy—New Approach Methodologies (NAMs)—is no longer optional. NAMs are rapidly becoming central to the future of drug development and chemical safety assessment. Recent accelerants—the FDA Modernization Act (2022)2, and the very recent FDA ‘Roadmap to Reducing Animal Testing in Preclinical Safety Studies’, NIH’s strategic initiatives, and the EPA’s roadmap for reducing animal testing—have aligned with a wave of public and policy pressures. Shortly after the FDA roadmap, the NIH announced their own initiative (OPRIVA, the Office of Research Innovation, Validation, and Application) to promote the use of NAMs to replace animal research. At the same time, breakthroughs in organoids, microphysiological systems (MPS), and AI/ML make it possible to capture human biology at rapidly increasing resolution. The convergence of need and capability sets the stage for a transformation. The growth of research using NAMs has been extraordinary: while in 1981–1984 fewer than 0.05% of NIH grants supported animal alternatives, by 2024 that number had surged to nearly 8%—a >150-fold increase that signals NAMs’ rise from fringe concept to mainstream engine of biomedical discovery. 1. Definition of New Approach Methodologies (NAMs) NAMs is a collective term used by regulatory agencies and the scientific community to describe human-relevant approaches that can reduce, refine, or replace the use of animals in biomedical research, toxicology, and drug development (Box 1). Unlike traditional models that rely on evolutionary proxies, NAMs aim to directly interrogate human biology with tools that are mechanistically precise, scalable, and computationally integrative. They are not defined by a single technology, but by a philosophy: end-to-end human systems that can credibly inform regulatory, clinical, and translational decision-making. The essential characteristics of a NAM are (Figure 1): 1. End-to-end human — the biological materials are, to the extent possible, derived from humans (e.g., primary cells, iPSCs, organoids, engineered tissues, or human-relevant biochemical systems). 2. Cohort-level measurement — the platforms enable in-depth, reproducible measurement of clinically relevant phenotypes across statistically meaningful sample sizes. 3. Computational integration — advanced data processing, analysis, and integration are embedded, using artificial intelligence (AI) and systems modeling to capture emergent properties and predict outcomes. BOX 1: Defining New Approach Methodologies (NAMs) Core characteristics of NAMs: •End-to-end human: Biological materials derived from humans (primary cells, iPSCs, organoids, engineered tissues, human-relevant biochemical systems). •Cohort-level measurement: Ability to deliver reproducible, clinically relevant phenotypes across statistically meaningful sample sizes. •Computational integration: Built-in use of AI, ML, and systems modeling to analyze and predict emergent outcomes. Modalities included under NAMs: •3D culture systems: Organoids, tumoroids, organ-on-chip. •Cell-free (in chemico) assays: Biochemical or molecular interaction platforms. •In silico models: Algorithmic and computational simulations of disease and therapeutic response. Key takeaway: NAMs are not a technology class, but a human-centric philosophy for building regulator-ready models. Currently, NAMs are deployed only in pilots, not in pivotal programs, because they lack the Phase 3–grade metrics that anchor regulatory approval.
5 | Page While many NAMs are associated with 3D culture systems such as organoids, tumoroids, or organ-on-chip models, the category is broader. It also encompasses cell-free (in chemico) systems that test biochemical or molecular interactions directly, as well as algorithmic (in silico) models that simulate disease processes and therapeutic responses. The unifying principle is not form, but function: NAMs provide human-centric, mechanistically faithful abstractions that can not only serve as foundational vessels for translation but also scale toward regulatory-grade decision-making. 2. The Current State of NAMs: A Brief History The movement toward human-relevant models began more than four decades ago at Johns Hopkins University with the launch of the Center for Alternatives to Animal Testing (CAAT) in 1981. Its mission, the 3Rs: reduction, refinement, and replacement of animal research, marked the first coordinated global effort to rethink preclinical science. At that time, modeling human biology outside the body was barely conceivable. Three decades later, breakthroughs in adult stem-cell culture redefined what was possible. It is noteworthy that CAAT was formed at a time when human model research was not a legitimate science. 3D epithelial organoid approaches were pioneered by Mina Bissell3. We improved this Matrigel-driven organoid culturing approach4, first in mice and then in humans, such that tissue stem-cell–derived epithelial structures were genetically stable, expandable, cryopreservable, and could faithfully recapitulate native tissue architecture. Simultaneously, Yoshiki Sasai developed approaches to create central nervous system organoids (e.g., brain, retina) by directed development of pluripotent stem cells, such as iPSCs5. The commercial and academic momentum that followed was explosive. New media formulations, scaffolds, and coculture systems transformed organoids into versatile new approach methodologies (NAMs) for disease modeling, efficacy testing, and personalized medicine. Today, NAMs-- encompassing organoids, microphysiological systems, in chemico, and in silico models—have evolved into a core pillar of human-relevant science. Organoids, organ-on-chip systems, high-content imaging, and AI/ML platforms have transformed the technical landscape, while OECD frameworks, FDA qualification pathways, and NIH crossagency programs have begun to set standards. The FDA Modernization Act 2.0 eliminated the requirement for animal testing before IND submission, signaling openness to alternatives. The FDA’s 2025 roadmap to minimize animal aims to go further, mandating NAM deployment in therapeutic monoclonal antibody programs and incentivizing sponsors who submit NAM data in parallel with animal data by offering “regulatory relief,” such as smaller required animal cohorts or reduced primate toxicology studies. Pilot programs such as ISTAND [Innovative Science and Technology Approaches for New Drugs] are designed to accelerate qualification pathways, while NIH’s NCATS is driving cross-agency investments to expand validation pipelines. Yet, despite remarkable progress, the field still lacks a universal standard for benchmarking fidelity, scalability, and mechanistic predictivity; nor does it have Phase 3–grade metrics that anchor regulatory approval, a gap the PICASSO framework aims to fill. NAMs remain incomplete, inconsistently validated, rarely have detailed demographic information about the source person, and poorly standardized, risking producing insights that may be elegant in design but fragile in translation. There is a gathering consensus that without population-scale datasets, reproducibility across labs, or accepted risk frameworks, NAMs cannot yet serve as the bedrock of regulatory confidence. Figure 1. Essential characteristics of ‘ideal’ NAMs. The workflow must begin [START] with Phase 3–sized human cohorts, which are mined to distill disease-driving essentials, stripping away superfluous complexity to reveal the processes and targets worth therapeutic
6 | Page reversal. These essentials are then modeled in (preferably) prospective “living” human biorepositories (organoids, MPS, primary cells, and patient-derived microbes). While these models use small-n cohorts, they must remain clinically annotated to allow outcome tracking and personalized testing. Deep phenotyping and multi-omics, anchored by objective gene-expression scores, generate actionable insights that form the center of the workflow. These insights must be iteratively refined through AI/ML training to sharpen relevance to clinical endpoints; ideally, they should also enable insilico prospective trials. In doing so, validation studies in small-n NAMs would always be anchored back to cohort-level conclusions to maintain relevance to clinical endpoints and ensure rigor and reproducibility. Anchoring must be achieved through continuous and objective evaluation of “model vs. disease match,” (using parameters that must avoid reproducibility crisis of annual upgrades that are subject to knowledge bias, e.g., gene ontology) with refinement at every stage—discovery, validation, and prediction. Validation platforms—including in chemico assays, ex vivo patientderived organoids (PDOs), microphysiological systems (MPS), and precision-cut tissue slices (PCTS), as well as in vivo animal studies—feed data back into computational systems models to project translational impact. Such a closed-loop, human-centric framework should enable the reproducible translation of foundational cellular and molecular insights into reliable, relevant and actionable therapeutic targets and companion biomarkers, thereby establishing NAMs as a true end-to-end foundation for human-first drug discovery.
7 | Page 3. The Inconvenient Truth of Tradeoffs: Complexity Without Relevance, Benchmarking Without Objectivity Biomedical science has been seduced by the false idol of complexity. For decades, we have equated “more” with “truth”: more cell types, more intricate matrices, more omics layers. But this dogma has not delivered translation, it has delivered fragility. Complexity without abstraction is not fidelity; it is noise disguised as rigor. Organoid systems are often engineered with every possible feature piled on, in the hope of mimicking native physiology. Yet every layer added without purpose is a liability, each one a potential source of irreproducibility and drift. Missing cues are not inherently fatal; their absence often sharpens the focus on what is essential. By contrast, bloated models obscure signal, risk replicating the irrelevant, and lull us into believing we have captured “the real thing.” The result is superficial biomimicry: models that look right but say little about disease drivers. A similar trap exists in gene expression analysis. For decades, our knowledge of a gene function has been treated as gold standard, even though only 8.2% of the human genome is functionally constrained, and only ~1.5% (estimated to be in the range of 19,587–20,2456-8) encodes proteins that are responsible for all cellular functions9. While AI has predicted many structures10-12, function remains largely uncharted-- 30% of all proteins are poorly understood and for much of the rest, knowledge is partial. Benchmarking NAMs with ontology-based (e.g., Gene Ontology13 [GO], Panther14, DAVID15) or pathway enrichment tools (Gene Set Enrichment Analysis16 [GSEA], Gene Set Variation Analysis17 [GSVA] and [ssGSEA18], PandaOmics19, Reactome20, KEGG20, Ingenuity [IPA; http://www.ingenuity.com], Enrichr21) risks irreproducibility because function-biased frameworks are incomplete, dynamic, and misleading, shifting longitudinally as blind spots are filled22. The way forward is disciplined restraint. Model only what matters—and where models are built, apply objective, unbiased, and mathematically precise frameworks that strip away noise and anchor findings to invariant, human-relevant drivers. This is not minimalism for its own sake; it is a deliberate strategy for fidelity, reproducibility, and clinical relevance. For example, applying SARS-CoV-2 to fetal lung organoids23 created models that lacked clinical premise (lungs of children were remarkably spared by the virus) and lacked relevance, assessed using objective gene signatures from the abstracted essentials of adult disease24. While adults developed distinct cytopathies which bore semblance to interstitial lung diseases25, children developed systemic immune storm that largely spared the lungs26. The principle of “less is more” must guide NAM design. Every exogenous addition risks artifacts that may suppress cell-autonomous niche signals, or misguided outputs unless parametrized by essentials abstracted from cohort-level comparisons. Figure 2. Multidimensional tradeoffs in NAM-based research and development. NAMs must balance realism with reduction, innovation with reproducibility.
8 | Page A. At the center, a balanced scale symbolizes the overarching theme of “tradeoffs.” Four domains surround it, each with distinct iconography and paired tensions: (1) Scientific Tradeoffs: - Complexity vs. Interpretability: Increasing cellular and matrix diversity may enhance physiological resemblance but introduces noise that can obscure mechanistic insights. - Fidelity vs. Scalability: High-fidelity organoids often require complex protocols, limiting throughput and translational utility. - Superficial Biomimicry vs. Functionality: Visual resemblance to native tissue does not guarantee functional relevance, especially in disease modeling. (2) Technical Tradeoffs: - Standardization vs. Personalization: Patient-derived models offer individualized insights but challenge reproducibility across platforms. - 2D vs. 3D Cultures: While 3D systems better mimic tissue architecture, they are less tractable for imaging and manipulation. - Longevity vs. Stability: Extended culture durations enable chronic modeling but risk phenotypic drift and instability. (3) Ethical & Economic Tradeoffs: - Commercialization vs. Donor Trust: The monetization of organoid technologies raises concerns about donor consent and ethical stewardship27. - Innovation vs. Accessibility: Advanced platforms may be cost-prohibitive, limiting equitable access across research environments. (4) Regulatory Tradeoffs: - Validation vs. Speed: Regulatory rigor ensures safety and reliability but may delay adoption of organoid-based alternatives to animal models. B. Together, the tradeoff domains in A underscore the need for identifying an optimal zone for innovation at scale, wherein intentional design and disciplined simplification in organoid and MPS systems will favor models that are not merely complex, but strategically constructed to yield interpretable, scalable, and clinically relevant insights. If the goal if for these pre-clinical models to yield better results in clinical trials, tradeoffs in ‘TRIALS’ (Y axis) must be optimized with radical simplification in complexity, content and cost (X axis), through disciplined reduction to essentials that can be scaled, standardized, and trusted. The tradeoffs are clear (Figure 2A): complexity enhances physiological resemblance but undermines interpretability, scalability, and reproducibility, while oversimplification sacrifices depth for throughput. This is the inconvenient truth: we are wasting resources, perpetuating translational failures, and widening the gap between model predictions and clinical outcomes, risking an era of reproducibility crisis worse than the decades of animal models. We believe that the future of NAMs depends not on multiplying complexity, but on our ability to identify first and subsequently converge upon (through consensus) on an ‘optimal zone’ for innovation at scale (Figure 2B). That path to optimization may require radical simplification—disciplined reduction to essentials that can be scaled, standardized, and trusted. 4. The Case for Abstraction: Simplify, but not oversimplify Abstraction is often misunderstood as oversimplification. In reality, abstraction is the art of deliberate reduction: stripping away what is superfluous or irrelevant to expose or arrive at what is essential. Precision Medicine is exactly this: identify the primary disease driver (Abelson Tyrosine Kinase in CML) and target it (the story of Gleevec). Mathematics, engineering, and even art embody this principle. Pablo Picasso’s sketches of a bull showed that stripping away detail can crystallize essence rather than erase it. In biomedical science, abstraction becomes powerful when grounded in mathematics and machine learning. Models built this way are reproducible—the answer today will be the same tomorrow—yet adaptive, updating as new data emerge. This ensures stability without rigidity. Applied to NAMs, abstraction delivers three advantages: (1) it strips away irrelevant complexity with its attendant consumption of time and resources while preserving exactness, (2) it enables iterative refinement as datasets grow, and (3) it anchors models—whether in silico or organoid-based—to disease-driving features tied to clinical outcomes. Machine learning illustrates this principle: from thousands of variables in tumor transcriptomes, it distills a handful of driver phenotypes with greater fidelity than bulk tumor models overloaded with noise that inevitably drift over time.
9 | Page Abstraction is not reductionism; it is disciplined precision. It transforms NAMs from elaborate replicas into streamlined, regulator-ready tools that focus on what truly changes outcomes. Simply put, NAMs under abstraction could be likened to “compressed algorithms” that are also expected to outperform bloated replicas. 5. The PICASSO™ Framework for Abstraction: The what and the how We propose PICASSO™: Phenotype-Informed Clinical Abstraction for Systematic Simulation and Outcomes (Box 2; Figure 3), a disciplined, algorithmic end-to-end human framework for making NAMs clinically credible. Inspired by the artistic principle that less can reveal more, PICASSO™ applies deliberate reduction—through four steps of integrated computational and systems approaches (Figure 4A) --to ensure models capture only what truly drives disease and therapy. Figure 3. From noise to nuance: PICASSO™ transforms NAMs into human-first, regulatory-ready tools for drug discovery. The PICASSO™ framework for abstracting complexity and prioritizing essential, human-relevant insights in biomedical research. Inspired by Pablo Picasso’s artistic reduction, the panel shows the progression from anatomically detailed to minimalist representations of a bull, symbolizing the distillation of biological complexity into essential, actionable signals. PICASSO codifies abstraction to clinical anchoring (large diverse cohorts of patients) to enforce Phase 3 rigor and predictive fidelity using a simple 4-step approach (see Fig 4). (i) Identify essentials — Modeling what matters. To identify invariant, disease-driving features within complex omic datasets, PICASSO™ begins with algorithmic abstraction, distilling biological complexity into logical essentials that remain stable across tissues, species, and disease states. This stage applies machine learning to large-scale human datasets spanning health-to-disease continua, capturing
16 | Page interventions, and recovery. PICASSO™ mandates that such data be analyzed with the same statistical rigor required in Phase 3 trials—including hazard ratios, Kaplan–Meier survival curves, quartile shifts, and multivariate regression models. Frameworks such as CANDiT62 and FORWARD63 exemplify how these standards operationalize predictive modeling by aligning molecular phenotypes with longitudinal outcomes. Anchor molecular phenotypes to real-world trajectories: Biobanked NAMs must not only reproduce disease features but track them over time. PICASSO™ enforces standardized pipelines where molecular phenotypes are systematically linked to clinical outcomes (e.g., remission, relapse, resistance or other relevant endpoints64) and validated against real-world cohort distributions. This anchoring transforms NAMs into quantitative mirrors of population diversity, enabling consistent cross-study benchmarking and predictive extrapolation to human populations. Digital twins and multi-scale computational models as regulatory assets: Multi-scale computational frameworks—including Boolean foundational models55, integrated end to end through digital biomarkers (now made accessible to biologists via the web-based intuitive interface of COMPASS™53), and dynamic network simulations—function as in silico digital twins that extend NAMs beyond the petri dish. These digital twins can: • Project therapeutic and diagnostic outcomes virtually, forecasting efficacy, toxicity, and resistance before human exposure. • Integrate cross-scale data (molecular → tissue → organ → system → population) using standardized, ontologyfree logic models. • Benchmark virtual predictions against real-world clinical data to quantify predictive accuracy, reliability, and reproducibility. Standardization principles for simulation and digital twins: • Relevance: Digital models must replicate clinically validated endpoints (e.g., survival, remission, organ function). • Reliability: Simulations must demonstrate reproducible outputs under identical parameter inputs across sites and software versions. • Robustness: Virtual models must tolerate biological and computational perturbations, parameter noise, and timeseries drift. • Traceability: Every model iteration must be version-controlled, auditable, and annotated with input data provenance and quality metrics. • Validation pipeline: Simulation outputs must undergo the same multi-step qualification as experimental NAMs— analytical validation → biological validation → clinical correlation. Hybrid Human–Digital Integration: PICASSO™ envisions NAMs and digital twins as reciprocal mirrors: • NAMs inform digital twins by providing real biological constraints and human-derived causal logic. • Digital twins guide NAM design by predicting optimal perturbations and experimental endpoints. This hybrid feedback loop creates a self-correcting, scalable system where predictions are continuously updated against real-world data—advancing a new gold standard of closed-loop translational modeling.
17 | Page Regulatory standardization statement: Digital twins, like NAMs, must be validated as fit-for-purpose predictive instruments. Their credibility is earned through documented reproducibility, defined performance thresholds, and transparent, explainable logic. Regulators increasingly treat validated digital twins as Phase-0 evidence engines, capable of supporting adaptive trial design, label expansion, and mechanism-based safety qualification. This four-step discipline transforms NAMs from ad hoc proxies into regulator-ready abstractions—minimal in form, maximal in clinical relevance. It creates a continuous, human-centric discovery engine where data informs models, models inform measurements, and measurements feed predictive systems, anchoring them across small ‘n’ and large ‘n’ datasets (Figure 4B). It also creates groundwork for NAM-first human trials. Integrate Essentials — The Regulatory Roadmap for Human Logic NAMs PICASSO™ integrates modeling, measurement, and simulation into a unified framework that converts NAMs from experimental artifacts into regulatory-ready digital and biological evidence engines (Figure 3). This integration enforces causality, comparability, and compliance across every layer, from petri dish to patient population, ensuring that predictive models are not just innovative but auditable, standardized, and clinically anchored. Unification of Logic and Measurement: The Boolean framework-based foundational model provides the logic, explainability and therefore, confidence in what NAMs can predict; COMPASS™ provides the quantification53. Together, they define what must be measured and how it must be measured. BoNE’s Boolean implication networks transform biological relationships into rule-based logic, while COMPASS translates that logic into deterministic activity scores with fixed thresholds and reproducible variance estimates. The combination ensures that NAM outcomes are both mechanistically transparent and statistically rigorous, thereby bridging the historical divide between mechanism and metric. Integration across Scales and Modalities: PICASSO™ enforces multi-scale alignment across four dimensions of translation: 1. Biological: Models capture cellular and molecular drivers with retained epigenetic and mechanical fidelity. 2. Functional: Readouts quantify organ-level performance—barrier function, metabolism, electrophysiology. 3. Clinical: Outcomes map to patient-relevant endpoints—remission, regression, recovery. 4. Computational: Digital twins integrate these layers into reproducible, explainable simulations. All outputs must conform to standardized metadata schemas and QC panels, enabling inter-laboratory comparability and regulatory traceability. Standardization Principles for Integration PICASSO™ mandates that integrated NAM-digital ecosystems adhere to the following regulatory criteria: • Relevance: Each component must demonstrate linkage to a clinically validated endpoint. • Reliability: Cross-platform reproducibility and inter-site concordance are required for qualification. • Robustness: Systems must withstand perturbations, temporal drift, and environmental variability. • Traceability: All data, algorithms, and experimental conditions must be version-controlled, annotated, and accessible for audit.
18 | Page • Transparency: Models must expose their causal structure and performance boundaries; no black boxes. Cross-Validation between Biological and Digital Systems PICASSO™ formalizes bi-directional validation: • NAMs supply empirical constraints and biological priors to digital twins. • Digital twins predict perturbations, dose ranges, and response dynamics to refine NAM experiments. This closed-loop ensures that computational inference is continuously checked against empirical truth, producing a harmonized continuum of evidence acceptable to regulators and industry alike. Regulatory Convergence and Qualification Pathway. Integrated NAM systems are evaluated through a three-tier qualification pipeline: 1. Analytical validation — confirm reproducibility and instrument fidelity. 2. Biological validation — confirm mechanistic alignment with human pathways. 3. Clinical correlation — confirm predictive alignment with real-world outcomes. Each tier requires documented SOPs, statistical performance metrics, and cohort anchoring to human data. Once validated, NAMs and their digital twins can be accepted as Phase-0 regulatory assays for mechanism-based approvals, adaptive trial design, or post-marketing surveillance. Hybrid Evidence Generation. Integration also redefines how evidence is created: NAMs generate mechanistic evidence; digital twins generate predictive evidence. Together, they fulfill the evidentiary requirements for explainable AI and data-driven regulatory science. PICASSO™ thus embodies the “3R-plus-1” standard of Relevance, Reliability, Robustness, and Reproducibility as a quantifiable, interoperable, and transparent benchmark for the next generation of human-logic medicine. Through integration, PICASSO™ transforms NAMs, organoids, and digital twins into a single, standardized continuum of discovery and validation. It delivers a regulatory roadmap where logic replaces opacity, reproducibility replaces assumption, and mechanistic truth replaces correlation. In this future, NAMs no longer imitate life, but they compute it, with clinical fidelity, mathematical precision, and regulatory confidence. A few therapeutic areas have already exemplified this process, albeit at early stages. A recent study tackled the long-elusive cancer stem cell (CSC), long implicated as the drivers of relapse and therapeutic resistance62. Using a machine learning framework—CANDiT (Cancer Associated Nodes for Differentiation Targeting)—the authors identified therapeutically exploitable transcriptomic networks capable of selectively inducing CSC differentiation and death. Anchored in Phase 3– sized human tumor cohorts, this effort began with a single “seed” gene, rigorously validated across >16,000 patients in 33 independent studies, and expanded into a network reproducibly confirmed in ~4,600 additional patients. Abstracted into PDO-based NAMs (n = 26), these networks yielded response signatures that, when simulated across 10 independent cohorts (~2,110 patients) for endpoints of clinical relevance, predicted a ~50% reduction in relapse and recurrence. The selectivity of this approach—killing CSCs while sparing normal stem and differentiated tumor cells—was explained by molecular mechanisms uncovered through cell-based and in chemico approaches a decade ago65, 66. This work shows how even small PDO numbers, when embedded in the PICASSO™ framework, can fuel reproducible, clinically anchored discoveries that bridge mechanism to outcome.
19 | Page Similarly, the elusive gut barrier state, long responsible for Phase 3 failures in Inflammatory bowel diseases (IBD), was distilled from ~1,600 datasets into a simple core defect: bioenergetic collapse driven by H₂S toxicity, the result of failed detoxification machinery63. This essential insight emerged from F.O.R.W.A.R.D (Framework for Outcome-based Research and Drug Development), a network-based target prioritization platform that ties molecular states to clinical outcomes. Built and validated initially as a map of continuum states in IBD on diverse datasets from ~1500 patients47, and subsequently trained on seven prospective randomized trials across four biologics (n = 332 patients), F.O.R.W.A.R.D defined remission at the molecular level and predicted, with network connectivity, the likelihood that targeting a given molecule would induce remission. Benchmarking against 210 completed trials across 52 targets (~200,000 patients), it achieved 100% predictive accuracy despite heterogeneity in drug mechanisms and trial designs. Single-cell RNA-seq and a prospective PDO biobank (n = 29 subjects, 40 PDO lines57) confirmed the remission signature as epithelium-specific and predictive of poor outcomes. By anchoring small-scale NAMs to Phase 3–sized cohorts, F.O.R.W.A.R.D enables in silico Phase 0 trials: de-risking development, reviving shelved drugs, and guiding trial design or early termination. It exemplifies how abstraction strips away noise, isolates essentials, and transforms compact models into reproducible, outcome-anchored discovery engines. In summary, PICASSO™ is inclusive but discerning: It does not discard models—algorithmic, human, or animal— but rather elevates them to a higher evidentiary standard, applied only when they pass the humanness test. Within its framework, every experimental or computational system earns its place by demonstrating measurable alignment with human biology, clinical outcomes, and causal logic. PICASSO codifies the five pillars of the regulatory quitet [Relevance, Reliability, Robustness, Reproducibility, and Traceability] as the foundation of regulatory-grade NAMs and digital twins. These parameters collectively define the qualification roadmap, from analytical validation to biological verification to clinical correlation—standardizing how NAMs, organoids, and digital twins advance toward regulatory acceptance. Because regulators approve interventions grounded in mechanisms, not mysteries, PICASSO™ enforces explainable predictivity: models that expose their causal logic, not conceal it in black boxes. By embedding Boolean causality (via BoNE) and quantitative anchoring (via COMPASS™), NAMs become evidence engines that are transparent, reproducible, and futureproof for integration with explainable AI. Through these essentials, PICASSO™ converts NAMs from experimental artifacts into logical, quantitative, and regulatory-ready surrogates of human biology. In short, PICASSO™ ensures that every model earns, not declares, its fidelity to human truth. 6. Why Standardization Matters: Making Biology as Reproducible as Mathematics and Engineering Despite decades of progress, NAMs remain fragmented. Organoids, microphysiological systems, and AI models differ not just in composition and readouts but in what they claim to represent. Without harmonized principles, NAMs risk becoming arbitrary proxies, demonstrations of technical ingenuity rather than engines of clinical translation. Standardization enables interoperability: labs, institutions, and companies working from the same playbook can compare results, pool datasets, and build trust. With a common language of essentials, NAMs evolve from boutique experiments to reproducible, regulator-ready platforms. The result is efficiency, credibility, and acceleration toward human outcomes. Mathematical abstraction is ‘connective tissue’: Disease is reduced to a set of essentials, identified mathematically, reconstructed biologically, and simulated computationally. The precision of engineering ensures fidelity; the discipline of mathematics ensures rigor. Together, they transform biology into a predictive science. PICASSO™ provides the framework for standardization NAMs need. By codifying abstraction, it ensures that models are deliberate reconstructions of disease-driving essentials, not open-ended exercises in complexity. This approach aligns with the FDA Modernization Act and NIH initiatives, which encourage the use of NAMs but demand rigor,
20 | Page reproducibility, and relevance. It also aligns with the changing times in which biology and medicine are increasingly branches of mathematics and engineering. The future of translational science lies not in descriptive cataloging of complexity but in mathematical inference and engineering precision. PICASSO™ operationalizes this vision. It provides a framework where every model is both an abstraction and an engineering construct, bridging the gap between biological noise and clinical signal. 7. Call to Action: Preparing academia, biotech and pharma The biomedical sciences now stand at a crossroads. One path chases complexity for its own sake, producing models that drown in noise. The other embraces abstraction as a disciplined art—distilling only what matters to deliver relevance. PICASSO™ charts this path (Figure 3C). • Academia must innovate and embrace regulatory-grade rigor, embed computational and regulatory expertise into discovery, and train a new generation fluent in both biology and translational science. • Biotech must design platforms for scalability, robustness, and reproducibility, not boutique demonstrations. Early engagement with FDA and participation in qualification programs (e.g., ISTAND) will be key. • Pharma must partner across sectors to ensure NAMs are population-representative and aligned with regulatory endpoints. • Regulators, in turn, must incentivize and support sponsors deploying NAMs, helping to overcome the understandable fear or caution around regulatory change. • Everyone involved in NAM creation must recommit to respecting patients’ and donors’ autonomy, ensuring informed consent for all research samples. With PICASSO™, NAMs evolve from proxies to standardized, clinically anchored abstractions capable of transforming drug discovery into a human-first, regulator-ready science. 8. Future Directions NAMs should be positioned as complementary extensions of human clinical biology, not siloed stand-ins. Integration, not replacement: NAMs extend what clinical samples or animal models cannot—longitudinal access, mechanistic dissection, and experimental control—while remaining anchored to human outcomes. Big-data NAMs: High-throughput organoid and MPS platforms, paired with AI, can generate population-scale, mechanistically rich datasets. These will reduce dependence on noisy, inconsistent human self-reporting (diet, sleep, lifestyle) while capturing the full range of genetic and physiological variation. BOX 3: Why standardization matters Challenges NAMs face today: •Lack of large-scale validated datasets •Lack of standardization of culture protocols and conditions •Poor reproducibility across labs •Fragmented SOPs and reference standards •Limited regulatory trust What standardization enables: •Cross-lab harmonization → reproducibility •Population-scale relevance → statistical power •Regulatory alignment → accelerated adoption •Reduced noise → clearer signal for clinical translation Key takeaway: Without standardization, NAMs remain boutique science; with it, they become the backbone of translational medicine.
21 | Page Standardization: Reference materials, harmonized SOPs, and interoperable data-sharing consortia will be essential to credibility and regulatory uptake. Case studies as proving grounds: Oncology (PDO-guided drug response67), liver toxicology68 (FDA’s adoption of MPSbased prediction of injury), and IBD (molecular subtyping and therapeutic efficacy assessments with organoids47, 57, 63) provide immediate testbeds to build trust and regulatory confidence. The trajectory is clear: NAMs must scale from boutique, artisanal models into robust, interoperable infrastructures that regulators and industry alike can rely on. 9. Closing Vision NAMs now stand at a tipping point. Regulatory momentum, technological maturity, and societal demand for human-relevant science have aligned. The winners in this landscape will not be those who showcase novelty alone, but those who bridge novelty with credibility, scaling NAMs into reproducible, regulator-ready tools. The challenge is scalability with simplicity. “Simple can be harder than complex”, said Steve Jobs “but it is worth it in the end because once you get there, you can move mountains”. Mark Twain once apologized to a correspondent for sending a long letter saying, “I didn’t have enough time to write short one”. But the art of abstraction demands it. PICASSO™ reframes disease not as an overwhelming mosaic of signals but as a handful of critical events NAMs must capture with rigor. Success will come not from producing more analytically impressive data, but from tracking the few signals that truly shift patient outcomes. In the future, biology and medicine will become predictive sciences--branches of mathematics and engineering. NAMs, standardized through PICASSO™™, will evolve from boutique experiments to mainstream practice, defining a new era of human-first drug discovery. Minimal in components, maximal in relevance—that is the essence of abstraction. That is the future PICASSO™ paints. PICASSO™ distills disease into the few decisive events NAMs must track—clinically anchored phenotypes that truly move the needle for patients. In doing so, PICASSO disciplines NAMs to focus only on outcome-defining events.
22 | Page Acknowledgments The authors wish to apologize for not citing many important publications on this topic due to the limit of references allowed. P.G. is supported by NIH grants R01-AI141630 and R01-AI55696. PG was also supported by the Leona M. and Harry B. Helmsley Charitable Trust and the Propel a Cure Foundation. S.S. was supported through The American Association of Immunologists (AAI) Intersect Fellowship Program for Computational Scientists and Immunologists. The views expressed are those of the authors and do not represent those of their employers or funders. Declaration of interests The authors declare no conflicts of interest.
23 | Page REFERENCES: 1. Butler D. Translational research: crossing the valley of death. Nature. 2008;453(7197):840-2. doi: 10.1038/453840a. PubMed PMID: 18548043. 2. Adashi EY, O'Mahony DP, Cohen IG. The FDA Modernization Act 2.0: Drug Testing in Animals is Rendered Optional. Am J Med. 2023;136(9):853-4. Epub 20230418. doi: 10.1016/j.amjmed.2023.03.033. PubMed PMID: 37080328. 3. Lee GY, Kenny PA, Lee EH, Bissell MJ. Three-dimensional culture models of normal and malignant breast epithelial cells. Nat Methods. 2007;4(4):359-65. doi: 10.1038/nmeth1015. PubMed PMID: 17396127; PMCID: PMC2933182. 4. Sato T, Vries RG, Snippert HJ, van de Wetering M, Barker N, Stange DE, van Es JH, Abo A, Kujala P, Peters PJ, Clevers H. Single Lgr5 stem cells build crypt-villus structures in vitro without a mesenchymal niche. Nature. 2009;459(7244):262-5. Epub 20090329. doi: 10.1038/nature07935. PubMed PMID: 19329995. 5. Mariani J, Vaccarino FM. Breakthrough Moments: Yoshiki Sasai's Discoveries in the Third Dimension. Cell Stem Cell. 2019;24(6):837-8. doi: 10.1016/j.stem.2019.05.007. PubMed PMID: 31173711; PMCID: PMC7085937. 6. Gaudet P, Michel PA, Zahn-Zabal M, Britan A, Cusin I, Domagalski M, Duek PD, Gateau A, Gleizes A, Hinard V, Rech de Laval V, Lin J, Nikitin F, Schaeffer M, Teixeira D, Lane L, Bairoch A. The neXtProt knowledgebase on human proteins: 2017 update. Nucleic Acids Res. 2017;45(D1):D177-D82. Epub 20161129. doi: 10.1093/nar/gkw1062. PubMed PMID: 27899619; PMCID: PMC5210547. 7. Aken BL, Achuthan P, Akanni W, Amode MR, Bernsdorff F, Bhai J, Billis K, Carvalho-Silva D, Cummins C, Clapham P, Gil L, Girón CG, Gordon L, Hourlier T, Hunt SE, Janacek SH, Juettemann T, Keenan S, Laird MR, ..., Flicek P. Ensembl 2017. Nucleic Acids Res. 2017;45(D1):D635-D42. Epub 20161128. doi: 10.1093/nar/gkw1104. PubMed PMID: 27899575; PMCID: PMC5210575. 8. The UniProt Consortium. UniProt: the universal protein knowledgebase. Nucleic Acids Res. 2017;45(D1):D158D69. Epub 20161129. doi: 10.1093/nar/gkw1099. PubMed PMID: 27899622; PMCID: PMC5210571. 9. Rands CM, Meader S, Ponting CP, Lunter G. 8.2% of the Human genome is constrained: variation in rates of turnover across functional element classes in the human lineage. PLoS Genet. 2014;10(7):e1004525. Epub 20140724. doi: 10.1371/journal.pgen.1004525. PubMed PMID: 25057982; PMCID: PMC4109858. 10. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, Bridgland A, Meyer C, Kohl SAA, Ballard AJ, Cowie A, Romera-Paredes B, Nikolov S, Jain R, Adler J, ..., Hassabis D. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583-9. Epub 20210715. doi: 10.1038/s41586-021-03819-2. PubMed PMID: 34265844; PMCID: PMC8371605. 11. Baek M, DiMaio F, Anishchenko I, Dauparas J, Ovchinnikov S, Lee GR, Wang J, Cong Q, Kinch LN, Schaeffer RD, Millán C, Park H, Adams C, Glassman CR, DeGiovanni A, Pereira JH, Rodrigues AV, van Dijk AA, Ebrecht AC, ..., Baker D. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 2021;373(6557):871-6. Epub 20210715. doi: 10.1126/science.abj8754. PubMed PMID: 34282049; PMCID: PMC7612213. 12. Binder JL, Berendzen J, Stevens AO, He Y, Wang J, Dokholyan NV, Oprea TI. AlphaFold illuminates half of the dark human proteins. Curr Opin Struct Biol. 2022;74:102372. Epub 20220416. doi: 10.1016/j.sbi.2022.102372. PubMed PMID: 35439658; PMCID: PMC10669925. 13. Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA, Hill DP, Issel-Tarver L, Kasarskis A, Lewis S, Matese JC, Richardson JE, Ringwald M, Rubin GM, Sherlock G. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nat Genet. 2000;25(1):25-9. doi: 10.1038/75556. PubMed PMID: 10802651; PMCID: PMC3037419. 14. Mi H, Muruganujan A, Casagrande JT, Thomas PD. Large-scale gene function analysis with the PANTHER classification system. Nat Protoc. 2013;8(8):1551-66. Epub 20130718. doi: 10.1038/nprot.2013.092. PubMed PMID: 23868073; PMCID: PMC6519453. 15. Sherman BT, Hao M, Qiu J, Jiao X, Baseler MW, Lane HC, Imamichi T, Chang W. DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res. 2022;50(W1):W216-W21. doi: 10.1093/nar/gkac194. PubMed PMID: 35325185; PMCID: PMC9252805. 16. Subramanian A, Tamayo P, Mootha VK, Mukherjee S, Ebert BL, Gillette MA, Paulovich A, Pomeroy SL, Golub TR, Lander ES, Mesirov JP. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci U S A. 2005;102(43):15545-50. Epub 20050930. doi: 10.1073/pnas.0506580102. PubMed PMID: 16199517; PMCID: PMC1239896. 17. Hänzelmann S, Castelo R, Guinney J. GSVA: gene set variation analysis for microarray and RNA-seq data. BMC Bioinformatics. 2013;14:7. Epub 20130116. doi: 10.1186/1471-2105-14-7. PubMed PMID: 23323831; PMCID: PMC3618321. 18. Chen Y, Feng Y, Yan F, Zhao Y, Zhao H, Guo Y. A Novel Immune-Related Gene Signature to Identify the Tumor Microenvironment and Prognose Disease Among Patients With Oral Squamous Cell Carcinoma Patients Using ssGSEA: A Bioinformatics and Biological Validation Study. Front Immunol. 2022;13:922195. Epub 20220706. doi: 10.3389/fimmu.2022.922195. PubMed PMID: 35935989; PMCID: PMC9351622.
24 | Page 19. Croft D, O'Kelly G, Wu G, Haw R, Gillespie M, Matthews L, Caudy M, Garapati P, Gopinath G, Jassal B, Jupe S, Kalatskaya I, Mahajan S, May B, Ndegwa N, Schmidt E, Shamovsky V, Yung C, Birney E, ..., Stein L. Reactome: a database of reactions, pathways and biological processes. Nucleic Acids Res. 2011;39(Database issue):D691-7. Epub 20101109. doi: 10.1093/nar/gkq1018. PubMed PMID: 21067998; PMCID: PMC3013646. 20. Kanehisa M, Goto S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 2000;28(1):27-30. doi: 10.1093/nar/28.1.27. PubMed PMID: 10592173; PMCID: PMC102409. 21. Kuleshov MV, Jones MR, Rouillard AD, Fernandez NF, Duan Q, Wang Z, Koplev S, Jenkins SL, Jagodnik KM, Lachmann A, McDermott MG, Monteiro CD, Gundersen GW, Ma'ayan A. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res. 2016;44(W1):W90-7. Epub 20160503. doi: 10.1093/nar/gkw377. PubMed PMID: 27141961; PMCID: PMC4987924. 22. Phan A, Joshi P, Kadelka C, Friedberg I. A longitudinal analysis of function annotations of the human proteome reveals consistently high biases. Database (Oxford). 2025;2025. doi: 10.1093/database/baaf036. PubMed PMID: 40338520; PMCID: PMC12060720. 23. Lamers MM, van der Vaart J, Knoops K, Riesebosch S, Breugem TI, Mykytyn AZ, Beumer J, Schipper D, Bezstarosti K, Koopman CD, Groen N, Ravelli RBG, Duimel HQ, Demmers JAA, Verjans GMGM, Koopmans MPG, Muraro MJ, Peters PJ, Clevers H, Haagmans BL. An organoid-derived bronchioalveolar model for SARS-CoV-2 infection of human alveolar type II-like cells. EMBO J. 2021;40(5):e105912. Epub 20210111. doi: 10.15252/embj.2020105912. PubMed PMID: 33283287; PMCID: PMC7883112. 24. Tindle C, Fuller M, Fonseca A, Taheri S, Ibeawuchi SR, Beutler N, Katkar GD, Claire A, Castillo V, Hernandez M, Russo H, Duran J, Crotty Alexander LE, Tipps A, Lin G, Thistlethwaite PA, Chattopadhyay R, Rogers TF, Sahoo D, ..., Das S. Adult stem cell-derived complete lung organoid models emulate lung disease in COVID-19. Elife. 2021;10. Epub 20210813. doi: 10.7554/eLife.66417. PubMed PMID: 34463615; PMCID: PMC8463074. 25. Sinha S, Castillo V, Espinoza CR, Tindle C, Fonseca AG, Dan JM, Katkar GD, Das S, Sahoo D, Ghosh P. COVID-19 lung disease shares driver AT2 cytopathic features with Idiopathic pulmonary fibrosis. EBioMedicine. 2022;82:104185. Epub 20220720. doi: 10.1016/j.ebiom.2022.104185. PubMed PMID: 35870428; PMCID: PMC9297827. 26. Ghosh P, Katkar GD, Shimizu C, Kim J, Khandelwal S, Tremoulet AH, Kanegaye JT, Group PEMKDR. An Artificial Intelligence-guided signature reveals the shared host immune response in MIS-C and Kawasaki disease. Nat Commun. 2022;13(1):2687. Epub 20220516. doi: 10.1038/s41467-022-30357-w. PubMed PMID: 35577777; PMCID: PMC9110726. 27. de Groot H, van Daal M, Hofland RW, Bronsveld I, Jongsma KR, Ten Ham RMT. The ethics and economics of organoid commercialization: potential donors' perspectives. BMC Med Ethics. 2025;26(1):109. Epub 20250729. doi: 10.1186/s12910-025-01269-3. PubMed PMID: 40730998; PMCID: PMC12309163. 28. Sinha S, Ghosh P. From Complex to Simple and Back: Mathematical Abstraction of Life's Logic Reveals DiseaseDriving Essentials. 2025. 29. Curtius K, Wright NA, Graham TA. Evolution of Premalignant Disease. Cold Spring Harb Perspect Med. 2017;7(12). Epub 20171201. doi: 10.1101/cshperspect.a026542. PubMed PMID: 28490542; PMCID: PMC5710095. 30. Gerstung M, Jolly C, Leshchiner I, Dentro SC, Gonzalez S, Rosebrock D, Mitchell TJ, Rubanova Y, Anur P, Yu K, Tarabichi M, Deshwar A, Wintersinger J, Kleinheinz K, Vázquez-García I, Haase K, Jerman L, Sengupta S, Macintyre G, ..., Consortium P. The evolutionary history of 2,658 cancers. Nature. 2020;578(7793):122-8. Epub 20200206. doi: 10.1038/s41586-019-1907-7. PubMed PMID: 32025013; PMCID: PMC7054212. 31. Beason-Held LL, Goh JO, An Y, Kraut MA, O'Brien RJ, Ferrucci L, Resnick SM. Changes in brain function occur years before the onset of cognitive impairment. J Neurosci. 2013;33(46):18008-14. doi: 10.1523/JNEUROSCI.140213.2013. PubMed PMID: 24227712; PMCID: PMC3828456. 32. Caselli RJ, Langlais BT, Dueck AC, Chen Y, Su Y, Locke DEC, Woodruff BK, Reiman EM. Neuropsychological decline up to 20 years before incident mild cognitive impairment. Alzheimers Dement. 2020;16(3):512-23. Epub 20200106. doi: 10.1016/j.jalz.2019.09.085. PubMed PMID: 31787561; PMCID: PMC7067658. 33. Vestergaard MV, Allin KH, Poulsen GJ, Lee JC, Jess T. Characterizing the pre-clinical phase of inflammatory bowel disease. Cell Rep Med. 2023;4(11):101263. Epub 20231107. doi: 10.1016/j.xcrm.2023.101263. PubMed PMID: 37939713; PMCID: PMC10694632. 34. Evans-Molina C, Sims EK, DiMeglio LA, Ismail HM, Steck AK, Palmer JP, Krischer JP, Geyer S, Xu P, Sosenko JM, Group TDTS. β Cell dysfunction exists more than 5 years before type 1 diabetes diagnosis. JCI Insight. 2018;3(15). Epub 20180809. doi: 10.1172/jci.insight.120877. PubMed PMID: 30089716; PMCID: PMC6129118. 35. Timmons JA, Szkop KJ, Gallagher IJ. Multiple sources of bias confound functional enrichment analysis of global - omics data. Genome Biol. 2015;16(1):186. Epub 20150907. doi: 10.1186/s13059-015-0761-7. PubMed PMID: 26346307; PMCID: PMC4561415. 36. Wijesooriya K, Jadaan SA, Perera KL, Kaur T, Ziemann M. Urgent need for consistent standards in functional enrichment analysis. PLoS Comput Biol. 2022;18(3):e1009935. Epub 20220309. doi: 10.1371/journal.pcbi.1009935. PubMed PMID: 35263338; PMCID: PMC8936487. 37. Haynes WA, Tomczak A, Khatri P. Gene annotation bias impedes biomedical research. Sci Rep. 2018;8(1):1362. Epub 20180122. doi: 10.1038/s41598-018-19333-x. PubMed PMID: 29358745; PMCID: PMC5778030.
25 | Page 38. Schnoes AM, Ream DC, Thorman AW, Babbitt PC, Friedberg I. Biases in the experimental annotations of protein function and their effect on our understanding of protein function space. PLoS Comput Biol. 2013;9(5):e1003063. Epub 20130530. doi: 10.1371/journal.pcbi.1003063. PubMed PMID: 23737737; PMCID: PMC3667760. 39. Brunet JP, Tamayo P, Golub TR, Mesirov JP. Metagenes and molecular pattern discovery using matrix factorization. Proc Natl Acad Sci U S A. 2004;101(12):4164-9. Epub 20040311. doi: 10.1073/pnas.0308531101. PubMed PMID: 15016911; PMCID: PMC384712. 40. Devarajan K. Nonnegative matrix factorization: an analytical and interpretive tool in computational biology. PLoS Comput Biol. 2008;4(7):e1000029. Epub 20080725. doi: 10.1371/journal.pcbi.1000029. PubMed PMID: 18654623; PMCID: PMC2447881. 41. Berglund AE, Welsh EA, Eschrich SA. Characteristics and Validation Techniques for PCA-Based GeneExpression Signatures. Int J Genomics. 2017;2017:2354564. Epub 20170206. doi: 10.1155/2017/2354564. PubMed PMID: 28265563; PMCID: PMC5317117. 42. Alkaabi AM, Abdallah AK. Portfolio practices in the principal evaluation process: A qualitative case study. Heliyon. 2024;10(21):e39467. Epub 20241018. doi: 10.1016/j.heliyon.2024.e39467. PubMed PMID: 39524733; PMCID: PMC11546155. 43. Kong W, Vanderburg CR, Gunshin H, Rogers JT, Huang X. A review of independent component analysis application to microarray gene expression data. Biotechniques. 2008;45(5):501-20. doi: 10.2144/000112950. PubMed PMID: 19007336; PMCID: PMC3005719. 44. Argelaguet R, Arnol D, Bredikhin D, Deloro Y, Velten B, Marioni JC, Stegle O. MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol. 2020;21(1):111. Epub 20200511. doi: 10.1186/s13059-020-02015-1. PubMed PMID: 32393329; PMCID: PMC7212577. 45. Argelaguet R, Velten B, Arnol D, Dietrich S, Zenz T, Marioni JC, Buettner F, Huber W, Stegle O. Multi-Omics Factor Analysis-a framework for unsupervised integration of multi-omics data sets. Mol Syst Biol. 2018;14(6):e8124. Epub 20180620. doi: 10.15252/msb.20178124. PubMed PMID: 29925568; PMCID: PMC6010767. 46. Aibar S, González-Blas CB, Moerman T, Huynh-Thu VA, Imrichova H, Hulselmans G, Rambow F, Marine JC, Geurts P, Aerts J, van den Oord J, Atak ZK, Wouters J, Aerts S. SCENIC: single-cell regulatory network inference and clustering. Nat Methods. 2017;14(11):1083-6. Epub 20171009. doi: 10.1038/nmeth.4463. PubMed PMID: 28991892; PMCID: PMC5937676. 47. Sahoo D, Swanson L, Sayed IM, Katkar GD, Ibeawuchi SR, Mittal Y, Pranadinata RF, Tindle C, Fuller M, Stec DL, Chang JT, Sandborn WJ, Das S, Ghosh P. Artificial intelligence guided discovery of a barrier-protective therapy in inflammatory bowel disease. Nat Commun. 2021;12(1):4246. Epub 20210712. doi: 10.1038/s41467-021-24470-5. PubMed PMID: 34253728; PMCID: PMC8275683. 48. Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. Deep generative modeling for single-cell transcriptomics. Nat Methods. 2018;15(12):1053-8. Epub 20181130. doi: 10.1038/s41592-018-0229-2. PubMed PMID: 30504886; PMCID: PMC6289068. 49. Way GP, Greene CS. Extracting a biologically relevant latent space from cancer transcriptomes with variational autoencoders. Pac Symp Biocomput. 2018;23:80-91. PubMed PMID: 29218871; PMCID: PMC5728678. 50. Tian Y, Chen X, Ganguli S, editors. Understanding self-supervised Learning Dynamics without Contrastive Pairs. International Conference on Machine Learning; 2021. 51. Taroni JN, Grayson PC, Hu Q, Eddy S, Kretzler M, Merkel PA, Greene CS. MultiPLIER: A Transfer Learning Framework for Transcriptomics Reveals Systemic Features of Rare Disease. Cell Syst. 2019;8(5):380-94.e4. doi: 10.1016/j.cels.2019.04.003. PubMed PMID: 31121115; PMCID: PMC6538307. 52. Lin Y, Dong H, Wang H, Zhang T, editors. Bayesian Invariant Risk Minimization. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 18-24 June 2022. 53. Sinha S, Ghosh P. COMPASS: A Web-Based COMPosite Activity Scoring System to Navigate Health and Disease Through Deterministic Digital Biomarkers. doi.org/10.5281/zenodo.17556273. Zenodo 2025. 54. Sahoo D. The power of boolean implication networks. Front Physiol. 2012;3:276. Epub 20120723. doi: 10.3389/fphys.2012.00276. PubMed PMID: 22934030; PMCID: PMC3429050. 55. Sinha S, Ghosh P. A Foundational Model for Biological Logic and Disease. Doi: 10.5281/zenodo.17459815 . 2025. 56. Chrisnandy A, Lutolf MP. An extracellular matrix niche secreted by epithelial cells drives intestinal organoid formation. Dev Cell. 2025. Epub 20250716. doi: 10.1016/j.devcel.2025.06.026. PubMed PMID: 40680738. 57. Tindle C, Fonseca AG, Taheri S, Katkar GD, Lee J, Maity P, Sayed IM, Ibeawuchi SR, Vidales E, Pranadinata RF, Fuller M, Stec DL, Anandachar MS, Perry K, Le HN, Ear J, Boland BS, Sandborn WJ, Sahoo D, ..., Ghosh P. A living organoid biobank of patients with Crohn's disease reveals molecular subtypes for personalized therapeutics. Cell Rep Med. 2024;5(10):101748. Epub 20240926. doi: 10.1016/j.xcrm.2024.101748. PubMed PMID: 39332415; PMCID: PMC11513829. 58. Aversano S, Caiazza C, Caiazzo M. Induced pluripotent stem cell-derived and directly reprogrammed neurons to study neurodegenerative diseases: The impact of aging signatures. Front Aging Neurosci. 2022;14:1069482. Epub 20221220. doi: 10.3389/fnagi.2022.1069482. PubMed PMID: 36620769; PMCID: PMC9810544.