scieee AI-readable full text Open interactive document viewer

On the importance of computational reproducibility in fostering Open and FAIR Science

Šimko, Tibor

Abstract

In this talk we propose to survey the computational reproducibility practices, opportunities and challenges in view of fostering Open and FAIR Science in research communities. We discuss several thinking models regarding computational reproducibility, focusing on the broader knowledge preservation and reuse aspects rather than on the raw computing evolution aspects. Building upon several use cases from experimental particle physics and related scientific disciplines, we discuss the variety of sociological and technological challenges inherent in making the research innately reproducible and reusable. From the researcher point of view, we argue how "preproducibility" should come early in the scientific process in order to ensure its future reusability. From the data infrastructure point of view, we argue how the data repository services benefit from accompanying "analysis engines" to ensure the correctness of data curation procedures of the validity of data usage recipes. The ultimate goal of the Open Science and Data Preservation efforts is to facilitate future reuse and reinterpretation of scientific data by new generation of researchers. A strong focus on the computational reproducibility of original data analyses provides a way to facilitate the reuse and reinterpretation of Open and FAIR data even many years after the original publication. This talk is heavily inspired, but not limited to, the experiences and lessons learnt from the past ten years of running the CERN Open Data portal and the REANA reproducible analysis platform for the particle physics community.

Full text

On the importance of computational reproducibility in fostering Open and FAIR Science Tibor Šimko CERN The reproducibility problem 2 3 I have divers times in cases, where the Experiments seem’d like to be thought strange, or to be distrusted, set down several Trials of the same thing, that they might mutually support and confirm one another. – Robert Boyle (1627-1691) Robert Boyle by Johann Kerseboom (1689) https://en.wikipedia.org/wiki/Robert_Boyle 4 Survey of 1,576 researchers on reproducibility M. Baker (2016) https://doi.org/10.1038/533452a Half of scientists cannot reproduce their own results Reproduce? What’s in a name? 5 6 The Turing Way model https://book.the-turing-way.org/reproducible-research/overview/overview-definitions.html 7 Reproducible? Replicable? Repeatable? H.E. Plesser (2018) https://doi.org/10.3389/fninf.2017.00076 8 From “reproducible” to “reusable” analyses C. Diaconu, U. Schwickerath (2025) CERN Courier Experimental particle physics data are being analysed decades after data taking DPHEP (2012) https://arxiv.org/abs/1205.4667 9 Experimental physics is done in large collaborations CMS collaboration: over 4000 particle physicists, engineers, computer scientists, technicians and students from around 240 institutes and universities from more than 50 countries. https://cms.cern/collaboration Producing robust / reusable code is expensive 16 Product (solid program) System Product (solid and reusable) Program (individual) System (reusable components) 3x 3x It pays to develop reusable code if you reuse it at least thrice. F. Brooks (1975) 17 Survey of 1008 researchers at the NIPS conference V. Stodden (2010) https://dx.doi.org/10.2139/ssrn.1550193 Researchers mostly worry about time; less so about ideas being scooped 18 Researchers are more likely to reuse data than code V. Stodden (2010) https://dx.doi.org/10.2139/ssrn.1550193 Preserve to reuse 19 20 Preserve-to-reuse: 1. Data CERN Open Data portal https://opendata.cern Trusted digital repositories can preserve data beyond experiment lifetimes 21 Preserve-to-reuse: 2. Code Trusted digital repositories can preserve code beyond version control lifetimes https://guides.github.com/activities/citable-code 22 ● Software changes (Freesurfer 4.3.1, 4.5.0, 5.0.0): 8.8±6.6% (volume); 2.8±1.3% (thickness) ● Operating system changes (macOS 10.5, 10.6): “about factor two smaller” 23 Preserve-to-reuse: 3. Computing environment https://hub.docker.com/u/atlas https://hub.docker.com/u/cmssw Container technology helps to encapsulate the original computing environment 24 Preserve-to-reuse: 3. Computing environment Computing environments may interact with other runtime services such as databases; these need “state encapsulation” too in order to allow future reuse Condition database snapshots for CMS open data 25 Preserve-to-reuse: 4. Computational workflows Declarative workflow languages can express complex computational worklfows CWL Snakemake Yadage 32 “Preproducible” science P. Stark (2018) https://doi.org/10.1038/d41586-018-05256-0 33 Continuous analyses Driving preproducibility via “continuous integration” of analyses T. Šimko et al (2021) https://doi.org/10.3389/fdata.2021.661501 34 Continuous reuse Periodical execution of data usage examples helps to catch troubles early M. Donadoni et al (2021) https://doi.org/10.5281/zenodo.10263203 Scenario: Workspace content When the workflow is finished Then the workspace should contain "njets.png" Scenario: Workspace size When the workflow is finished Then the workspace size should be less than 75 MiB Scenario: Log content When the workflow is finished Then the job logs of the step "skimming" should contain "Event has good muons: pass=36921" Scenario: Run duration When the workflow is finished Then the workflow run duration should be less than 25 minutes "adaptable software examples [are] the most efficient way to pass on the knowledge needed for research-level studies on these data" — CMS A holistic point of view 35 36 Funders: Is the grant money well spent? A. Mullard (2022) https://doi.org/10.1038/d41573-022-00012-6 37 Publishers: The fraud is growing R. Richardson et al (2025) https://doi.org/10.1073/pnas.2420092122 38 International Committee of Future Accelerators (ICFA) S. Campana et al (2025) https://arxiv.org/abs/2508.18892 https://icfa-data-best-practices-demo.app.cern.ch/ Best practices addressing a large variety of stakeholders Conclusions 39 40 ● Data + Code + Environment + Workflow → Reusable Analyses ● Technological challenges: large containers, complex workflows ● Sociological challenges: carving out time in publish-or-perish culture ● Close collaboration between researchers and computer scientists ● Driving future reusability through early preproducibility ● Synergies across scientific disciplines (astronomy, life sciences, physics) Conclusions → See also Clemens Lange’s talk this afternoon “Nudging Scientists into adopting Open Science Practices” https://indico.cern.ch/event/1484392/contributions/6523967/ https://opendata.cern https://www.reana.io