scieee AI-readable full text Open interactive document viewer

Reproducibility and reuse

Kleiner, Brian; Morgan de Paula, Emilie; Furrer, Eva; Schütz, Frédéric; Held, Leonhard; Molo, Fabio; Fraga Gonzalez, Gorka

Abstract

This slideset has been used for a lesson of the CAS in Data Stewardship, University of Lausanne, edition 2024-2025 (Module RDM: background, general information and legal framework, Lesson on Reproducibility and Reuse).

Full text

M1 Research Data Management M2 Visibility of the activity and networking Orientation module M4 Advice and technical support M3 CAS in DATA STEWARDSHIP REPRODUCIBILITY AND REUSE Legal framework and good scientific practices Brian Kleiner (FORS), Emilie Morgan de Paula (FORS), Eva Furrer (CRS UZH), Frédéric Schutz (SIB), Leonhard Held (CRS UZH), Fabio Molo (CRS UZH), Gorka Fraga Gonzalez (CRS UZH) 2 ABOUT THIS PRESENTATION This presentation is released under a CC-BY 4.0 license, which means that you are free to reuse, distribute, remix, adapt, and build upon the material in any medium or format only so long as attribution is given to the creator. If you remix, adapt, or build upon the material, we highly recommend to license the modified material under identical terms. This course is part of the CAS in Data Stewardship. Authors Brian Kleiner, Emilie Morgan de Paula, Eva Furrer, Frédéric Schutz, Leonhard Held, Fabio Molo, Gorka Fraga Gonzalez Providers FORS, CRS, SIB Title Reproducibility and reuse Education level Graduates Language English License CC-BY 4.0 Estimate total time 4 hours Version v20241008 How to attribute KLEINER, Brian, MORGAN DE PAULA, Emilie, FURRER, Eva, SCHUTZ, Frédéric, HELD, Leonhard, MOLO, Fabio & FRAGA GONZALEZ, Gorka. (FORS, SIB, CRS), 2024. Reproducibility and reuse [Lausanne]. CAS Data Stewardship UNIL. 01 November 2024. 3 YOUR HOSTS Fabio Molo Managing Director Center for Reproducible Science University of Zurich Gorka Fraga González Scientific staff Center for Reproducible Science University of Zurich Brian Kleiner Head of Data Management and Archive Services FORS – Swiss Centre of Expertise in the Social Sciences Emilie Morgan de Paula Scientific staff FORS –Swiss Centre of Expertise in the Social Sciences Frédéric Schutz Head of Biostatistics SIB –Swiss institute of Bioinformatics 4 PEDAGOGICAL OBJECTIVES In this lesson we will work on these professional skills: •PS11 – Advice on reproducibility and reuse for existing data •PS12 –Understand the practices that contribute to data quality, completeness and consistency •PS20 –Be familiar with best practices and the specifics of long-term data preservation (formats, standards, infrastructures, data preparation and selection, etc.) •PS22 –Advise researchers on how to make the most of their data, in particular by publishing it for re-use 5 LESSON PLAN 1. Introduction and warm-up –FORS (30’) 2. Ensuring reproducibility: navigating roles and responsibilities –FORS (30’) 3. Reproducibility and replicability: concepts and implementation with dynamic reporting –CRS (30’) 4. Reproducibility and versionning : concepts and implementation with Git –SIB (45’) 5. Reproducibility in practice: Hands-on session –CRS & SIB (45’) 6. Conclusion: The role of data stewards in relation to reproducible research –FORS (30’) Estimate total time = 210 minutes 6 WARM-UP Drawing a monster CAS in Data Stewardship UNIL 2024-2025 Individual work: Take a pen, a piece of paper and follow the instructions below. The goal is to follow these instructions as precisely as possible. Do not add any extra details or change any of the instructions. Time: 5’ ICEBREAKER: DRAW A MONSTER 7 Draw a large oval shape in the center of the page. This will be the monster's body. Below the oval, draw 2 short, thick legs. Each leg should end in a foot with three toes. On top of the oval, between the eyes, draw two curved horns that point upwards. 01 06 07 At the top of the oval, draw 3 small circles for the eyes. The eyes should be 0,5 cm apart. Inside each eye, draw a smaller circle to represent the pupils. 02 03 On each side, draw 2 curved, long lines for arms, ending in hands with three fingers each. Below the eyes, in the middle of the oval, draw a wide smile with a row of 5 sharp, triangular teeth. 0405 Draw 7 small dots randomly across the monster's body to represent spots. Give your monster a tail on the bottom right side of the oval, curving upwards with a pointy end. 09 08 8 DRAWING A MONSTER We’d love to learn more about your academic background and experiences. Please display your monster and answer the following questions to introduce yourself: •Academic Background: What discipline are you from? Do you specialize in quantitative, qualitative, or mixed-methods research? •Experience with Syntax: How comfortable are you with writing and understanding syntax in research software or programming languages? •Preregistration Experience: Have you ever preregistered a study? If so, please share your experience with the process. •Replication Experience: Have you been involved in replication studies? What insights or challenges have you encountered? 9 INTRODUCTION Why are reproducibility and re-use important? General definitions CAS in Data Stewardship UNIL 2024-2025 16 •Reproducibility involves different actors in the scientific ecosystem, including: oresearchers and authors; ojournals publishers, peer-reviewers, and editors; ofunders and policy-makers; oand data repositories •Anetwork of change for open science: united action on research integrity ROLES AND RESPONSIBILITIES OF VARIOUS STAKEHOLDERS •Sharing underlying data •Proper documentation •Sharing of materials for reproducibility •Data citation ROLE OF RESEARCHERS AND AUTHORS 17 •Changing journal policies and practices regarding reproducibility •Tools •Preregistration and pre-prints •Registered reports •Data citation •Open science badges •Data availability statements ROLE OF JOURNALS, PUBLISHERS, PEER-REVIEWERS, AND EDITORS 18 •Setting obligations for open science (ex: SNSF data sharing requirements, DMPs) •Budgeting for open sciences practice ROLE OF FUNDERS AND POLICYMAKERS 19 •Storage of data and materials underlying articles •Preregistration and pre-prints •Registered reports •Providing recommended data citations and persistent identifiers ROLE OF REPOSITORIES 20 21 REPRODUCIBILITY AND REPLICABILITY: CONCEPTS Headlines and terminology CAS in Data Stewardship UNIL 2024-2025 •More than 70% of researchers have tried and failed to reproduce another scientist’s experiments •More than half have failed to reproduce their own experiments. IS THERE A REPRODUCIBILITY CRISIS? Source: Baker (2016), 1,500 scientists lift the lid on reproducibility, Nature, 533(7604). DOI 22 SOME RECENT HEADLINES Source: Gihavi et al. (2023), Major data analysis errors invalidate cancer microbiome findings, mBio, 14(5). DOI 23 SOME RECENT HEADLINES (2) Source: Sohn (2023), The reproducibility issues that haunt health-care AI, Nature, 613(7943). DOI 24 TERMINOLOGY Source: The Turing Way Community (2023), The Turing Way: A handbook for reproducible, ethical and collaborative research (1.0.2). DOI 25 OPEN DATA AND CODE Source: Peng (2011), Reproducible Research in Computational Science, Science, 334(6060). DOI Optional reading: •Heyard and Held (2022), When should data and code be made available, Significance, 19(2). DOI •Fraga González et al. (2024), Primer: Software containers for reproducible research, Zenodo. DOI 32 Reproducible Replicable Robust Generalisable THE SWISS REPRODUCIBILITY NETWORK Working groups: •Open Research Data •Preregistration and Registered Reports •Computational Reproducibility •Training •Research Assessment and Incentives https://www.swissrn.org/ 33 34 REPRODUCIBILITY AND REPLICABILITY: IMPLEMENTATION WITH DYNAMIC REPORTING Concepts CAS in Data Stewardship UNIL 2024-2025 •Based on literate programming: a programming paradigm in which executable code is embedded in descriptive text, human readable •The result is a single source document that can be read and rerun with identical results WHAT IS DYNAMIC REPORTING? Source: Knuth, D. E. (1984). Literate programming. The computer journal, 27(2), 97-111. 35 36 WHAT IS DYNAMIC REPORTING? 37 WHAT IS DYNAMIC REPORTING? 38 WHY SHOULD I USE IT? Source: The Turing Way project illustration by Scriberia. Used under a CC-BY 4.0 licence. DOI: 10.5281/zenodo.3332807. •Improves efficiency of report writing •Helps making analysis workflows easier to follow •Facilitates error detection •Facilitates sharing In sum: dynamic reporting increases transparency and reproducibility Language-specific •R Markdown •Matlab* live scripts and report generator Made for multiple languages (Python, R, Julia… ) •Jupyter Notebook / Jupyter Lab •Quarto POPULAR TOOLS FOR DYNAMIC REPORTING 39 *Not open-source 40 POPULAR TOOLS FOR DYNAMIC REPORTING •R Markdown •Matlab live scripts •Jupyter Notebook •Quarto Source (.rmd) file Rendered document R Markdown are interactive documents that can be exported in multiple formats (MS word, PDF, HTML, LaTeX) The source R Markdown files (.rmd) can be edited with any text editor. 41 POPULAR TOOLS FOR DYNAMIC REPORTING •R Markdown •Matlab live scripts •Jupyter Notebook •Quarto Live scripts are interactive documents that can be exported into PDF, MS Word, HTML, LaTeX, Markdown, Jupyter Notebooks The source file (.xml) should be opened with Matlab © 1994-2024 The MathWorks, Inc. (source access). Included on the basis of educational purposes QUARTO 48 Programming language Markup language Output format User knitr (.Rnw)RLaTeX pdf intermediate to advanced R Markdown (.Rmd) RMarkdown html, pdf, docx, pptx beginner to advanced Quarto (.qmd)R, Python, Julia Markdown html, pdf, docx, pptx, … beginner to advanced ●Quarto uses Knitr engine to execute R code, just like R markdown ●Quarto is developed by Posit: https://posit.co/ ●It is free (for academic purposes) and open-source COMPONENTS OF A DYNAMIC REPORT An option header uses YAML language to specify: -Metadata (e.g., author, title) -Settings (e.g., output format) 49 COMPONENTS OF A DYNAMIC REPORT Markdown is lightweight markup language with simple syntax to add formatting to text. Tables and images can be included here. 50 COMPONENTS OF A DYNAMIC REPORT Executable code in R, Python or Julia Many options can be specified for each code chunk (e.g., hide/display the code, warnings off, etc). 51 COMPONENTS OF A DYNAMIC REPORT The output of the code can be rendered in the same document. 52 COMPONENTS OF A DYNAMIC REPORT The text parts can have information generated by code, for example describing data or results. 53 COMPONENTS OF A DYNAMIC REPORT 54 Dynamic reporting can be kept simple or include many advanced features like: •Parameterized reports •Advanced interactivity •Execution from a software container ADVANCED FEATURES 55 •.qmd versus .Rmd files •Easy transition: •Very minor syntax changes in the way code options are defined •Quarto still reads .Rmd files and has more options •Quarto supports also Python and Julia •Quarto supports Documents, Presentations, Websites, Books and Dashboards Quarto is the new thing. But R Markdown will still be out there! QUARTO VERSUS R MARKDOWN 56 CRS Primer: Dynamic Reporting Quarto •Official guide to get started •Explore the gallery : you can see source code of most entries R Markdown •Good tutorial series from R Studio R Markdown tutorial •Free book R Markdown: The Definitive Guide •R Markdown cheatsheet ADDITIONAL RESOURCES 57 "TRACK CHANGES" IN YOUR WORD PROCESSOR 64 •Requires a master document •Does not scale (only a few participants) •Does not provide the complete history (just the latest proposed changes) ISSUES 65 ONLINE DOCUMENTS ON COLLABORATIVE OFFICE SUITES 66 •Not FAIR •Stored in a given company's cloud •Impossible to download (and use) the complete history •Does not work with all types of documents ISSUES 67 68 What is version control? Version control (also known as revision control, source control, and source code management) is the software engineering practice of controlling computer files and versions of files; primarily source code text files, but generally any type of file. WHAT IS VERSION CONTROL ACCORDING TO WIKIPEDIA? 69 Source: https://en.wikipedia.org/wiki/Version_control, version of 17 August 2024, 02:21 Version control (also known as revision control, source control, and source code management) is the software engineering practice of controlling computer files and versions of files; primarily source code text files, but generally any type of file. A version control system is a software tool that automates version control. Alternatively, version control is embedded as a feature of some systems such as word processors, spreadsheets, collaborative web docs, and content management systems, e.g., Wikipedia's page history. WHAT IS VERSION CONTROL ACCORDING TO WIKIPEDIA? 70 Source: https://en.wikipedia.org/wiki/Version_control, version of 17 August 2024, 02:21 Version control (also known as revision control, source control, and source code management) is the software engineering practice of controlling computer files and versions of files; primarily source code text files, but generally any type of file. A version control system is a software tool that automates version control. Alternatively, version control is embedded as a feature of some systems such as word processors, spreadsheets, collaborative web docs, and content management systems, e.g., Wikipedia's page history. Version control includes viewing old versions and enables reverting a file to a previous version. WHAT IS VERSION CONTROL ACCORDING TO WIKIPEDIA? 71 Source: https://en.wikipedia.org/wiki/Version_control, version of 17 August 2024, 02:21 •Provide an introduction to version control •Discuss the Git software tool •Discuss the Github and Gitlab systems •Show a demonstration of how to use these tools GOALS OF THIS PART 72 •A version control system created in 2005 by Linus Torvalds to support the development of the Linux kernel. •Quickly adopted by many software developers •Free software, available under most computing platforms (Windows, Mac, Linux, …) •Git could stand for "Global Information Tracker" WHAT IS GIT? 73 CASE STUDY 1: SOFTWARE OR DOCUMENTS UPDATED 80 Code and data Revision 1.2.2024 Code and data Revision 2.2.2024 Code and data Revision 20.3.2024 Initial version Adapt for new data Use new analysis method Realize that the new analysis method does not work with the new data R code Revision 1.2.2024 R code Revision 2.2.2024 R code Revision 20.3.2024 R code Revision 1.11.2024 R code "Paper" Revision 1.3.2024 … R code "Paper" Revision 1.11.2024 … Initial version Format results for journal Correct bug Correct bug CASE STUDY 2: PARALLEL VERSION (BRANCHES) 81 Output results Update code for new data R code Revision 1.2.2024 Alan R code Revision 2.2.2024 Brad R code Revision 20.3.2024 Alan Initial version Adapt for new data Use new analysis method R code Revision 2.2.2024 Alan Format output for paper CASE STUDY 3: MULTIPLE DEVELOPERS 82 83 Git basics •Create a file in an empty directory (or take an existing textfile): echo "First Git example" > README •Initialize a git repository: git init •Tell git to include the file README in the index (stage the file): git add README •Commit staged changes to the repository: git commit -m "Initial commit for first Git example" •Check history of repository in log file git log A FIRST EXAMPLE 84 •Working directory: the actual files you are currently working with in the project schutz@laptop:~/git$ ls README schutz@laptop:~/git$ ls -a . .. .git README •Repository: database with all information needed to manage history (including revisions) of a project schutz@laptop:~/git$ ls .git branches config HEAD index logs refs COMMIT_EDITMSG description hooks info objects BASIC CONCEPTS 85 schutz@laptop:~/git$ git commit -m "Initial commit for first Git example" [master (root-commit) 332ba12] Initial commit for first Git example Committer: Frederic Schutz <schutz@laptop> Your name and email address were configured automatically based on your username and hostname. Please check that they are accurate. You can suppress this message by setting them explicitly: git config --global user.name "Your Name" git config --global user.email [email protected] After doing this, you may fix the identity used for this commit with: git commit --amend --reset-author 1 file changed, 1 insertion(+) create mode 100644 README PART OF THE RESULTS FROM THE EXAMPLE 86 88 Using Git Locally •Initialize a Git repository: $ git init (or clone an existing repository) •Edit/add files •Add/stage files/changes: $ git add (update the index) •Check state of index: $ git status (files can be tracked, ignored or untracked) •Commit staged changes: $ git commit USING GIT: MAIN STEPS 89 FILE STATUS LIFECYCLE 90 Source: Figure 2-1 from "Pro Git" We move now to some practical examples www.gitlab.uzh.ch/crsuzh/workshop-dynamic-reporting HANDS-ON 97 98 CONCLUSION The role of data stewards in relation to reproducible research and their area of support CAS in Data Stewardship UNIL 2024-2025 •Data stewards play a vital role in supporting researchers to achieve greater reproducibility, especially in the context of open science. •Their expertise in data management, curation, and sharing ensures that research data and related materials are properly organized, accessible, and reusable by others. ROLE OF DATA STEWARDS IN REPRODUCIBILITY AND RE-USE 99 •Data management planning •Documentation and metadata •Data integrity and quality assurance •Data sharing and licensing •Training and capacity building •Ethical and legal compliance DATA STEWARDS SUPPORT AREAS 100 •Guidance on best practices: Data stewards assist researchers in developing comprehensive data management plans (DMPs) at the start of a project, ensuring that data are well-organized, properly annotated, and structured for reproducibility. •Standardization of formats: They help researchers choose and use standardized file formats, metadata schemas, and ontologies that make it easier for others to understand, reproduce, and re-use the data. DATA MANAGEMENT PLANNING 101 •Ensure proper data documentation: Data stewards help researchers create detailed metadata, documenting how data were collected, processed, and analyzed. This ensures that others can reproduce findings and replicate studies using the same methods. •Promote FAIR principles: By encouraging researchers to make their data Findable, Accessible, Interoperable, and Reusable, data stewards improve the likelihood that research data can be discovered, understood, and reused by others, fostering reproducibility. DOCUMENTATION AND METADATA 102 •Implement data quality control: Data stewards may work with researchers to implement data validation and quality control measures, ensuring the accuracy and integrity of datasets throughout the research lifecycle. •Help with version control: They may assist in maintaining version control of datasets and code, making it clear how data have evolved and helping others reproduce findings based on the same version of the dataset. DATA INTEGRITY AND QUALITY ASSURANCE 103 •Support open data sharing: Data stewards guide researchers through the process of making their data publicly available, while ensuring compliance with ethical guidelines and legal restrictions (e.g., data privacy). •Advise on Licensing: They help researchers choose appropriate open licenses (e.g., Creative Commons, Open Data Commons) that clarify how others can reuse the data, promoting broader access and re-use without ambiguity. DATA SHARING AND LICENSING 104 •Provide training on data management tools: Data stewards offer workshops, tools, and resources that help researchers efficiently manage, share, and publish their data in ways that facilitate reproducibility. •Educate on open science practices: By training researchers in open science principles and tools (e.g., pre-registration, open data, open code), they foster a research culture that values transparency and reproducibility. TRAINING AND CAPACITY BUILDING 105 •Ensure compliance with regulations: Data stewards ensure that data sharing and reuse practices align with ethical standards, institutional policies, and legal frameworks (e.g., GDPR), enabling safe and reproducible data use. •Balance openness with privacy: They help researchers balance the openness of data with the need to protect sensitive information, providing guidance on anonymization or data protection strategies. ETHICAL AND LEGAL COMPLIANCE 106