scieee AI-readable full text Open interactive document viewer

Stronger together: Advancing the data repository ecosystem through strategic coopetition

Buys, Matt; Chandramouliswaran, Ishwar; Hahnel, Mark; Holmes, Kristi; Van Gulick, Ana

Abstract

This presentation summarizes the achievements and future directions of the NIH Generalist Repository Ecosystem Initiative (GREI), a collaboration launched in 2022 between seven major generalist repositories: Dataverse, Dryad, Figshare, Mendeley Data, OSF, Vivli, and Zenodo. Operating under a "coopetition" model, the initiative balances cooperation on foundational standards with competition on unique repository features. The session highlights four years of progress in enhancing the FAIR data ecosystem, focusing on the successful implementation of a common metadata schema based on DataCite standards and the adoption of standardized metrics for data usage and citation through Make Data Count principles. These efforts have improved data discovery and interoperability across the repository landscape. The presentation also details the evolution of GREI's extensive community engagement and demonstrates its tangible impact on data sharing and scientific discovery through real-world user stories. Looking ahead, the initiative's future direction will be significantly guided by the recommendations issued by the the NIH Council of Councils’ GREI Working Group in 2025. Key priorities for the future include: Shifting focus from policy compliance to more robustly supporting and measuring the impact of biomedical data reuse. Expanding direct engagement with biomedical researchers to better meet their needs. Exploring new technologies to further strengthen the data repository ecosystem and enhance interoperability.

Full text

Acknowledgement of Country We acknowledge and celebrate the First Australians on whose traditional lands we meet, and we pay our respect to their elders past and present. Stronger together: Advancing the data repository ecosystem through strategic coopetition Presented by the NIH Generalist Repository Ecosystem Initiative (GREI) October 13, 2025 Session agenda ●Laying the Foundation: The Strategic Imperative and Vision of GREI ●A Balancing Act: Navigating Competition and Collaboration in the GREI Ecosystem ●Highlights & impacts from 4 years of GREI ○Building a Cohesive Data Ecosystem: Evolving Metadata Standards ○Measuring What Matters: Standardizing Metrics for Data Impact ○The Evolution of GREI Community Engagement: Building Capacity and Fostering Trust ○From Infrastructure to Impact: Real-World Stories of Data Sharing and Discovery ●Charting the Course: The Future of the GREI Data Ecosystem ●Facilitated discussion and presenter Q&A Join our Slido session to share your feedback! slido.com join code: 3682553 Presenters Ishwar Chandramouliswaran Program Director, NIH Office of Data Science Strategy Mark Hahnel VP of Open Research, Digital Science (Figshare) Matt Buys Executive Director, DataCite Kristi Holmes Associate Dean for Knowledge Management and Strategy Northwestern University Ana Van Gulick Head of Customer Engagement for Government, Funders & Nonprofits, Digital Science (Figshare) Laying the Foundation: The Strategic Imperative and Vision of GREI Join our Slido session to share your feedback! slido.com join code: 3682553 A strong repository landscape provides researchers with multiple locations to share their data. The inclusion of established generalist repositories in the data sharing ecosystem is helping to catalyze the flexible sharing of data and other research outputs in any format that cannot be shared elsewhere. Domainspecific Repositories Generalist Repositories* Institutional Repositories �� �� �� �� �� �� Data Repository Landscape 🌲🌲 �� �� �� �� �� �� *Some generalist repositories serve as backbone infrastructure for institutional repositories. ●Data type ●Format ●Content ●Disciplinary focus ●Methodology Software Presentation Image Video Article Poster Data Generalist Repositories Accept Data Regardless of: In support of the planned launch of the NIH Policy for Data Management and Sharing, ODSS launched the Generalist Repository Ecosystem Initiative (GREI) in January 2022. Adapted from NIH Office of Science Policy 20042003 2007 2008 Model Organism Policy Genome-wide Association (GWAS) Policy 2012 NIH Public Access Policy (Publications) Big Data to Knowledge (BD2K) Initiative NIH Data Sharing Policy 2023 2013 Modernization of NIH Clinical Trials 2014 Genomic Data Sharing (GDS) Policy NIH Public Access Plan 2015 Cancer Moonshot 2016 All of Us Research Program 2017 Request for Information on Data Sharing NIH Data Science Strategic Plan 2018 NIH Data Commons Pilot NLM Strategic Plan Generalist Repositories Workshop 2020 2022 HHS Rule and NIH Policy on Clinical Trial Results Dissemination NIH Policy for Data Management and Sharing (DMS) White House Initiative (“OSTP Memo”) NIH’s Culture of Data Sharing A Balancing Act: Navigating Competition and Collaboration in the GREI Ecosystem Join our Slido session to share your feedback! slido.com join code: 3682553 Coopetition: Balancing Cooperation & Competition (Bengtsson & Kock,1999) Cooperation Competition Collaborate on common standards, features, and activities (e.g., metadata, metrics, identifiers, etc.) Compete on unique features: data visualization and analytics tool integrations curation strategies etc. VALUE LINE Inside GREI Repositories Distinct Attributes ●Organization: ○Nonprofit, academic, and commercial ●Infrastructure: ○Open source and proprietary ●Offer features such as: ○Data visualization ○File types and sizes ○Curation ○Licenses ○Controlled access Similarities ●FAIR data sharing across disciplines ●Adhere to repository best practices ●Leverage community metadata standards and PIDs Coopetition Model: Supporting the NIH Strategic Plan for Data Science Goal: Improve Capabilities to Sustain the NIH Policy for Data Management and Sharing Objective: Strengthen NIH Data Repository and Knowledgebase Ecosystem Enhance infrastructure and interoperability ●Develop shared standards ●Modernize data repository ecosystem Foster innovation and flexibility ●Encourage technological advancement ●Support repository growth Strengthen community engagement ●Promote shared learning ●Encourage the secondary use of data Coopetition in Action: GREI’s Collaborative Framework Governance Framework Streamline decision-making Encourage innovation Shared deliverables to enhance functionality and services Strategic Partnerships Shared training and outreach Collaborate with: DataCite, ROR, Make Data Count, The Carpentries Cross-repository consistency Implement shared metadata standards (e.g., DataCite) Align metrics using Make Data Count best practices Enable cross-repository linking via PIDs User-Focused Approach Harmonize metadata across repositories Incorporate community feedback loops Develop GREI training resources Driving Shared Progress: GREI Coopetition 2 3 4 5 1 Foster Data Sharing & Reuse Culture Through Collaboration ●GREI Code of Conduct Enhance Community Engagement & Build Capacity ●Webinars, Workshops, Training resources ●Community feedback Document Use Cases ●Real-world, cross-disciplinary use cases Develop & Implement Open Common Metrics ●Make Data Count standards for data usage tracking Advance Metadata Standards & Interoperability ●GREI Metadata Recommendations 1 2 3 4 5 Coopetition Drives Innovation ●Enhance governance efficiency ●Streamline workflows ●Continuously improve operational performance ●Foster a partnership-oriented mindset ●Accelerate innovation through knowledge sharing ●Engage the community to increase transparency and accountability VALUE LINE https://bit.ly/GREIcoopetition PREPRINT: Coopetition as a Model to Advance a FAIR Data Ecosystem Strategic Goals for Future Collaboration Measure Impact with Meaningful Metrics Support Data Reuse Scalable and impactful biomedical data reuse. Expand Community Engagement Inclusive development through community collaboration. Robust metrics to measure data reuse frequency and scientific impact. Highlights & impacts from 4 years of GREI Building a Cohesive Data Ecosystem: Evolving Metadata Standards in GREI Join our Slido session to share your feedback! slido.com join code: 3682553 A Collaborative Effort to Standardize Data Sharing One of the founding objectives goals of the GREI program has been to implement consistent metadata models to enable discovery of NIH funded data in generalist repositories. All GREI metadata work is based on the DataCite metadata schema. Year 1 Common Fields: ●Assess metadata schemas across GREI repositories ●Establish common baseline set of metadata fields ●Agree to use DataCite as common schema Year 2 & 3 Consistent Schema Version 1: ●Identify metadata that meets use cases & enables common metrics ●Review quality of metadata captured across GREI ●Publish & implement metadata recommendation Version 1 to to standardize the use of optional DataCite properties. Year 4+ Consistent Schema Version 2: ●Refine and publish metadata recommendation version 2 based on community needs, repository implementation experiences, and emerging best practices ●Implementation to resolve gaps between internal repository schema and Version 2 GREI Metadata Recommendation Version 1 (doi.org/10.5281/zenodo.8101956) and Version 2 (doi.org/10.5281/zenodo.16953588) Looking Ahead: Continuing to Evolve with Community Needs The NIH Council of Councils recommended that GREI shifts focus from data sharing for policy compliance towards prioritizing and enabling data reuse. To meet this recommendation, we plan to: Explore AI and automation to improve quality and efficiency of metadata creation and enhancement. Identify and develop support for high-priority data types that do not have a domain-specific repository home. Add capabilities that allow datasets to be searchable and findable by different parameters. Implement product updates to resolve gaps between repository metadata and Version 2 recommendations. Highlights & impacts from 4 years of GREI Measuring What Matters: Standardizing Metrics for Data Impact in GREI Join our Slido session to share your feedback! slido.com join code: 3682553 GREI Common Metrics | Year 4 Focus GREI objective: Implement open metrics ●Provide repositories with information about interactions with the data they host ●Increase visibility about data value and encourage data reuse ●Lay groundwork to better understand the use of NIH-funded data The GREI repositories adopted Make Data Count recommendations for the implementation of data usage indicators to help achieve these goals. Year 4 activities: ●All-GREI Metrics Dashboard (Task 4.3) ●Metadata Completeness Dashboard (Tasks 2.5 & 2.6) ●Usage Metrics via Make Data Count integration ●Goal: consistent, comparable, actionable insights for repositories, GREI, and NIH Metrics Dashboard ●Purpose: Establish consistent set of cross-GREI metrics for NIH-funded data. ●MVP Features : ○“Funded By” facet (using ROR + Crossref IDs, incl. child orgs) ○Repository facet (single + future multi-select) ○Widgets: citations, views, downloads ○Stable, shareable links + API access ●Next Steps: Iterative build and test with GREI repositories. Metadata Dashboard ●Purpose: Help repositories assess and improve metadata quality . ●Features: ○Completeness overview by repository, consortium, org ○Metrics for 20+ schema properties + sub-properties ○Filtering by resource type, registration year, etc. ○Guidance on high-impact improvements ●Feedback: Align with GREI Metadata & Search Subcommittee; refine based on community input. Usage Metrics Integration ●Based on Make Data Count (MDC) ●GREI repositories now collect & share normalized usage counts ●Benefits: ○Insights into NIH-funded data use & impact Comparable usage measures across repositories ○Trends monitoring & outlier detection ○Adds recognition for dataset creators ●Next Steps: ○Ensure usage stats surface on repository pages ○Add FAQ & implementation support Standardizing Data Citation to Measure and Incentivize Reuse Moving beyond simple usage counts, GREI collaborated with Make Data Count to develop the GREI Data citation best practices for repositories to provide clear evidence of data reuse and credit for researchers. GREI Best Practices recommendations: ●Collect: Harvest citations from community sources (e.g., DataCite, Crossref, Europe PMC) and user submissions. ●Standardize: Use consistent DataCite metadata fields (e.g., relatedIdentifier, relationType) to define the citation link. ●Display: Clearly show the number of citations and their sources on every dataset landing page for transparency. ●Aggregate: Contribute all collected citations back to the community via DataCite to build a global Data Citation Corpus. Outcome: A More Connected Ecosystem For Researchers: Easier to receive academic credit for their data. For Funders: Provides a reliable, cross-repository metric for evaluating the impact of funded research. For the Community: Creates a richer, more transparent network of linked research objects, accelerating discovery for everyone. GREI Data citation best practices for repositories: https://doi.org/10.5281/zenodo.10562428 The collaboration between the GREI repositories, the DataCite team and Make Data Count has made it possible to implement the collection and sharing of normalized metrics. Benefits Insights into the use and impact of funded data in ways not possible when the GREI initiative started. These usage counts will feed into the GREI metrics dashboard to be completed in Year 4. Normalized usage counts that the repositories can add to data records to showcase the value of datasets and give recognition to its creators. Insights to monitor repository use trends over time and to identify outliers or odd behaviours. Better understanding of implementation needs for the usage tracker. Mechanisms to work with repositories to showcase dataset usage and to signal if there are implementation needs or unexpected trends. Delivering on GREI metrics through collaboration Looking Ahead: Evolving Impact Metrics The NIH Council of Councils recommended that metrics development be focused specifically on demonstrating the impact of biomedical data reuse, moving beyond general usage statistics like page views and downloads. To meet this recommendation, we plan to: Explore partnership strategies around the development and repository implementation of the S-Index Focus on metrics illustrating the impact of data reuse and connecting the metrics back to real world user stories. Continue to develop and enhance the DataCite All GREI metadata & metrics dashboards Highlights & impacts from 4 years of GREI The Evolution of GREI Community Engagement: Building Capacity and Fostering Trust Join our Slido session to share your feedback! slido.com join code: 3682553 A Multi-Faceted Approach to Direct Community Engagement Broadening Reach Through Partner Webinars Publicly offered GREI Collaborative webinars as well as partner webinars with FASEB, NNLM, and NIH Library, covering topics from creating Data Management and Sharing Plans to understanding repository features, reaching a global audience. Guiding Development with Community Feedback Community feedback calls and interactive workshops and conference sessions (e.g., Spring 2024 Calls, Jan 2023 Workshop) to solicit direct input on GREI resources, ensuring our tools and guidance meet the evolving needs of the research community. Strengthening the Broader Data Ecosystem Domain-specific repository community engagement to foster a complementary and interconnected, rather than competitive, relationship through presentations at the NIH DRKB Network Program Meeting in both 2024 and 2025 to share progress on common metadata and metric standards. To: Evolution of Engagement: From Broad Awareness to Targeted Support Moving from: Broad Webinars General Guidance One-Way Communication Librarian Focus Targeted, Thematic Workshops & User Stories Practical, Actionable Resources (Checklists, Flowcharts) Active, Transparent Feedback Loops Expanded Outreach Directly to Biomedical Researchers GREI's community engagement has successfully raised awareness through hosting or co-hosting 15+ webinars, presenting our work at over 20 conferences, creating essential resources (over 35,000+ downloads), and building a responsive framework to support the research community in navigating data sharing. Looking Ahead: Doubling Down on Researcher Engagement The NIH Council of Councils recommended that GREI expand engagement with biomedical researchers and tailoring outreach activities to this specific audiences. To meet this recommendation, we plan to: Develop a conference toolkit for gathering strategic researcher feedback. Launch a scientific advisory board to provide strategic guidance on GREI priorities, outreach, and long-term strategy. Identify collaborations with domain-specific repositories. Focus on the researcher audience through targeted outreach activities. Highlights & impacts from 4 years of GREI From Infrastructure to Impact: Real-World Stories of Data Sharing and Discovery Join our Slido session to share your feedback! slido.com join code: 3682553 GREI Real-World User Stories To demonstrate the real impact of the GREI repositories, each repository has begun collecting stories from real repository users about their experiences sharing and reusing biomedical data in generalist repositories. Through these stories we have learned how researchers have, for example: GREI Real-World User Stories can be found at: https://bit.ly/GREIUserStories Shared complex, multimodal datasets that have been reused by other labs, leading to new discoveries and collaborations. Combined and reused data from separate clinical trials to create new disease models and directly inform national clinical practice guidelines. Published analysis code, software, and workflows, improving reproducibility and accelerating research in computationally intensive fields like multi-omics and biomedical informatics. Sharing the Brain Genomics Superstruct Project (GSP) Researcher: Dr. Buckner, Randy L.(Harvard, MGH, HMS) The Research: The Brain Genomics Superstruct Project Open Access Data Release exposes a carefully vetted collection of neuroimaging, behavior, cognitive, and personality data for over 1,500 human participants. Repository Use: Dr. Buckner chose the Harvard Dataverse Repository due to the General support and excitement over a Harvard open data repository. The reuse dataset identified in this real world user story is is also shared on Harvard Dataverse. Data Sharing and Reuse User Stories “It’s been amazing to watch the field change. When you ask about sharing standards in our field, I think most people — especially if they receive NIH funding — follow them. But more broadly, because foundations…run by the Coalition for Aligning Science all encourage open, transparent data sharing, it has just now become the new standard. And it wasn’t 30 years ago.” https://doi.org/10.5281/zenodo.17081413 Exploring data sharing practices through Dryad: A researcher’s insight Data Sharing and Reuse User Stories "Data sharing becomes extremely important for multi-institution collaborations, as well as training activities we have with scholars from other laboratories.” https://doi.org/10.5281/zenodo.17038533 Researcher: Dr. Troy D. Wood, University at Buffalo The Research: The data benefits researchers studying soybean metabolomics and stress-adaptive phenotypes, and offers valuable examples for mass spectrometry practitioners applying the Kendrick mass defect in complex biological samples. Repository Use: Dr. Wood chose Dryad to host the dataset due to his previous positive experience publishing data on Dryad and the platform’s expert team of curators—“real people who provide hands-on support for every dataset to ensure quality, consistency, and usability.” Open Science in Neuroscience Researcher: Oliver Contier, Max Planck Institute for Human Cognitive and Brain Sciences The Research: The THINGS-data project provides a massive collection of brain and behavioral data to shed light on how we see and distinguish between different objects. Repository Use: The research team chose Figshare as the central repository for THINGS-data due to its flexibility in organizing large, multimodal collections. To further enhance accessibility and meet community standards, THINGS neuroimaging data was shared on OpenNeuro.org, while behavioral data was shared via the Open Science Framework (OSF). Data Sharing and Reuse User Stories “Reusing open data can be particularly beneficial for early-career researchers, helping to minimize risks when resources are limited…and sharing one's own data can increase visibility and help develop valuable technical skills.” https://doi.org/10.5281/zenodo.16970812 Advancing Health Equity: Sharing Obesity Prevention Policy Data Researcher: Renee Parks, MS, Senior Research Manager, Prevention Research Center, Washington University in St. Louis The Research: Surveyed 195 U.S. practitioners, policymakers, and researchers to identify the most impactful and feasible local policy actions for advancing health equity in obesity prevention, and shared de-identified survey responses to support transparency, reproducibility, and further research. Repository Use: Used WashU’s Digital Commons Data@Becker, powered by Mendeley Data, to share both open access and restricted datasets, ensuring ethical data sharing and compliance with funder requirements. Benefited from strong support for metadata, documentation, and access request workflows, making the process smooth and supportive. Data Sharing and Reuse User Stories “We believe our science should be made available to the public, especially since most of our research projects have been funded by federal funding. It allows others to explore the data from different angles, apply it to new questions, or use it in educational settings.” User Story Pending Publication Pattern Glare Sensitivity Distinguishes Subclinical Autism and Schizotypy Researchers: Wendy A. Torrens, Jenna N. Pablo, Marian Berryhill, Sarah M. Haigh, Department of Psychology and Institute for Neuroscience, University of Nevada, Reno, USA The Research: This study behaviorally identified early visual processing differences between non-clinical individuals with autism-like traits and those with schizotypy traits. Schizophrenia and autism symptoms substantially overlap clinically, which heightens risk of misdiagnosis. Repository Use: The research team found sharing data through OSF to be convenient, simple, and streamlined compared to other repositories they’ve used. They find OSF’s data storing system easy to use, with an intuitive organizational structure. Data Sharing and Reuse User Stories “Sharing data has harbored a culture of accountability by pushing us as researchers to have well-documented and readable code, as well as easy-to-use ReadMe files, etc. It has also facilitated new connections and collaborations across the globe.“ https://doi.org/10.5281/zenodo.17046374 The Working Group recommended focusing on: ●Expand engagement with biomedical researchers ●Engage with others in repository landscape (domain-specific repos) ●Contribute to FAIR-enabling methods and approaches ●Establish a steering committee from the community How GREI plans to respond: ●Engage biomedical researchers through focused outreach events ●Design a feedback framework for engaging researchers at conferences ●Collaborate with domain-specific repositories ●Launch a scientific advisory board Expand Community Engagement Connect with GREI Stay in touch with the GREI repositories - ask questions, provide feedback, get updates, learn about future events: ●Join the GREI mailing list: groups.google.com/g/contactgrei ●Read the GREI blog: medium.com/@blog-grei ●GREI GitHub Repository: github.com/NIH-GREI/grei-pm ●GREI GitHub Community Discussion Board: github.com/NIH-GREI/grei-pm/discussions bit.ly/GREIZenodo Access all published GREI resources via the GREI Zenodo Community → Audience discussion and presenter Q&A Join our Slido session to share your feedback! slido.com join code: 3682553 1. What work should GREI prioritize to encourage researchers to share and reuse data? 2. What work should GREI prioritize to enhance the discoverability and impact measurement of data and to promote interoperability across the data repository landscape? 3. How can GREI successfully engage researchers? 4. Which organization would you most like to see GREI partner with and why? 5. With your primary role in mind, what is one “killer feature/functionality” that would improve the use of a data repository for YOU? Audience members - we welcome live participation at this time but if you prefer to submit a response via text, please submit via Slido. Thank you! Access all GREI resources via the NIH ODSS GREI landing page → bit.ly/ODSSGREI →