scieee AI-readable full text Open interactive document viewer

find.software: Foundations for Interdisciplinary Discovery of (Research) Software

Gey, Ronny; Mietchen, Daniel; Karras, Oliver; Wittenborg, Tim; Schubotz, Moritz; Bumberger, Jan

Abstract

Across essentially all fields of research, many aspects of the respective research processes – whether experimental, theoretical, empirical or outright computational – are closely related to software. Yet the process of finding software that is directly suitable or at least a good starting point for a given research task is cumbersome.This project aims to develop a community-driven system that provides potential users of research software with a diversity of pathways towards actually finding software that closely matches their research needs if such software exists. Conversely, it will provide software developers with mechanisms to make their software findable for research-related tasks and it will highlight mismatches between software supply and demand for specific tasks.To this end, we will document how various stakeholders of the research landscape have been searching for – or stumbling upon – research software so far, identify variables associated with successful search outcomes and build workflows that assist in describing software and associated concepts in a standardised fashion. These descriptions will then be aligned across various sources of relevant information and integrated into Wikidata, the knowledge graph that anyone can edit and that already contains considerable breadth and depth of information related to research, software and their interactions.While keeping an eye on similar approaches to software discovery that might work in parts of the research ecosystem, existing Wikidata content and workflows will be reviewed and built upon. Additional documentation, tooling and workflows will be developed to enrich, expand, curate, query and explore this content, both for specific use cases and with ongoing engagement of the communities involved in research software, open data or collaborative curation. Within its three years, the project seeks to establish a dedicated community overseeing a well-documented and smoothly running infrastructure for software discovery and to devise a plan for how this can be sustained for the longer term.

Full text

Research Ideas and Outcomes 11: e179253 doi: 10.3897/rio.11.e179253 Reviewable v 1 Grant Proposal find.software: Foundations for Interdisciplinary Discovery of (Research) Software Ronny Gey , Daniel Mietchen , Oliver Karras , Tim Wittenborg , Moritz Schubotz , Jan Bumberger ‡ Helmholtz Centre for Environmental Research (UFZ), Leipzig, Germany § FIZ Karlsruhe — Leibniz Institute for Information Infrastructure, Berlin, Germany | Institute for Globally Distributed Open Research and Education, Jena, Germany ¶ TIB - Leibniz Information Centre for Science and Technology, Hannover, Germany # L3S Research Center, Leibniz University Hanover, Hannover, Germany ¤ German Centre for Integrative Biodiversity Research (iDiv), Halle-Jena-Leipzig, Germany Corresponding author: Ronny Gey ([email protected]), Daniel Mietchen ([email protected]), Jan Bumberger ([email protected]) Received: 21 Nov 2025 | Published: 03 Dec 2025 Citation: Gey R, Mietchen D, Karras O, Wittenborg T, Schubotz M, Bumberger J (2025) find.software: Foundations for Interdisciplinary Discovery of (Research) Software. Research Ideas and Outcomes 11: e179253. https://doi.org/10.3897/rio.11.e179253 Abstract Across essentially all fields of research, many aspects of the respective research processes – whether experimental, theoretical, empirical or outright computational – are closely related to software. Yet the process of finding software that is directly suitable or at least a good starting point for a given research task is cumbersome. This project aims to develop a community-driven system that provides potential users of research software with a diversity of pathways towards actually finding software that closely matches their research needs if such software exists. Conversely, it will provide software developers with mechanisms to make their software findable for researchrelated tasks and it will highlight mismatches between software supply and demand for specific tasks. To this end, we will document how various stakeholders of the research landscape have been searching for – or stumbling upon – research software so far, identify variables associated with successful search outcomes and build workflows that assist in describing software and associated concepts in a standardised fashion. These descriptions will then ‡ §,| ¶ # § ‡,¤ © Gey R et al. This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. be aligned across various sources of relevant information and integrated into Wikidata, the knowledge graph that anyone can edit and that already contains considerable breadth and depth of information related to research, software and their interactions. While keeping an eye on similar approaches to software discovery that might work in parts of the research ecosystem, existing Wikidata content and workflows will be reviewed and built upon. Additional documentation, tooling and workflows will be developed to enrich, expand, curate, query and explore this content, both for specific use cases and with ongoing engagement of the communities involved in research software, open data or collaborative curation. Within its three years, the project seeks to establish a dedicated community overseeing a well-documented and smoothly running infrastructure for software discovery and to devise a plan for how this can be sustained for the longer term. Keywords software discoverability, research software, software discovery pathways, Wikidata Starting point State of the art Software has become an integral part of the research landscape. According to DOI registrar DataCite, there were 608,539 DOI registrations for software as of February 2025* . This includes 132,168 software DOIs that have been registered in 2024 alone and the trend continues upwards. The amount and diversity of software developed and used in research projects are so high that it is becoming increasingly challenging for participants in the research process – be they researchers, research software engineers (RSEs, cf. Goth et al. (2025)), research administrators, research funders or students – to make informed decisions when searching and selecting research software (RS) for specific research-related tasks. Challenge 1 – Search inefficiency: The process of searching for – and ultimately finding – relevant RS is often inefficient (Hucka and Graham 2018): the search is cumbersome, there is no guarantee that the software found is suitable for an intended task or that software suitable for a given task will be found by those looking for it. While software discovery has received less attention than discovering other types of resources – for example, literature (Kang et al. 2023) or data (Contaxis et al. 2022) – some insights from the study of the latter ones are transferable to the former. For instance, the concept of discovery pathways (Thagard 2022) as a series of cognitive steps towards a discovery goal can be applied to software discovery, as can existing typologies in this space (Nishikawa-Pacher 2021) or ideas around workflows that extend beyond papers as the main starting points for discovery (Kang et al. 2023). There is, indeed, a wide range of available discovery pathways, which vary greatly depending on the context (Wittenborg 1 2Gey R et al et al. 2025). Searching for software may, for instance, entail exploration of various code or publication repositories, search engines or package managers, each of which having its own search functions and metadata standards. Furthermore, the type of software being searched for, such as data analysis tools, visualisation libraries or simulation frameworks, influences the choice of search strategy and platform. Parameters, such as the scientific domain, familiarity with programming languages, licence considerations and specific use cases, can also influence the search strategy. For example, a bioinformatician might use bio.tools* to search for suitable software, while physicists might prefer physics.tools* , and humanities' scholars the Social Sciences and Humanities Open Marketplace* . In addition, there are approaches at the institutional (e.g. Helmholtz Research Software Directory - cf. Maassen (2023)) or national level (e.g. NFDI.software - cf. Castro et al. (2025)). This complex network of discovery pathways leads to researchers having difficulties finding relevant software that meets their research needs, which ultimately hinders the progress of their work. Challenge 2 – Lack of annotations: RS is often insufficiently annotated regarding its functionality or potential scope of usage, making it challenging to understand its capabilities, compatibilities, limitations and actual or potential applications. Many software packages lack comprehensive documentation detailing their algorithms and associated assumptions, input/output formats or usage examples. As a result, researchers and other potential users face difficulties determining whether a particular software, tool or workflow meets their specific needs. This lack of clarity also hinders research reproducibility, since users may not be able to accurately replicate results without a clear understanding of the software’s functionality, parameters, dependencies or interactions with other components of relevant research workflows. Despite promising initiatives for standardised software descriptions (Jones et al. 2017, Druskat et al. 2021, Garijo et al. 2021), their consistent implementation in practice remains lacking, and those descriptions focus on some discovery pathways, while not covering others. Challenge 3 – Software sustainability: The development of RS, the many initiatives in the area of research software engineering (RSE) and a variety of RSE services that are perceived as infrastructure often rely on a very limited group of people (Hettrick et al. 2022). These individuals are crucial to the development, maintenance and support of the tools and platforms that underpin RS, but compete for attention with their own research or professional obligations. This dependency on a few individuals jeopardises the long-term sustainability of these infrastructures and initiatives. More stable and sustainable models are needed to support the continued development, deployment and maintenance of reliable and effective RS infrastructures, in general and for RS discovery, in particular. Challenge 4 – Research silos: RS infrastructures, such as the European Open Science Cloud* or Germany’s National Research Data Infrastructure (NFDI)* , are often built within national or institutional boundaries, making them less accessible to researchers outside those boundaries. These infrastructures are designed to provide researchers with essential resources and tools for managing, sharing and analysing data. However, the geographic and organisational restrictions create barriers to international collaboration, as researchers from different countries or institutions may not be able to access these 2 3 4 5 6 find.software: Foundations for Interdisciplinary Discovery of (Research) ... 3 infrastructures, resulting in fragmented data and tools making it difficult to share data and knowledge across borders. This highlights the need for more inclusive and globally accessible RS infrastructures to support a truly open and collaborative scientific community. Initiatives like swMATH – which annotates mathematical publications indexed in zbMATH Open with the underlying software (Azzouz-Thuderoz et al. 2022) – can help overcome such administration-related silos. We will work closely with swMATH as well as with MaRDI – the mathematical NFDI consortium (Schubotz et al. 2023) – and the NFDI Basic service nfdi.software* to extend the functionality of swMATH and to use it as a starting point to enable interlinking between software and other research resources also for other disciplines Needs analysis As RS discovery has not received much attention so far (see above) and the associated needs are not well circumscribed. We thus decided to conduct an initial analysis of user groups for RS discovery in the form of personas, as well as to identify a total of 12 scenarios for a search for RS (Gey et al. 2025). Based on these findings, we conducted a workshop with representatives of the RSE community (Wittenborg et al. 2025), which led to the following conclusions: • One of the greatest challenges is the lack of a coherent mechanism for researchers to find relevant software. Such a mechanism serves as the basis for a single, unified platform and would be particularly beneficial for users who are less experienced in dealing with software and do not have the necessary knowledge of the different discovery pathways. • Infrastructure for discovering RS should be publicly funded and administered. This would benefit the stability of the infrastructures and their independence from commercial interests. • Infrastructure for discovering research software should be supported and steered by the community, especially with regard to content. This would ensure that it is aligned with the needs of users and content curation is open and up to date. • Software developers can make a significant contribution to the findability of their software by providing clean and enriched metadata using suitable metadata schemas and documentation. Developers should be able to use easily accessible, collaborative tools for annotation. The results of the workshop are confirmed by a position paper of the German association for research software, deRSE e.V. (Anzt et al. 2021). In the paper, the authors outlined challenges for sustainable research software in Germany and consider RS discovery as a prerequisite for enabling the reuse of existing software and avoiding redundant development. However, insufficient information retrieval strategies and a lack of knowledge about relevant repositories are crucial factors hindering a sustainable RS discovery (Anzt et al. 2021). 7 4Gey R et al Environmental analysis The discovery of RS is facilitated by various platforms, tools and workflows that can be grouped into different categories, of which we will briefly outline some of the most commonly used ones (a detailed environmental analysis was recently conducted by the authors - cf. Gey et al. (2025)). Code repositories such as GitHub, GitLab and Codeberg provide management and collaboration features, while mixed content publication platforms, such as Zenodo, allow researchers to publish their research outcomes more generally, including software. Software repositories or package repositories are centralised locations for software packages and may be specific to programming languages or operating systems. Examples include PyPi for Python, CRAN for R or the App Store for iOS. Software archives, such as Software Heritage, preserve code for future generations, while catalogues, including domain-specific, institutional and national/supranational catalogues, collect metadata and point to persistent RS repositories. Curated lists, such as the Awesome Lists on GitHub, are collections of high-quality content on specific topics and are often used to search for topic-specific software. Search engines, including general-purpose search engines like Google and science-orientated competitors like BASE, provide comprehensive strategies for aggregating content and search functions. Specialised search engines, such as Betty’s ReSearch Engine, specialise in research software. Software journals, such as JORS and JOSS, publish text articles on software publications, making software searchable. Social networks can be a valuable resource for RS discovery, too, as researchers often rely on their academic network for information (Hucka and Graham 2018, Murphy-Hill et al. 2015). Furthermore, large language models can be used for RS discovery to identify relevant tools and provide information about their functions and capabilities. Last, but not least, knowledge graphs provide a structured approach to organising and linking diverse pieces of information (including from all the other categories mentioned before), facilitating semantic search capabilities and discovery of RS tailored to specific research purposes. They can also support recommendation systems, provide contextual information and lead to engaging visualisations. A key feature of knowledge graphs in the context of research software discovery is that many discovery pathways can be supported, with every node potentially serving as an entry point or as a seed for more specific subgraphs and every edge as a bridge to other parts of the graph or even to other graphs, databases or further resources. The Knowledge Graph Infrastructure KGI4NFDI* is an initiative to review the usage of scholarly knowledge graphs within NFDI and to support disciplinary communities in choosing and using suitable knowledge graph approaches. We are involved and we will closely monitor KGI4NFDI activities from the perspective of RS discovery, consulting with related initiatives like NFDI.software as appropriate. There have been prior efforts regarding knowledge graphs for software. They can be classified in several ways, one of which would be whether they are based on stand-alone (Schubotz et al. 2023, Kelley and Garijo 2021, Samuel and Mietchen 2024a) or community-driven infrastructures, particularly Wikidata (Thornton et al. 2018, Rasberry et al. 2022). 8 find.software: Foundations for Interdisciplinary Discovery of (Research) ... 5 Wikidata For this initial phase of the implementation of find.software, we have chosen to concentrate on Wikidata, for several reasons. First, Wikidata is open, transdisciplinary and multilingual, which forms a good foundation for a potentially large user base. Second, Wikidata runs on Wikibase – a tried and tested backend for collaborative curation of structured data – that can also be used in a stand-alone fashion, as is the Figure 1. find.software - A knowledge graph for research software. This project aims to create a platform – find.software – that provides its users with support for a large and diverse array of research software discovery pathways. It is based on Wikidata and has community, technical and data aspects that it weaves together through data integration from various sources, data curation via MediaWiki-based tools, data visualisation via Scholia and a rich set of interactions at the interfaces between science, software and Wikidata.  6Gey R et al case with MaRDI (Schubotz et al. 2023). Third, a worldwide community has already formed around Wikidata and Wikibase, from whose experiences and organisational structures we want to benefit and to which we want to contribute with find.software for the research software domain. Fourth, the barriers for the participation of individuals are as low as possible, which facilitates broad participation in the design of find.software. Fifth, thanks to diverse community efforts, the Wikidata database already contains many software-related items (e.g. 13000 direct instances of software* , over 5000 subclasses of software* and millions of instances of software or any of its subclasses* ), along with over 100 Wikidata properties related to software* , of which over 60 have the term software in their English label, description or aliases* . This forms a good basis for supporting a range of discovery pathways for software in general and – in light of the strong presence of research-related concepts in Wikidata (Nielsen et al. 2017, Waagmeester et al. 2020) and efforts to establish links between them and software (Rasberry and Mietchen 2022) – for research software in particular. In addition, MaRDI – the mathematical consortium within Germany’s National Research Data Infrastructure (NFDI) – uses a stand-alone Wikibase instance as a productive backend for its own portal for research data management in mathematics (Schubotz et al. 2023), which is integrated with swMATH to facilitate software discovery in the field. This software stack can be reused by find.software and serve as a starting point, especially since our project team already has considerable experience in working with Wikidata and Wikibase. find.software - A knowledge graph for research software Taking into account the identified needs of the target audience and the identified existing alternatives for software discovery (both previous section), we would like to propose find.software as a solution to enable cross-domain and community-maintained software discovery. With find.software, we are setting up a functional prototype (cf. Fig. 1) and thus assign ourselves to Phase 1 (Set-up and testing) of the DFG funding programme Research Software Infrastructures. Technical aspects Key considerations in knowledge graph construction are the kinds of entities that should be represented there (nodes in graph terminology, subjects in terms of semantic webs or items in Wikidata terms) and what kinds of relationships (edges, predicates and properties, respectively) should exist within the graph and with external pieces of information. For specific types of nodes, suitable data models need to be considered as well, ideally in a way that allows for alignment with external identifiers for the respective concepts, especially in ontologies. As mentioned above in the environmental analysis, Wikidata already has considerable coverage of both software and research. However, their intersections – software research and research software – are not systematically covered there. Before designing data models for research software-related concepts, we would thus document the existing data models (whether defined in some way or inferred from usage practice) and their alignment with external resources like the Software Ontology (Malone et al. 2014) or software categories (Hasselbring et al. 2025). On that 9 10 11 12 13 find.software: Foundations for Interdisciplinary Discovery of (Research) ... 7 basis, we would identify potential improvements relevant to research software, propose corresponding changes to the Wikidata community (e.g. through the software section of WikiProject Informatics* or through the property proposal process* ) and work with it towards practical and implementable solutions. Once the data model for research software in Wikidata is stable and applied consistently (Thornton et al. 2019, Turki et al. 2022) in the realm of research software, our focus will shift towards identifying and documenting gaps in Wikidata's coverage of research software-related matters, paying special attention to the use cases identified in the environmental analysis above. These gaps will be mostly related to missing statements on existing Wikidata items or properties, missing references for existing statements or missing Wikidata items. Filling the gaps would require finding suitable sources for the relevant information. Suitability here has multiple dimensions, including technical ones (fit to the Wikidata data model, availability of batch processing workflows), legal ones (especially licensing – as a public-domain database, Wikidata can only import public-domain materials) and community ones (e.g. whether the data are notable in Wikidata terms, of sufficient interest to the Wikidata community and maintainable by it). It is quite possible that a given source fulfils these criteria only partially: for example, the FAIR Jupyter knowledge graph (Samuel and Mietchen 2024a) has reproducibility data about individual Jupyter notebooks and even individual code cells therein, both of which would not meet Wikidata’s notability criteria, but the more coarse-grained information in which an individual scholarly article describes the use of Jupyter notebooks would fit in. Once suitable sources are identified, data have to be collected from them. This will be an iterative process, for which existing tools like the Software Metadata Extraction Framework (SOMEF) (Garijo et al. 2021) will be explored. The collected data needs to be normalised for the Wikidata data model and deduplicated before ingestion (Turki et al. 2022). In combination with other workflows that continuously curate Wikidata content, research software-related entries will then be enriched in an ongoing fashion. As a result, the content related to research software can be explored in ever more detail using tools like Scholia (Rasberry et al. 2022, Nielsen et al. 2017). Community aspects With find.software, we are responding to the call, already expressed in the renewal of the Budapest Open Access Initiative (BOAI20 Steering Group 2022) and recently taken up by RFII (RfII - German Council for Scientific Information Infrastructures 2024) or DFG (Deutsche Forschungsgemeinschaft. Ausschuss für Wissenschaftliche Bibliotheken und Informationssysteme 2025 to base research work on open, community-controlled infrastructure, to promote cooperation and shared responsibility for it and, thus, to maximise the distribution of knowledge and to minimise access inequalities to knowledge infrastructures. We will involve both the scientific and the open knowledge community in our endeavour from the very beginning, if possible even before the actual project start. Therefore, in WP3, we will network with scientific (NFDIs, deRSE) and open knowledge communities (Wikidata/ Wikibase communities) in the project environment and establish governance and decision-making mechanisms for find.software that involve representatives from these communities. Such community involvement will be further supported by joint workshops, 14 15 8Gey R et al a find.software roadshow, user and developer documentation as well as scientific communication. Preliminary work FIZ Karlsruhe FIZ provides scientific information, infrastructure and services in support of research and innovation. It offers access to databases, journals and patents, particularly in STEM fields. The institute conducts research in information science, data management and knowledge discovery, with a focus on knowledge graphs to structure and analyze complex data and develops research software for text mining, semantic search and data analysis. FIZ promotes open access to scientific knowledge and manages key resources for disciplinary communities, including zbMATH Open (Bär 2024, Rahkooy et al. 2025) – a comprehensive database for mathematical publications jointly operated with the Heidelberg Academy of Sciences and the European Mathematical Society – and swMATH, a database for software cited in zbMATH-indexed publications. Furthermore, FIZ contributes to Germany’s National Research Data Infrastructure (NFDI) through multiple disciplinary consortia – for example, NFDI4Chem (Steinbeck et al. 2020), NFDI4Culture (Altenhöner et al. 2020) or MaRDI (The MaRDI consortium 2022) – where its role often involves managing data, including via knowledge graphs (Schubotz et al. 2023, Conrad et al. 2024). Activities with a close relationship to the find.software project comprise contributions to the development of scholarly content in Wikidata (Waagmeester et al. 2020), to the discoverability of clinical trials via Wikidata (Rasberry et al. 2022) and to the development of the Wikidata frontend Scholia (Nielsen et al. 2017), including adaptations for software (Rasberry and Mietchen 2022). They also include the analysis of computational reproducibility of research-related Jupyter notebooks (Samuel and Mietchen 2023, Samuel and Mietchen 2024a, Samuel and Mietchen 2024b) and systematic archiving of research software (Ramy-Badr et al. 2024), as well as the analysis of deep dependencies of software cited from research publications (Nesbitt et al. 2024) and the analysis of GitHub repositories with respect to the use of Wikidata and Wikipedia (Turki et al. 2024) or concerning the demographics of contributors to open-source repositories (Levitskaya et al. 2022). TIB The ERC Consolidator Grant ScienceGraph (829536) developed the Open Research Knowledge Graph (ORKG) (cf. http://orkg.org), a platform for scholarly communication using a semantically rich, interlinked knowledge graph. ORKG facilitates the representation, analysis and augmentation of scientific publications and related artefacts, especially (research) software. ORKG includes over 35,000 interconnected scholarly resources (1600 of which describe (research) software) and is permanently operated by TIB as a service for the scientific community. The BMBF SCINEXT (01IS22070) project uses artificial intelligence to revolutionise scholarly knowledge discovery and automate find.software: Foundations for Interdisciplinary Discovery of (Research) ... 9 WP 3.3 User and developer documentation. We are creating comprehensive user and developer documentation for find.software. The documentation will cover both technical and non-technical aspects and provide clear instructions for using the system, contributing metadata and interacting with the platform. The developer manuals will include technical specifications, APIs and integration instructions, while the user documentation will explain how to search, retrieve and analyse metadata. WP 3.4 Science communication. Science communication activities are being developed to raise awareness of the project and its goals. Short videos and blog posts will explain the project’s purpose, its benefits for the RS community and how users can become involved. The material will be disseminated via online platforms, such as social media, project websites and academic networks to ensure broad visibility. The videos are intended to appeal to a wide audience, from researchers to the general public. WP4 (month 1 - 36): Project management WP 4.1 Initialisation. This sub-package focuses on laying the foundations for the research project. Activities include defining the scope of the project, setting key objectives and assembling the core project team. Project tools and communication channels are also defined. The preliminary technical requirements for the Wikibase instance are also evaluated. WP 4.2 Coordination. This sub-package includes the ongoing management and synchronisation of project activities across all work packages. Regular project meetings are organised and documented. A comprehensive communication strategy is implemented. Coordination efforts also include risk management, monitoring the schedule (work packages, deliverables, reports) and ensuring that tasks and deliverables are completed on time. WP 4.3 Research data management. We ensure that metadata is curated, stored and shared according to best practice. In this sub-package, the storage, versioning and backup of data and the underlying workflows are realised. Ethical considerations and compliance with relevant data protection regulations are addressed. The project ensures that RS metadata is freely accessible to the global research community in line with the principles of open science. All project relevant data will be published at the end of the project (Deliverable D4.1) under a common open access licence. WP 4.4 Reporting. This sub-package focuses on the systematic documentation and reporting of project progress and outcomes. We track achievements, challenges and adjustments to the project plan. The preparation of deliverables and research findings will be managed within this sub-package. The main focus of the package, however, is to write the DFG Final Report under Infrastructure Funding (DFG form 12.02 - 10/24), (Deliverable D4.2). 16 Gey R et al Supplementary information on the project context General ethical aspects We do not anticipate that our research will result in any risks and/or harm to individuals or groups and/or the potential for other negative impacts. Some potential discovery pathways for research software involve information about people, such as authorship or co-authorship of papers or software, yet we will only process such information if it is already public and such ethical dimensions will be taken into account when we prioritise which discovery pathways to support and how (cf. WP 1.1 and WP 1.4). Considerations on aspects of ecological sustainability in the planning and implementation of the project We will avoid air travel as part of the project implementation* . When organising workshops, care will always be taken to enable hybrid participation and, thus, minimise the cost of travel to and from the event. When collecting, processing and presenting software metadata in find.software, we will try to take ecological aspects (carbon footprint, ecological benchmarks) of the assessed research software into account (Lannelongue et al. 2021) and offer specially tailored services (filter options, ecological assessments). Measures to meet funding requirements and handle project results We commit ourselves to making all project results (text, data and software publications, documentation and software code) openly available to the general public via permissive licences* . Data suitable for Wikidata will be shared there. The project results will conform to the FAIR and FAIR4RS principles. The applicants are supported in research data management and compliance with the FAIR and FAIR4RS criteria by the responsible structural units at their institutions* . Formal assurances The applicant institutions each guarantee the 10% of their own contributions required for Phase 1 of the Research Software Infrastructures programme. All results will be Open Access and comply with FAIR and FAIR4RS principles. Software developed will be open source. Developer documentation and user documentation will be created and made publicly available alongside the respective software and data model. In general, software will be modularised and constructed in a generally applicable manner to allow for easy reuse and transfer in/to other projects. For long-term archiving, all research artefacts will be published and made available free of charge. Source code and documentation will be made available under free open source licences. Scientific publications and tech reports will be made available in (green or gold) open access. In terms of sustainability, all partners commit to using the software developed in this project internally also outside of the scope of the project. 16 17 18 find.software: Foundations for Interdisciplinary Discovery of (Research) ... 17 Institutions or researchers in Germany with whom the partners have agreed to cooperate on this project NFDI4Earth. The Consortium for the establishment of a National Research Infrastructure (NFDI) for Earth System Sciences (NFDI4Earth) addresses the digital needs of Earth System Sciences. NFDI4Earth is willing to collaborate with find.software in areas such as the collection, processing and the provision of metadata on research software, the organisation of joint workshops, as well as in providing joint education, documentation and training materials. Wikimedia e.V. Wikimedia Deutschland – Gesellschaft zur Förderung Freien Wissens e.V. is a non-profit organisation which establishes and promotes the creation, collection and distribution of free knowledge in all parts of society. The find.software idea of creating an open and globally accessible platform for research software discovery – especially through Wikidata – is a good fit to that and their activities in terms of community engagement, software development and legal advice, as well as their broad reach into several of our target communities, align well with the project's goals. Use of AI tools declaration The authors declare they have used artificial intelligence tools (DeepL, ChatGPT) to assist in the writing of this proposal, which includes grammar checks, language enhancements summarisation and improved clarity. Funding program The find.software project is funded under grant number 567156310 through the Research Software Infrastructure funding programme within the Scientific Library Services and Information Systems programme of the German Research Foundation (DFG). Grant title find.software Conflicts of interest The authors have declared that no competing interests exist. Disclaimer: This article is (co-)authored by any of the Editors-in-Chief, Managing Editors or their deputies in this journal. 18 Gey R et al References • Altenhöner R, Blümel I, Boehm F, Bove J, Bicher K, Bracht C, Brand O, Dieckmann L, Effinger M, Hagener M, Hammes A, Heller L, Kailus A, Kohle H, Ludwig J, Münzmay A, Pittroff S, Razum M, Röwenstrunk D, Sack H, Simon H, Schmidt D, Schrade T, Walzel A, Wiermann B (2020) NFDI4Culture - Consortium for research data on material and immaterial cultural heritage. Research Ideas and Outcomes 6 https://doi.org/10.3897/rio. 6.e57036 • Anzt H, Bach F, Druskat S, Löffler F, Loewe A, Renard B, Seemann G, Struck A, Achhammer E, Aggarwal P, Appel F, Bader M, Brusch L, Busse C, Chourdakis G, Dabrowski PW, Ebert P, Flemisch B, Friedl S, Fritzsch B, Funk M, Gast V, Goth F, Grad J, Hegewald J, Hermann S, Hohmann F, Janosch S, Kutra D, Linxweiler J, Muth T, Peters-Kottig W, Rack F, Raters FC, Rave S, Reina G, Reißig M, Ropinski T, Schaarschmidt J, Seibold H, Thiele J, Uekermann B, Unger S, Weeber R (2021) An environment for sustainable research software in Germany and beyond: current state, open challenges, and call for action. F1000Research 9 https://doi.org/10.12688/ f1000research.23224.2 • Auer S, Barone DC, Bartz C, Cortes E, Jaradeh MY, Karras O, Koubarakis M, Mouromtsev D, Pliukhin D, Radyush D, Shilin I, Stocker M, Tsalapati E (2023) The SciQA Scientific Question Answering Benchmark for Scholarly Knowledge. Scientific Reports 13 (1). https://doi.org/10.1038/s41598-023-33607-z • Azzouz-Thuderoz M, Schubotz M, Teschke O (2022) Sustaining the swMATH project: Integration into zbMATH Open interface and Open Data perspectives. European Mathematical Society Magazine 126: 62‑64. https://doi.org/10.4171/mag/118 • Bär C (2024) zbMATH Open – Where do we stand and what’s next? European Mathematical Society Magazine 133: 66‑68. https://doi.org/10.4171/mag/214 • BOAI20 Steering Group (2022) The Budapest Open Access Initiative: 20th Anniversary Recommendations. https://www.budapestopenaccessinitiative.org/boai20/. Accessed on: 2025-11-11. • Bumberger J, Gey R, Kollai H, Rosenow D, Schnicke T (2024) Principles for the responsible handling of research data at the Helmholtz Centre for Environmental Research - UFZ. Helmholtz Centre for Environmental Research https://doi.org/10.57699/ ny5g-gd55 • Bumberger J, Abbrent M, Brinckmann N, Hemmen J, Kunkel R, Lorenz C, Lünenschloss P, Palm B, Schnicke T, Schulz C, van der Schaaf H, Schäfer D (2025) Digital ecosystem for FAIR time series data management in environmental system science. SoftwareX 29 https://doi.org/10.1016/j.softx.2025.102038 • Castro LJ, Bleier A, Baum R, Hagemeier B, Michels T, Pempe W, Reinhardt M, Voß J, Venkatesh S, Windeck J, Zacharias M (2025) Initial concept for metadata exchange between nfdi.software and other Base4NFDI services. Zenodo https://doi.org/10.5281/ zenodo.15676105 • Conrad TF, Ferrer E, Mietchen D, Pusch L, Stegmüller J, Schubotz M (2024) Making Mathematical Research Data FAIR: Pathways to Improved Data Sharing. Scientific Data 11 (1). https://doi.org/10.1038/s41597-024-03480-0 • Contaxis N, Clark J, Dellureficio A, Gonzales S, Mannheimer S, Oxley P, Ratajeski M, Surkis A, Yarnell A, Yee M, Holmes K (2022) Ten simple rules for improving research find.software: Foundations for Interdisciplinary Discovery of (Research) ... 19 data discovery. PLOS Computational Biology 18 (2). https://doi.org/10.1371/journal.pcbi. 1009768 • Deutsche Forschungsgemeinschaft. Ausschuss für Wissenschaftliche Bibliotheken und Informationssysteme (2025) Digitale Forschungspraxis und kooperative Informationsinfrastrukturen. Ein Diskussionspapier der Deutschen Forschungsgemeinschaft zu Förderung und Finanzierung wissenschaftlicher Informationsinfrastrukturen. Deutsche Forschungsgemeinschaft (DFG) https://doi.org/ 10.5281/zenodo.14621979 • Druskat S, Spaaks J, Chue Hong N, Haines R, Baker J, Bliven S, Willighagen E, PérezSuárez D, Konovalov A (2021) Citation File Format. Zenodo https://doi.org/10.5281/ zenodo.5171937 • Garijo D, Ratnakar V, Gil Y, Khider D (2021) The Software Description Ontology, Revision 1.9.0. Open Knowledge Network URL: https://w3id.org/okn/o/sd/1.9.0 • Gey R, Wittenborg T, Mietchen D, Struck A, Karras O (2025) Seek and You Shall Find - Or Not! Why Can't We Find the Research Software We Really Need? Zenodo https:// doi.org/10.5281/zenodo.14878608 • Goth F, Alves R, Braun M, Castro LJ, Chourdakis G, Christ S, Cohen J, Druskat S, Erxleben F, Grad J, Hagdorn M, Hodges T, Juckeland G, Kempf D, Lamprecht A, Linxweiler J, Löffler F, Martone M, Schwarzmeier M, Seibold H, Thiele JP, von Waldow H, Wittke S (2025) Foundational Competencies and Responsibilities of a Research Software Engineer: Current State and Suggestions for Future Directions. F1000Research 13 https://doi.org/10.12688/f1000research.157778.2 • Harpke A, Brünecke J, Bohring H, Grescho V, Haase K, Kühn E, Kuhnert T, Musche M, Petruschke S, Schnicke T, Sielaff D, Strätling J, Voigt M, Bumberger J (2024) BioMe - The butterfly monitoring Germany usecase. Zenodo https://doi.org/10.5281/zenodo. 13387849 • Hasselbring W, Druskat S, Bernoth J, Betker P, Felderer M, Ferenz S, Hermann B, Lamprecht A, Linxweiler J, Prat A, Rumpe B, Schöning-Stierand K, Yang S (2025) MultiDimensional Categorization of Research Software with Examples. Zenodo https://doi.org/ 10.5281/zenodo.14082553 • Hettrick S, Bast R, Crouch S, Wyatt C, Philippe O, Botzki A, Carver J, Cosden I, D'Andrea F, Dasgupta A, Godoy W, Gonzalez-Beltran A, Hamster U, Henwood S, Holmvall P, Janosch S, Lestang T, May N, Philips J, Poonawala-Lohani N, Richmond P, Sinha M, Thiery F, Werkhoven B, Zhang Q (2022) International RSE Survey 2022. Zenodo https://doi.org/10.5281/zenodo.7015772 • Hucka M, Graham MJ (2018) Software search is not a science, even among scientists: A survey of how scientists and engineers find software. Journal of Systems and Software 141: 171‑191. https://doi.org/10.1016/j.jss.2018.03.047 • Jaradeh MY, Oelen A, Farfar KE, Prinz M, D'Souza J, Kismihók G, Stocker M, Auer S (2019) Open Research Knowledge Graph. Proceedings of the 10th International Conference on Knowledge Capture243‑246. https://doi.org/10.1145/3360901.3364435 • Jones MB, Boettiger C, Mayes AC, Smith A, Slaughter P, Niemeyer K, Gil Y, Fenner M, Nowak K, Hahnel M, Coy L, Allen A, Crosas M, Sands A, Hong NC, Cruse P, Katz D, Goble C (2017) CodeMeta: an exchange schema for software metadata. KNB Data Repository https://doi.org/10.5063/schema/codemeta-2.0 20 Gey R et al • Kabongo S, D’Souza J, Auer S (2023) ORKG-Leaderboards: a systematic workflow for mining leaderboards as a knowledge graph. International Journal on Digital Libraries 25 (1): 41‑54. https://doi.org/10.1007/s00799-023-00366-1 • Kang HB, Soliman N, Latzke M, Chang JC, Bragg J (2023) ComLittee: Literature Discovery with Personal Elected Author Committees. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems1‑20. https://doi.org/ 10.1145/3544548.3581371 • Kelley A, Garijo D (2021) A framework for creating knowledge graphs of scientific software metadata. Quantitative Science Studies 2 (4): 1423‑1446. https://doi.org/10.1162/ qss_a_00167 • Kuckertz P, Göpfert J, Karras O, Neuroth D, Schönau J, Pueblas R, Ferenz S, Engel F, Pflugradt N, Weinand J, Nieße A, Auer S, Stolten D (2024) DataDesc: A framework for creating and sharing technical metadata for research software interfaces. Patterns 5 (11). https://doi.org/10.1016/j.patter.2024.101064 • Lannelongue L, Grealey J, Inouye M (2021) Green Algorithms: Quantifying the Carbon Footprint of Computation. Advanced Science 8 (12). https://doi.org/10.1002/advs. 202100707 • Levitskaya E, Korkmaz G, Mietchen D, Rasberry L (2022) Analysis of Linked GitHub and Wikidata. Zenodo. Release date: 2022-12-15. URL: https://doi.org/10.5281/zenodo. 7443338 • Maassen J (2023) The Research Software Directory - NFDI Software Marketplace Workshop. Zenodo https://doi.org/10.5281/zenodo.7781935 • Malone J, Brown A, Lister AL, Ison J, Hull D, Parkinson H, Stevens R (2014) The Software Ontology (SWO): a resource for reproducibility in biomedical data analysis, curation and digital preservation. Journal of Biomedical Semantics 5 (1). https://doi.org/ 10.1186/2041-1480-5-25 • Messner F, Rosenow D, Zacharias S, Bumberger J (2025) Managing the transformation from small data to big data. Wissenschaftsmanagement (January 2025). URL: https:// www.wissenschaftsmanagement.de/news/managing-transformation-small-data-big-data • Murphy-Hill E, Lee DY, Murphy G, McGrenere J (2015) How Do Users Discover New Tools in Software Development and Beyond? Computer Supported Cooperative Work (CSCW) 24 (5): 389‑422. https://doi.org/10.1007/s10606-015-9230-9 • Nesbitt A, Veytsman B, Mietchen D, Brown EM, Howison J, Pimentel JF, Hébert-Dufresne L, Druskat S (2024) Biomedical Open Source Software: Crucial Packages and Hidden Heroes. arXiv https://doi.org/10.48550/arxiv.2404.06672 • Nielsen FÅ, Mietchen D, Willighagen E (2017) Scholia, Scientometrics and Wikidata. Lecture Notes in Computer Science237‑259. https://doi.org/10.1007/978-3-31970407-4_36 • Nishikawa-Pacher A (2021) A typology of research discovery tools. Journal of Information Science 49 (4): 1086‑1095. https://doi.org/10.1177/01655515211040654 • Oelen A, Jaradeh MY, Auer S (2024) ORKG ASK: a Neuro-symbolic Scholarly Search and Exploration System. arXiv https://doi.org/10.48550/arxiv.2412.04977 • Rahkooy H, Schubotz M, Teschke O, Fuhrmann M, Roy N, Mietchen D (2025) Wikipedia and zbMATH Open: Connecting several layers of mathematical information. European Mathematical Society Magazine 136: 59‑63. https://doi.org/10.4171/mag/252 • Ramy-Badr A, Schubotz M, Mietchen D, Samuel S (2024) Software Heritage Archival IDs of FAIR Jupyter GitHub repos. Zenodo https://doi.org/10.5281/zenodo.12806151 find.software: Foundations for Interdisciplinary Discovery of (Research) ... 21 • Rasberry L, Mietchen D (2022) Scholia for Software. Research Ideas and Outcomes 8 https://doi.org/10.3897/rio.8.e94771 • Rasberry L, Tibbs S, Hoos W, Westermann A, Keefer J, Baskauf SJ, Anderson C, Walker P, Kwok C, Mietchen D (2022) WikiProject Clinical Trials for Wikidata. medRxiv https:// doi.org/10.1101/2022.04.01.22273328 • RfII - German Council for Scientific Information Infrastructures (2024) Federated Data Infrastructures for Scientific Use. NFDI, EOSC, Gaia-X, and the European Data Spaces: Comparison and Recommendations for a Committed Engagement to Shape the European Research Data Ecosystem. RfII, Göttingen, 103 pp. [In English]. • Samuel S, Mietchen D (2023) Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications. Zenodo https://doi.org/10.5281/zenodo. 8226725 • Samuel S, Mietchen D (2024a) FAIR Jupyter: A Knowledge Graph Approach to Semantic Sharing and Granular Exploration of a Computational Notebook Reproducibility Dataset. Transactions on Graph Data and Knowledge 2 (2): 4:1‑4:24. https://doi.org/10.4230/ TGDK.2.2.4 • Samuel S, Mietchen D (2024b) Computational reproducibility of Jupyter notebooks from biomedical publications. GigaScience 13 https://doi.org/10.1093/gigascience/giad113 • Schmidt L, Schäfer D, Geller J, Lünenschloss P, Palm B, Rinke K, Rebmann C, Rode M, Bumberger J (2023) System for automated Quality Control (SaQC) to enable traceable and reproducible data streams in environmental science. Environmental Modelling & Software 169 https://doi.org/10.1016/j.envsoft.2023.105809 • Schubotz M, Ferrer E, Stegmüller J, Mietchen D, Teschke O, Pusch L, Conrad TF (2023) Bravo MaRDI: A Wikibase Powered Knowledge Graph on Mathematics. arXiv https:// doi.org/10.48550/arxiv.2309.11484 • Schulz C, Lange R, Gonzalez Pestana N, Panda M, Schnicke T, Bumberger J (2025) spatial.IO - An integrated cloud-ready geospatial data management system. Zenodo https://doi.org/10.5281/zenodo.10391523 • Steinbeck C, Koepler O, Bach F, Herres-Pawlis S, Jung N, Liermann J, Neumann S, Razum M, Baldauf C, Biedermann F, Bocklitz T, Boehm F, Broda F, Czodrowski P, Engel T, Hicks M, Kast S, Kettner C, Koch W, Lanza G, Link A, Mata R, Nagel W, Porzel A, Schlörer N, Schulze T, Weinig H, Wenzel W, Wessjohann L, Wulle S (2020) NFDI4Chem - Towards a National Research Data Infrastructure for Chemistry in Germany. Research Ideas and Outcomes 6 https://doi.org/10.3897/rio.6.e55852 • Thagard P (2022) Pathways to Biomedical Discovery. Philosophy of Science 70 (2): 235‑254. https://doi.org/10.1086/375465 • The MaRDI consortium (2022) MaRDI: Mathematical Research Data Initiative Proposal. Zenodo. https://doi.org/10.5281/zenodo.6552436 • Thornton K, Seals-Nutt K, Cochrane E, Wilson C (2018) Wikidata For Digital Preservation. Zenodo https://doi.org/10.5281/zenodo.1214319 • Thornton K, Solbrig H, Stupp G, Labra Gayo JE, Mietchen D, Prud’hommeaux E, Waagmeester A (2019) Using Shape Expressions (ShEx) to Share RDF Data Models and to Guide Curation with Rigorous Validation. Lecture Notes in Computer Science606‑620. https://doi.org/10.1007/978-3-030-21348-0_39 • Turki H, Hadj Taieb MA, Shafee T, Lubiana T, Jemielniak D, Aouicha MB, Labra Gayo JE, Youngstrom E, Banat M, Das D, Mietchen D, on behalf of WikiProject COVID- (2022) 22 Gey R et al *1 *2 *3 *4 *5 *6 *7 *8 *9 *10 *11 *12 *13 *14 *15 *16 *17 *18 Representing COVID-19 information in collaborative knowledge graphs: The case of Wikidata. Semantic Web 13 (2): 233‑264. https://doi.org/10.3233/sw-210444 • Turki H, Hadj Taieb MA, Ben Aouicha M, Rasberry L, Mietchen D (2024) Preregistration: Comparing the use of Wikidata and Wikipedia by open-source software programmers on GitHub repositories. In: Wikidata Workshop 2023 Organizers (Ed.) Proceedings of the Wikidata Workshop 2023. URL: https://openreview.net/forum?id=brCdQnsIQG • Waagmeester A, Stupp G, Burgstaller-Muehlbacher S, Good BM, Griffith M, Griffith OL, Hanspers K, Hermjakob H, Hudson TS, Hybiske K, Keating SM, Manske M, Mayers M, Mietchen D, Mitraka E, Pico AR, Putman T, Riutta A, Queralt-Rosinach N, Schriml LM, Shafee T, Slenter D, Stephan R, Thornton K, Tsueng G, Tu R, Ul-Hasan S, Willighagen E, Wu C, Su AI (2020) Wikidata as a knowledge graph for the life sciences. eLife 9 https:// doi.org/10.7554/elife.52614 • Wittenborg T, Gey R, Karras O, Mietchen D, Struck A (2025) Research Software Discovery: How do we Want to Search Research Software and Where do we Want to Find it? Zenodo https://doi.org/10.5281/zenodo.14878664 Endnotes https://commons.datacite.org/doi.org?query=+&resource-type=software https://bio.tools https://physics.tools https://marketplace.sshopencloud.eu/ https://open-science-cloud.ec.europa.eu/ https://nfdi.de/ https://base4nfdi.de/projects/nfdi-software https://base4nfdi.de/projects/kgi4nfdi4 https://w.wiki/DEK8 https://w.wiki/DEK2 https://w.wiki/DFSo https://w.wiki/DG5r https://w.wiki/DFUR https://www.wikidata.org/wiki/Wikidata:WikiProject_Informatics/Software https://www.wikidata.org/wiki/Wikidata:Property_proposal https://noflyclimatesci.org/biographies/daniel-mietchen https://github.com/Daniel-Mietchen/pledges FIZ: https://www.fiz-karlsruhe.de/en/forschung/forschung; TIB: https://www.tib.eu/de/ die-tib/policies/open-access-policy; UFZ: https://www.ufz.de/rdm and Bumberger et al. (2024) find.software: Foundations for Interdisciplinary Discovery of (Research) ... 23