scieee AI-readable full text Open interactive document viewer

Interview Evaluation

Martens, Claudia

Abstract

During the Joint Plenary of NFDI4Earth and NFDI4Biodiversity on September 22-25 2025 at Marum in Bremen we conducted 25 short interviews to examine the need of semantic services across divergent institutions. Therefore we divided our interviews into two sections: while the questionnaire for Scientists [A] focussed on questions regarding publishing and reusing data, the one for Data Managers [B] referred to workflows for metadata management. Even though this study is not representative, some conclusions can be drawn about needs and gaps around metadata management.

Full text

1 cite as: Martens, C. (2025). Interview Evaluation. NFDI4Biodiversity meets NFDI4Earth - Joint Plenary 2025, Bremen, Marum. Zenodo. https://doi.org/10.5281/zenodo.17630162 Authors: Claudia Martens, Anette Ganske, Andrea Lammert, Alexander Wolodkin, Ivonne Anders Interview Evaluation During the Joint Plenary of NFDI4Earth and NFDI4Biodiversity on September 22-25 2025 at Marum in Bremen we conducted 25 short interviews to examine the need of semantic services across divergent institutions. Therefore, we divided our interviews into two sections: while the questionnaire for Scientists [A] focussed on questions regarding publishing and reusing data, the one for Data Managers [B] referred to workflows for metadata management. Even though this study is not representative, some conclusions can be drawn about needs and gaps around metadata management. 2 Study Design & Outcome Overall, we asked 25 people of which 9 were Scientists and 14 classified themselves as Data Managers (in a broad sense, including repository staff and data stewards). One PhD candidate was attributed to the Scientists and one PI to the Data Managers which gave us a reasonable distribution of the different questionnaire sections. Part A We asked the Scientists [A1] whether they publish data and if so, where and if not, why not respectively. Out of 10 people only 2 did not publish data, the reason being ‘no time’ and ‘no need currently’. Repositories used for data publishing are widely spread (multiple responses were possible) as are the metadata standards used (DataCite, ISO 19115/19139, INSPIRE, OGC sensor, GeoNames, Darwin Core, ABCD, CF, ops4MIPs, archiveSS, and project specific metadata). The question [A2] “Could automated metadata enrichment help you to publish your data as FAIR data?” was answered by all participants with “yes”, including the ones that currently do not publish data. Many participants emphasized their agreement with statements such as “yes, that would be really cool!” or “of course, give me tools for it!”. However, other remarks referred to the reliability of services (“yes that would help, but it 3 must be maintained”, “only if it is reliable”), those using Zenodo as publisher restrained their statements (“only if Zenodo would implement it, I have no time”). Our last question [A3] “Do you (re)use data from others?” was answered with “yes” by 6 participants whereas again all respondents agreed that richer metadata would make data better understandable. The answers for the follow-up question “Which metadata would help you mostly” highlights the diverse landscape and a lack of interoperability options: “specific information about experiments”, “harmonisation of NCDF format”, “CF conformity”, “Provenance: who completed the data, original sources, applied data transformation”, “spatial/temporal coverage, parameter, resolution”, and “more metadata for transparency, provenance, governance information” were mentioned. Part B For Data Managers, the distribution across institutions/repositories is widespread as well. For our first question [B1] “Would tools for metadata validation against controlled vocabularies simplify the data ingestion and the curation workflows?” 12 participants responded with “yes”, the other 3 persons claimed to have already automated workflows. However, here too, the answer “yes” was sometimes emphatically emphasized (“of course”, “certainly”). The second question [B2] “Are you looking for ways to automate your data quality assurance?” was answered by two-thirds of respondents with “yes”. One-third responded with “no” but of those 3 were claiming that workflows for automated data quality assurance already exist. Interestingly, within the group of “yes” some participants stated that even though they already have data quality mechanisms those could and should be improved (which is why they answered with “yes”). The range for the follow-up question in which areas data quality assurance could be improved spans from “tools for term definition and standards”, “crossquality checks”, and “statistics and error dealing” to “tools for extracting metadata from papers”. Our final question [B3] “Are you concerned about the lack of standardisation in your repository / data lake asset?” shows no clear message: even though most participants do not 4 recognise a lack of standardisation we got many restricting answers such as “no but curation is needed”, “not for published data but definitely for internal data”, “not for my repo, but certainly in general”, “no but curation and maintenance are an issue", or “no but there is much room for improvement”. Interestingly, one comment referred to concerns about the standardisation of training data for AI because "there are no clear criteria what 'good' training data is and a very fast development in this area of AI". We also received comments for already existing services, e.g., that "in EarthPortal there is a dropdown menu for variables but if you look into it, the definitions of these variables are different". Apart from those remarks about metadata standardisation all participants agreed that richer metadata would make data better understandable - which we see as an inducement to continuously supporting the automatic enrichment of metadata with semantic context. 5