scieee AI-readable full text Open interactive document viewer

D2.2 Updated Report on the H2IOSC Landscapes

Monachini, Monica; Quochi, Valeria; Luzietti, Roberta Bianca; Spadi, Alessia; Mancuso, Giacomo; D'Eredità, Antonio; Caravale, Alessandra; Giampietro, Nicola; Caradonna, Marta; Melaccio, Daniele

Abstract

This deliverable presents the updated results of the Landscaping and Building Communities activities carried out within WP2 of the H2IOSC project. Through a mixed-methods approach—combining questionnaires, interviews, focus groups, and internal and external scouting—the report provides a systematic overview of the research landscape in the Humanities, Linguistics, and Heritage Science domains in Italy. It maps existing digital resources, tools, services, community practices, FAIRness levels, training needs, and infrastructural gaps across the communities served by the four participating Research Infrastructures (CLARIN, DARIAH, E-RIHS, OPERAS). The findings feed directly into key components of the project, including the DHeLO landscaping platform and the H2IOSC Observatory, supporting resource integration, FAIRification, service development, capacity-building, and pilot innovation activities. Overall, the deliverable offers the strategic evidence base necessary for guiding the technical and community-oriented work of subsequent WPs.

Full text

1 1 CALL Ministry of University and Research (MUR), DirectorateGeneral for Internationalization and Communication, Notice D.D. 3264 dated 28/12/2021. Mission 4 "Education and Research" - Component 2 "From research to enterprise" - Investment Line 3.1 "Fund for the realization of an integrated system of research and innovation infrastructure", Action 3.1.1 "Creation of new IRs or strengthening of existing ones that contribute to Horizon Europe's Scientific Excellence objectives and networking" PROJECT TITLE Humanities and cultural Heritage Open Science Cloud PROJECT ACRONYM H2IOSC PROJECT ID IR0000029 PROJECT WEBSITE https://www.h2iosc.cnr.it CUP B63C22000730005 WP 2 - LANDSCAPING AND BUILDING COMMUNITIES OU OU 8 ILC-PI DELIVERABLE NUMBER 2.2 DELIVERABLE TITLE Updated Report on the H2IOSC Landscapes EXPECTED DELIVERY DATE 30/04/2025 ACTUAL DELIVERY DATE 31/10/2025 AUTHORS Monica Monachini, Valeria Quochi, Roberta Bianca Luzietti, Alessia Spadi, Giacomo Mancuso, Antonio D’Eredità, Alessandra Caravale, Nicola Giampietro, Marta Caradonna, Daniele Melaccio This project is funded by the European Union - NextGenerationEU. The views and opinions expressed are only those of the authors and do not necessarily reflect those of the European Union or the European Commission. Neither the European Union nor the European Commission can be held responsible for them. Due to project-level deliverable submission deadlines, the information in this document is updated as of M36 (October 31, 2025). However, the authors intend to also publish the updates to the activities as soon as they are available. This will ensure that the most up-to- 2 2 date and comprehensive information is made available, supporting continuous improvement and strategic alignment across the cluster. 3 3 Table of contents Sommario Table of contents ............................................................................................................................... 3 List of acronyms ................................................................................................................................ 5 List of Figures ..................................................................................................................................... 7 Executive Summary ........................................................................................................................ 10 1. Introduction: the H2IOSC project and WP2 ........................................................................ 12 1.1 Landscaping and Building Communities .................................................................................. 12 1.2 Synergies with other project’s activities .................................................................................. 14 2. Overall Methodology ................................................................................................................. 16 3. Questionnaire-based Survey ................................................................................................... 21 3.1 Survey Methodology ........................................................................................................................ 21 3.2 Results and analysis ......................................................................................................................... 22 4. Interviews and other community-involving activities.................................................... 32 4.1. Structure of interviews and results .......................................................................................... 32 4.2 Other community-targeted activities: format and results ................................................ 35 5. Focus Groups ................................................................................................................................ 38 5.1 Focus Group Methodology ............................................................................................................. 38 5.2 Features of planned focus groups. .............................................................................................. 39 5.3 Results and analysis ......................................................................................................................... 42 6. Landscape analysis of participating RIs ............................................................................. 47 6.1 RIs Landscape strategy ................................................................................................................... 47 6.2 Knowledge Organization Schema ............................................................................................... 52 6.3 Unified H2IOSC Landscape criteria and methods ................................................................. 55 7. Digital Heritage Landscaping PlatfOrm: DHeLO .............................................................. 60 7.1 DHeLO 1.0. A technical overview of the data model ............................................................ 61 7.2 DHeLO Data collection process .................................................................................................... 63 7.3 DHeLO 2.0. - A technical overview .............................................................................................. 64 7.4 DHeLO 2.0 - Adjusting the data model ...................................................................................... 66 7.5 Results overview ............................................................................................................................... 68 7.6 BiDiAr - Bibliography of Digital Archaeology ......................................................................... 70 8. H2IOSC Observatory .................................................................................................................. 72 8.1 Concept and general aims .............................................................................................................. 72 8.2 Structure, scope, and positioning ............................................................................................... 72 4 4 8.3 Methodological approach and system design ........................................................................ 73 8.4 Information architecture and content organisation ........................................................... 74 8.5 Functionalities and interaction with the H2IOSC Marketplace ....................................... 76 8.6 Sustainability, adaptability, and long-term vision ............................................................... 77 8.7 Current development status ......................................................................................................... 78 9. Conclusions ................................................................................................................................... 79 References ......................................................................................................................................... 81 Annex 1: Questionnaire structure and layout........................................................................ 83 Annex 2: Focus groups preparation materials ................................................................... 102 5 5 List of acronyms H2IOSC Humanities and Heritage Open Science Cloud RI Research Infrastructure CLARIN-IT CLARIN Italian node (Common Language Resources and Technology Infrastructure) DARIAH-IT DARIAH Italian node (Digital Research Infrastructure for the Arts and Humanities) E-RIHS E-RIHS Italian node (European Heritage Science Research Infrastructure) OPERAS OPERAS Italian node (Open scholarly communication in the European research area for social sciences and humanities) SSH Social Sciences and Humanities SSHOC Social Sciences & Humanities Open Cloud FAIR Findable, Accessible, Interoperable and Reusable ERIC European Research Infrastructures Consortium EOSC European Open Science Cloud ESFRI European Strategy Forum on Research Infrastructures WP Work Package DHeLO Digital Heritage Landscaping PlatfOrm ERA European Research Area PiD Persistent Identifier GIS Geographic Information System CC BY NC Creative Commons Attribution Non Commercial 6 6 OJS Open Journal System DPH Digital Philology Hub FG Focus Group VLO Virtual Language Observatory BiDiAr Bibliography of Digital Archaeology TEI Text Encoding Initiative CIDOC CRM International Committee for Documentation Conceptual Reference Model DCMI Dublin Core Metadata Initiative LOD Linked Open Data FOAF Friend Of A Friend DA Digital Archaeology CH Cultural Heritage HS Heritage Science 7 7 List of Figures Some figures and illustrations in this deliverable retain their original Italian captions, as they are directly taken from internal documentation, pilot materials, or project outputs produced in their original linguistic context. The choice to preserve the original language aims to maintain consistency with the source data and the integrity of the visual materials. Figure 1: Map of the H2IOSC Cloud Network organization ...................................................... 12 Figure 2. Synergies with other WPs ......................................................................................... 14 Figure 3: Schema illustrating the Mixed Methods approach employed in the H2IOSC surveying activity. .................................................................................................................... 19 Figure 4: summary of provenance of respondents ................................................................. 23 Figure 5: institutions of provenance of respondents .............................................................. 23 Figure 6: career level of respondents ...................................................................................... 24 Figure 7: disciplinary fields of respondents ............................................................................. 24 Figure 8: summary of resources, tools and services mostly used by respondents ................. 25 Figure 9: summary of resources, tools and services created by respondents ........................ 26 Figure 10: knowledge on the existence and characteristics of Ris by respondents ................ 27 Figure 11: professional profiles of respondents unaware of the existence and characteristics of RIs......................................................................................................................................... 27 Figure 12: list of Ris known by respondents ............................................................................ 28 Figure 13: respondents use and consultation of Ris services .................................................. 28 Figure 14: respondents' future expectations on RIs ................................................................ 28 Figure 15: respondent’s interest and availability in depositing their data and research results in Ris repositories ..................................................................................................................... 29 Figure 16: respondents' awareness of the H2IOSC project ...................................................... 29 Figure 17: respondents’ involvements in the H2IOSC project ................................................. 29 Figure 18: respondents’ publication practices ........................................................................ 30 Figure 19: respondents’ types of publications in Open Access ............................................... 31 Figure 20: most common licenses used by respondents ......................................................... 31 Figure 21Expected conceptual positioning of the two focus groups. ................................... 41 Figure 22: Composition by biological sex of the four focus groups......................................... 42 Figure 23: Composition by study area of the four focus groups ............................................. 43 Figure 24: Positioning of the four focus groups obtained after the discussions. .................... 45 Figure 25: Screenshot of the DHeLO 1.0 interface. ................................................................. 61 Figure 26: Overview of the DHeLO table schema structure .................................................... 62 Figure 27: Project page on DHeLO 2.0 ..................................................................................... 65 Figure 28: Entity relationship schema for DHeLO 2.0. ............................................................. 68 Figure 29: Quantities of data collected in DHeLO, broken down by type of entity. ............... 69 Figure 30: Geographic distribution of the resources collected within DHeLO, visualized through the platform’s mapping interface. ............................................................................. 70 Figure 31: The H2IOSC Observatory home page ..................................................................... 73 Figure 32: The H2IOSC Observatory section about focus groups ............................................ 75 Figure 33: An example of interactive data visualisation from the questionnaires section ..... 76 Figure 34: Markeplace-Observatyory interaction: statistics about published resources ....... 77 Figure 35: introduction and presentation of structure and purpose of the questionnaire together with privacy and policy statement............................................................................ 83 8 8 Figure 36: information on how the respondents received the notification to answer the questionnaire ........................................................................................................................... 84 Figure 37: personal information .............................................................................................. 85 Figure 38: information on the disciplinary sector of respondents (ERC) ................................ 86 Figure 39: information on experience in using data resources, tools and services ................ 87 Figure 40: information on experience in creating data resources, tools and services ........... 88 Figure 41: information on the (level of) knowledge of Research Infrastructures (in general)90 Figure 42: information on the (level of) knowledge of the H2IOSC project ............................ 90 Figure 43: information on training preferences ...................................................................... 91 Figure 44: information on publication practices ..................................................................... 93 Figure 45: contact information ................................................................................................ 94 Figure 46: information on how to proceed with the second questionnaire ........................... 95 Figure 47: final comments section ........................................................................................... 95 Figure 48: introductory information on the structure, purpose and time required to answer the questionnaire ..................................................................................................................... 96 Figure 49: e-mail request to match responses in previous questionnaire .............................. 96 Figure 50: information regarding data resources created ...................................................... 97 Figure 51: information regarding data resources used ........................................................... 98 Figure 52: information regarding tools and software created ................................................ 99 Figure 53: information regarding tools and software used ................................................... 100 Figure 54: final comments section ......................................................................................... 101 9 9 List of Tables Table 1: Targeted Research Infrastructures ............................................................................ 18 Table 2: Distribution of interviews participants ...................................................................... 33 Table 3: Number and characteristics of the focus group JUNIOR 1 and 2 .............................. 40 Table 4: Number and characteristics of the focus group SENIOR 1 and 2 .............................. 40 Table 5: CLARIN-IT Landscape criteria ..................................................................................... 48 Table 6: DARIAH-IT Landscape criteria .................................................................................... 49 Table 7: E-RIHS.it Landscape criteria ....................................................................................... 50 Table 8: Fields and Metadata of tables of Datasets and Tools ................................................ 54 Table 9: DHeLO metadata overview ........................................................................................ 63 Table 10: Metadata fields in the DHeLO platform ................................................................... 69 Table 1: Targeted Research Infrastructures ............................................................................ 18 Table 2: Distribution of interviews participants ...................................................................... 33 Table 3: Number and characteristics of the focus group JUNIOR 1 and 2 .............................. 40 Table 4: Number and characteristics of the focus group SENIOR 1 and 2 .............................. 40 Table 5: CLARIN-IT Landscape criteria ..................................................................................... 48 Table 6: DARIAH-IT Landscape criteria .................................................................................... 49 Table 7: E-RIHS.it Landscape criteria ....................................................................................... 50 Table 8: Fields and Metadata of tables of Datasets and Tools ................................................ 54 Table 9: DHeLO metadata overview ........................................................................................ 63 Table 10: Metadata fields in the DHeLO platform ................................................................... 69 16 16 2. Overall Methodology The approach is conceived to be cyclically repeatable over time, serving both as a sort of health check of the four infrastructures and as a strategic compass to monitor developments, detect emerging needs, and inform future directions. A key feature of this work was the complexity derived from the diverse nature of the four Research Infrastructures (RIs) involved in the project (CLARIN, DARIAH, E-RIHS, and OPERAS), each operating within distinct disciplinary fields and engages with different types of data. In fact, despite operating under the broader umbrella of Humanities and Cultural Heritage, each infrastructure prioritizes a different set of practices, services, and research communities. CLARIN places emphasis on the development and dissemination of high-quality linguistic resources that are technically robust, well-documented, and interoperable, supporting researchers in fields such as linguistics, literature, history, sociology, and other languagerelated disciplines. DARIAH, is oriented toward the wider digital humanities landscape, with particular attention to areas including philology, lexicography, archival science, Latin and Romance languages, and art history. E-RIHS operates primarily in the cultural heritage domain, offering a range of repositories, services, and tools designed for the management, sharing, and preservation of heritage data, and promoting standards and best practices across the field. Finally, OPERAS focuses on advancing Open Science within the Humanities, emphasizing FAIR principles and supporting scholarly communication through platforms, tools, and services tailored to the needs of humanities researchers. In the table below (Table 1) a brief description of each participating RI is provided, according to Part 3 of the ESFRI Roadmap 2021 1 : RI Name ESFRI Roadmap Description CLARIN The Common Language Resources and Technology Infrastructure (CLARIN) is a distributed Research Infrastructure that provides easy and sustainable access for scholars in the Humanities and Social Sciences to FAIR digital language data − in written, spoken or multimodal form − and advanced tools to discover, explore, exploit, annotate, analyse or combine them, independent of their location. CLARIN is building a networked federation of language data repositories, service centres and centres of expertise, with single sign-on access for all members of the academic community in all participating countries. Tools and data from different centres are interoperable, so that data collections can be combined and tools from different sources can be chained to perform complex operations to support researchers in their work. Entered in the ESFRI Roadmap 2016, CLARIN became a European Research Infrastructure Consortium (ERIC) in 2012. Since then several countries have joined either as full Member, or as Observer. The ultimate goal is to include all European countries as well as any interested third countries in or outside Europe. The majority of operations, services and centres of the CLARIN infrastructure is provided and funded by CLARIN Members (and Observers). They set up a national consortium, typically consisting of universities, research 1 https://roadmap2021.esfri.eu/projects-and-landmarks/ 17 17 institutions, libraries and public archives, of which at least one has the status of CLARIN Centre which is expected to create and provide access to digital language data collections, and digital tools and expertise for researchers to work with them. DARIAH The Digital Research Infrastructure for the Arts and Humanities (DARIAH) is a distributed Research Infrastructure to enhance and support digitally enabled research and teaching for Arts and Humanities. DARIAH is a network of people, expertise, information, knowledge, content, methods, tools and technologies from its member countries. It develops, maintains and operates an Infrastructure that sustains researchers in building, analysing and interpreting digital resources. By working with communities of practice, DARIAH brings together state-of-the-art digital arts and humanities activities and scales their results to a European level. It preserves, provides access to and disseminates research that stems from these collaborations and ensures that best practices, methodological and technical standards are followed. Entered in the ESFRI Roadmap 2006, DARIAH was established as a European Research Infrastructure Consortium (ERIC) in 2014. DARIAH was awarded Landmark Status in 2016 as a Research Infrastructure that reached its Implementation Phase and was considered a pan-European hub of scientific excellence. Currently, DARIAH has 20 Members, one Observer and several Cooperating Partners in six non-member countries. Structurally, DARIAH operates through the Europe-wide networks of the Virtual Competency Centres (VCCs) and their constituent Working Groups. Each of the four VCCs is cross-disciplinary, multiinstitutional, international and centred on a specific area of expertise. Within this structure, DARIAH has over 20 dynamic Working Groups to integrate national services under specific operational categories. E-RIHS The European Research Infrastructure for Heritage Science (E-RIHS) is a distributed Research Infrastructure to support research on heritage interpretation, preservation, documentation and management. E-RIHS delivers integrated access to expertise, data and technologies through a standardized approach, and integrate world-leading European facilities into an organization with a clear identity and a strong cohesive role within the global heritage science community. Through interdisciplinary access to the four platforms – E-RIHS ARCHLAB, E-RIHS DIGILAB, E-RIHS FIXLAB, E-RIHS MOLAB – E-RIHS supports a wide variety of research, from smaller object-focused case studies to largescale and longer-term collaborative projects, and stimulates innovation in instrumentation, portable technologies and data science. The long‐-term tradition of this field of research, the ability to combine science with innovation, and the support provided by EU-funded projects and integrating activities such as EUARTECH, CHARISMA, IPERION CH and IPERION HS in conservation science, and ARIADNE in archaeology, represent the background of E-RIHS. Entered in the ESFRI Roadmap 2016, E-RIHS was established as a European Research Infrastructure Consortium (ERIC) in 2025. It offers access to a wide range of fixed and mobile instruments in national facilities of recognized excellence, physically accessible collections/archives 18 18 and virtually accessible heritage repositories for standardized data storage, analysis and interpretation. OPERAS The Open Scholarly Communication in the European Research Area for Social Sciences and Humanities (OPERAS) is a distributed RI to enable Open Science and upgrade scholarly communication practices in the Social Sciences and Humanities in line with the EOSC. OPERAS pools resources and offers services to enable all SSH stakeholders to streamline their activities and maximize the societal impact, in an interdisciplinary, mission-driven approach. OPERAS fosters the co-creation and adoption of scholarly communication services addressing research needs in terms of discovery, content creation, quality assurance, dissemination, outreach, and evaluation of outputs. It catalyses knowledge and know-how sharing, practices adoption, and increases return on socio-economic investments. The OPERAS Concept and Design Phase dates back to 20122018 when the coordinated actions to establish the network and shaping the project began. This effort lead to the recognition of OPERAS as a project addressing High strategic potential area for research in SCI in the ESFRI Roadmap 2018. OPERAS started its Preparation Phase in 2019 by developing the business plan and governance model and promoting services creation and alignment for EOSC catalogue and transnational access. In 2019, OPERAS established the status of International non-profit Association under Belgian law (AISBL). Entered the ESFRI Roadmap 2021, OPERAS is striving to efficiently guarantee an efficient move to the Operation Phase with a coherent approach to technical, administrative, and financial issues and the establishment of OPERAS ERIC. Table 1: Targeted Research Infrastructures To address the complexity of the research landscape and the diversity of the communities involved, a Mixed Methods approach is adopted (Figure 2). This combines qualitative and quantitative strategies to: • Provide a comprehensive overview of the current landscape of resources, tools and services, along with the needs expressed by the research communities • Support related activities within the H2IOSC federation, particularly those focused on infrastructure construction and enhancement, service integration, user engagement, and training • Allow for data-driven analysis and forecasting of user needs • Elaborate a long-term strategy for the implementation and development of the H2IOSC Observatory. 19 19 Figure 3: Schema illustrating the Mixed Methods approach employed in the H2IOSC surveying activity. Since the design of an apparatus capable of accounting for the existing projects, resources, tools, communities, best practices, and standards, in relation to each RI community involved in the project, requires the elaboration of a composite strategy, four primary instruments have been put together to collect structured information: 1. an online questionnaire-based survey designed to engage the stakeholder communities at large, and used to collect data on the use, production, and awareness of digital resources, tools, and Open Science practices, as well as training needs and alignment with FAIR principles; 2. focus group meetings, intended to complement quantitative data, are conducted with selected representatives of the target communities, including students, researchers, senior scholars and experts pertaining to the different disciplinary areas represented by the RIs. These sessions gather new insights on expectations, perceived gaps, and specific needs of the target communities in the use of digital infrastructures. Focus groups also foster interest in the project’s activities. Participants are selected to ensure disciplinary and career-stage diversity across the four RIs; 3. a mapping matrix used to catalogue systematize data retrieved from the different sources on projects, datasets, tools, protocols and standards, across different disciplines; 4. a database for storing and facilitating navigation of all the collected data. It facilitates longitudinal analysis and supports the development of dashboards and observatory functions. Strategic starting points for this activity are existing repositories and 20 20 catalogues pertaining to two RIs involved in this project, which served as valuable initial sources of information (e.g. ILC4CLARIN). Quantitative data (obtained primarily from the questionnaires and integrated with data derived from interviews and focus groups) is analysed using descriptive statistics to identify usage patterns, levels of awareness, and potential gaps. Qualitative data from focus groups is, on the other hand, thematically coded to extract recurrent concerns, needs, and suggestions. As a result, the triangulation of methods ensures robustness and consistency between findings and strategic recommendations. The obtained insights are expected to address several key areas: • Prioritization: identifying the most critical data resources, tools, and services that require urgent integration into the RIs and the H2IOSC Marketplace; • Enhancements, FAIRification, and servification: determining which resources need improvement, FAIRification, or transformation into services; • Gaps identification: detecting the absence of crucial resources, tools, and services within existing RIs, and addressing the specific needs of the Italian research communities; • Training needs: identifying gaps in knowledge, skills, and competencies within the community to guide the development of targeted training programs, materials, and initiatives. 21 21 3. Questionnaire-based Survey In this section we dig into the main landscaping activity carried out in the WP, the questionnaire-based survey. 3.1 Survey Methodology Among the core surveying activity there is a designed and comprehensive online questionnaire aimed at directly engaging with the target research communities to gather essential information on key aspects, including the usage and needs related to data resources and technologies, awareness and current utilization of existing research infrastructure services, desiderata for new services or offerings, and training needs as perceived by community members (students, researchers, and experts pertaining to the different disciplinary areas represented by the Ris) . The latter focused particularly on competencies and skills identified throughout the survey, which are crucial for the development of effective training programs, materials, and campaigns. The decision to create a single questionnaire applicable across all RI communities was initially made to prevent data fragmentation and mitigate the risk of "community overload" [Error! Reference source not found.], given the interdisciplinary nature of many researchers’ fields of specialization, which can often span multiple RIs. The elaboration process began with identifying the expected information to be collected [Error! Reference source not found.], which encompassed several outcomes: the identification of stakeholders, user needs, training needs, priorities, gaps in available tools and services, existing resources both within and outside the RIs, the degree of FAIR compliance, the digitization of new resources, and best practices and standards. A series of targeted questions were then formulated to acquire the necessary information for each of these objectives. For instance, to determine if a resource indicated by a respondent was FAIR-compliant, four additional questions were included to gather details on its location, cataloguing, persistent identifier (PID), and licensing. Each question was accompanied by explanatory notes to provide further clarification and to minimize potential misunderstandings. Additionally, definitions and references for potentially ambiguous acronyms or technical terms, such as FAIR and PID, were included to help participants understand the questions and respond accurately. This supplementary information also aimed to spark curiosity and encourage engagement. Each question was carefully designed and assigned a specific response format, including open-ended, single-choice, and multiplechoice options. The questionnaire was assembled using LimeSurvey (Annex 1) and was structured into the following sections: 1. Personal Information & Privacy Policy: was aimed at collecting general information about the respondents age, career level, ERC research field, and institutional affiliation. Respondents were provided with a written informed consent form at the beginning of the questionnaire, before answering any questions, ensuring that they were fully aware of the purpose and scope of the data collection. All personal data collected follow the requirements of Article 13 of Regulation EU 2016/679 (General Data Protection Regulation), and responses have been analyzed and stored in anonymized and aggregated form. 22 22 2. Data Resources, Software, Tools, and Technologies: served to gather information about existing or newly created data resources, tools, etc. The goals here were multiple, from acquiring information on the state-of-the-art material within the Italian panorama, users’ knowledge, expectations, and needs, as well as identifying gaps that should be addressed by other project’s activities (e.g., points 3 and 7 in this list). 3. Projects: focused on collecting information on projects that have developed or are currently developing data resources, digital tools, software, or other technologies relevant to the respondents’ disciplines, such as linguistics, archaeology, and philology. 4. Training needs: aimed at collecting respondents’ expectations and comments regarding their personal training needs, with a particular focus on competencies and skills they deemed essential for future professional development. 5. Prior knowledge of RIs: was included to assess respondents’ awareness of and level of engagement with the partnering RIs, which was considered critical for designing outreach and training actions to increase participation and engagement. 6. Publications: was aimed to explore how respondents and their institutions publish scientific articles. Questions focused on publication practices, awareness of institutional policies, and the extent of commitment to open-access publishing. 7. Contacts: e-mail addresses were collected at the end of the survey for contact purposes only and only after the respondents’ informed consent. 8. Feedback: to share impressions and suggestions to the questionnaire and provide insights on how they learned about the survey. Regarding the deployment of the questionnaire, a first draft was initially distributed to a selected group of colleagues representing various subcommunities of interest for the investigation; these colleagues acted as a test group and provided valuable first-hand feedback. Following this preliminary phase, the revised first official version of the questionnaire was forwarded to a control group composed of members from key Linguistics, Digital Humanities, and Cultural Heritage associations, for which the exact numerosity was known. This step allowed for an estimation of the potential participation of the broader community in the subsequent stages of the investigation. As a result of incorporating all the important feedback acquired from the first communities involved, we produced a second version of the questionnaire 2 that was divided in two parts: the first comprising the set of questions regarding the profile of respondents, training needs, knowledge of the RIs and publications, while the second part 3 was dedicated to the description of resources, tools and services ever used and created by respondents. The final version of the questionnaire was widely disseminated across all target community members, associations, and mailing lists through various channels, including social networks, project and RI websites, conference presentations, and other outreach events (for deals see D2.1 and Annex 1). 2 The questionnaire-based survey is available at https://survey.cnr.it/index.php?r=survey/index&sid=825638&lang=it 3 To complete this second part of the questionnaire, respondents were asked at the end of the first part to indicate their preference for providing information either via a second online questionnaire or through a guided interview. 23 23 3.2 Results and analysis The analysis conducted on the two versions of the questionnaire provided valuable insights into the participation, awareness, and use of Research Infrastructures (RIs), as well as the respondents’ training needs and publication habits. • Participation: overall, the first version of the questionnaire recorded 396 views with a completion rate of 22% (86 responses), while the second version had a higher relative participation, with 65% (97 responses) of participants completing the first part and 56% (29 responses) completing the second. The responses primarily came from researchers affiliated with the National Research Council (CNR), with 106 respondents, and universities, with 70 respondents. Figure 4: summary of provenance of respondents Figure 5: institutions of provenance of respondents • Career and disciplinary field: the distribution of respondents by career level shows a higher presence of senior researchers, with 99 participants, followed by 44 mid-career researchers, 17 early-stage researchers, and 14 students. Most respondents belong to diverse disciplinary fields, with a predominance in the social sciences and humanities, particularly in fields such as SH4_9 (theoretical and computational linguistics), SH5_8 (cultural studies, cultural identities and memories, cultural heritage), and SH6_3 (archaeology of early literate societies and early civilizations). 0 5 10 15 20 25 30 35 CNR Università 24 24 Figure 6: career level of respondents Figure 7: disciplinary fields of respondents • Use and creation of resources and tools: relational databases are the most frequently used, with 66 responses, followed by linguistic databases and written language corpora, both with 54 responses. GIS resources and geographic datasets are used by 49 participants, while the creation of new resources is most common in relational databases, with 56 responses, followed by written language corpora with 36 responses, and linguistic databases with 28. Less frequently used technologies include artificial intelligence applications and 3D/BIM modeling. Seventy-nine percent of respondents reported being aware of research infrastructures, with CLARIN, DARIAH, and E-RIHS being the most well-known, with 84, 100, and 104 responses, respectively. However, 20% of participants are unfamiliar with the available resources, while 73% of those who do not currently use them express an interest in exploring them further. Regarding the sharing of research data, 109 respondents expressed their willingness to deposit their materials, while 22 have already done so. The reasons cited by those who have not yet deposited data include the perception of outdated platforms, Studenti Early stage Mid carreer Senior SH5_8 SH6_3 SH4_9 SH6_2 SH4_11 SH4_8 SH4_10 SH5_12 SH6_1 SH5_3 PE_6_11 PE6_3 PE6_7 SH5_1 PE6_6 SH6_6 PE6_1 PE4_2 PE6_4 SH4_1 SH5_2 PE8_3 SH6_4 SH6_5 SH5_7 SH6_13 SH5_11 SH2 SH5_9 0 5 10 15 20 25 30 25 25 complexities in managing data ownership, and the need for more information on sharing procedures. Figure 8: summary of resources, tools and services mostly used by respondents Database relazionali - inventari, cataloghi, … Banca dati linguistici Corpus di lingua scritta GIS / Dataset geografici Software per annotazione Rilievo 2D / 3D - rilievo fotogrammetrico, laser … Archivio Dataset di analisi - diagnostiche, prospezioni … Prodotto multimediale / virtuale - video, dataset … Corpus di lingua parlata Piattaforma per fruizione di un servizio specifico Wordnet Tokenizer Corpus multimodale Software per elaborazione dati Thesauri Corpus management building information modelling, building … Non ho usato/creato nulla Reference manager Software per elaborazione immagini Software per analisi statistiche Remote sensing Siti (es. Limes) Applicazioni AI (deep learning) Modellazione 3D/BIM Mappa interattiva georeferenziata con … Interviste, questionari 0 10 20 30 40 50 60 70 32 32 4. Interviews and other community-involving activities To complement the results of the questionnaire and gain a more detailed understanding of community needs, a series of semi-structured interviews was conducted with a selection of researchers operating within the SH6, SH5 and SH4 disciplinary sectors. The aim was to explore in greater depth how members of the scholarly community relate to digital data — in terms of usage, creation and dissemination — and to investigate their expectations, needs, concerns, and habits regarding data management and sharing practices. As part of the broader strategy for landscape assessment in the H2IOSC project, each infrastructure has curated a series of community-facing initiatives to gather expert feedback, foster dialogue, and align research infrastructure services with the needs and practices of scholarly communities. 4.1. Structure of interviews and results The structure of the interviews generally followed this outline: • Research field o Briefly define your primary and secondary research areas using keywords. • Use of digital tools in research o Use of already digital data ▪ How frequently do you use data that is already digital (excluding bibliographic data)? ▪ What types of data do you most often search for? o Do you consider that research infrastructures—such as data aggregators— offer added value to your research or that of your collaborators? • Digitisation of data o How frequently do you use tools for digitizing data? o In both cases, do you tend to rely on systems you're already familiar with, or do you explore new resources to obtain data? • Processing of digital data o Do you process data using digital or analogue methods? o What categories of software do you use most frequently? o What digital tools do you consider essential for the advancement of your discipline? • Availability of digital tools o In your current research context, do you feel that the available tools meet your needs? o Is there anything important that is missing or particularly difficult to use? o Would you be willing to pay for a service that fills these gaps? • Data production and management o Do you or your team produce digital data in the course of your research? o When selecting software tools to invest in (for yourself or your research group), how important is it that they are open source? o If the results were the same, would you prefer open source tools? 33 33 o If the results were different, would you prioritize output quality or methodological openness and data shareability? • Data sharing and dissemination o How important is sharing and publishing your data outside of traditional publication formats? o In what ways would you like your research data to be made accessible? Due to the informal nature of the interviews and the lack of consent from participants for both audio recording and name disclosure, it was not possible to produce full transcripts or attribute statements to specific individuals. Instead, structured notes were taken during each conversation. This respect for privacy created a more open and relaxed atmosphere, allowing participants to speak freely and raise critical points. The insights gathered in this way have been analysed and synthesised in the results section 4 . The sample, although limited in size, comprised 27 individuals selected to reflect a range of professional experiences and academic profiles. Due to privacy agreement, it was decided to maintain the identity of the participant anonymous; nevertheless, we would like to deeply thank all the individuals who generously agreed to be interviewed and shared their experiences and reflections. Particular attention was given to ensuring a balance across different career stages, following the same logic adopted for the questionnaire. The distribution of respondents can be defined as follows (Table 2): Definition Number Role Students 7 Master & Phd Students Early-career researchers 8 • University researchers • Postdoctoral fellows (assegnisti di ricerca) • Fixed-term researchers (CNR & University) • CNR Researchers – Level III Mid-Career researchers 5 • CNR Researchers – Level I • Associate Professors (University) • Ministry officials Senior Researchers 7 • CNR Research Directors • CNR Senior Research • Full Professors (University) Table 2: Distribution of interviews participants While the relatively small size of the sample may limit the generalizability of the findings, efforts were made to ensure diversity and representativeness across career stages and institutional backgrounds. This exploratory phase was primarily intended to provide a qualitative foundation for identifying recurring issues and needs, rather than to produce statistically generalisable results. 4 Preliminary insights on the interviews were published in (Mancuso et al. 2025; Mancuso 2025). 34 34 The interviews provided a very complex, albeit partial, snapshot of the practices, perceptions, and challenges that characterise the relationship between the research community, especially in archaeology, and digital data. One of the clearest findings is the strong demand, expressed across all career levels, for more structured training across the field. This need was spontaneously mentioned by 14 participants, mainly in the earlyand mid-career groups, and received unanimous agreement when addressed directly during the interviews. The type of training envisioned spans a wide range of topics, from data management to digital documentation. In two cases, interviewees explicitly invoked the need not just for a wider ‘digital literacy’, but for a broader data culture and critical awareness of the informational potential embedded in the data that the community itself produces. A second significant theme concerns the widespread and transversal engagement, at all levels, with both the production and the use of digital data, confirming what was already inferable by the questionnaire results itself. In terms of usage, participants frequently (88%) mentioned popular archaeological databases and online services (such as Beazley Archive) as well as wellknown WebGIS (such as the Geoportale Nazionale per l’Archeologia). Types of data produced ranged from inventories and lists to relational databases, GIS systems, statistical datasets, 3D surveys and virtual reconstructions, and even interactive notebooks or specific software or plugins. However, a structural issue, already inferred but now clearly confirmed, also emerged: most of this work begins from scratch, without building on existing datasets or infrastructures. This isolated approach appears to correlate with a widespread lack of familiarity with metadata standards, which several interviewees (ca. 50%) identified as a priority area for targeted and conscious training, as well as for broader thematic aggregation strategies. Some interviewees (28%) also pointed out that infrastructures themselves could play a central role in this process, not only as providers of training, but as reference points for defining and actively disseminating best practices. This could happen, some suggested, through the creation of guidelines (12%) or through the provision of ready-to-use and customisable tools and services (16%) aligned with community needs and standards for data management. Particularly telling is the variety of definitions attributed to the very concept of ‘data’. Across the interviews, the term was used to refer to a wide range of materials: from basic lists of artefacts to raw excavation records, from remote sensing outputs to unprocessed 3D scans. This semantic ambiguity also informs attitudes towards sharing practices: many respondents equated open access publication of written outputs with the sharing of data. While this position is perfectly legitimate, it coexists with other, more data-centric views that emphasise the need to publish structured, reusable datasets as independent research products (16%). In few cases, there was even a clear resistance to the idea of sharing raw or unprocessed data (20%). One particularly thoughtful observation pointed out that relying mostly on raw data sharing is not always sustainable, given that datasets often represent significant investments of time, training and financial resources. At least 36% of the interviewees, primarily earlyand 35 35 mid-career researchers, as well as some PhD students, raised concerns about the current lack of formal recognition for dataset publication in academic evaluation. Compared to the established value of journal articles or monographs based on the same data, datasets are still perceived as having little or no weight in career progression. This reinforces the central role of written publications, both in print and digital formats, which continue to be perceived and used as the primary vehicles for recognition and the preferred medium for communicating research outputs. In this respect, all interviewees acknowledged the importance of open access as a mode of scholarly dissemination. Nevertheless, in 15 cases, this support was accompanied by an awareness of the costs associated with open access publishing. In this context, digital data is often viewed as an interactive complement to a publication – a kind of navigable illustration – rather than as an autonomous object with its own life cycle, capable of being circulated, updated, or reinterpreted over time by other actors. In summary, what emerges from this evaluation is the complex image of a field that is, on the one hand, remarkably innovative in its experimentation with digital data and its potential, but on the other, still deeply rooted in traditional scholarly formats of transmission. Digital technologies are widely present, used, and valued, but rarely experienced as a native or structuring dimension of collaborative research practice. At present, it seems that digital infrastructures tend to be perceived more as functional supports than as fully embedded ecosystems within everyday scholarly routines. 4.2 Other community-targeted activities: format and results In line with the project’s landscape assessment goals, CLARIN-IT and DARIAH-IT carried out dedicated actions to engage with scholarly communities, promote interaction, and ensure that services evolve in response to concrete research practices and needs. The CLARIN community contributed through its series of monthly online “Cafés” — informal and interactive discussion sessions that bring together researchers, lecturers, students, infrastructure staff, and language resource experts to exchange ideas and address emerging topics in digital humanities and linguistic data infrastructures. Between November 2022 and May 2025, the CLARIN-IT team at ILC participated in and supported a wide range of CLARIN Cafés organised by CLARIN ERIC, each typically involving around 30 participants from academia and research institutions. These sessions addressed a broad spectrum of subjects — from text and data mining exceptions, multilingual corpora, and digital tools for online learning, to copyright in AIgenerated content, lexical semantic change, and pedagogical corpora containing sensitive language. More recent Cafés have expanded to cover translation technologies and workflows for SSH researchers, which attracted a broader interdisciplinary audience including translators, technical experts, and professionals working with multilingual data. Through its continued engagement, CLARIN-IT has contributed to knowledge exchange and 36 36 capacity building within the H2IOSC network, strengthening the link between the Italian and European research communities and ensuring that user feedback informs infrastructure development. In each Café session, participants share experiences, highlight infrastructural gaps, propose workflows, and reflect on organisational and technical requirements. Through open dialogue and moderated exchange, the CLARIN Café served as a mechanism for bottom-up feedback into the infrastructure ecosystem. These meetings addressed various challenges such as multilingual and annotated corpora, pedagogical resource creation, rights management in language data, automated transcription or annotation workflows, and user-driven infrastructure design. For example, the July 2024 Café entitled “Of Users and Infrastructures” explored how language researchers engage with resources and what infrastructures must provide in a post-pandemic digital research landscape. DARIAH-IT launched its own (i) series of community-oriented events and (ii) co-design working sessions. Between December 2023 and June 2025, a total of 20 dedicated meetings were held, including five public seminars under the title “I giovedì di H2IOSC” and 15 working sessions devoted to the design and implementation of the Digital Philology Hub (DPH) pilot (Activity 7.3). The “giovedì di H2IOSC” seminars served as a forum for targeted discussions on topics such as the lemmatization of ancient Italian, digital library models, AI-supported transcription workflows, and sustainability in digital research infrastructure. These events gathered feedback from scholars, technical experts, and institutional stakeholders, providing valuable insights into research practices, infrastructural needs, and methodological challenges. The feedback collected directly informed the mapping activity, helping to define key metadata fields such as data standards, technical readiness levels, and the role of responsible institutions. Moreover, these sessions helped surface community priorities for reusable workflows, openness, and technical interoperability. Complementing this public dialogue, the DPH working sessions supported a structured codesign process aimed at shaping the DARIAH-IT Pilot. Held monthly, these meetings enabled continuous feedback loops between researchers and infrastructure providers. As a result, core requirements for the pilot were collaboratively defined. This iterative process ensured the alignment of pilot components with both scholarly workflows and infrastructure capabilities. In the context of H2IOSC, the interplay between the CLARIN Cafés and the “Giovedì di H2IOSC” series broadens the scope of WP2’s landscape mapping activities. While the CLARIN Cafés emphasise language-resource infrastructures, multilingual issues, and cross-domain digital humanities practices, the “Giovedì” seminars focus on heritage sciences, digital philology, and 37 37 the Italian node ecosystem. Together, they create shared touchpoints between linguistic, heritage, and digital-humanities infrastructures, ensuring that service design, metadata documentation, and workflow development reflect diverse disciplinary needs. Collectively, these engagement activities contribute to the overall community mapping strategy. They clarify the categories of resources to be identified—including not only software and services but also workflows, data types, and user support mechanisms—and validate key criteria such as usability, openness, standard compliance, and integration potential. Ultimately, the interviews and participatory events serve as both a diagnostic and developmental layer of landscape analysis, grounding the mapping methodology in real SSH research practices and community expectations. In Deliverable 1.3 “H2IOSC Training Activities Coordination and Management Procedures” a complete list of engagement and dissemination activities is provided for further information. 38 38 5. Focus Groups 5.1 Focus Group Methodology Focus groups are a widely used research technique in the humanities and social sciences. They are used to collect rich data derived from verbal interactions among a group of people brought together specifically to discuss a particular topic, through a series of stimuli provided by a researcher (Caradonna et al., 2025). A typical focus group (FG) involves 6–8 participants, a researcher who proposes discussion topics and moderates verbal interactions, and an assistant who follows the discussion and notes particularly relevant points and topics. Several variations of this basic setup exist: the number of participants can be increased to a maximum of 12–13, and the assistant can be replaced by a second researcher acting as a critical friend. This researcher assists the main moderator in managing interactions and covering the topics to ensure all knowledge objectives are achieved. The most common methodology involves FG meetings taking place in person in a suitably equipped room with furnishings that allow for easy interaction. According to traditional methodology, the discussion takes place in a dedicated environment — usually a room equipped with audio-visual recording equipment — and lasts between one and a half and two hours. This phase is very important as the significance of the data is directly related to the discussion. The discussion depends on the stimuli provided, the researcher's ability to encourage communication and interaction between participants, and the characteristics of the participants invited. Depending on the research problem and the questions to be answered, the focus group follows a standard procedure. 1. The characteristics that participants must meet are identified. 2. The interviewees are identified and invited to attend on a given day, at a given place and time. 3. The discussion takes place, which is normally audio-video recorded. 4. The discussion is transcribed. 5. The data from the discussions is analysed. Of the above stages, the first and fourth (transcription of the discussion) are undoubtedly the most important. In the first case, it is crucial that the participants possess the necessary knowledge and experience; otherwise, they are unable to respond to the stimuli, and any silence or inappropriate responses could hinder or divert the discussion. In the second case, the transcription should retain as much non-verbal communication, 'fillers' and other forms of expression as possible, since, as communication theory tells us, explicit speech constitutes only part of a person's communicative message. Fillers, tone of voice, gestures, and all other forms of non-verbal communication (e.g. nodding or shaking one's head, gesturing with one's hands) are of great importance. 39 39 In the analysis of the data, it is important to look at all communications. From the very beginning and then including factors such as the group atmosphere. For example, a participant's statement could trigger a series of consents that are only manifested through non-verbal communication (gestures of assent, activation of the body through the posture of the torso, etc.) but do not produce verbal communication. If this involves all participants, it creates a climate of consensus, even if only momentary, on a statement, so much so that it can then be taken up in the data analysis as “the whole group agrees with what ... said”. What has just been expressed represents the cognitive power that can be gained from a group meeting. Obviously, the difficulty lies in the fact that when presenting the results, it is necessary to provide specific interpretations, otherwise the impression could be given that the analysis does not perfectly match the data. In recent decades, this classic approach has been complemented by group meetings held via videoconference using web applications. Web-based technologies have made it possible to expand the opportunities for using the focus group technique, with obvious advantages for research activities such as those offered by WP2, for example. In WP2, landscaping activities were carried out using a mixed method approach, i.e. surveys with quantitative and qualitative tools, and the focus groups planned as part of the project activities were conducted via video call. This made it possible to involve a very large audience of participants, even though they were in different areas of the peninsula, thereby substantially reducing the costs of the activity. Below, we describe the phase of defining the meetings. The WP research group planned four focus groups, two of which were conducted between July and December 2024, and the remaining two in March 2025. First, the characteristics decided during the planning phase are listed, followed by the characteristics of the participants who took part in the meetings. 5.2 Features of planned focus groups. During the planning phase, it was decided that each group would consist of eight participants. The following variables were monitored for these participants: biological sex, age, membership of the four macro research areas to which the infrastructures confederated in H2IOSC belong, and orientation towards innovation (Tables 3 and 4). 40 40 Table 3: Number and characteristics of the focus group JUNIOR 1 and 2 Women Men Age researchers Senior Researchers/Univers ity professors Senior researchers Researchers/Univers ity professors Total 33 - 40 2 2 2 2 8 41 + 2 2 2 2 8 Total 4 4 4 4 16 Table 4: Number and characteristics of the focus group SENIOR 1 and 2 Suggestions for infrastructures regarding the identification of group participants regarding orientation towards innovation. The identification of these figures was delegated to the four infrastructures. It was asked to invite PhD students, post-doctoral researchers/junior research fellows and senior researchers/university professors that were: • open and extroverted individuals who allow for free and open discussion of the possibilities for innovative and digital development in the SSH; • open-minded scholars with a strong focus on the future; • participants with solid approaches to both “classical” and “innovative” SSH studies; • open and extroverted researchers who can discuss the possibilities for future development of the SSH but who are not ideologically biased or overly enthusiastic; • Scholars who also express critical but constructive views on the digital future. The four groups were named: FGSENIOR1, FGSENIOR2 and FGJUNIOR1, FGJUNIOR2. From a theoretical standpoint (see Figure 1), the research team resolved to proceed with two distinct groups. The initial group (FGSENIOR1) comprised individuals who exhibited a marked propensity towards exploring novel avenues of research, utilizing innovative or experimental digital resources, and possessing substantial research experience. The second group (FGJUNIOR1) consisted of individuals who, despite possessing a high level of education, lacked extensive research experience. The former group was characterized by a strong focus on the Women Men Age PhD students Post-doctoral researchers/Junior research fellows PhD students Post-doctoral researchers/Junior research fellows Total 25 - 28 2 2 2 2 8 29 - 32 2 2 2 2 8 Total 4 4 4 4 16 41 41 present of learning (FGJUNIOR1), coupled with a positive outlook towards the future, driven by a desire to pursue a research career (FGSENIOR1). The research group adopted this theoretical position subsequent to the observation that the digital dimension exhibited a divergent geometry between the two groups. Senior researchers and professors belong to a generation (Generazione X) 5 where the digital dimension has grown in importance, while doctoral students and junior researchers belong to a generation (Generazione Z) 6 in which digital technology was already an integral part of both everyday life and study and research activities. This discrepancy, when considered in conjunction with the varying degrees of research experience, gives rise to disparate expectations regarding infrastructure orientation amongst users of differing ages. Figure 21Expected conceptual positioning of the two focus groups. The participants of the four groups have been invited to deliberate on their expectations regarding the outcomes, future developments, and long-term sustainability of the project. Furthermore, researchers and students have been assigned the following tasks: 5 In sociology, Generation X is a term widely used in the Western world to describe the generation born in different years depending on the source (1964-1979 or 1965-1980) who lived through a crucial historical transition, marking the shift from the analogue to the digital age[1]. The term was coined by Canadian writer Douglas Coupland, https://it.wikipedia.org/wiki/Generazione_X 6 Generation Z (or Centennials, Digitarians, Gen Z, iGen, Plurals, Post-Millennials, Zoomers) is the generation of people born between the second half of the 1990s[1][2][3][4][5] and the first half of the 2010s[6][7] (although the exact numbers vary depending on the different definitions presented),[8][9][10][11] and whose members are generally the children of Generation X (1965-1979) and the last baby boomers (1946-1964), https://it.wikipedia.org/wiki/Generazione_Z 48 48 complementarity. Tables 5, 6, and 7 present a comparative overview of the needs, priorities, and perspectives expressed by each RI as part of this coordinated strategy. RI CLARIN-IT LANDSCAPING OBJECTIVES • strengthening of the infrastructure (general): enrich CLARIN resource catalogues; • Identifying gaps in data and software resources, in order to guide actions for their onboarding and deposition into CLARINIT certified repositories • mapping the needs of the target communities • reaching out to resource stakeholders LANDSCAPING CRITERIA • Digital(ised) language resources related to language technology, linguistics, NLP or applications in the Digital Humanities • resources produced or primarily used by domain experts/scholars • Deposited in a CLARIN certified repository (F) • clear license (A) • format/standard (I) • in/for the Italian language and/or classical and ancient indoeuropean languages (R) LANDSCAPING SOURCES • CLARIN VLO • Sector-specific conferences and symposia: esp. AIUCD, CLIC-IT • Interviews and dedicated events LANDSCAPING TOOLS • spreadsheet for internal landscaping • Search into conference proceedings • H2IOSC Landscaping Questionnaire RESOURCES INFORMATION • authors • status (created/updated/used) • publications related to the resource • year of creation/integration • reference project • type of resource • type of data (oral/written/video) Table 5: CLARIN-IT Landscape criteria RI DARIAH-IT LANDSCAPING OBJECTIVES • creation of a resource catalogue to be made available to other WPs • evaluation of the FAIRness of the resources • mapping the needs of the target communities • identification of available resources/tools in relation to their domain and use; mapping of unmet needs in line with the previous objective 49 49 • reaching out to resource stakeholders LANDSCAPING CRITERIA • resources related to texts, literature, and the language of Old Italian • resources produced or primarily used by domain experts • resources involving DARIAH participation • relevance to the objectives of the Pilots (WP7) • resources related to the following academic-disciplinary sectors: L-FIL-LET/09, L-FIL-LET/10, L-FIL-LET/12, L-FIL-LET/13 LANDSCAPING SOURCES • DARIAH Registry • DARIAH Tools and Services Catalogue • Opera del Vocabolario Italiano (OVI-CNR): http://www.ovi.cnr.it/Progetti.html • SSHOC Marketplace • Sector-specific conferences and symposia • Recommendations from domain experts and partners • Journal of Cultural Heritage • Umanistica Digitale • Digital Humanities Quarterly • Annali d'Italianistica • Digital Medievalist • Digital Scholarship in the Humanities • Humanist Studies and the Digital Age • International Journal of Heritage in the Digital Era • Studies in Digital Heritage • Scientific reports and sector-specific articles on Zenodo (used as a repository for several DARIAH-affiliated projects) LANDSCAPING TOOLS • spreadsheet for internal landscaping • research in journals and other sources using keywords or thematic search • consultation with domain experts and project partners • collaborations with partners and other institutes RESOURCES INFORMATION • Lead RI/Institute • Description • Current status • Curator • Link to the webpage and source code • Standard/format Table 6: DARIAH-IT Landscape criteria RI E-RIHS.it LANDSCAPING OBJECTIVES • survey of existing digital resources for Heritage Science (HS) and Cultural Heritage (CH) 50 50 • survey of the main services (tools) developed for / used by the research community • identification of the research community’s needs • survey of the main research centers actively involved in creating digital resources for HS and CH • implementation of the dataspace LANDSCAPING CRITERIA • thematic relevance: digital datasets related to HS or CH • thematic relevance: tools developed by / used by the research community • thematic relevance: open access journals and their citation network over the last 5 years LANDSCAPING SOURCES • IperionHS • ARIADNEplus • Europeana • Zenodo Communities • Dataspace | ISPC • scientific articles and conference proceedings published in open access: Archeologia e Calcolatori journal repository • responses to the Landscaping questionnaire • scientific articles in open access journals: Scopus repository. Search terms include "Heritage Science", "Cultural Heritage", and "Digital Cultural Heritage" • scientific articles in open access journals: Journal of Cultural Heritage • scientific articles and sector-specific conference proceedings: Computer Applications in Archaeology • sector-specific conference proceedings: ArchaeoFOSS LANDSCAPING TOOLS • H2IOSC Landscaping Questionnaire • Spreadsheet for data entry, management, and retrieval • Keyword-based search of source materials • Web App: DHeLO - Digital Heritage Landscaping Platform • Zotero BiDiAr collection RESOURCES INFORMATION • Authors • Research institutions • Research projects • Datasets linked to research projects • Tools for the creation, visualization, management, and reprocessing of data Table 7: E-RIHS.it Landscape criteria Landscape Sources SSHOC MarketPlace 51 51 The SSH Open Marketplace is a key resource for identifying and contextualizing tools, datasets, training materials, services, workflows, and publications relevant to the Social Sciences and Humanities (SSH) research community. It provides an accessible, user-oriented portal where researchers can find digital solutions to support all phases of the research data lifecycle in SSH. The SSH Open Marketplace is built around five main content types: • Tools and services • Training materials • Workflows • Datasets • Publications CLARIN VLO The CLARIN Virtual Language Observatory (VLO) is a discovery platform that provides unified access to a wide range of language resources, including corpora, lexical databases, metadata records, and language tools deposited in the various CLARIN certified repositories distributed across Europe. Built on a metadata harvesting and aggregation framework, the VLO allows researchers to search, browse, and filter resources according to multiple facets such as language, resource type, modality, or license. By exposing interoperable metadata compliant with CLARIN standards (Melaccio et al., 2025), the VLO facilitates cross-resource comparability and supports the findability and accessibility dimensions of the FAIR principles. As a central access point, the VLO plays a strategic role in enabling scholars from diverse research (sub-)communities to discover and reuse resources beyond disciplinary and geographic boundaries. DARIAH-EU Tool and Services Catalogue In September 2023, DARIAH officially launched the DARIAH Tools and Services Catalogue, now accessible through both the DARIAH website and the SSH Open Marketplace. The catalogue features over 200 resources and distinguishes between two main types of services: Core Services and Community Services. Core Services are owned and managed by DARIAH-ERIC and delivered through its national nodes; they represent the infrastructure backbone and strategic mission. Community Services, on the other hand, are maintained by one or more DARIAH partner institutions and typically support local capacity building and research instrumentation. The DARIAH Tools and Services Catalogue serves as a central access point for exploring the rich landscape of digital resources offered by the DARIAH infrastructure. The catalogue does not represent the full internal portfolio of DARIAH services but is rather a curated selection of those deemed suitable for public listing. It offers a user-friendly interface that allows scholars to discover, access, and evaluate available digital resources. As of now, the catalogue 52 52 includes 29 contributions from DARIAH-IT, further illustrating the active involvement of national nodes in the provision of reusable and open-access research tools. 6.2 Knowledge Organization Schema The adopted methodology focused on developing a structured system for organizing and describing resources, with attention to their typology, stage within the data lifecycle, data formats, applicable standards, and disciplinary characteristics. While devising the mapping strategy, attention was also placed on their alignment with the FAIR principles: how easily resources can be found, accessed, integrated, and reused across contexts. This activity involved a systematic collection of information on both datasets and digital tools, drawing on inputs provided by the participating Work Packages and affiliated institutions. Resources were described using a shared metadata framework to enable comparative analysis and consistent evaluation. In the case of datasets, emphasis was placed on aspects such as content type, disciplinary scope, availability, data stewardship, and adherence to recognized standards. For digital tools and services, the analysis focused on their core functionalities, technical attributes, stage of development, and their ability to interact with or enhance relevant datasets. The metadata framework incorporates essential descriptive elements, including basic resource information, technical characteristics (such as supported input/output formats and compliance with standards), institutional ownership, and possible links to other Pilots or research infrastructures. This mapping exercise offers a structured and integrated perspective on both existing and emerging digital resources, while also highlighting external assets that may be suitable for integration or alignment with the project's strategic objectives (Table 8). Field Description Datasets Tools ID Identifier for internal reference X X Name full name (expanded acronym) X X Acronym Or short name of the project X X Title To be added to the name if provided X X Type of dataset classification of the dataset (e.g., corpus, glossary, image repository, edition, etc.) X Type of tool classification of the tool (e.g., annotation tool, data management system, learning system, visualization platform, etc.) X Description type of information conveyed; content included in the datasets / main features, details about the context of use X X Native data access tool how are the data accessed? If available online, specify the name of the tool used X 53 53 to access them (preferably also include the tool in the tools list) Link (to web instance or website) link to the download page if available, or to the online instance X Link to the tool link to the download page if available, or to the online instance X Last update (optional) Date format X Offered services / functionalities details on the type(s) of supported services, e.g., DBMS, KB, LMS, ML, etc. X Relevant disciplines Relevant discipline: academic or research domain(s) to which the dataset or tool is related (e.g., linguistics, philology, archaeology, lexicography, art history, heritage science, digital humanities, etc.) X X Documentation Documentation: link to user manuals, technical documentation, or descriptive materials related to the dataset or tool X Involved Research Infrastructure (RI) infrastructure that contributed to the development X X Responsible institution/center name of the institution, research center, or organization in charge of maintaining or overseeing the dataset or tool X X Curator name of the individual or institution responsible for the curation, maintenance, or scholarly oversight of the dataset X Contact person if different from the curator / whom to contact for clarifications; not necessarily the developer X File format (extension) format or extension of the files stored in the dataset (e.g., .xml, .csv, .json, .xlsx, .txt, .rdf, etc.) X Data model/standard conceptual model or formal standard used to structure the data (e.g., CIDOC CRM, TEI, Dublin Core, EAD, METS, etc.) X Programming language (optional) programming language(s) used to develop or implement the tool (e.g., Python, Java, PHP, JavaScript, etc.) X Link to source code (if available) e.g., on GitHub X Developer / Responsible center individual, institution, company, or development team responsible for creating the tool X Developer contact (optional) contact information for the developer or development team responsible for the tool X Input data format(s) (file extension) file extension of the input (e.g., xml, csv, xlsx, dtd, txt, etc.) X 54 54 Required data standards standard followed by the input data (e.g., TEI, ICCD, EAD, EAC, etc.) X Status developed or under development X Usable datasets datasets that can be processed, accessed, or integrated by the tool (include names, links, or references where applicable) X Output data format file extension of the output (e.g., xml, csv, xlsx, dtd, txt, etc.) X Provided data standards standard followed by the output data (e.g., TEI, ICCD, EAD, EAC, etc.) X Link to Pilot (for internal H2IOSC use) Reference to the H2IOSC Pilot the resource can be compared to X X Notes additional remarks, contextual information, or clarifications relevant to the dataset or tool that do not fit into other fields X X Priority relevance or urgency level assigned to the dataset or tool within the project context (low, medium, high) X X Table 8: Fields and Metadata of tables of Datasets and Tools To develop a unified framework for resource mapping within the H2IOSC project, WP2 began analyzing the respective needs to develop a unified strategy. While each RIs differed in scope and disciplinary focus, they revealed several common themes, including the need to document resource status, development context, disciplinary relevance, and FAIRness aspects. The Knowledge Organization Schema provides the conceptual and structural basis for organising the landscape tables, allowing for the systematic classification of resources according to their typology, stage within the data lifecycle, technical characteristics, and degree of FAIR alignment. The mapping and analysis activities carried out within WP2 laid the groundwork for understanding the heterogeneous landscape of research infrastructures, resources, and user needs across the domains of humanities, language sciences, and heritage science in Italy. The process has highlighted the diversity of data types, access modalities, and disciplinary cultures, prompting an initial harmonisation effort to align perspectives and expectations across domains. The results of WP2 didn’t remain confined to internal reporting but have actively fed into the development of key strategic assets of the H2IOSC infrastructure. The collected data and reports have been incorporated into the design of the H2IOSC Marketplace, where they inform the classification, onboarding, and presentation of resources and services in a way that reflects the real needs and practices of the research communities. Furthermore, they constitute the informational foundation of the H2IOSC Observatory, a permanent knowledge space dedicated to monitoring communities, identifying trends, and tracking evolving training 55 55 and technological demands. WP2 Permanent Observatory (§7) acts as an interface between analysis and infrastructure, ensuring that strategic foresight remains embedded in platform development. A concrete outcome of the landscaping process is the development of DHeLO - Digital Humanities Enhanced Learning Observatory, a platform that leverages the knowledge base generated by WP2 to support exploration, discovery, and visualisation of disciplinary resources and user needs. DHeLO exemplifies how landscaping results can be transformed into operational tools supporting both the community and platform designers. WP2 also established robust interdependencies with the other Project’s Work Packages, sharing data, frameworks, and survey results. It collaborated with WP6 in identifying and describing available services; contributed to WP4 in the definition of the resources to be included in the semantic framework; cooperated with WP5 in shaping onboarding strategies for the Marketplace; and worked with WP7 in aligning the landscaping results with pilot implementation. This demonstrates the central role of WP2 as both an analytical and integrative component of the project, enabling the operational reuse of its outputs across the entire H2IOSC ecosystem. 6.3 Unified H2IOSC Landscape criteria and methods These standardized metadata fields were then organized into three dedicated spreadsheets: • Pilot Requirements, aligned with WP7, capture the functional needs of the Pilots and their expected tools or services. • Landscaping Tools, used by WP6, WP7, and WP8, lists software, services, and platforms relevant to the project. • Landscaping Datasets, aligned with WP6, and WP7, focuses on structured data resources used or produced by the community. Each sheet is structured into thematic sections: General Information, Development, Input/Output Data, Curation, and Interconnections, with fields derived from the original RI criteria. For example, mentions of "status," "developer," and "programming language" were formalized under a Development section, while references to "disciplines," "projects," and "repository presence" informed the General Information and Interconnections fields. The adoption of this structured template allowed WP2 to harmonise metadata, thus ensuring that the information provided by the various RIs and WPs remained comparable, reusable, and interoperable. This common schema supports both the mapping effort and the broader goals of WP2 to achieve landscaping of the whole Federation. 56 56 The organization of this activity is detailed in the following paragraphs, which explain the sheets prepared to carry out the WP2 mapping. Pilot Requirements This sheet gathers the requirements of the Pilots from the four infrastructures, as described by WP7, that are of interest for WP2. The objective of this analysis is to create a comprehensive list of requirements, useful for identifying shared development goals and highlighting interconnections and overlaps that may arise during the project. Fields Section: General Information includes: • Reference Pilot: the name of the reference pilot (selected from a drop-down menu); • Task: the related task (drop-down menu); • RI: the reference research infrastructure (drop-down menu); • Description: a brief description of the Pilot’s requirement; • Status: the development status (drop-down menu), which can be “Developed” (indicating a product already available), “Under development” (a product currently in progress but not yet in production), or “Not developed” (a product required by the Pilot but whose development has not yet started). Section: If Developed or Under Development (to be completed if the “Status” field is “Developed” or “Under development”): • Development percentage: approximate percentage of the tool’s development completion; • Developer: the institute, university, or company responsible for the tool’s development; • Link to the tool: webpage link where the tool is available; • Link to the code (if available): link to the GitHub/GOGS/Jupyter repository containing the tool’s code. Section: If Not Developed (to be completed if “Status” is “Not developed”): • Development approach: specifies whether development is handled internally within the project or outsourced through an external tender. Section: Reports and Interconnections includes: • Link to other Pilots: indicating whether the tool may be of interest to other Pilots beyond the one that produced it; 57 57 • Contact person: contact details of the person responsible for the entry; • Report of already developed products: indication of external tools (outside the project) known to satisfy the Pilot’s requirement; • Notes: field for comments or remarks. Landscaping Tools This sheet allows the relevant Work Packages (WP2, WP6, WP7, WP8) to report tools and services relevant to the objectives of each WP and of the H2IOSC project. The entries may refer to products developed within the Operative Units (OUs) involved in the project, as well as those developed externally. Fields Section: General Information includes: • ID: a sequential number that uniquely identifies the resource; • Name: name of the resource; • Description: brief description of the resource; • Link to the tool: webpage link where the resource is available; • Last update: date of the most recent update; • Offered services: types of services provided (e.g. data analysis, NLP, NER, controlled vocabularies, etc.); • Relevant disciplines: research areas for which the resource was designed (e.g. linguistics, archaeology, lexicography, legal studies, etc.); • Reference WP: the Work Package responsible for entering the resource into the sheet; • RI: whether the resource was developed within one of the research infrastructures involved in the project or is otherwise available in the corresponding marketplaces. Section: Development includes: • Status: development status (drop-down menu: Developed, Under development, Not developed); • Development percentage: approximate percentage of completion; • Developer: institute, university, company, or individual responsible for development; • Developer contact: email contact of the developer; • Programming language: programming language(s) used to develop the resource; • Link to the code (if available): GitHub/GOGS repository link containing the tool’s code. Section: Input Data includes: • Input data format: format of the data accepted by the resource; 64 64 Heritage», while the HS domain was explored through «Heritage Science» and «Heritage» journals. Additionally sector-specific proceeding series (e.g. Computer Application in Archaeology, MetroArchaeo) within the last five years were considered. During this process, due to H2IOSC Project requirements, priority was given to Italian research products, with consideration also given to those developed abroad, but when conducted by Italian researchers or research groups. 7.3 DHeLO 2.0. - A technical overview During the project, reflections stemming from the analysis of the questionnaire data and the interviews made it increasingly clear that DHeLO could assume a more active role within the broader landscape of research infrastructures. An idea that gradually took shape during the platform’s development phase, and was later reinforced through several interviews, was the potential value of a system capable of aggregating and thematically, geographically, and chronologically indexing existing datasets, interactive resources, and research projects. It became evident that DHeLO could move beyond functioning solely as a passive virtual observatory operating behind the scenes, evolving instead into an active service for the community. The envisioned system would not require contributors to adhere to a rigid data model; rather, it would focus on collecting and exposing digital resources already available online, in diverse formats, and making them accessible through a unified, searchable interface, thus addressing a gap identified within the current infrastructure ecosystem (see further §4 on survey results). The transition from a relational database to a service-oriented infrastructure designed for the broader research community called for a partial reconfiguration of DHeLO. Rather than a full restructuring, this shift marked a progressive alignment with LOD principles, aiming to improve interoperability, discoverability, and semantic integration (https://www.w3.org/DesignIssues/LinkedData.html). This new perspective made it easier to map and contextualise resources that are often available online with only minimal or generic metadata or whose descriptive elements are accessible exclusively through associated scholarly literature. Structuring the platform around LOD-compatible logics offered a way to capture and represent such resources more effectively, while maintaining a lightweight, inclusive approach to aggregation. This conceptual transition was accompanied by a substantial rethinking of the technical infrastructure supporting DHeLO (Fig. 27). 65 65 Figure 27: Project page on DHeLO 2.0 After evaluating a number of alternatives (such as Arches, used in the configuration of the ISPC dataspace), it was decided to migrate the system to Omeka S, an open-source web publishing platform developed by the Digital Scholar project at the Roy Rosenzweig Center for History and New Media, based at George Mason University (https://rrchnm.org/). Omeka S (https://omeka.org/s/) is the evolution of the original Omeka Classic software (https://omeka.org/classic/), and was designed to support multi-site environments, semantic interoperability, and the management of digital objects across multiple contexts; functioning as a content management system (CMS) and allowing for both structured data handling and visual customisation via templates and CSS. At its core, Omeka S provides a powerful yet user-friendly backend for managing structured metadata, as well as a flexible publishing layer for creating public-facing sites that present digital collections in curated and accessible ways. Notably, the platform can be managed effectively by users with no specific technical background, making it especially suitable for collaborative and multidisciplinary environments. While commonly adopted in the digital humanities and museum sectors, its use in DA project is still limited, though steadily increasing. One of the key advantages that Omeka S provided to DHeLO lies in its modular 66 66 architecture. Through an ecosystem of plugins and extensions, it was possible to integrate advanced functionalities without compromising the system’s core structure. Among these add-ons, one of the features that most strongly influenced the choice toward this software was the Collecting module, which allows users to submit metadata for content to be ingested into the system. This capability is particularly significant, as it enables broader community participation while contributing to the long-term sustainability of the platform. This functionality envisions a future where DHeLO could become partially sustained through usersubmitted content, curated and published with minimal editorial oversight by a restricted editorial board. This orientation aligns DHeLO with emerging models of participatory infrastructures, where communities are not only users but also contributors to the growth, curation, and sustainability of shared digital ecosystems. Another compelling reason for choosing Omeka S is its close conceptual and institutional affinity with Zotero, the widely used reference management tool (www.zotero.org). Both platforms were developed within the same institutional framework, and share a commitment to openness, usability, and sustainability. In the context of the H2IOSC project, this synergy was particularly valuable given the relationship between DHeLO and BiDiAr, a bibliographic database originally built with Zotero (see further), and soon to be imported into Omeka S. The possibility of cross-platform integration, particularly in terms of metadata reuse and synchronisation, positioned Omeka S as a natural solution for supporting the long-term development of both platforms within a shared ecosystem. Additionally, Omeka S includes built-in RESTful APIs, which facilitate programmatic access to the data and support interoperability with external systems and services. This feature aligns closely with current recommendations on infrastructural openness and machine-actionability and allowed to exploit DHeLO’s data for the Open Digital Archaeology Hub (see further). Equally important in the selection process was Omeka S’s native support for controlled vocabularies and ontology-based metadata structures. The platform is designed to work natively with RDFcompliant ontologies and supports the import of external vocabularies expressed in SKOS or OWL. It allows seamless alignment with widely adopted semantic standards such as Dublin Core, FOAF, and schema.org, and – especially relevant for the archaeological domain – with CIDOC CRM, which is increasingly recognised as a reference model for the representation of CH data. In this respect, the contribution of the H-Setis platform, developed within the project to collect and share semantic resources and ontologies within the HS an CH domain, was instrumental in defining a conceptual and technical framework of the new DHeLO infrastructure, compliant with disciplinary domain standards. 7.4 DHeLO 2.0 - Adjusting the data model As previously mentioned (§7.3) the shift to a LOD infrastructure brought with it a partial reorganisation of DHeLO’s underlying data model. A close analysis of the existing entities revealed a high degree of correspondence with classes defined in two widely adopted ontologies: FOAF and the DCTerms. These vocabularies, while relatively lightweight, provide 67 67 a robust and widely compatible foundation for describing agents, organisations, intellectual outputs, and related entities. Accordingly, the core tables were mapped to the following classes: foaf:Person for individuals; foaf:Organization for institutions; dctype:Software and dctype:Dataset for tools and research products, respectively; and foaf:Project for research initiatives. One important refinement emerged during this reconfiguration: the introduction of dctype:InteractiveResource as a distinct class to represent digital assets that serve as interfaces for accessing data — such as visualisation platforms, WebGIS portals, or multimedia exploration tools. In the previous model, these resources had been grouped under the general category of ‘tools’, a choice that inadvertently reduced their intellectual specificity and made it difficult to establish precise relationships with other entities in the system. By assigning them to a dedicated semantic class, it became possible to capture their dual nature as both access mechanisms and scholarly outputs, and to situate them more meaningfully within the larger graph of research activities and digital production. Importantly, this semantic distinction also enabled these resources to benefit from the same thematic, geographic, and chronological indexing strategies applied to datasets — a crucial step in enhancing their visibility and interoperability within the platform. The thematic classification system adopted in DHeLO remained consistent with the previous version of the platform. It continues to draw on the taxonomy developed by the journal Archeologia e Calcolatori, which offers a well-established and domain-specific framework for categorising research products in the field of archaeological computing. As for spatial and temporal metadata, the system had already adopted a gazetteer-based approach in its earlier implementation, and this strategy was retained in the current version. However, it was observed that in order to enable certain advanced visualisation features, such as interactive maps and timelines, it would be preferable to structure spatial and temporal references as independent items within the data model. For this purpose, two additional classes were introduced: dcterms:Location and dcterms:PeriodOfTime, which serve as containers for geographic and chronological metadata imported from external gazetteers. On the temporal side, this approach ensured continuity with the use of the PeriodO gazetteer. On the spatial side, it opened the possibility of integrating and combining data from multiple gazetteers, thereby addressing a potential limitation in relying solely on a single source, such as Pleiades, where certain locations might not be represented. The introduction of dedicated items for places and periods not only improved the precision of metadata but also enhanced the platform’s ability to support semantic linking, filtering, and visual exploration of resources across time and space. 68 68 Figure 28: Entity relationship schema for DHeLO 2.0. Finally, the bibliographic references associated with the various entities are drawn from the BiDiAr system and include links to their corresponding Zotero records, allowing for rapid consultation and reuse. The following table lists the metadata fields used for the main resource classes structured within the DHeLO platform (Table 10). Each class is described according to a set of standardized properties that ensure semantic clarity and interoperability. For a complete overview of the DHeLO data model and its implementation, please refer to the official documentation (see also: Annex 3.2). 7.5 Results overview The activities carried out during the reporting period have contributed to consolidating DHeLO as a key infrastructure for the collection, organization, and dissemination of digital resources in the wider domains of Digital Cultural Heritage and Heritage Science. Beyond the technical improvements, the platform has been enriched with a diversified corpus of data, thus providing researchers and practitioners with a broader and more representative basis for their work. Importantly, DHeLO itself should also be regarded as one of the main outputs of the landscaping process undertaken within WP2. By systematically gathering, harmonizing, and exposing resources, the platform not only documents the current landscape of digital practices but also enables their further development in terms of interoperability, accessibility, and reuse. The platform, already accessible at https://dhelo.cnr.it, is fully operational and available to the community. This ensures that the outcomes of the landscaping activity are not confined to an internal testing phase but are instead offered as an open tool for immediate consultation and use by researchers and institutions. In particular, the platform has been populated with a wide range of resources, reflecting both the heterogeneity of the cultural heritage record and the multiple needs of the research community. The data currently available include (fig. 29; 30): • Dataset (175) 69 69 • Interactive Resources (100) • Projects (183) • Softwares (81) • Locations (386) • Periods (101) • Organizations (298) • People (1074) Figure 29: Quantities of data collected in DHeLO, broken down by type of entity. 0 200 400 600 800 1000 1200 Serie 1 70 70 Figure 30: Geographic distribution of the resources collected within DHeLO, visualized through the platform’s mapping interface. Taken as a whole, the results achieved demonstrate the capacity of WP2 to generate both conceptual and technical tools in support of the Digital Cultural Heritage and Heritage Science communities. They also provide a foundation for further developments in the framework of H2IOSC and beyond, in line with the principles of Open Science and FAIR data management. In this perspective, the data collected through DHeLO will also be harvested and made available within the WP2 Permanent Observatory (§8), ensuring their long-term visibility and integration into a broader monitoring environment. 7.6 BiDiAr - Bibliography of Digital Archaeology In recognition of the continuing centrality of bibliographic resources as the primary means of scholarly communication within the research community – as emerged from the interviews – particular importance was assigned on the management of scientific bibliography in the fields of DA, CH, and HS. This attention led to the creation of BiDiAr, conceived both as a tool for deepening the understanding of disciplinary dynamics and as a service-oriented resource for the broader research community. BiDiAr was designed as a database, aimed at systematically collecting, organizing, and thematically indexing bibliographic metadata of sector-relevant publications. The platform’s primary objective is to ensure accessibility, thematic classification, and long-term preservation of references critical to the development of these disciplinary areas. The ambition behind BiDiAr is not to create a bibliometric tool, but rather to develop a platform that facilitates the sharing, use, and reuse of bibliographic resources. By accelerating the production of scholarly articles and supporting the thematic exploration 71 71 of existing literature, BiDiAr aims to foster the discovery of new connections and research directions within the DA and CH communities. Analogous to the iterative development process of DHeLO, the design of BiDiAr evolved during the project, adapting to the needs that emerged both from community feedback and from the evolving specifications of the infrastructure. The first implementation of BiDiAr was based on Zotero, a widely adopted open-source bibliographic management tool selected for its collaborative features, harvesting capabilities, and increasing acceptance within the scholarly publishing ecosystem. The choice aligns with methodologies adopted by other European research infrastructures, such as DARIAH, and reflects practices increasingly common in projects like EAGLE or PATHS. Zotero’s flexibility, together with its capacity to generate unique URLs for each record, has been particularly advantageous for bibliographic management within a LOD approach. Currently, BiDiAr aggregates bibliographic records across four principal categories: 1) articles published in Archeologia e Calcolatori between 1990 and 2024, including both regular issues and Supplements (over 1,300 records); 2) all bibliographic references cited within the articles published in A&C volumes from 2019 to 2024, totalling approximately 10,000 entries; 3) references associated with datasets indexed in DHeLO, incorporated as external links; and 4) references related to the thematic collections included in the Open Digital Archaeology Hub (see further). The close integration between BiDiAr and DHeLO is prompting a significant evolution of the system. A migration process is currently underway to incorporate BiDiAr into DHeLO’s Omeka S environment, with metadata structured according to the Bibliographic Ontology (BIBO). This transition, although technically demanding, aims to achieve a unified management of research data within the infrastructure and to align the bibliographic repository with semantic web standards. The adoption of a knowledge graph model will allow not only for enhanced retrieval and exploration of bibliographic resources, but also for the dynamic integration of metadata across research outputs, projects, and datasets. Through this transformation, BiDiAr is positioned to evolve from a traditional bibliographic archive into an active component of an interconnected, semantically enriched research ecosystem. This approach, cantered on the semantic enrichment and interconnection of research outputs, also underpins the design philosophy of the ArchaeoHub (Mancuso 2025; Caravale et al. 2025), which aims to further consolidate and expand the networked knowledge environment with a focus on DA. 72 72 8. H2IOSC Observatory 8.1 Concept and general aims The H2IOSC Permanent Observatory is one of the flagship components of H2IOSC, designed to monitor, analyse, and valorise the ecosystem of research infrastructures, services, and user practices that converge within the H2IOSC federation. Conceived as a digital and methodological instrument, the Observatory complements the H2IOSC Marketplace, acting as an evidence-based interface for understanding how data, tools, and communities interact within an open science environment (Sichera et al., 2025). It is both a scientific output and an operational tool that informs the design of policies, infrastructures, and services aligned with the principles of openness, interoperability, and sustainability. The Observatory emerges from the integrated mixed-methods approach developed in Work Package 2, which combines quantitative surveys, one-to-one interviews, focus groups, and semi-automatic data collection protocols. The results and insights produced through these instruments are made publicly accessible through a dynamic and continuously updated web application, providing a structured and transparent environment for the dissemination of evidence-based analyses. In this sense, the Observatory operates as a permanent and reflexive mechanism that supports the strategic evolution of the H2IOSC infrastructure over time. 8.2 Structure, scope, and positioning Like the European Open Science Cloud (EOSC) Observatory, but with a more targeted and granular focus, the H2IOSC Observatory serves as an intelligence platform for the continuous monitoring of research practices, resource usage, and community needs within the H2IOSC ecosystem. It provides an integrated analytical perspective on how data resources, digital instruments, and users interact in practice, offering a multidimensional view that connects the technological, organisational, and cultural dimensions of research infrastructures. The Observatory includes a public interactive dashboard that allows users to explore visualisations and metrics concerning research data, tools, and community activities. These insights directly inform the development and uptake of H2IOSC nodes and services, helping identify emerging gaps, prioritise the FAIRification of resources, and design strategies for integrating new datasets and tools into the Marketplace. Unlike traditional landscape analyses or static observatories—where methodologies and data sources often remain opaque—the H2IOSC Observatory is grounded in the replicable and validated instruments tested in WP2, ensuring transparency, methodological rigour, and accountability to the communities that contribute data. 73 73 Figure 31: The H2IOSC Observatory home page 8.3 Methodological approach and system design The design of the Observatory follows a user-centered and evidence-driven approach, grounded in principles of accessibility, clarity, and methodological robustness. Its development involves close collaboration between research teams, data engineers, and UX/UI specialists, ensuring coherence between scientific objectives and technological implementation. The methodological backbone of the Observatory combines statistical analysis with qualitative interpretation, enabling the representation of complex, multidimensional realities in a comprehensible and visually coherent form. From a technical standpoint, the Observatory is implemented as a modular web application developed in synergy with the H2IOSC Marketplace. The two platforms are interconnected, establishing a bidirectional flow of information: while the Marketplace aggregates and exposes digital resources, the Observatory analyses their use and evolution, returning insights that guide further development. This dynamic relationship ensures that data collected from user interactions are continuously integrated into the federation’s broader strategic framework. The UX/UI design ensures usability across a diverse audience of researchers, infrastructure managers, and policy makers. The interface adopts a minimalist and institutional aesthetic, aligned with H2IOSC’s visual identity and compliant with the Web Content Accessibility Guidelines (WCAG). A modular architecture allows for independent management of each 80 80 participation and evidence-based analysis at the core of infrastructure development, H2IOSC is well positioned to act as a living cluster of infrastructures aligned with the evolving European Open Science ecosystem. 81 81 References All references were formatted using the Chicago 17th citation template. Marta Caradonna, Nicola Giampietro, Roberta Bianca Luzietti, Monica Monachini, Valeria Quochi, Emiliano Degl’Innocenti, Alessia Spadi, Alessandra Caravale, Antonio D’Eredità, Paola Moscati, Giacomo Mancuso, 2025, Il ruolo delle Infrastrutture nella costruzione di un ambiente di ricerca inclusivo. Un modello di buone pratiche, AIUCD 2025 Caravale, Alessandra, Antonio D’Eredità, Giacomo Mancuso, and Paola Moscati. 2025. "An open system for textual, visual, and bibliographic resources: the Open Digital Archaeology Hub." Archeologia e Calcolatori 36 (1): 455-468. https://doi.org/https://doi.org/10.19282/ac.36.1.2025.25. https://www.archcalc.cnr.it/journal/articles/1399. Caravale, Alessandra, Nicolau Duran-Silva, Berta Grimau, Paola Moscati, and Bernardo Rondelli. 2023. "Developing a digital archaeology classification system using Natural Language Processing and Machine Learning techniques." Archeologia e Calcolatori 34 (2): 9-32. https://doi.org/https://doi.org/10.19282/ac.34.2.2023.01. http://www.archcalc.cnr.it/journal/id.php?id=1259. Corbetta, Piergiorgio. 2014. Metodologia e tecniche della ricerca sociale.Strumenti. Bologna: Il Mulino. Luzietti, Roberta Bianca, Alessia Spadi, Nicola Giampietro, Giacomo Mancuso, Alessandra Caravale, Antonio D’Eredità, Marta Caradonna, Paola Moscati, Valeria Quochi, Monica Monachini, and Emiliano Degl’Innocenti. 2024. "Digital Humanities and Heritage Science: moving from landscaping to a dynamic research observatory in an Open Science Cloud." MeTe digitali. Mediterraneo in rete tra testi e contesti, Catania. Luzietti R. B., Spadi A., Giampietro N., Mancuso G., Caravale A., D'Eredità A., Caradonna M., Moscati P., Quochi V., Monachini M. e Degl'Innocenti E. (2025) “Digital Humanities and Heritage Science: moving from landscaping to a dynamic research observatory in an Open Science Cloud”, UMANISTICA DIGITALE, ISSN 2532-8816, vol. 2025 (20), pagg. 419-439. Mancuso, Giacomo. 2025. "Beyond monitoring. Reimagining DHeLO as a Linked Open Data infrastructure for Cultural Heritage research." Archeologia e Calcolatori 36 (1): 443454. https://doi.org/https://doi.org/10.19282/ac.36.1.2025.24. https://www.archcalc.cnr.it/journal/articles/1398. Mancuso, Giacomo, Alessandra Caravale, Antonio D’Eredità, and Paola Moscati. 2025. "Merging Knowledge and Tools in Heritage Science and Digital Archaeology. Practices from H2IOSC WP2." DIGITAL HERITAGE 2025. The premier global forum where culture meets cutting-edge technology, Siena, 2025. 82 82 Mancuso, Giacomo, and Antonio D’Eredità. 2024. "DHeLO and BiDiAr: new digital resources within the H2IOSC Project." Archeologia e Calcolatori 35 (1): 521-542. https://doi.org/https://doi,org/10.19282/ac.35.1.2024.31. http://www.archcalc.cnr.it/journal/id.php?id=1312. ---. in press. "Two FOSS resources for Cultural Heritage and Digital Archaeology in the H2IOSC project." ARCHEO.FOSS XVIII | 2024. International conference on Open software, hardware, processes, data and formats in archaeological research, Chieti, 2024. Melaccio D., Boschetti F., Monachini M. (2025): Interfacing CLARIN with H2IOSC: Metadata Interoperability through Ontology-based Mediation, in CLARIN Annual Conference Proceedings, 2025, ISSN 2773-2177, Grisot & Kontino eds., Vienna. Monachini, Monica, and Francesca Frontini. ‘Infrastrutture Digitali per Le Scienze Umane e Sociali’. In Digital Humanities. Metodi, Strumenti, Saperi, 197–213. Studi Superiori. Carocci Editore, 2023. Moscati, Paola. 1999. ""Archeologia e Calcolatori": dieci anni di contributi all'informatica archeologica." Archeologia e Calcolatori 10: 343-352. http://www.archcalc.cnr.it/journal/id.php?id=284. Pietro Sichera, Monica Monachini, Valeria Quochi, Nicola Giampietro, Vittoria Fabiani, Roberta Ottaviani and Roberta Bianca Luzietti (2025), Synergies between CLARIN-IT and OPERAS-IT within H2IOSC: Monitoring Communities and Orchestrating Digital Services, in CLARIN Annual Conference Proceedings, 2025, ISSN 2773-2177, Grisot & Kontino eds., Vienna. pagg. 26-30 83 83 Annex 1: Questionnaire structure and layout Figure 35: introduction and presentation of structure and purpose of the questionnaire together with privacy and policy statement 84 84 Figure 36: information on how the respondents received the notification to answer the questionnaire 85 85 Figure 37: personal information 86 86 Figure 38: information on the disciplinary sector of respondents (ERC) 87 87 Figure 39: information on experience in using data resources, tools and services 88 88 Figure 40: information on experience in creating data resources, tools and services 89 89 OR 96 96 Figure 48: introductory information on the structure, purpose and time required to answer the questionnaire Figure 49: e-mail request to match responses in previous questionnaire 97 97 Figure 50: information regarding data resources created 98 98 Figure 51: information regarding data resources used 99 99 Figure 52: information regarding tools and software created 100 100 Figure 53: information regarding tools and software used 101 101 Figure 54: final comments section 102 102 Annex 2: Focus groups preparation materials Form with selection criteria for focus group participants. 103 103 Focus group invitation e-mail 104 104 Form for recording the demographic and professional characteristics of guests 105 105 1 1 Annex 3: DHeLO Metadata A3.1 DHeLO 1.0 Metadata List Product table Field Explanation id Unique identifier of the product title Title of the digital product or resource description Short description of the content or purpose category General category (e.g., dataset, publication, software) subjectField Primary disciplinary field dataProvider Entity or project that provided the original data intermediateDataProvider Entity that transformed or curated the data before final publication isShownAt URL of a landing page or metadata record isShownBy Direct link to the actual digital resource (e.g., image, dataset) project Reference to the associated project doi Digital Object Identifier, if available resourceTypeTools Type of tool or format used to create or deliver the resource temporalCoverages Time periods covered by the content spatialCoverages Geographic coverage of the content ercSector Corresponding ERC scientific sector contributors People or entities who contributed to the product creators Main authors or creators of the product Project table id Unique project identifier title Project title alternateTitle Alternative or short title founder Funding body or founding institution description Project description or abstract temporalCoverageStart Start date of the project temporalCoverageEnd End date of the project website URL of the project website coordinators Main coordinators or principal investigators participants Additional team members or collaborating entities ercSector Related ERC scientific sector People table id Unique identifier for the person surname Last name name First name email Contact email address institution Affiliated institution or organization nationality Nationality of the person position Academic or professional role (e.g., researcher, curator) 2 2 orcidId ORCID identifier for the person Institution table id Unique identifier for the institution name Full institutional name department Department or subdivision, if applicable website Official website description Short institutional description Tools table id Unique tool identifier name Tool name alternateName Other names or acronyms type Tool type (e.g., software, service, library) isPartOf Larger software suite or environment the tool belongs to license Usage license (e.g., GPL, CC BY) softwarehouse Developer or institution responsible for the tool isShownAt Landing page for the tool developers People or institutions involved in development A3.2 DHeLO 2.0 Metadata List FOAF:ORGANIZATION DCTERMS:TITLE The official name of the organization, ensuring standardized identification. DCTERMS:ALTERNATIVETITLE Any alternative names or abbreviations commonly used for the institution. DCTERMS:DESCRIPTION A brief overview of the organization, including research focus and contributions to CH. DCTERMS:ISPARTOF Indicates if the organization is part of a larger entity (e.g., a university or research center). DCTERMS:IDENTIFIER Unique identifier for the organization (preferably from ROR) for interoperability. FOAF:HOMEPAGE URL linking to the organization’s official website. FOAF:PERSON DCTERMS:TITLE Full name of the person (given name + surname). FOAF:NAME Given name of the person. FOAF:SURNAME Surname of the person. DCTERMS:IDENTIFIER ORCID identifier linking the person to global research infrastructure. FOAF:PROJECT DCTERMS:TITLE Official title of the project. DCTERMS:ALTERNATIVETITLE Alternative names, acronyms, or abbreviations. DCTERMS:DESCRIPTION Overview of objectives, methods, and expected impact. DCTERMS:DATE Timeframe of the project (start and end dates). BIBO:STATUS Current status (e.g., ongoing, completed). 3 3 DCTERMS:REPLACES Links to a previous project that this one continues or replaces. DCTERMS:ISREPLACEDBY Refers to a subsequent project that continues or extends this one. DCTERMS:ISPARTOF Indicates whether the project is part of a broader framework, such as a consortium or research program. FOAF:HOMEPAGE URL of the project’s official website. DCTERMS:IDENTIFIER DCTERMS: DATASETer for the project (can be internal or from external repositories). DCTERMS:CREATOR Coordinator(s) — individual(s) or institution(s) responsible for the project. DCTERMS:CONTRIBUTOR Participant(s) — contributors to the project beyond the main coordinator(s). DCTERMS:INTERACTIVE RESOURCE DCTERMS:TITLE Title of the interactive resource. DCTERMS:ALTERNATIVE Alternate name used for the resource, if applicable. DCTERMS:DESCRIPTION Description of the interactive resource (e.g., its purpose or audience). DCTERMS:TYPE Type of the interactive content (e.g., exhibition, game, visualization). DCTERMS:CREATOR Main person(s) or organization(s) responsible for creating the resource. DCTERMS:CONTRIBUTOR Additional contributors who helped in the creation. DCTERMS:PUBLISHER Entity that published or hosts the interactive resource. DCTERMS:DATE Date of creation or publication. DCTERMS:LANGUAGE Language(s) used in the interface or content. DCTERMS:SUBJECT Thematic subject(s) of the resource. DCTERMS:SPATIAL Spatial coverage or geographic focus of the content. DCTERMS:TEMPORAL Temporal coverage (historical period represented or addressed). DCTERMS:FORMAT File format or delivery format (e.g., HTML5, WebGL). DCTERMS:LICENSE Usage license for the resource (e.g., CC BY-SA, proprietary). DCTERMS:RIGHTSHOLDER Person or entity holding intellectual property rights. FOAF:HOMEPAGE Main access point or URL to the resource. DCTERMS:ISPARTOF Indicates whether the resource belongs to a larger project or collection. DCTERMS:SOFTWARE DCTERMS:TITLE Official name of the software tool. DCTERMS:ALTERNATIVE Alternative name or abbreviation used for the software. DCTERMS:DESCRIPTION Description of the software's functionality and scope. DCTERMS:TYPE Type of tool (e.g., library, application, plugin). DCTERMS:CREATOR Main developer(s) or institution(s) who created the software. DCTERMS:CONTRIBUTOR Individuals or organizations who contributed to the development. DCTERMS:PUBLISHER Organization or group that maintains or distributes the software. DCTERMS:DATE Release or publication date of the software. DCTERMS:LANGUAGE Programming language(s) or interface language(s). DCTERMS:SUBJECT Domain of application (e.g., epigraphy, GIS, 3D modeling). DCTERMS:FORMAT Distribution format or package type (e.g., ZIP, GitHub repo, executable). DCTERMS:LICENSE Usage license (e.g., MIT, GPL, proprietary). 4 4 DCTERMS:RIGHTSHOLDER Legal holder of the intellectual property rights. FOAF:HOMEPAGE Official website or main access point. DCTERMS:ISPARTOF If part of a larger software suite or infrastructure project. DCTERMS:DATASET DCTERMS:TITLE Title of the dataset. DCTERMS:DESCRIPTION A concise description of the dataset’s content and purpose. DCTERMS:TYPE Type or category of dataset (e.g., point cloud, tabular, image set). DCTERMS:CREATOR Main author(s) or institution(s) that created the dataset. DCTERMS:CONTRIBUTOR Additional contributors involved in the dataset’s creation. DCTERMS:PUBLISHER Entity responsible for publishing or hosting the dataset. DCTERMS:DATE Date of publication or creation of the dataset. DCTERMS:SUBJECT Thematic subjects covered by the dataset. DCTERMS:SPATIAL Geographic area(s) represented in the dataset. DCTERMS:TEMPORAL Time period(s) covered by the dataset. DCTERMS:FORMAT File format of the dataset (e.g., CSV, XML, GeoTIFF). DCTERMS:LANGUAGE Language(s) used in the content of the dataset. DCTERMS:LICENSE Usage license under which the dataset is distributed. DCTERMS:RIGHTS Legal holder of the intellectual property rights. DCTERMS:ISPARTOF If applicable, project or collection the dataset belongs to. DCTERMS:IDENTIFIER Persistent identifier (e.g., DOI, internal ID). DCTERMS:SOURCE Reference to the original source from which the dataset was derived. DCTERMS:RELATION Other datasets or resources related to this one. FOAF:HOMEPAGE Main URL or access point for the dataset.