Full text
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base European Cloud for Heritage Open Science Deliverable D6.1 - Data Strategy for the CH Knowledge Base HORIZON-CL2-2023-HERITAGE-ECCCH-01 Horizon Innovation action Responsible authors: Cyprus Institute (CYI) & Poznańskie Centrum Superkomputerowo-Sieciowe (PCSS) © 2025 by the authors, the ECHOES consortium. This work is licensed under a “CC BY 4.0” license. Project start 1 June 2024 Project duration 60 months Document Identifier 10.5281/zenodo.17751757
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 2 Deliverable Information Project Number 101157364 Acronym ECHOES Full title European Cloud for Heritage OpEn Science Project URL https://www.echoes-eccch.eu/ EU Project Officer Christian WILK Change log Date Version Author Change 29/04/2025 V1 All authors All chapters drafted 10/06/2025 V2 All authors All chapters extended 14/07/2025 V3 Andriana Sielli, Avgoustinos Avgousti (CYI), Joanna Kowalska (PCSS) Refactoring and new assignments added 16/08/2025 V4 All authors Required sections updated 15/09/2025 V5 Andriana Sielli, Avgoustinos Avgousti (CYI), Joanna Kowalska (PCSS) Second round of review update completed and new assignments added 17/10/2025 V6 All authors Required sections updated 31/10/2025 V7 Andriana Sielli, Avgoustinos Avgousti (CYI), Joanna Kowalska (PCSS) Third round of review update completed 31/10/2025 V7.1 Dimitris Kotzinos (CNRS-CYU) Full document review and comments Deliverable D6.1 Title Data Strategy for the CH Knowledge Base Work Package WP6 Title Setting the Cloud Environment Submission date 10/12/2025 Pages 71 Keywords Data Strategy, Knowledge Base, Metadata, Semantic Data, Heritage Digital Twin, Ontology, Federated Architecture
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 3 14/11/2025 V8 Andriana Sielli, Avgoustinos Avgousti (CYI) Joanna Kowalska (PCSS) Refactoring 31/11/2025 V9 Joanna Kowalska (PCSS) Final updates 31/11/2025 V10 Xavier Rodier (CNRS) Final review Lead Partner CYI, PCSS, FORTH Responsibles Authors CYI, PCSS Authors (Partner) Andriana Sielli (CYI), Avgoustinos Avgousti (CYI), Sorin Hermon (CYI), Paolo Cignoni (CNR), Emanuel Demetrescu (CNR), Joanna Kowalska (PCSS), Konrad Pawlikowski (PCSS), Anastasia Axaridou (FORTH), Carlos Andújar (UPC), Matej Ďurčo (OEAW), Evangelos Kritsotakis (FORTH), Maria Theodoridou (FORTH), Elias Tzortzakakis (FORTH), Marios Pitikakis (FORTH) Contributors Dimitrios Kotzinos (CNRS-CYU), Antoine Isaac (EUROPEANA), Sarah Middle (York University)
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 4 Abstract This document, Deliverable D6.1 "Data Strategy for the CH Knowledge Base," establishes the foundational data and metadata management framework for the ECHOES project (European Cloud for Heritage Open Science). It outlines a comprehensive strategy to build a federated, semantically interoperable infrastructure that will support the creation, preservation, and reuse of Heritage Digital Twins (HDTs)—dynamic digital replicas of cultural heritage assets enriched with contextual knowledge. The strategy is designed to overcome critical challenges in the cultural heritage sector, including data fragmentation across institutions, semantic gaps, and scalability issues. It is grounded in core principles of semantic interoperability, a federated architecture that respects data ownership and provenance, and strict adherence to FAIR (Findable, Accessible, Interoperable, and Reusable) data practices. This approach ensures that cultural heritage institutions maintain control over their data while actively contributing to a shared, collaborative ecosystem. Key components of the strategy include: • Technical Architecture: A detailed design for the Knowledge Base (KB), comprising distributed and federated triplestores (with both public and private content), an administrative repository, REST APIs, and user interfaces to facilitate data ingestion, management, and querying. • Governance & Lifecycle: A robust governance model defining roles, responsibilities, and policies for secure, ethical, and compliant data sharing, covering the entire data lifecycle from raw acquisition to interpreted knowledge. • Strategic Alignment: Direct support for ECHOES's core objectives, particularly enabling a sustainable legal entity (SO2), integrating diverse outcomes, from past, present and future projects (SO4), and fostering collaborative knowledge co-creation around Digital Commons (SO5). This deliverable defines the functional and non-functional requirements for the system, illustrated through user personas and operational use cases. It also establishes mechanisms for monitoring, evaluation, and quality assurance to ensure the infrastructure meets its goals. Crucially, this Data Strategy lays the essential groundwork for the subsequent deliverable "D6.2 Interoperability Requirements and Guidelines," which will expand upon this design to support advanced features such as distributed computing, automated annotation, and integration with Cascading Grants and Vertical Applications. This document is a critical step in realizing the vision of a unified, open, and sustainable European Collaborative Cloud for Cultural Heritage (ECCCH). Disclaimer: Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 5 Table of Contents Table of Contents ...................................................................................................................................... 5 List of abbreviations ................................................................................................................................. 7 Definitions .................................................................................................................................................. 8 1.Introduction .......................................................................................................................................... 11 1.1 Purpose and Scope of the Document .................................................................................................. 11 1.2 Knowledge Base in the Context of the CH Cloud ................................................................................ 11 1.3 Alignment with ECHOES Goals .............................................................................................................. 12 Strategic Objective 1: Open-Source Infrastructure for Community-Driven Applications ..................... 13 Strategic Objective 2: Holistic Digital Transformation Approach ............................................................ 13 Strategic Objective 3: Community Integration and Cohesion ................................................................. 13 Strategic Objective 4: Integration of ECCCH-Related Projects ................................................................. 14 Strategic Objective 5: Enabling Collaborative Co-Creation around Digital Commons .......................... 14 Strategic Objective 6: Collaborative Co-Creation around Digital Commons .......................................... 14 2. State of the Art .................................................................................................................................... 15 2.1 Current Practices in Cultural Heritage Data Strategy ......................................................................... 15 2.2 Gaps and Challenges ............................................................................................................................. 15 2.3 Heritage Digital Twins: A New Paradigm ............................................................................................. 17 2.4 Metadata Transformation and Ontologies .......................................................................................... 17 2.5 Federated Architectures and FAIR Data ............................................................................................... 18 2.6 Positioning of the ECHOES Knowledge Base ....................................................................................... 19 3. Understanding the Data Ecosystem ................................................................................................. 20 3.1 Data Processing Stages and Levels of Interpretation ........................................................................ 20 3.1.1 Example 1: Heritage science-based art historical investigation of the Derynia icon ................... 23 3.1.2 Example 2: Ayios Ioannis Lampadistis Monastery, Cyprus ............................................................. 25 3.1.3 Example 3: Andrea Pisano's pulpit in the Church of Sant’Andrea in Pistoia (Italy) ....................... 28 3.2 Data Sources ........................................................................................................................................... 33 3.2.1 Category 1: Structured databases ..................................................................................................... 33
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 6 3.2.2 Category 2: Semi-structured databases ............................................................................................ 34 3.2.3 Category 3: No/minimal-structured databases ................................................................................ 35 3.3 Ontology Models .................................................................................................................................... 35 4. Data Management Requirements ..................................................................................................... 36 4.1 Functional Requirements ...................................................................................................................... 36 4.1.1 End User Profiles ................................................................................................................................. 37 4.1.2 Data Governance Roles ...................................................................................................................... 41 4.1.3 User Scenarios ..................................................................................................................................... 42 4.2 Non-Functional Requirements .............................................................................................................. 44 4.2.1 FAIR Principles in ECHOES .................................................................................................................. 45 4.2.2 System Quality Characteristics in ECHOES ....................................................................................... 46 5. Data Architecture Design ................................................................................................................... 48 5.1 Architecture Overview ........................................................................................................................... 48 5.2 Federated Architecture and Data Lifecycle Management .................................................................. 49 5.2.1. Federated Knowledge Strategy ........................................................................................................ 49 5.2.2 User Personal Space ........................................................................................................................... 56 5.2.3 Semantic Metadata Persistence Strategy ......................................................................................... 56 5.2.4 Potential Issues and Solutions ........................................................................................................... 57 5.3 Data Storage Solutions .......................................................................................................................... 58 5.4 Replication / Backup Strategy ............................................................................................................... 59 5.5 Core Platform Services and Administrative Subsystem ..................................................................... 60 5.5.1 Access to ECHOES - SSO ...................................................................................................................... 61 5.5.2 High-Level SSO process ...................................................................................................................... 62 5.5.3 Administrative Repository .................................................................................................................. 63 5.5.4 Back-End - Web Services ..................................................................................................................... 64 5.5.5 Management Console ......................................................................................................................... 65 Conclusion ................................................................................................................................................ 66 References ............................................................................................................................................... 67
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 7 List of abbreviations AAI Authentication and Authorization Infrastructure API Application Programming Interface CH Cultural Heritage CPU Central Processing Unit CRM Conceptual Reference Model CRUD Create, Read, Update, Delete DB Database DMP Data Management Plan DT Digital Twin ECCCH European Collaborative Cloud for Cultural Heritage EOSC European Open Science Cloud ETL Extract, Transform, Load FAIR Findable, Accessible, Interoperable, and Reusable GLAM Galleries, Libraries, Archives, and Museums GPU Graphics Processing Unit HDT Heritage Digital Twin HDTO Heritage Digital Twin Ontology KB Knowledge Base LOD Linked Open Data OKD Origin Kubernetes Distribution PCA Principal Component Analysis PID Persistent Identifier RDF Resource Description Framework RDM Research Data Management RML RDF Mapping Language SSO Single Sign-On UI User Interface VA Vertical Application
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 8 Definitions Heritage Digital Twin (HDT) The Heritage Digital Twin (HDT) is the state-of-the-art digital information about a real-world heritage asset, encompassing its tangible and intangible components. It describes the asset’s properties and captures its space-time-culture identity, all formally organized within a semantic framework described by the CIDOC Conceptual Reference Model (CRM)-based ontology (HDTO). The HDT enables access to data, their interpretation, and the workflows that generated them. Additionally, associated Knowledge Graphs represent inferences and reasoning processes related to statements about heritage assets. Digital Commons Digital Commons are holistic social institutions that govern the (re)production, sharing, and stewardship of digital resources—such as data, information, culture, and knowledge—through collaborative and interdisciplinary processes. These resources are created and/or maintained online and shaped by interrelated legal, socio-cultural, economic, and institutional dimensions. Digital Commons function within a sustainability-oriented framework, supported by interconnected components: • Digital Twins (for acquisition, analysis, conservation, dissemination), • Innovation (e.g. data handling, integration, thematic groups), • Digital Ecosystem (platforms, marketplaces, socio-technical systems), • Communities (end-user interfaces, training, participation), • Digital Continuum (workflows, assessments), and • Knowledge (semantic value, AI, pilots and vertical apps). Digital Continuum The ability to maintain digital information in such a way that it will continue to be available, as needed, despite changes in digital technologies. This means ensuring that digital content remains complete, available, and usable over time. Activities involved in supporting the digital continuum include: • Information management and risk assessment • Managing technical environments • File format conversion • Long-term preservation The full glossary including all related ECHOES terms is being developed separately and its current version can be find in this document . For the purpose of Data Strategy the most relevant definitions have been copied and extended here.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 9 Data In ECHOES, data refers to the actual content or information collected, created, or used in the project. This includes images, texts, audio recordings, 3D models, measurements, and other digital representations of cultural heritage objects or phenomena. Metadata Metadata is “data about data.” In ECHOES, it describes and organizes the main data — for example, who created it, when, where, what it represents, and how it can be used. Metadata helps make the data searchable, understandable, and reusable, especially across different systems and institutions. Paradata Paradata is “data about the process of creating, processing, or interpreting data.” In ECHOES, paradata captures the context, methods, and provenance of how cultural heritage data was acquired, processed, refined, or interpreted. EGI Check-in EGI Check-in is a proxy service that operates as a central hub to connect federated Identity Providers (IdPs) with EGI service providers – more details can be found on this website. providers. Heritage Digital Twin Ontology A Heritage Digital Twin ontology is a structured framework used to capture, organize, and interrelate all relevant knowledge about a heritage asset within a HDT. (Niccolucci et al., 2022). Later on, in this document, term E-HDTo will be also used as ECHOES HDTo. FAIR Data The FAIR Data Principles provide a framework for enhancing scientific data management by ensuring data is Findable, Accessible, Interoperable, and Reusable. These principles emphasize machineactionable metadata, enabling both researchers and automated systems to efficiently discover, access, and reuse data. By following FAIR guidelines, researchers improve data sharing, reproducibility, and cross-disciplinary collaboration, increasing the long-term value of their work. Widely adopted by funding agencies, publishers, and research institutions, FAIR principles promote better data stewardship in an increasingly digital research landscape. (Wilkinson et al., 2016)
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 16 they provide frequently lack the essential functionality required for seamless collaboration, ultimately reinforcing existing institutional and geographic boundaries instead of bridging them. Risk to Long-Term Preservation: Many legacy systems lack essential capabilities for digital stewardship, such as robust provenance tracking, version control, and support for semantic updates. This critical gap significantly increases the risk of data loss, format obsolescence, and the decontextualization of cultural heritage assets over time. The challenge is further exacerbated because preservation systems and collection management platforms often operate in separate silos. Built upon different technical standards and underlying data models, these disconnected systems disrupt metadata continuity and actively impede unified, comprehensive stewardship of the digital cultural record. Semantic Gaps Undermine Access and Context: Traditional systems lack the mechanisms to dynamically link entities—such as people, events, places, and artifacts—across time, space, and collections. Without richer semantic frameworks like knowledge graphs and ontologies, cultural heritage data remains fundamentally fragmented. This severely limits its contextual richness and ultimate research value. The problem is exacerbated when preservation systems and collection management platforms operate on distinct metadata models, further isolating critical contextual knowledge. Bridging this semantic gap requires a systemic commitment to semantic alignment across the entire data stewardship pipeline, from initial data creation to long-term preservation. Lack of True Federated Solutions: The challenges of collaboration and scalability are further constrained by the absence of robust, genuinely federated architectures. These would allow distinct data providers to participate in a shared environment while maintaining control over their data. A key limitation of current approaches is the inability to query a distributed knowledge repository as a single, unified system capable of supporting integrated analysis across all connected datasets. Existing infrastructures often cannot be queried in an integrated manner, while other solutions merely implement parallel named graphs within a centralized knowledge base; neither approach fulfils the requirements of true interoperability and federation. Consequently, a paramount challenge for crossinstitutional research is to enable both seamless data integration and discoverability, as well as the capacity to perform comprehensive semantic queries across isolated sources with highly diverse content. Lack of Process-Aware Data Documentation: Current cultural heritage platforms excel at storing structured data but remain rooted in traditional paradigms, acting as repositories for static information - mere "still pictures" of data - while failing to capture the dynamic, iterative, and context-rich processes that generated it. Consequently, the scientific workflows, interpretive decisions, and methodological sequences behind data creation are often lost. This gap critically hinders reproducibility, interpretability, and collaborative knowledge building. Without access to the rationale, tools, and decision points behind data generation, it becomes difficult to assess quality, integrate results, or reuse information. A fundamental shift is needed: from storing only data outputs to formally capturing the entire knowledge creation workflow. This requires
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 17 recording provenance, methodological pipelines, and decision-making logic in structured, semantic formats that allow for dynamic re-contextualization. Existing initiatives like the Italian DIGILAB platform (E-RIHS.eu) explicitly address this by adopting a three-pillar model (Data, Process, and Knowledge Management). It extends beyond a semantic repository to include a workflow editor and engine that formally describe and execute scientific processes. This enables both the reproducibility of known procedures and the dynamic composition of new workflows, thereby bridging the critical divide between static data storage and active knowledge generation. ECHOES will function as a unifying system, designed to aggregate and streamline similar efforts. 2.3 Heritage Digital Twins: A New Paradigm Heritage Digital Twins (HDTs) represent a transformative paradigm for the documentation, management, and interpretation of cultural heritage assets. HDT serves as the nexus of digital representations and their associated interpretations, enabling dynamic, evolving, and multiperspective engagement with the cultural artifact over time. Within the ECHOES platform, HDTs are realized as interoperable entities grounded in semantic ontologies, notably HDTO. They are hosted within the KB and linked to a broader digital continuum, a design that fundamentally promotes transparency, reusability, and cultural equity. 2.4 Metadata Transformation and Ontologies Supporting a semantic infrastructure for cultural heritage data necessitates the transformation of legacy, heterogeneous datasets into structured, interoperable formats. This is increasingly achieved by adopting formal ontologies like CIDOC CRM, which provide a robust conceptual framework for consistently representing complex relationships between events, actors, and objects. These semantic models are essential for reconciling disparate data sources into unified, machine-readable knowledge structures. Within the ECHOES project, this foundation is extended through the ECCCH Heritage Digital Twin Ontology (HDTO), a CIDOC CRM-compatible extension specifically designed to serve as the core semantic layer of the European Collaborative Cloud for Cultural Heritage (ECCCH). E-HDTo builds upon and refines the original Heritage Digital Twin ontology [NMTF23] by introducing dedicated classes and properties—such as HC1 Heritage Entity, HC2 Heritage Digital Twin, HC9 Study, HC10 Heritage Valuation, and HC12 Heritage Declaration Event—to explicitly model the dynamic, socially conferred nature of heritage value, the provenance of digital knowledge production, and the evolving relationship between real-world assets and their digital counterparts. This enables the systematic aggregation of multimodal data (3D models, scientific analyses, restoration records, narratives) into versioned, semantically coherent Digital Commons that capture the full space-time-culture identity of heritage assets.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 18 The CIDOC CRM Special Interest Group has defined SYNERGY, a reference model designed to standardize data provisioning and aggregation processes within the cultural heritage sector. SYNERGY defines a consistent set of business processes, user roles, generic software components, and open interfaces. Its core rationale addresses a fundamental sector-wide need: while institutions document collections in vastly different ways due to varying disciplines, objectives, and languages, handling this metadata as a unified whole is vital for advancing humanities research, enabling knowledgeable information retrieval, and facilitating seamless data exchange. E-HDTo fully aligns with SYNERGY principles by providing a standardized ontological target for metadata transformation pipelines, ensuring that institution-specific schemas can be mapped, enriched, and integrated into a shared, extensible knowledge graph. This transformation process is inherently complex and resource intensive. It requires powerful tools to align, map, and convert existing schemas and datasets into E-HDTo compliant structures. Successful implementation demands close collaboration among domain experts, ontologists, data engineers, and cultural heritage curators to ensure both semantic fidelity and practical usability. Within ECHOES, this interdisciplinary effort is operationalized through dedicated work packages that develop mapping templates, validation workflows, and training resources - ultimately enabling institutions of all sizes to contribute to and benefit from the ECCCH’s interoperable, collaborative digital ecosystem. 2.5 Federated Architectures and FAIR Data A key limitation of current semantic approaches is the absence of a unified querying layer that can seamlessly operate across multiple institutional endpoints. In practice, each institution exposes its data differently - through distinct SPARQL endpoints, APIs, or partial exports - requiring users to manually locate, access, and harmonize results. This places a heavy burden on researchers and prevents the execution of cross-collection queries that require consistent reasoning across distributed datasets. As a result, federated knowledge networks often behave as a set of parallel repositories rather than a truly integrated ecosystem. Addressing this gap is essential for ECHOES, where unified discovery, cross-institutional analytics, and collaborative enrichment depend on reliable, coherent querying mechanisms across the entire federation The FAIR principles form a cornerstone of modern CH data management and are a core requirement within the ECHOES framework. Their systematic implementation is essential for maximizing the visibility, usability, and long-term value of CH datasets. In ECHOES, these principles are operationalized through a federated architecture. Unlike centralized repositories, this model maintains data at its source institution while connecting it via a shared ontology, standardized APIs and interoperability protocols to enable machine-actionable access. This is underpinned by persistent identifiers and robust provenance tracking to guarantee data integrity and citation reliability. This approach offers key advantages: it preserves institutional autonomy, safeguards the provenance and authenticity essential for scholarly trust, and supports responsible governance of sensitive or
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 19 restricted information in alignment with legal, ethical, and intellectual property constraints. Through federated search and semantic integration, the ECHOES Knowledge Base enables controlled discovery of data across domains and jurisdictions. This balanced model fosters cross-institutional collaboration, broadens data reuse, and facilitates the participation of both large infrastructures and smaller, resource-limited institutions in a common knowledge space. By uniting the distributed resilience of federation and application of the FAIR framework, ECHOES ensures cultural heritage data remains technically sustainable, socially inclusive, and semantically coherent. 2.6 Positioning of the ECHOES Knowledge Base The ECHOES Cultural Heritage Knowledge Base (KB) is positioned to become a central hub within the evolving, interconnected landscape of cultural heritage data. Built on semantic web technologies and rigorously grounded in FAIR principles, it provides a unified, standards-aligned platform for the creation, discovery, and reuse of Heritage Digital Twins (HDTs). The ECHOES KB represents more than a technical system; it embodies a vision for the dynamic preservation and continuous contextual enrichment of cultural heritage. By bridging institutional, disciplinary, and technological divides, it enables a future where Europe’s cultural legacy is not only safeguarded but also actively reinterpreted and reimagined, ensuring its vitality, accessibility, and relevance for generations to come. Architecturally, the KB is designed as a modular, federated system that supports distributed contributions and shared governance while ensuring strict semantic consistency across diverse sources. It builds upon well-established models like the CIDOC CRM and leverages trusted linked data vocabularies to deliver reliable and scalable interoperability across institutions and domains. Open access repositories like Zenodo (more information can be found here) are commonly used for sharing CH research outputs. While Zenodo supports FAIR principles and structured metadata, it is a general-purpose platform that cannot be customized to support specialized CH needs, such as semantic annotations or tailored APIs. Moreover, Zenodo is offering the ability to use datasets through a download only facility, so access (e.g. querying) specific data points is not possible. The CH domain is thus shifting toward more collaborative, open data practices. Initiatives like Europeana have been instrumental in building frameworks for sharing high-quality, reusable data, with principles integrated into models like the Europeana Data Model (EDM). ECHOES KB will be a starting point to search for information on Cultural Heritage and work as integration layer between different data repositories. Federated strategies are gaining traction as they allow institutions to maintain control over their data and infrastructure while participating in a shared ecosystem built on common standards. This approach supports incremental ECHOES partnership by moving through different levels of integration, scalability, and local autonomy without requiring centralization.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 20 Semantic federation is typically implemented in one of two ways: as distinct institutional knowledge bases or as a collection of semantic named graphs within a central repository. However, the first approach often lacks unified querying mechanisms, while the second is not a true federation due to its centralized nature. Notable progress has been made within the European Open Science Cloud (EOSC) framework through projects such as: • FAIRsFAIR, which developed practical tools and policies for implementing FAIR principles. • DARIAH's Generic Search Service, which harvests metadata from distributed collections to create a unified, faceted search index. • SSH Open Marketplace, which harmonizes tools, services, and workflows for the research data lifecycle. Building on these foundations, the ECHOES project aims to establish a federated knowledge base aligned with EOSC and ECCCH principles. It will provide a FAIR-compliant KB infrastructure and metadata associated with it with rich APIs, tailored to the CH community. By leveraging existing expertise and advancing federation-based interoperability, ECHOES will create a distributed network of institutional knowledge bases connected through common standards and protocols, enhancing access, reuse, and long-term sustainability of CH data across Europe. 3. Understanding the Data Ecosystem 3.1 Data Processing Stages and Levels of Interpretation This section outlines the data lifecycle within the ECHOES project, a framework designed to address the challenges faced by the cultural heritage sector—such as fragmentation and heterogeneity—by accommodating data at different stages and levels of refinement. From primary observations to semantically enriched knowledge representations, this structured progression ensures inclusivity for institutions with varying levels of technical maturity and enables increasingly sophisticated analysis and integration over time. Within ECHOES, cultural heritage data evolves through a lifecycle characterized by progressive refinement, contextualization, and semantic enrichment (as presented in the Figure 1). Identifying these distinct stages is essential for ensuring interoperability, promoting reuse, and supporting the The document mentions several standards (e.g. RDF, Dublin Core, SKOS, IIIF, OAIPMH), but it does not yet provide: • a consolidated list of data/metadata formats, • a distinction between mandatory and recommended standards, • a mapping to specific contexts (e.g. 3D data, images). D6.1 is more strategic in nature and technical details will be covered by D6.2.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 21 long-term preservation of cultural heritage knowledge. The lifecycle comprises the following key stages: • Raw Data Raw data refers to the unprocessed outputs collected directly from data acquisition instruments. These initial measurements - such as energy spectra or imaging data - are structured but have not yet undergone computational processing or semantic enrichment. o Raw data must be accompanied by both descriptive metadata (e.g., unique identifier, object name, date of creation) and acquisition paradata (e.g., instrument settings, environmental conditions, operator). This is essential for establishing transparency and enabling accurate interpretation. • Processed Data Processed data is derived from raw data through computational or analytical techniques that improve its quality, clarity, and utility for specific research questions (e.g., elemental mapping, spectral peak fitting, Principal Component Analysis). o Processed data must be rigorously linked to processing paradata, which documents the algorithms, software, and parameters used. This ensures full traceability, guarantees reproducibility, and defines the dataset's fitness-for-purpose. • Post-Processed Data Post-processed data results from targeted refinement of processed outputs to optimize accuracy, consistency, and interoperability (e.g., through noise reduction, artifact correction, normalization). o The refinement paradata details the rationale and methods for these corrective actions. This stage yields a validated, structured foundation ready for semantic annotation and integration. • Interpreted Data Interpreted data integrates expert-derived knowledge and domain-specific insights, encoding higher-order relationships—such as temporal sequences or causal links—into structured, computationally tractable forms using semantic ontologies. o This stage produces structured scholarly assertions. The provenance of these interpretations (including the expert, evidence, and rationale) is meticulously documented as paradata to ensure transparency and credibility. It enables advanced reasoning and drives knowledge discovery. These stages align directly with the DIKW (Data–Information–Knowledge–Wisdom) model shown in Figure 1. Raw and processed inputs form the Data layer, providing the basic measurements collected from instruments. After validation and refinement, they become Information, suitable for reliable reuse and comparison. When experts interpret these validated outputs and map them to shared ontologies, they create Knowledge that is structured and semantically meaningful. Finally, connecting these knowledge elements across datasets and institutions produces Wisdom in the form of integrated
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 22 knowledge graphs. This progression ensures that ECHOES can handle heterogeneous inputs while still enabling increasingly rich and interoperable cultural heritage insights. Synthesis: Towards a Semantic Knowledge Graph Metadata will be connected to the HDTO. This consistent foundational layer facilitates the construction of rich, interoperable knowledge graphs that represent complex heritage relationships, thereby fulfilling the core ECHOES objective of transforming isolated data into a connected and federated, meaningful Digital Commons. The ECHOES infrastructure will be specifically designed to support the creation, management, and querying of these evolving knowledge graphs. Figure 1: A DIWK pyramid that illustrates the project's data lifecycle framework
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 23 3.1.1 Example 1: Heritage science-based art historical investigation of the Derynia icon This example illustrates the progression of cultural heritage data through its lifecycle stages by examining a pigment analysis of the Derynia icon, which is a specific religious icon from a church in the village of Derynia, Cyprus. The artwork was investigated using advanced imaging techniques, including Macro X-Ray Fluorescence (MA-XRF), Reflectance Imaging Spectroscopy (RIS), and Luminescence Imaging Spectroscopy (LIS). This case study demonstrates how raw instrumental signals are transformed into semantically rich, interpreted knowledge that supports both scientific and art historical research. Raw Data The raw data consists of unprocessed signals directly acquired from the analytical instruments. This foundational layer includes: • XRF energy spectra • RIS reflectance spectra • LIS luminescence signals Acquisition Metadata & Acquisition Paradata This raw data is accompanied by detailed acquisition metadata (e.g., unique sample ID, artwork name, date of analysis) and acquisition paradata, which documents the essential instrumental settings and context required for reproducibility and interpretation: • XRF settings: voltage (40 kV), current (500 µA), dwell time (80 ms/pixel) • Spectral ranges: RIS (470–820 nm), LIS excitation wavelengths (365 nm and 655 nm) • Scan details: duration (3 h 30 min), area (25 × 17 cm²), resolution (498 × 340 pixels) Processed Data At this stage, raw signals are transformed through computational procedures to improve clarity and utility. This involves: • Converting XRF spectra into elemental distribution maps (e.g., Hg-L for mercury, Pb-L for lead). • Processing RIS and LIS data using spectral angle mapping (SAM) and NFINDER endmember extraction to isolate distinct pigment signatures.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 24 Processed Paradata The derivation of processed data is rigorously documented through processing paradata, which captures the computational environment and choices: • Software tools used: PyMCA, pystools, and SPECTRONON. Post-Processed Data Processed outputs are further refined to optimize accuracy and support interoperability. This stage creates composite visualizations for enhanced analysis: • Generation of RGB correlation maps. • Creation of multi-element fusion images. These products standardize the data, providing a validated foundation for integration and expert interpretation. Interpreted Data Domain experts integrate the analytical results with art-historical and material knowledge to generate higher-order insights. This stage produces structured scholarly assertions that encode meaning: • Identification of specific pigments (e.g., mercury sulfide signifies cinnabar/vermilion; lead compounds indicate lead white). This entire pipeline, from raw signals to interpreted knowledge, demonstrates how the ECHOES lifecycle facilitates the transformation of isolated technical measurements into semantically rich, interoperable, and trustworthy knowledge for the broader cultural heritage community. Figure 2: Derynia Icon
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 25 3.1.2 Example 2: Ayios Ioannis Lampadistis Monastery, Cyprus This example describes the 3D-based data workflow for the digital documentation and preservation of the Ayios Ioannis Lampadistis Monastery, which is a UNESCO World Heritage site in the Troodos Mountains of Cyprus, located in Kalopanayiotis village. Raw Data The initial data acquisition combines multiple sources to capture both geometric detail and historical context (Figure 3 and Figure 4). This includes high-resolution laser scanning point clouds and photogrammetric images documenting the site’s precise geometry and surface appearance. These are complemented by topographic measurements, as well as archival drawings and historical documents provided by the Cyprus Department of Antiquities. Figure 3: A 2D panoramic image (raw data of the scanner) of the Kalopanayiotis monastery Figure 4: An aerial image of Kalopanayiotis monastery taken by the drone.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 32 dissemination activities. Administrative metadata also record the responsible institution, data custodianship, and temporal scope of the acquisition campaign, ensuring traceability and compliance with institutional data policies • Descriptive Metadata: Detail content-based metadata describing what the digital resources represent (subject matter, historical context, relationships to other objects, | Linked Data: Relationships to other objects using ontology links referencing CIDOC CRM classes such as E22 Man-Made Object or E55 Type). Data are documented with descriptive metadata defining their cultural and contextual meaning. Metadata fields about the asset include title, subject, artist/author, date, typology, materials, dimensions, and iconographic themes.Their relationships are expressed using CIDOC-CDM classes and properties. • Structural Metadata: Information about how data components are organized and related to each other (e.g. file hierarchies, data dependencies). A formal structural metadata schema will be created during the pilot to define the organisation and relationships between the project's heterogeneous datasets. • Technical Metadata: Information about the digital files' technical characteristics (file formats, size, creation date, software used). Technical metadata relating to the 3D survey, modelling and monitoring of the Pulpit are also available, although they are not yet systematically organised. They will be organised, together with those that will be generated for all new digital resources created as part of the pilot project. • Paradata: There are no defined standards for collecting paradata relating to the surveying and modelling of 3D data. During data collection and management, paradata was collected in unstructured form (currently being organised according to the CIDOC-CRM schema) which will be used and implemented in the project, including: Acquisition methodologies (photogrammetry, laser scanning) and processing pipeline. Sources: historical inventories, floor plans, descriptive texts, historical photographs. • Data Quality Metadata: Information about accuracy, completeness, consistency of the data. The quality and accuracy of the geometric data will be certified by the metadata contained in the survey reports and 3D modelling. The quality of the information content relating to the works of art described will be certified by its origin from the catalogue records of the Italian Ministry of Culture. • Usage History: Documentation of how the data has been used previously, by whom, and for what purposes. • The data has been delivered to stakeholders and partly used for previous studies and scientific publications.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 33 Figure 10: Orthophoto of the 3D model 3.2 Data Sources The ECHOES Knowledge Base integrates and standardizes cultural heritage data from diverse sources, ensuring compliance with FAIR principles and semantic interoperability. This federated ecosystem draws from institutional archives, public repositories, and collaborative platforms, each contributing unique datasets to enrich the collective knowledge base. 3.2.1 Category 1: Structured databases Institutional Databases • Description: Structured datasets from museums, libraries, archives, and archaeological institutions • Examples: The Louvre Database • Key Features: o Advanced search capabilities o Metadata standards compliance o API support for external integrations A comprehensive list of data sources can be found in the D3.1 Integration Strategy document.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 34 o Robust database management systems Public / Open Datasets • Description: Freely accessible datasets provided by governments, cultural organisations, and museums, often under Creative Commons or open data licenses and very often derived from institutional databases • Examples: The datasets from the common European data space for cultural heritage, made available at Europeana.eu, and national heritage databases. • Key features: o Standardized Licensing: Clear usage rights (e.g., CC0, CC-BY, Open Data Commons) o Bulk Access & APIs: Enables large-scale downloads and programmatic integration o Multilingual Metadata: Often includes cross-lingual (semantic) annotations for broader accessibility o Community-Enhanced: Some platforms allow crowdsourced corrections or contributions o Pre-processed Harmonization: Many datasets already align with common vocabularies (e.g., Schema.org, Europeana Data Model) o Global Reach: Aggregates data from multiple countries/institutions into unified interfaces 3.2.2 Category 2: Semi-structured databases Research Infrastructures • Description: High-value, scholarly datasets produced by EU-funded projects (e.g., Horizon Europe) and academic consortia (e.g., DARIAH, CLARIN), featuring rigorous curation and rich contextual metadata • Examples: o A multilingual epigraphy corpus developed under DARIAH o CLARIN’s linguistic datasets with TEI/XML markup o Archaeological field reports from Horizon Europe projects • Key Features: • Scholarly-Grade Data: Outputs from digitization projects, archaeological surveys, historical corpora, or linguistic documentation • Enhanced Metadata: Includes: o Multilingual annotations o Provenance and attribution tracking o Temporal/spatial references (e.g., period-specific ontologies) o Alignment with CIDOC CRM, SKOS, or domain-specific standards o Project-Backed Curation: Structured for long-term preservation and academic reuse (e.g., via institutional repositories)
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 35 3.2.3 Category 3: No/minimal-structured databases Fieldwork and Crowdsourced Data Description: Data generated through on-site cultural heritage documentation (fieldwork) or contributed by the public via crowdsourcing platforms. These sources capture rich, often underrepresented information such as transcriptions, local knowledge, classifications, or geotagged imagery, enhancing institutional collections. Examples: o Public-contributed artifact transcriptions o Geotagged historical photos from community mapping projects o Ethnographic fieldwork recordings with indigenous knowledge Key Features: o Diverse Data Types: Includes text transcriptions, images, audio recordings, and geospatial data o Community Participation: Leverages public contributions for broader coverage and perspectives o Complementary Value: Fills gaps in institutional records with local/non-institutional knowledge o Metadata Challenges: Often needs standardization due to varied contributor expertise Legacy/Digitized Collections Description: Historical materials (manuscripts, photographs, maps, objects) digitized from physical archives or older digital formats, held by museums, libraries, or archives. Often requires normalization to meet modern standards. Examples: • Scanned excavation notebooks from a national archive • Digitized microfilm images of colonial-era documents • Legacy museum catalogues converted from spreadsheets Key Features: • Format Diversity: Supports multiple file types and obsolete digital formats • Preservation Focus: Digitizes fragile physical materials (e.g., aged paper, photographs) to prevent deterioration • Metadata Variability: Often lacks standardized metadata, requiring manual or automated enrichment Processing Needs: Requires conversion of image-based text into machine-readable formats. 3.3 Ontology Models HDTo (Heritage Digital Twin ontology) is the core semantic framework of the European Collaborative Cloud for Cultural Heritage (ECCCH), a CIDOC CRM-compatible extension that models the dynamic
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 36 relationship between HC1 Heritage Entity (real-world tangible, intangible, or born-digital assets with socially conferred value) and HC2 Heritage Digital Twin (evolving propositional objects aggregating all state-of-the-art digital representations, studies, and provenance metadata). HDTo enables the creation of versioned, collaborative Digital Commons capturing the full space-time-culture identity of heritage assets. ECHOES uses other standardized ontologies and vocabularies to ensure interoperability: • CIDOC CRM: Museum and artifact modeling (ISO 21127:2023). • CIDOC CRM compatible models: CRMinf: the argumentation (inference) model CRMsci: the scientific observation model CRMdig: the model for provenance metadata and digitization processes CRMarchaeo: the excavation model CRMtex: the model for the study of ancient texts CRMact: the model for activity planning and execution • LRMoo: the library reference model (FRBR-aligned). • Dublin Core: Descriptive metadata. • Schema.org: Web schema and compatibility. • FOAF: Personal entities (creators, contributors). • GeoSPARQL: Geospatial data and queries. HDTo serves as the integrative layer, harmonizing contributions from all listed models into a unified, extensible knowledge graph that supports cross-institutional collaboration, long-term preservation, and advanced research within the ECCCH. 4. Data Management Requirements 4.1 Functional Requirements To ensure effective data management and user engagement, the ECHOES Knowledge Base must address the practical needs of its stakeholders. This section outlines key usage scenarios through representative user profiles, reflecting real-world challenges and informing system design. It needs to be emphasized that a detailed and comprehensive description of the functional requirements will be included in D4.1 Report on identification of communities and their needs, which is due in Month 24 of the project. Therefore, in this document, we present the current shape of the requirements, assuming that they will evolve as needed throughout the project.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 37 This functional-needs analysis serves as a blueprint for the explicit and implicit expectations of stakeholders, encompassing a broad spectrum of project objectives. It includes a description of the proposed system's intended capabilities, user interactions, and output requirements to ensure that it will be designed to meet its intended purpose efficiently, thereby bridging the gap between user expectations and technical execution. In this context, we identify the following: • Key Stakeholders: categorization of the system’s diverse user groups with an emphasis on their distinct roles and interactions, their primary needs and challenges, and the ECHOES functional response and impact. • High-Level Functionalities and Requirements: establishment of the fundamental capabilities, criteria and guidelines that steer the system's development, value proposition and user engagement. This ensures that every aspect of the system architecture and design converges towards fulfilling its intended purpose and user expectations. • Operational scenarios and representative use cases: determination of how the system will integrate into the existing or intended operational environment, ensuring both efficiency and effectiveness in real-world applications. • Expected outcomes: summarizing the expected outcomes and exploitable results, while providing information about the relevant stakeholders and the corresponding inherent requirements. Below, there is a section that categorizes users based on their digital and technical advancement. Even though, for each of those categories a set of ECHOES functional responses is listed, it needs to be taken into account that needs of communities vary. The categories were made to present diversity of Cultural Heritage ecosystem and lines between those categories are blurry. That means ECHOES will offer its capabilities to all kinds of users, for them to select the most applicable ones. 4.1.1 End User Profiles User Profile A: Independent Data Owners (Unstructured Assets) Profile: • Independent archaeologists, grassroots heritage NGOs, small private collectors. • Generate unstructured data (photos, GPS tracks, field notes, spreadsheets, AV recordings). • Store data across personal devices or consumer cloud services with minimal backup/versioning.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 38 Challenges: • Unstandardized Metadata – Poor interoperability due to inconsistent schemas/vocabularies. • No Persistent Identifiers – Limits citation and long-term referencing. • High Data Loss Risk – Reliance on personal devices (laptops, USB drives) without structured backups. • Limited Infrastructure Support – No institutional repositories or IT assistance for preservation. • Low Technical Capacity – Users lack time/expertise for complex data publishing workflows. Primary Needs: • Safe, durable storage solutions to preserve digital assets • Low-barrier tools for organising and enriching data • Mechanisms to make data findable, accessible, and citable • Integration into broader cultural heritage data ecosystems without needing specialist skills ECHOES Functional Response To address these needs, ECHOES will provide the following capabilities tailored for User Profile A: • User-Friendly Upload Interface: A guided, web-based tool for uploading files in various formats (images, CSV, GPX, PDFs) with automatic metadata extraction where applicable and linking it to the KB. • Lightweight Metadata Capture: Customisable templates with dropdowns and minimal required fields, aligned with CH standards (e.g., Dublin Core, DCAT, or LIDO), supporting later enrichment. • Automatic PID Assignment: Persistent identifiers are generated automatically upon dataset submission, ensuring data can be reliably cited in publications. • Basic Access Controls: Data owners can designate assets as private, restricted, or public and manage access accordingly. • Public Discoverability: Indexed in the ECHOES Knowledge Base, enabling external researchers and institutions to find and reuse datasets (with attribution). • Minimal Technical Requirements: All tools and services will be accessible via a standard web browser, requiring no installation or technical configuration.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 39 Impact: By lowering the barriers to participation in digital heritage ecosystems, ECHOES empowers small-scale actors to preserve and share their data responsibly, enhancing long-term value for scholarship, education, and community heritage preservation. User Profile B: Small-to-Medium GLAM Institutions (Structured Data, No Durable Storage) Profile: • Regional museums, galleries, libraries, and archives (GLAMs) • Maintain organized catalogues in legacy systems. • Operate with limited digital infrastructure and staff capacity Challenges: • Data Loss Risk - Outdated hardware and lack of backup systems • Preservation Barriers - Inability to ensure long-term digital preservation • Interoperability Issues - Limited connection to external platforms • Technical Limitations - Lack of capacity to migrate to modern standards • Compliance Pressure - Need to meet FAIR principles and open data requirements Primary Needs: • Reliable long-term hosting for structured digital records • Support for standard data models • Tools to enrich and link existing data • Solutions that integrate with existing institutional workflows ECHOES Functional Response: • Legacy data import - Supports access with auto-mapping to standards • Standards-compliant repository - Linked data models and semantic enrichment • Customizable access controls - Manage internal/restricted/open permissions • Automated preservation - Cloud storage with backups and PID assignment • Curator-friendly interfaces - Web dashboards for non-technical users • Federated publishing both via APIs and web dashboards • Automated metadata harvesting and standards mapping • Synchronized PID and access policy management • Built-in FAIR compliance monitoring
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 40 Impact: ECHOES enables GLAM institutions to sustainably manage and share their collections, bridging legacy systems with modern digital infrastructures while maintaining institutional control over cultural assets. User Profile C: Scientific Community/Academics/Research Infrastructures, Large GLAM Institutions (Structured & Internally Hosted Data) Profile: • Universities, national laboratories, digital humanities centres • Large-scale research consortia with advanced data infrastructures • Operate sophisticated internal platforms Challenges: • Duplication of Effort - Redundant dissemination across multiple platforms • Limited Visibility - Low reuse beyond immediate academic networks • Fragmented Exposure - No unified EU-wide data dissemination layer • Maintenance Overhead - High cost of supporting multiple dissemination systems • Compliance Pressure - Need to meet evolving standards Primary Needs: • Federation capabilities for EU-wide integration • Centralized dissemination without data migration • Persistent identifiers and metadata propagation • Semantic web and linked data interoperability • Compliance tools for FAIR/EOSC requirements ECHOES Functional Response • Federated publishing directly via APIs • Automated metadata harvesting and standards mapping • No-migration integration model preserving data sovereignty • Synchronized PID and access policy management • Built-in FAIR compliance monitoring Impact: ECHOES serves as a trusted intermediary, amplifying research data visibility across Europe while maintaining institutional infrastructure autonomy, enabling compliance and broader participation in cultural heritage knowledge commons.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 41 4.1.2 Data Governance Roles Effective data governance within the ECHOES project relies on the collaboration of multiple actors, each with clearly defined roles and responsibilities. This section outlines the key functions of data providers, infrastructure operators, data stewards, data curators, and end users. Their coordinated efforts ensure that data is accurate, ethically managed, semantically aligned, and technically robust - supporting both immediate research needs and long-term sustainability. Data Providers Data Providers are responsible for delivering accurate, high-quality data while adhering to all relevant policies and standards. Key Responsibilities: • Submitting data accompanied by comprehensive documentation and detailed metadata to ensure clarity and usability. • Complying with data privacy regulations, ethical guidelines, and GDPR requirements to protect sensitive information. • Collaborating closely with infrastructure operators and data curators to uphold data integrity, quality, and long-term usability. • Respect and align with semantic models and ontologies endorsed by the project to facilitate integration and discovery. ECHOES will provide mapping tools for those purposes. Infrastructure Operators Responsible for managing and maintaining the technical environment and ensuring the security of the infrastructure. Key Responsibilities: • Implement enforce role-based access control (RBAC), encryption, and federated identity management systems to safeguard data confidentiality, integrity, and availability. • Monitoring service health, system performance and data flows to prevent data loss, corruption, or unauthorized access. • Supporting data provenance tracking and maintaining comprehensive audit trails to ensure transparency. • Ensuring interoperability by adhering to established standards such as CIDOC CRM. Governance roles across the federation have to be confirmed. Below, the foundation for data governance has been described and may change when project progresses. Details on a consistent governance model and streamline decision-making will be defined by WP10: Cloud Governance in the related WP10 deliverables.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 48 • Traditional Knowledge and Indigenous Rights: Consideration of cultural sensitivities, communal ownership, and consent for use and dissemination of culturally significant materials. • Ethical Use of Data: Protection against misuse or misrepresentation of cultural heritage materials; responsible, respectful engagement with source communities. 5. Data Architecture Design 5.1 Architecture Overview The ECHOES architecture is being designed around a layered approach that supports multiple repositories while ensuring seamless integration within a federated environment. This allows the platform to balance flexibility and openness with the requirements of security, access governance, and accountability. In this document we focus on the layers describing the KB design. The Knowledge Base Management System (KBMS) is the central component of this architecture. It acts as a facilitator between users, external applications, and the cloud infrastructure. The KBMS is composed of several modules that support both administrative tasks and user operations. It manages querying, publishing, and editing of semantic data, while also providing monitoring and auditing functionalities to ensure traceability and compliance. The end-users' access to the system is provided through a single-entry point based on Single Sign-On mechanism that allows a user to authenticate once and then access all authorized services without repeated login attempts. The process is handled through EGI Check-in, which functions as an identity broker and supports open standards such as OIDC and SAML. This ensures interoperability with a wide range of external identity providers, including academic federations, institutional systems, and social login services. This deliverable outlines interoperability in general terms (federation, private/public nodes) but does not define maturity levels or stages of integration. This will be covered by D6.2. A complete and fully detailed system architecture will be delivered in June 2026, as defined in the ECHOES project timeline. The present document provides only a preliminary architectural outline based on the information, requirements, and design assumptions available at this stage of the project. As ongoing work in WP6 and other Work Packages continues to refine functional needs, interoperability specifications, and governance rules, the architecture described here should be understood as an evolving blueprint rather than a final technical design. The forthcoming architecture deliverable (D6.3) will consolidate these developments into a comprehensive, validated, and implementation-ready specification.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 49 The backbone of data management relies on triplestores, which operate as semantic repositories supporting cultural heritage digital twins. The architecture distinguishes between i) ECHOES knowledge-sharing repositories, with open or restricted access, which serve as the official and versioned resources with strong provenance guarantees, ii) external repositories and iii) privatespaces repositories, which provide flexible collaborative spaces for individuals and teams. Semantic data prepared in private-spaces repositories can later be published into knowledge-sharing repositories, enabling a workflow that moves smoothly from iterative editing to formal release. A dedicated administrative repository complements this architecture by storing auxiliary information such as user roles, descriptions and locations of triplestores, ownership records, and workspace definitions. This ensures consistent access management and makes it possible to monitor the health and availability of all nodes across the federation. Low-level, technical, interaction with the system takes place through a user-facing web interface designed to go in the detail of the repository. The interface provides simple mechanisms for querying, editing, and inspecting data, while offering additional administrative features to authorized users, such as node registration and role assignment. To serve diverse communities, from researchers and curators to administrators and external stakeholders in a more specific way we envision the use of dedicated vertical applications. The service layer of the architecture implements the business logic of the KBMS. It manages queries, registry functions, monitoring of federated nodes, and secure communications between components. All services are exposed via REST APIs, making it straightforward to integrate third-party vertical applications and external tools into the platform. 5.2 Federated Architecture and Data Lifecycle Management 5.2.1. Federated Knowledge Strategy As described in section 2, above, a federated architecture enables searching for information across distributed data sources without consolidating them in a central repository. Instead of gathering all data into a single location, a federated Knowledge Graph (KG) leaves data in its original silos and dynamically constructs responses through federated queries. This approach is commonly used in portals that retrieve data from multiple databases while maintaining decentralized storage. This section describes how a federation knowledge strategy is designed for the ECHOES requirements. The main key features of this strategy are discussed in the next paragraphs. Types of federation nodes Triplestores constitute the core storage layer of the Knowledge Base (KB), hosting the semantic RDF content, serving as the federated nodes of the infrastructure. These nodes may reside either within the ECHOES main infrastructure or be hosted by external providers. Their availability is monitored
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 50 regularly to ensure the integrity and reliability of query results. Semantic representations of the HDTs are stored within these triplestores, enabling their seamless integration into the federated Knowledge Graph (KG). The federation engages three main types of triplestores for knowledge aggregation and manipulation: ECHOES KG/knowledge-sharing nodes (for simplicity ECHOES-nodes), authoritative triplestores of the system where semantic data can be shared either openly with all the users, or with restrictions only to authorized ECHOES users. The content of these nodes is built and structured according to the ECHOES federation rules using the provided KB REST API. They support ingest, snapshot versioning and provenance mechanisms, ensuring that their content is traceable, auditable, and immutable over time. These capabilities are critical for maintaining data integrity, reproducibility, and compliance with provenance requirements. External KG/knowledge-sharing nodes (for simplicity External-nodes) associated queryable-only triplestores provided by other projects or organizations that want to join the ECHOES ecosystem. These triplestores publish semantic data to be shared within the ECHOES, but their content has been developed outside the ECHOES infrastructure. No snapshot versioning support or provenance mechanisms must be applied by the system. Auxiliary nodes, serving internal KB mechanisms: • private-spaces nodes, dedicated triplestores used to implement collaborative editing workspaces, where individuals or groups can prepare and enrich HDT representations and finally publish and share that work on the ECHOES-nodes. These spaces allow for concurrent updates and, therefore, exclude features such as history snapshots, provenance provisioning, etc. to prevent synchronization conflicts when multiple users edit the same content simultaneously. This setup facilitates flexible and iterative data editing workflows. • recycle-bin nodes, used to preserve the replaced or deleted content from the federated ECHOES-nodes, for archival, snapshot versioning, and rollback purposes. The ECHOES KB federation can integrate multiple distributed ECHOES nodes to support the creation and sharing of new knowledge within the ecosystem. It may also include External nodes, which are developed and maintained by external stakeholders. All knowledge-sharing nodes must be registered to participate in the federation. The ECHOES-nodes are governed by rules and mechanisms applied by the KBMS. Named graphs (NGs) are an essential mechanism for partitioning RDF content in these triplestores for flexible management and snapshot versioning of semantic data. NG is the fundamental unit of knowledge exchanged and tracked within the KB, allowing for monitoring the evolution of the repository in time. Since the ECHOES-nodes support HDT knowledge sharing and enrichment provided in the ECHOES ecosystem, activity based on user access, regarding ingestion, deletion, and querying of the semantic
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 51 content, is manipulated via the KB RESTful API that controls and guarantees reliable operation for each knowledge-sharing node and the total federated KG as well. Proposed technology to be used: Virtuoso (Reference: https://virtuoso.openlinksw.com/) Information about the decision of using Virtuoso can be found in the Appendix Joining and leaving the federation Data providers can join the federation in two ways: as full ECHOES Nodes and as External Nodes. Full ECHOES Nodes Repositories of this type are considered ECHOES nodes and are used for sharing knowledge within the ecosystem. Access to these repositories—either open or restricted is managed through the KBMS. This mode of participation requires the deployment of a Virtuoso triplestore, whose management is governed by the ECHOES rules via the KB REST API. During the registration process, the repository owner and privileged users authorized to modify the repository must be defined. Each triplestore registration must include a URL endpoint for REST access and a Virtuoso-specific JDBC connection string (JDBC URL, port, and user credentials). Furthermore, repositories providing HDT data are expected to comply with the HDT ontology; otherwise, their content may lack discoverability. To safeguard intellectual property rights (IPR) on RDF content, restricted access policies should be applied at the triplestore level, controlling who can view or modify the data and thus maintaining data confidentiality and security. However, other solutions that combine both open-access and restricted content within the same repository will also be explored. External Nodes Repositories of this type also contribute to knowledge sharing within the ecosystem but are not governed by the ECHOES rules. This mode of participation requires only the registration of an openaccess SPARQL REST endpoint, regardless of the underlying triplestore engine. Such endpoints can be integrated into federated queries across the ecosystem. Data providers requiring restricted access to their endpoints must implement appropriate protection mechanisms offered by their respective triplestore vendors. A typical example of an external repository includes large GLAM institutions that participate using their own managed SPARQL endpoints. Leaving the Federation A repository can leave the federation by invoking the ECHOES unregister process. This action notifies the Administrative Repository through the corresponding KB REST API service. However, unregistering - particularly in the case of full ECHOES nodes - is strongly discouraged, as it may lead to inconsistencies and integrity issues within the federated KG.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 52 Querying the federation Queries over the federated Knowledge Graph (KG) can be executed through dedicated services of the KB REST API in two modes: federated SPARQL queries and parallel execution of SPARQL queries. Both methods are designed to handle service unavailability and can provide information on how each endpoint contributes to the result set. Federated SPARQL Queries These queries adhere to the SPARQL 1.1 standard, utilizing the SERVICE operator inline to specify the target endpoints. The query is executed on a triplestore that supports federated query processing, enabling seamless data retrieval across multiple sources. For queries that search for the same graph pattern across the federation, in other words to execute the same query on all KG nodes, the KB REST API offers a service that automatically builds the federated query and then executes it on a target triplestore. The target triplestore then manages the interaction with the participating federation nodes. Parallel execution of SPARQL Queries SPARQL queries can be executed directly and concurrently on target nodes within the ECHOES federation. The individual query results are then synchronized and merged to produce a unified response. A timeout parameter can be defined for each endpoint to control the synchronization process. Central Management and Access Control All the federated nodes are centrally managed by the ECHOES KBMS. The system handles the triplestores’ registration for joining the federation, applying user access control, and privacy policies. New data provider nodes can be registered to the federation by an infrastructure operator user interacting with the Administrative Repository of the KBMS. Users may query the shared content but can only edit or manage semantic content for which they have controlled authorization. Also, groups of users will be enabled under common authorization to interact with the triplestores of the federation. During the registration of a new triplestore, an ECHOES user can be assigned as the Owner, granting them authorization to enrich, update, and delete RDF content, as well as to designate maintainers who share these privileges for that triplestore. Finally, unregistering any triplestore from the federation is possible. However, this operation may affect consistency and data integrity if other knowledge-sharing nodes reference data hosted by the removed node. RDF Data Upload Ingestion of RDF data into a target ECHOES-node triplestore is performed exclusively through distinct RDF file uploads. Each upload results in the creation of a newly named graph within the triplestore to store the uploaded RDF content and triggers the generation of metadata documenting the upload
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 53 event, thereby supporting traceability and provenance. Since the semantic data is ingested into a named graph, it becomes available for querying by ECHOES users. Collaborative Editing in Private Spaces of the KG The system supports collaborative editing of semantic data through the creation of “temporary” Private Spaces. These are named graphs hosted on dedicated auxiliary triplestores facilitating easy isolation from shared content. These private named graphs allow for collaborating editing by individuals or groups to prepare and enrich HDT representations which can then be published and shared on the ECHOES-nodes. To ensure good performance and prevent synchronization conflicts during concurrent updates, these private spaces exclude features such as history snapshots and provenance provisioning. Private Spaces are accessible only to authorized users. They are suitable for either working with completely new data or modifying a working copy of the KB shared data. Once editing is completed in a Private Space, the resulting content can be published as a newly created named graph on a target ECHOES-node. The system preserves the content in a new RDF file and records the publishing event, generating provenance metadata for the new upload. Within Private Spaces, users can ingest any content from RDF files or perform SPARQL-Update operations directly on their named graphs. As already mentioned, the system does not track changes during this private editing process. RDF Data Update and Versioning Once a NG is created on an ECHOES-node, it cannot be modified with new content. Authorized users, including the NG owner, may replace an existing NG by uploading a new one via the KB REST API. When an NG is replaced, the following strategic actions take place: • The existing NG is moved to a Recycle-bin node to preserve it for rollback and audit purposes. • A new NG is created through the upload of a new RDF file. • Metadata recording the replacement event is generated within the system. The same strategy applies when an NG needs to be completely removed without replacement. Whenever an authorized system user deletes an NG, it is moved to a Recycle-bin node, and appropriate metadata records the deletion event within the system. In contrast, triplestores dedicated to hosting Private Spaces do not support versioning or provisioning. Unlike the knowledge-sharing nodes, RDF provenance metadata is not enforced, allowing users to iteratively insert and update content with minimal overhead. Ownership is managed by the KBMS at the level of named graphs, which function as useror group-specific workspaces. The NG update strategy described here supports snapshot versioning, enabling the capture of the “active” RDF content at any point in time as needed. A snapshot corresponding to a specific date or
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 54 timespan for a given HDT can be obtained by querying the knowledge-sharing and recycle-bin triplestores to retrieve the named graph URIs and/or their content that were active during the desired period. Aggregated Knowledge and Discovery Within the federated system, information about HDTs and other CH entities is distributed across named graphs. A single HDT may be referenced by multiple named graphs hosted on different triplestores. Each HDT maintains semantic links with a set of named graphs that capture these relationships within the context of each graph. Queries across knowledge-sharing nodes—and across private spaces for authorized users—enable semantic discovery by identifying where digital twins are mentioned and in what context. Intellectual Property Rights To protect semantic content from public access on account of IP rights limitations, a dedicated triplestore can be deployed for each provider to host the protected content, with access restrictions enforced at the node level. This approach ensures the highest level of content protection. Additionally, solutions combining both open-access and restricted content within the same repository will also be explored. RDF File Persistence and Provenance All uploaded RDF files, either directly ingested by the users or published from Private Spaces are stored for provenance and reconstruction purposes. To meet this requirement, the system generates, alongside each RDF data file (e.g., filename1.nt), a companion metadata file (e.g., filename1.nt.metadata) containing audit information such as URI of the created named graph, upload time, user identity, etc. HDT Registration Policy To prevent duplicate references to the same HDT or to the physical object it represents, it is strongly recommended that creators verify the existence of an existing URI before instantiating a new HDT. To support this process, the system maintains a catalog of all HDTs created within ECHOES, which users can search to identify reusable URIs. When the creation of a new HDT URI is required, a dedicated service of the KB REST API is invoked to generate and register the new URI. If duplicates are nevertheless detected within the Knowledge Base (KB), an authorized repair workflow can be initiated. During this process, redundant URIs can be replaced with the correct ones, the HDT registry is updated, and the corresponding named graphs are replaced with new versions. All modifications are documented through metadata to ensure that the KB remains consistent, traceable, and fully operational.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 55 URI Generation Policy The system applies to a ruled-based URI policy to ensure that each semantic resource has a persistent, unique, and context-independent identifier across the federation. This policy adopts the following principles: Namespaces allocation. Namespaces used in URIs are systematically allocated based on resource types, enabling robust dereferencing across the federation. Human-readable elements in URIs. Human-readable elements may be included in URIs when they are derived from stable, context-independent attributes of the resource. Where available and consistently generated, concise descriptive labels may be incorporated to improve readability. These enhance clarity, assist users and developers in identifying resource types, and support debugging and review during semantic data transformation and browsing. Avoidance of volatile elements. Transient or context-specific information (e.g., temporary labels, locations, or user-specific data) should be excluded to ensure URIs remain valid across different systems, use cases, and timeframes. Use of classification attributes: Faceted or hierarchical classifications that indicate the resource type or category may be embedded in the URI to support semantic consistency. Reuse of existing Identifiers. If a resource is already identified in a trusted third-party repository, its existing URI should be reused. If no URI exists, but a stable identifier (e.g., inventory number or catalogue ID) is available; it may be incorporated into the new URI. Use of timestamps. For time-sensitive resources, creation of timestamps may be used to distinguish versions or establish temporal uniqueness. Inclusion of provenance Information. Elements such as the data provider, archival source, or originating system may be included if they help uniquely and persistently identify the resource. The URI policy applies to the generation of URIs in the following cases: • HDTs: System created URIs under a common namespace. • Physical Objects that are represented by HDTs • Media/Data files • Named Graphs: System produced, formatted like e.g. [echoes-domain]/graph/user or organization/upload-timestamp • ECHOES Actors: o Institutions / VOs (EGI VirtualOrganizations) o System Users o Authors • Places: o Geographic: URIs aligned with Wikidata or geospatial registries. o Non-Geographic: Institutional places, areas, or local site-specific locations. • Other Entities: Events, activities, and additional entity types follow domain-specific conventions.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 56 5.2.2 User Personal Space The User Personal Space (UPS) should be implemented as a dedicated logical space within the ECHOES infrastructure, isolated from the public-facing federated graph but fully integrated with the platform’s core services. At a high level, this could be achieved through: • Authentication and Authorization: Integration with the ECHOES AAI (e.g., EGI Check-in) to ensure secure user identification, role management, and access control. • Data Storage Layer: Private triplestore or RDF graph segment for unpublished data, with associated binary object storage for files and media assets. This private storage should be logically separated from the federated public knowledge graph, but interoperable for later controlled publishing. • Editing and Enrichment Tools: Web-based user interface for HDT creation and editing, metadata management, and ontology alignment, reusing the same toolset available in the public space but operating in a sandboxed environment. • Version Control and Provenance: Built-in mechanisms for tracking all changes, maintaining full revision history, and capturing provenance metadata even before publication. • Publishing Pipeline: Configurable workflows to transition selected resources from the UPS to the public KB or shared collaborative spaces, with automated validation against metadata completeness, interoperability standards, and governance policies. • Security and Compliance: data encryption at rest and in transit (so how we make sure that data is both secured when stored and transferred safely), regular backups, and strict enforcement of GDPR and IPR requirements, ensuring that sensitive or restricted content remains protected until publication. 5.2.3 Semantic Metadata Persistence Strategy A key architectural consideration for federated knowledge graphs is ensuring that essential semantic information remains accessible and consistent across distributed environments. In cultural heritage contexts, where data may be replicated, shared, or migrated across institutions over time, the persistence of contextual metadata becomes critical for maintaining data integrity and legal compliance. This section outlines strategies for embedding core semantic metadata directly within RDF datasets while maintaining operational efficiency through targeted separation of concerns. D6.1 does not yet provide a layered view of interoperability requirements. It outlines general principles (e.g., federation, public/private nodes) but does not distinguish levels of integration maturity or set priorities (mandatory vs. later-phase elements). D6.2 will introduce these aspects more explicitly.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 57 The ECHOES KB uses a hybrid approach for metadata management, separating semantic metadata (which belongs in the RDF graph) from operational metadata (which stays in the Administrative Repository for performance). Semantic Metadata in RDF Core semantic metadata should live directly in the RDF datasets, following the principle that "metadata travels with data" - similar to how modern file systems like ZFS embed ownership and permissions with the actual files. Here are some indicative examples of metadata that should go in RDF: • Authorship: dcterms:creator, foaf:maker • Licensing: dcterms:license, dcterms:rights • Provenance: prov:wasGeneratedBy, prov:generatedAtTime • Versioning: pav:version, dcat:version This practice ensures that when data moves across federated nodes, its context and legal framework move with it. This separation maintains operational efficiency while keeping semantic context in RDF. When federated SPARQL queries run across multiple triplestores, the semantic metadata automatically comes with the results. This ensures that users receive proper attribution and licensing information without requiring additional database lookups. 5.2.4 Potential Issues and Solutions Consistency and Integrity Issues Any removal of parts of the federated Knowledge Graph (KG), whether at the federation node level or at the named graph level, can lead to inconsistencies and integrity issues within the federation. Additionally, integrating an external repository into the federation may introduce inconsistencies. Solution: Consistency checks using appropriate queries can detect such issues, allowing for timely repairs. Redundant RDF Content and Performance Issues The requirement to include minimal RDF data for each resource to ensure integrity may result in redundant information, as multiple named graphs can contain identical data. This redundancy complicates updates and corrections across the distributed data.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 64 While user authentication and role management are handled by an external identity management service, the Administrative Repository stores complementary metadata necessary to define userspecific permissions within the KB. This includes references to authenticated users, their ownership and maintenance assignments, and their associations with specific triplestores of the ECHOES ecosystem. It also maintains descriptive and operational metadata for each registered in the federation triplestore, such as its identifier, URL, availability status, and configuration details. By centralizing this information, the Administrative Repository contributes to the governance and integrity of the federated KB, ensuring that user access rights are consistently enforced and that KB repository interactions remain traceable and auditable. Its document-oriented structure supports scalability and extensibility, enabling the inclusion of additional metadata types or policy mechanisms as the federation evolves. In this way, the Administrative Repository plays a key role in maintaining secure, coherent, and well-governed access to the distributed resources of the KB. A document-oriented database (MongoDB) supports the storage of this metadata, providing flexibility for evolving access control structures and integrating seamlessly with the external user and role management infrastructure. Proposed technology to be used: MongoDB (Reference: https://www.mongodb.com/) 5.5.4 Back-End - Web Services This is the core component of the KB architecture, responsible for implementing all essential business logic. It provides: 1. A KB REST API for third-party applications and individual stakeholders (i.e., Vertical Applications). 2. RESTful microservices that facilitate communication with the front-end. The back end interacts with all components described above and manages security by integrating with the SSO & RBAC system, which is part of the Single-Entry Point—a separate component outside the KB architecture itself. At a high level, this component, together with the Administrative Repository utilizes metadata to deliver the following functionality: • Web services to support basic CRUD operations on the federated KB; • Registry services to maintain a catalog of all connected triplestores and their associated ownership and maintenance assignments; • Monitoring and tracking of federated node availability; • Administrative functionality for managing users and KB nodes; • Administrative functionality for auditing KB RDF content enrichment evolution • Security features, including the protection of all REST API interactions, ensuring that user access rights are consistently enforced.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 65 • A KB REST API for third-party applications and individual stakeholders (i.e., Vertical Applications). Proposed technology to be used: Spring Boot (Reference: https://spring.io/projects/spring-boot) API Endpoints and Operations The KB REST API for the basic CRUD operations (for triplestore interaction) and registry services is presented in the Appendix section API Endpoints and Operations. The completed API for all the provided services will be published in the ECHOES GitHub repository. 5.5.5 Management Console The KB Management Console is a web-based user interface designed to support administrative operations within the federated KB environment. It serves as a central access point for managing knowledge-sharing and private-spaces repositories, configuring user access, and handling digital twin registrations. The interface aims to provide a unified and intuitive environment for administrators and authorized users to perform governance-related actions without requiring direct interaction with lowlevel APIs or backend configurations. Through this console, users will be able to perform operations such as registering new triple stores and digital twins, constructing and managing working groups, and assigning ownership and maintenance roles to users. It will also offer a simple interface for executing basic Create, Read, Update, and Delete (CRUD) operations over the triple stores, leveraging the underlying KB Management REST API for secure and standardized interactions. Beyond these core functions, the console will emphasize usability, transparency, and traceability. Administrative actions—such as repository registration and access control modification—will be logged to ensure accountability and facilitate auditability. The interface will include mechanisms for visualizing repository status, such as availability indicators, providing administrators with a clear overview of the operational landscape of the federated KB. In alignment with the overall architectural principles, the console will act primarily as a thin client layer, with all logic and data manipulation handled by the backend services. This separation of concerns ensures scalability, maintainability, and consistent enforcement of access control policies across the system. The design will also allow for future extensions, such as integrating analytics dashboards, monitoring services, or collaborative management tools, supporting the evolving needs of the Knowledge Base infrastructure. Proposed technology to be used: ReactJS (Reference: https://react.dev/)
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 66 Conclusion Deliverable D6.1 presents the foundational Data Strategy for the ECHOES Cultural Heritage Knowledge Base (KB), outlining the principles, requirements, and preliminary architectural approach for building a federated, semantically interoperable data ecosystem within the European Collaborative Cloud for Cultural Heritage (ECCCH). The document establishes how cultural heritage data - from raw scientific measurements to interpreted knowledge -will be managed, integrated, enriched, and made reusable across institutions, domains, and countries. At its core, the strategy addresses long-standing challenges in the cultural heritage sector, including fragmentation of data across legacy systems, lack of harmonized metadata, insufficient interoperability, and limited support for collaborative research workflows. It highlights the importance of semantic technologies, ontology alignment (notably CIDOC CRM and the ECHOES Heritage Digital Twin Ontology), and FAIR principles as the foundation for a shared European knowledge environment. The document also identifies critical gaps in existing infrastructures, such as the absence of unified querying across distributed repositories, and positions ECHOES as a unique opportunity to overcome these limitations through a federated but semantically coherent architecture. D6.1 defines the conceptual data lifecycle used in ECHOES - from raw data to processed, postprocessed and interpreted datasets - and demonstrates it through detailed examples from heritage science, 3D documentation, and conservation research. It categorizes data sources expected to be integrated into the KB and describes the roles and responsibilities of data providers, institutions, data stewards, curators, and infrastructure operators within a multi-level governance model. The deliverable outlines high-level functional and non-functional requirements based on user profiles ranging from independent researchers to large research infrastructures. It presents example user scenarios illustrating how ECHOES will support data ingestion, transformation, semantic enrichment, publication, preservation, and federated analytics. It also offers an initial sketch of the technical architecture, including triplestores, metadata pipelines, APIs, identity and access management, and federation mechanisms, noting that the complete, final architecture will be published in June 2026. Overall, D6.1 establishes the strategic, conceptual, and methodological framework for managing cultural heritage data in ECHOES. It prepares the ground for subsequent deliverables, particularly D6.2 on interoperability requirements and guidelines and D6.3 on cloud architecture, where the technical specifications, standards, and implementation details will be further formalized and extended.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 67 References • Niccolucci, F., Felicetti, A. and Hermon, S., Populating the Data Space for Cultural Heritage with Heritage Digital Twins. Data, 7(8), 105 (2022). • Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016). https://doi.org/10.1038/sdata.2016.18 • Fernandez, George & Zhao, Liping & Wijegunaratne, Inji. (2003). Patterns for Federated Architecture.. Journal of Object Technology. 2. 135-149. 10.5381/jot.2003.2.3.a4. • Niccolucci, F., Markhoff, B., Theodoridou, M., Felicetti, A., & Hermon, S. (2023). The Heritage Digital Twin: A bicycle made for two. The integration of digital methodologies into cultural heritage research. arXiv preprint arXiv:2302.07138. https://doi.org/10.48550/arXiv.2302.07138
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 68 Appendix API Endpoints and Operations The final KB REST API - currently in development - will be published in the ECHOES GitHub repository. The following table outlines several core operations. All the requests of the API must include a valid JWT in the Authorization header. Endpoint Method Description graph/insert POST Imports an RDF file on a given triplestore under certain named graph which is automatically generated (multipart/form-data) Parameters: RDF File – The triples file to be validated and located in the triplestore - multipart/form-data; The content type of the file to be uploaded - Text Triplestore ID - Text; Response: A JSON object holding A succeed/failed Boolean value; The automatically generated named graph used; The appropriate message. graph/replace POST Replaces an existing named graph with a new Parameters: RDF File – The triples file to be validated and located in the triplestore - File The content type of the file to be uploaded - Text Triplestore ID - Text; Named graph old - Text; Named graph new - Text; Response: A JSON object holding A succeed/failed Boolean value; The appropriate message. graph/download GET Downloads the contents of specific named graph Parameters:
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 69 The content type of the output file - Text; The content disposition – Text Triplestore ID - Text; The named graph - Text; Response: The raw binary (or text) content of the file itself. hdt/download GET Downloads the RDF content of specific digital twin as a file Parameters: The content type of the output file - Text; The content disposition – Text; Triplestore ID – Text; The digital twin ID - Text; Response: The raw binary (or text) content of the file itself. triplestore/query/download GET Downloads the contents from specific query as a file Parameters: The content type of the output file - Text; The content disposition – Text Triple SPARQL query - Text; Response: The raw binary (or text) content of the file itself. graph/content GET Returns the RDF contents of specific named graph Parameters: Triplestore ID – Text; The named graph - Text; Response: The text content of the named graph. hdt/content GET Returns the RDF content of specific digital twin Parameters: Triplestore ID - Text; The digital twin ID - Text; Response: The text content of the digital twin.
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 70 triplestore/query/result GET Returns the text result after applying specific query Parameters: The SPARQL query - Text; Response: The text content of the query’s results. graph/delete DELETE Deletes a named graph from specific triplestore Parameters: Triplestore ID - Text; Named graph - Text; Response: A JSON object holding A succeed/failed Boolean value; The appropriate message. triplestore/catalog/register POST Registers or updates metadata of a triplestore in the Administrative Repository Parameters: URL – Text; Name – Text; Description – Text; Response: A JSON object holding A succeed/failed Boolean value; The appropriate message. triplestore/catalog/get GET Returns all metadata of a triplestorefrom the Administrative Repository triplestore/search POST Searches for triplestore based on filters Parameters: Creator – Text; Triplestore name; Creation Date; Response: A JSON object holding
ECHOES Project 101157364 – Deliverable D6.1 Data Strategy for the CH Knowledge Base 71 A succeed/failed Boolean value; The appropriate message. Hdt/search POST Searches for digital twins Parameters: Digital Twin’s name – Text; Creator – Text; Creation Date; Response: A JSON object holding A succeed/failed Boolean value; The appropriate message. Selecting Triplestore for the ECHOES The decision for the adopted selection of SPARQL Triplestore for the ECHOES project is documented in 241204_TripleStoreComparison.docx.url