Draft Catalogue of Innovative Workflows for Data Management and Connection to the SSH KG
Abstract
This milestone presents the preliminary catalogue of innovative workflows for data management and connection to the SSH KG developed within Task 3.2 of the GRAPHIA project. The catalogue will aggregate and describe ESFRI’s workflows, aiming to improve KG-ready or compliant data for WP3 partners serving as Data Providers. This task focuses on the connection of heterogeneous and legacy data to the SSH KG, complementary to T2.4. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the Agency. Neither the European Union nor the granting authority can be held responsible for them.
Full text
GRAPHIA Knowledge Graphs, AI Services and Next Generation Instrumentation for R&D in Social Sciences and Humanities Work Package WP3 Innovative Solutions and Instruments for the SSH KG Milestone n. 4 Draft Catalogue of innovative workflows for data management and connection to the SSH KG Funding Instrument: Horizon Europe Call: HORIZON-INFRA-2024-TECH-01 Call Topic: R&D for the next generation of scientific instrumentation, tools, methods, solutions for RI upgrade Project Start: 2025-01-01 Project Duration: 36 months Document Identifier: 10.5281/zenodo.17941071 Funded by the European Union. Grant Agreement number 101188018. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.
Deliverable Information WP Number: 3 WP Title: Innovative Solutions and Instruments for the SSH KG Milestone Number: 4 Milestone Full Title: Definition of total number of innovative workflows Deliverable Short Title: Workflows draft Document Identifier: GRAPHIA_WP3_M4_Workflows draft-vn.n Beneficiary in Charge: CNRS MAP Report Version: v1.0 Report Submission Date: 2025-12-15 Dissemination Level: PU Nature: Catalogue Lead Author(s): Anthony Pamart (CNRS MAP) Co-author(s): Julien Homo (Foxcub), Luca De Santis (Net7), Maeo Greco (CNR ISPC), Giovanni Pescarmona (CNR ISPC), Valentina Vassallo (CYI) Keywords: Workflows, Connection, W7 Status: Submied Change Log Date Version Author/Editor Summary of Changes made 2025/11/25 v0.1 Anthony Pamart Initial draft 2025/12/12 v0.2 Anthony Pamart Modified draft 2025/12/14 v1.0 Giovanni Pescarmona Final draft 2 of 22
Table of Contents Deliverable Information 2 Change Log 2 List of Figures 4 List of Tables 4 List of Abbreviations 4 Executive Summary 6 1. Introduction 8 1.1 Purpose and Scope of the Document 8 1.2 Structure of the Document 9 2. Concepts definitions 10 2.1 Definition of catalogue 10 2.2 Definition of workflow (WF) 10 2.3 Tentative definition of Catalogue of innovative workflows 11 3. Workflows definitions 14 3.4 Tentative methodology for workflows description to be managed in the Catalogue 18 3.4.1 Draft template for workflow description 19 3.4.2 Generic workflow description of [MAP-CNRS] 20 4. Conclusion 21 4.1 Roadmap for the implementation of the Catalogue of workflow 21 3 of 22
List of Figures Figure 1: Examples of realistic workflow nodes organising and structuring time-oriented relations between digital activities: a) chain, b) parallel sequences, c) knot, d) iterative sequences, e) repetitive sequences (credits: Dudek, 2023). Figure 2. Screenshot of Quasimodo platform showing a graph-based W7 provenance exploration of Notre-Dame de Paris KG-ready HDT (credits: De Luca, 2025). Figure 3. Current overview of the ANAMNESIS platform (Credit: Pamart, 2025). Fig 4. Graphical overview of workflow management within ANAMNESIS, current state (up), ongoing (middle) and expected (boom). (Credit: Pamart 2025) List of Tables ND List of Abbreviations CH Cultural Heritage DoA Description of Action ESFRI European Strategy Forum on Research Infrastructures HMI Human Machine Interface HS Heritage Science ICT Information and Communication Technology IT Information Technology KG Knowledge Graph LoD Level of Detail SSH Social and Human Sciences SSHOC Acronym of the project Social Sciences & Humanities Open Cloud UC Use case(s) WF Workflow 4 of 22
5 of 22
Executive Summary This milestone presents the preliminary catalogue of innovative workflows for data management and connection to the SSH KG developed within Task 3.2 of the GRAPHIA project. The catalogue will aggregate and describe ESFRI’s workflows, aiming to improve KG-ready or compliant data for WP3 partners serving as Data Providers. This task focuses on the connection of heterogeneous and legacy data to the SSH KG, complementary to T2.4. The catalogue will establish and document protocols toward the creation of semantic-aware data to create KGs integrable in the GRAPHIA Knowledge Graph (SSH KG). This preliminary version will evolve toward Deliverable D3.2 (M24), which will present the FAIR-by-design workflows to feed GRAPHIA KG with data from the Catalogue of Instruments (T3.1) and Use-Cases (T3.3). In line with the Description of Action, this milestone clarifies that the catalogue developed in T3.2 includes not only W7-based Cultural Heritage workflows (documentation, provenance, and digital pipelines) but also their explicit relevance for the future connection of heterogeneous and legacy datasets to the SSH Knowledge Graph. While workflows rooted in CH practice remain valid and essential components of the catalogue, they must be articulated with the KG-connection perspective that structures WP3 and the entire GRAPHIA ecosystem. This framing ensures that workflows are not documented solely as operational procedures, but as enablers of semantic-ready, interoperable data flows aligned with the project's long-term objectives. T3.2 is therefore conceived as a complementary and collaborative eort with the architectural, semantic, and ingestion activities carried out in WP2. T3.2 and WP2 will work together to explore and identify the most eective approaches to achieve the project’s KG ingestion objectives. This milestone does not aim to define technical specifications for KG ingestion at this stage; instead, it provides a set of conceptual and methodological foundations that will support a shared and iterative alignment between T3.2 and WP2. As a framing document, it outlines the principles and 6 of 22
structures that will support the production of KG-ready or KG-compliant datasets, acknowledging that detailed mappings, controlled vocabularies, and API-based mechanisms will be formalised in the next phases of WP3. Deliverable D3.2 (M24) will build on this groundwork to provide a consolidated, fully interoperable catalogue aligned with WP2 specifications. 7 of 22
1. Introduction 1.1 Purpose and Scope of the Document This document will initiate the actions to be led within the context of T3.2 Innovative workflows for data management and connection to the SSH KG. The Description of Action will be achieved by the link to this task with other WP3 tasks, namely the T3.1 Catalogue of innovative instruments and T3.3 Use Cases implementation. Cross-WPs collaboration is also required; this document will also present some links between T3.2 and actions from WP2 SSH Knowledge Graph (T2.3 and T2.4) and WP4 Artificial Intelligence Solutions for the SSH KG (T4.1 and T4.3). This document adopts a high-level approach to introduce how innovative workflows can act as bridges between the operational practices of ESFRI data providers and the semantic requirements of the SSH KG. In particular, it frames W7-based workflow documentation as a conceptual layer that captures provenance, methodological context, and digital activity structures. Such a layer is essential to support future mappings to GoTriple-compliant metadata schemas and to enable the transformation of CH and HS workflows into KG-ingestible resources. By articulating workflows, W7 conceptual models, and the requirements of the SSH KG, this milestone clarifies the position of T3.2 within the GRAPHIA workflow-to-graph pipeline. Recognising that T3.2 is complementary to the work carried out in WP2, this milestone explicitly acknowledges the need to prepare workflows and metadata structures that can evolve toward KG-ready formats. While WP2 defines the semantic architecture, ingestion pipelines, and interoperability standards, T3.2 focuses on documenting and structuring the processes that produce data, ensuring that they can be connected to the SSH KG in a transparent, traceable, and FAIR-by-design manner. As such, this milestone positions T3.2 as a foundational activity for the broader GRAPHIA ecosystem and clarifies that more detailed technical specifications will be developed jointly with WP2 during the upcoming phases of WP3. 8 of 22
1.2 Structure of the Document This document is organised as follows: Section 2 describes the Workflow and Provenance topics, as necessary contextualization. Section 3 describes the workflow definition in a methodological open discussion and potential tools to support the task. Section 4 describes the future actions and strategies to develop the GRAPHIA catalogue of workflows. 9 of 22
Fig 2. Screenshot of Quasimodo platform showing a graph-based W7 provenance exploration of Notre-Dame de Paris KG-ready HDT (credits: De Luca, 2025)9 This W7-based structuring constitutes a crucial preparatory step before ingestion into a SSH Knowledge Graph. It provides a harmonised semantic surface from which richer ontological models (e.g. PROV-O, CIDOC-CRM, or project-specific vocabularies) can be derived in a systematic way. In practical terms, the W7 description functions as a pivot: it captures the essential meaning behind activities and data transformations, enabling consistent mapping to KG entities and relations, regardless of disciplinary vocabularies or local documentation standards. Thus, W7 enhances interoperability, improves the quality of the metadata entering the KG, and ensures that workflow-derived resources can be aligned, queried, and reused within the broader SSH semantic ecosystem. W7 and GoTriple are not competing models but complementary layers that capture dierent facets of the research process. W7 focuses on activities — how data are produced, by whom, with which tools, where and why — while GoTriple focuses on the resulting research resources and how they can be shared, discovered and reused in the SSH ecosystem. Because activities produce resources, a natural continuity exists between the approaches: W7 documents the context, provenance and meaning behind research actions, and GoTriple provides the semantic structure for exposing the resulting datasets, methods, software or publications. Using W7 as an intermediate layer therefore enriches GoTriple-compatible resources with clearer provenance and contextual grounding, while allowing workflows and practices to remain fully traceable. 3.3 The ANAMNESIS web-application for W7 metadata and paradata management of CH and HS documentation ANAMNESIS is an open-source platform for collaboratively creating and managing dynamic metadata and paradata for CH and HS instrumentation-based documentation workflow. It is included in GRAPHIA’s Catalogue of innovative instruments as a software compound of the HYPERMNESIA framework dedicated to workflow provenance for HS applications. It has been developed by an eponym 9 Source and credits (De Luca, 2025) : ERC Advanced Grant NDame Heritage 16 of 22
project funded by the French Foundation of Heritage Science (FSP) in 2024 to contribute to the French HS infrastructure project EQUIPEX+ ESPADON. The FSP is holding the national node of E-RIHS France, insofar as GRAPHIA is a logical extension to test and update the tool ANAMNESIS to a European scale. This prototype is currently hosted and maintained in the ESPADON distributed computing infrastructure, but is ready for new instance deployment (Docker-based). This innovative tool is currently being integrated into the E-RIHS Catalogue of Service as DIGILAB. ANAMNESIS, in its current stage of development, enables a user-based documentation of digital activities through the W7 structure. The W7 implementation introduces a small adaptation to the W7 scheme and states the following classes to define a Digital Activity: - WHAT: Digital Entity (i.e. a data resource in its digital or digitalized form) and a Material Entity (i.e. the physical object itself) - WHERE: Location - WHO: Agents (individual and or institution) - WHEN: Temporal information - HOW: Methods and techniques - WHICH: Instruments (i.e tools, software, devices, etc.) - WHY: Context (i.e projects, etc.) The activities belong to dierent stages of CH documentation from digitization practices with dedicated schemas. Schemas can be created, modified and shared by expert users with Admin roles directly within the interface. ANAMNESIS enables the stabilization of 5W with perennial semantic identifiers: - WHAT: DOI, Ark, Uri - WHERE: GeoNames, or WhatThreeWord - WHO: HAL API (to be linked to ORCID) - WHEN: ISO Standard or PeriodO - WHY: HAL API (linked to CORDIS) The last 2Ws (HOW and WHICH) interlink instruments and their specific modalities used in the activity are stabilized by user-defined concepts (OpenTheso) or by other perennial identifiers (URI, ARK, Handle). ANAMNESIS enables the creation of W7 structured metadata schemas to document digital activities that are exported in JSON format. The Import features enable the platform to be fed with new schemes and a conflict management system, which 17 of 22
keeps the database clean. The platform also enables users to create templates to support batch processing. The export feature includes a mapping functionality to transform the output to other metadata schemas (GoTriple, Dublin Core, EDM). Metadata descriptions could be gathered in user-customised Collections and could be shared using dierent visibility parameters (public, private, confidential) to be enriched dynamically and collaboratively. ANAMNESIS is already available in open-source (licence GNU AGPLV3)10 and a testing instance is available11. Fig 3. Current overview of the ANAMNESIS platform (Credit: Pamart, 2025). 11 hps://anamnesis.espadon.net/ 10 hps://gitlab.espadon.net/anamnesis/anamnesis_monorepo 18 of 22
3.4 Tentative methodology for workflows description to be managed in the Catalogue In this section, all partners involved in the task are invited to describe their most common workflow in natural language (or a matrix form to be discussed in 2026). The main idea is to evaluate the compliance or readiness level of workflows to be managed to tailor the best methodological and technological approach. The task leader and the partners will conduct a self-evaluation of the FAIRness of their workflow to be described by semantic layers (GoTriple, W7, EDM, PROV, and others). List of 3.2 participants (PM involvements) and expected inputs: CNRS MAP (18PM) - Workflows catalogue, WF description (expected numbers of WF = 3) CNR (9PM) - Workflows description (expected numbers of WF 1) UNIBO (9PM) - Workflows description (expected numbers of WF = 1) CYI (5PM) - Workflows description (expected numbers of WF = 1) Foxcub (5PM) - Interoperability assessment and development TIB (5PM) - Interoperability assessment and development Net7 (3PM) - Interoperability assessment and development KNAW (2PM) - Interoperability assessment and development 3.4.1 Draft template for workflow description Based on the tools of the catalogue of instruments (T3.1) in relation to the use cases (T3.3), try to summarize the most generic and conventional workflow, including data collection/acquisition, data processing and data analysis activities. Workflow textual description: The workflow of [Name of the Partners] starts with a data collection performing [Type of the Tools and Instruments] using [Type of Technique] generating [Type of input and Data source]. The input data sources are transformed by [single or multiple] processing step generating [Type of output and Data output] using [Type of the Tools and Instruments]. This output is [intermediary or final]. If intermediary, this intermediary data are transformed by [single or multiple] a new [Processing or Analysis] step consisting of [Data filtering, optimization, extraction, annotation, segmentation, enrichment, enhancement, fusion] generating [Type of output and Data output] using [Type of the Tools and Instruments]. Repeat this sentence until Final output to be integrated in GRAPHIA KG is obtained. 19 of 22
Type of application domain: CH documentation, Heritage Science, Geospatial, Digital Humanities and studies Type of workflows: Linear, cyclic or both Type of technique: Either generic (Real-based modeling / Image Capture), intermediary (Image-based modeling / 2D scanning) or specific (Photogrammetry / Photography) Type of the Tools and Instruments: Name of the instruments or tools (acquisition device, software, etc.) Type of input: Single, multiple or variable Type of Intermediary Data: Single, multiple or variable Type of output: Single, multiple or variable Type of Data source: Most used file type (text, image, audio, table, 3D model) and format Type of Data output: Most used file format (text, image, audio, table, 3D model) Metadata and paradata of input, intermediary and output: Yes or no, technical and/or descriptive. Metadata scheme or standard used: Example (EXIF, DublinCore, Europeana Data Model, etc.) Workflow formalization: Type of data collection, type of data processing, type of data analysis 3.4.2 Generic workflow description of [MAP-CNRS] Based on the tools of the catalogue of instruments (T3.1) in relation to the use cases (T3.3), MAP-CNRS summarizes its most conventional workflow, including data collection/acquisition, data processing and data analysis. Workflow textual description: The workflow of MAP-CNRS starts with a data collection performed by Real-based capture using Photogrammetry and generating multiple image sets as input. The input data sources are transformed by a single processing step, generating a pointcloud and spatialized images using commercial or FLOSS image-based modeling software. This output is intermediary. This intermediary data is transformed by multiple outputs by a new digital activity step consisting of annotation and segmentation, generating a semantically enriched 3D model composed of several file types using the web platform AIOLI. Occasionally, this process is iterated on dierent parts of CH objects at dierent temporalities to construct HDT, aiming to be integrated in GRAPHIA KG. Type of application domain: CH documentation and Heritage Science Type of workflows: Linear 20 of 22
Type of technique: Generic (Real-based modeling), intermediary (Image-based modeling) or specific (Photogrammetry) Type of the Tools and Instruments: SPIDER, STUDIO or ARCH photogrammetry rig combined with 3D modeling software (Metashape, MicMac, Meshroom) integrated in the AIOLI spatial annotation platform Type of input: Multiple Type of output: Multiple Type of Data source: Image set (RAW+JPG) Type of Intermediary Data: 3D Pointcloud (PLY) Type of Data output: Segmented 3D models (PLY), vectors (SVG), images (JPG), table (CSV), and text (JSON). Metadata and paradata of input, intermediary and output: Yes, technical and descriptive. Metadata scheme or standard used: EXIF and W7 Workflow formalization: Real-based modelling for data collection, photogrammetric processing, and semantic enrichment through annotation for data analysis. In this example, formalized from UC 003 (Data ingestion & enrichment for architectural digital twins), we will mobilize innovative instruments mentioned in UC-040 to UC-043 of the GRAPHIA SSH KG - Use cases Matrix (internal file). The workflow is constructed and semantically enriched from the data source to its KG connection. This data corpus is typical of KG-ready ESFRI’s source, requiring a metadata mapping to fit the minimal requirement of the GoTriple model to be integrable into SSH KG. 4. Conclusion As mentioned in the DoA, a W7 structured approach of CH workflows (documentation, provenance, pipelines) is a valid component of the catalogue in order to support KG connection perspectives. Starting in January 2026, T3.2 will focus on WP3 serving as the Data Provider to feed GRAPHIA KG by standardizing interoperable connections with T2.4. These further technical specifications regarding intraoperative processes will be developed during the next phases of WP3. 4.1 Roadmap for the implementation of the Catalogue of workflow Proposal of actions, subtask’s deadlines to formalize the catalogue of workflows expected to be delivered in December 2026. 21 of 22
● Early 2026, necessity of validation from WP leaders (mostly W7 compliance) internal meeting scheduled ● Early 2026, translation of the Front-End in English ● First semester 2026, WF description to HS instruments (T3.1) and UCs (T3.3) ● First quarter 2026, mapping of W7 metadata to GoTriple ● Second quarter 2026, ad-hoc WF activities schemes for GRAPHIA WF (LLM4SSH, etc ) ● Third quarter 2026, extension to SSH workflows ● October 2026, draft catalogue based on the existing user-friendly prototype tool ANAMNESIS (TRL similar to Quagga) ● Last quarter 2026, delivery of the Catalogue of Workflows and interconnection with GRAPHIA (API) In addition, a necessary task meeting has to be planned (summer or third quarter 2026) in physical or hybrid format in Marseille (MAP-CNRS) or Florence (Headquarters of E-RIHS) with the form of a datathon aiming for the co-construction of workflow and interoperability solving. Fig 4. Graphical overview of workflow management within ANAMNESIS, current state (up), ongoing (middle) and expected (boom) (Credit: Pamart 2025). 22 of 22