scieee AI-readable full text Open interactive document viewer

Supporting executable scientific workflows in a clustered Infrastructure: DARIAH-IT and H2IOSC

Degl'Innocenti, Emiliano; Pinna, Francesco; Spadi, Alessia; Spinelli, Federica

Abstract

Workflows have become essential in digital humanities, enabling the formalisation, automation and reproducibility of complex research processes. As the humanities increasingly adopt data-driven methodologies, workflows offer structured approaches to manage diverse data and tools while supporting transparency and collaboration. Recognising this need, DARIAH-IT has advanced research infrastructure development by focusing on workflow-based services within the H2IOSC (Humanities and Cultural Heritage Italian Open Science Cloud) project. DARIAH-IT leads the design and implementation of a national cloud system to support digital humanities research, ensuring interoperability, FAIR data practices and semantic integration across disciplines. Central to DARIAH-IT’s effort within H2IOSC is AEON (dAriah sErvice Oriented iNfrastructure), a platform that enables the creation, execution and management of scientific workflows. AEON integrates service provisioning, semantic validation and runtime orchestration, supporting reproducible research and collaborative practices.

Full text

Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC Emiliano Degl’Innocenti , Francesco Pinna , Alessia Spadi , Federica Spinelli Opera del Vocabolario Italiano, Consiglio Nazionale delle Ricerche, Florence, Italy Transformations, A DARIAH Journal Volume 1, 2025 https://transformations.episciences.org Dates Received: 15/11/2024 Accepted: 04/09/2025 Published: 14/10/2025 DOI: 10.46298/transformations.14776 ©The authors Creative Commons Attribution 4.0 International Abstract Workflows have become essential in digital humanities, enabling the formalisation, automation and reproducibility of complex research processes. As the humanities increasingly adopt data-driven methodologies, workflows offer structured approaches to manage diverse data and tools while supporting transparency and collaboration. Recognising this need, DARIAH-IT has advanced research infrastructure development by focusing on workflow-based services within the H2IOSC (Humanities and Cultural Heritage Italian Open Science Cloud) project. DARIAH-IT leads the design and implementation of a national cloud system to support digital humanities research, ensuring interoperability, FAIR data practices and semantic integration across disciplines. Central to DARIAH-IT’s effort within H2IOSC is AEON (dAriah sErvice Oriented iNfrastructure), a platform that enables the creation, execution and management of scientific workflows. AEON integrates service provisioning, semantic validation and runtime orchestration, supporting reproducibleresearchand collaborativepractices. Keywords: digital humanities, scientific workflows, research infrastructures, FAIR data principles, semantic interoperability To cite this article: E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli (2025). Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC. Transformations, A DARIAH Journal, vol. 1, p.1-19, DOI: 10.46298/transformations.14776 Eseguire workflow scientifici con il supporto di un cluster di infrastrutture: DARIAH-IT e H2IOSC Abstract I workflow sono diventati sempre più rilevanti nell’ambito Digital Humanities, poiché permettono di formalizzare, automatizzare e rendere riproducibili processi di ricerca complessi. Con la crescente adozione di metodologie basate sui dati nelle scienze umane e sociali, i workflow offrono approcci strutturati per gestire dati e strumenti eterogenei, promuovendo al contempo trasparenza e collaborazione. Riconoscendo questa esigenza, DARIAH-IT ha contribuito allo sviluppo di infrastrutture di ricerca focalizzandosi su servizi basati su workflow all’interno del progetto H2IOSC (Humanities and Cultural Heritage Italian Open Science Cloud). In questo contesto, DARIAH-IT ha come obiettivo l’implementazione di un sistema cloud nazionale a supporto della ricerca nelle digital humanities, garantendo interoperabilità, pratiche FAIR per i dati e integrazione semantica tra le risorse. Al centro di questo impegno di DARIAH-IT nell’ambito di H2IOSC si colloca AEON (dAriah sErvice Oriented iNfrastructure), una piattaforma che consente la creazione, l’esecuzione e la gestione di workflow scientifici. AEON integra l’erogazione di servizi, la validazione semantica e l’orchestrazione in fase di esecuzione, supportando la riproducibilità della ricerca e le pratiche collaborative. Parole chiave: digital humanities, workflow scientifici, infrastrutture di ricerca, principi FAIR, interoperabilità semantica Acknowledgements The H2IOSC – Humanities and Cultural Heritage Italian Open Science Cloud – project is funded, as part of the European Union Next Generation programme, by Italy’s National Recovery and Resilience Plan (NRRP) under Mission 4 “Education and Research”, Component 2 “From research to business”, Investment 3.1 “Fund for the realisation of an integrated system of research and innovation infrastructures”, Action 3.1.1“Creationof new researchinfrastructures strengthening of existing ones and their networking for Scientific Excellence under Horizon Europe”. (Project code IR0000029 – CUP B63C22000730005. Implementing Entity: CNR.) Conflict of Interest The authors do not have any conflict of interest to declare Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC Introduction Access to digital tools and data is now common, as are opportunities to deploy advanced technologies such as big data, artificial intelligence and large language models. The mature phase of humanities studies, supported by the widespread availability of digital data, also entails a transition from experimental approaches to standard practice 1 . Widespread digitisation has had a direct impact on scholarly methods in general and, more particularly, on the workflows adopted to perform research tasks. The definition of workflows in digital humanities (DH) follows the development of the discipline from the analogue era to the digital era. Humanists can be said to have already employed workflows, understood as a sequence of steps to produce a research result, in their endeavours; what is new is the application of digital tools to these steps. This transition enables scholars to analyse vast amounts of digital data, uncover new patterns and insights, and create innovative forms of scholarly expression (Biemann et al. 2014). This apparently simple transition 2 brings with it the need to reorganise well-established and successful procedures in humanities research, since the rooting of research methods in a consolidated tradition makes the application of digitally enabled workflows complex. A persistent challenge in the social sciences and humanities (SSH) and in cultural heritage (CH) sectors is the fragmentation of data, tools and practices. The diversity of languages, disciplinary traditions, metadata standards and technological systems often hinders collaboration, data reuse and large-scale analysis. This fragmentation not only limits interoperability among resources but also hinders the development of cohesive research environments. In this context, research infrastructures (RIs) 3 can play a leading role in SSH research to coordinate distributed knowledge systems 4 , promote standardisation and enable shared access to tools, services and datasets. They are uniquely positioned to reduce conceptual and technical silos, foster semantic continuity (i.e. ensuring that different data and tools can refer to and operate on shared or compatible meanings across systems and disciplines) and support interdisciplinary research by translating and aligning diverse resources. 1. Father Roberto Busa also recommended considering information technology not as a mere tool for efficiency but as a catalyst for innovation. By demanding new research strategies and a higher level of human engagement, intellectual advancement is propelled beyond the limitations of traditional methods: “Mi preme fare ai giovani una raccomandazione: non mettete vino vecchio in otri nuovi, e tenete conto che l’informatica non è per fare le stesse ricerche di prima con gli stessi metodi di prima ma solo più velocemente e magari con meno lavoro umano. L’informatica obbliga a due cose: primo, all’invenzione di nuove strategie di ricerca, proporzionate alla possibilità di questo strumento; e, secondo, impegna a un lavoro umano più intenso, più condensato, a livelli umani superiori.” See: http://circe.lett.unitn.it/attivita/eventi/pdf_eventi/busa.pdf (accessed 15 April 2024). 2. In his lectures, French philosopher Bruno Latour characterised this digital transition as the “screentoria” (Latour 2014,2005,1993), a modern-day equivalent of the scriptorium where mediaeval monks contemplated knowledge in solitude. The scriptorium has evolved into a digital screen, reflecting the shift in dimension and distribution. 3. For a definition of research infrastructure, see European Union 2013, Article 2 (6): "’research infrastructures’ mean facilities, resources and services that are used by the research communities to conduct research and foster innovation in their fields [...]." 4. Let us refer to a knowledge system as an organised structure of people, practices, data, tools and institutions involved in producing, managing and using knowledge. It includes both formal and informal mechanisms for knowledge creation, dissemination and validation within a domain. E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 3 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC DARIAH-IT, the Italian node of the Digital Research Infrastructure for the Arts and Humanities (DARIAH ERIC), advanced this mission within the H2IOSC (Humanities and Cultural Heritage Italian Open Science Cloud) project. 5 One of the objectives of DARIAH-IT within H2IOSC 6 is precisely to bridge the gap between traditionally separate domains, such as hard sciences and humanities (Snow 1961), enabling deeper, interdisciplinary research. This objective is further exemplified by DARIAHIT’s activities in developing integrated strategies for data management and semantic interoperability in the arts and humanities (Spadi, Degl’Innocenti, and Di Meo 2024). As part of this effort, within H2IOSC DARIAH-IT leads the development of integrated solutions that support workflow management, interoperability and FAIR data practices. To maximise this impact and foster interoperability and large-scale dataset analysis, it is essential to involve all stakeholders, including researchers, interested communities and beneficiaries. Building upon this broader perspective, the following section provides an overview of the specific context in which these efforts materialise, namely the AEON (dAriah sErvice Oriented iNfrastructure) platform, a cloud-based infrastructure developed within DARIAH-IT under the H2IOSC project. AEON’s primary goal is to support the execution and management of complex scientific workflows, thereby enhancing interoperability and reproducibility in digital humanities research, while also exposing the full catalogue of DARIAH-IT services. Context H2IOSC is a project funded by Italy’s National Recovery and Resilience Plan (NRRP) (Presidenza del Consiglio dei Ministri 2021), which aims to create a federated cluster comprising the Italian branches of four European research infrastructures (RIs) – CLARIN, DARIAH, E-RIHS and OPERAS – that operate in the social sciences and humanities domain of the European Strategy Forum for Research Infrastructures. This initiative aligns with broader European strategies promoting data-driven research and fostering interdisciplinary collaboration (ESFRI 2020; Wilkinson et al. 2016). DARIAH-IT is in charge of creating H2IOSC’s cloud system, where it leads architecture design, coordinates software and hardware implementation and manages the information organisation and representation framework within the cloud environment. This involves setting up the physical infrastructure, which includes designing and building eight data centres across different locations and then connecting them, as well as developing a shared semantic framework to manage project knowledge. This framework is basically a set of agreed-upon terms, relationships, and vocabularies ensuring that all the data gathered by the four RIs in H2IOSC can be understood and used more efficiently for research purposes. The federated cloud is needed to implement the primary objective of H2IOSC: enabling data-driven research activities and supporting the description and execution of scientific workflows, which are implemented through "scientific pilots". Within the H2IOSC project, scientific pilots are structured and replicable research experiments designed to 5. https://www.h2iosc.cnr.it/ 6. This is one of the main projects that DARIAH-IT currently participates in. A detailed presentation is provided at the beginning of section. E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 4 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC demonstrate the practical utility and robustness of workflows in addressing specific scholarly tasks. These pilots showcase the technological infrastructure capabilities in concrete scenarios, typically involving collaboration among researchers, data curators and developers, and they serve as practical benchmarks in validating the interoperability, reproducibility and scalability of digital tools and workflows. In particular, DARIAH-IT will implement scientific pilots supporting the execution of a digital philology workflow. To achieve this, DARIAH-IT will undertake a set of preparatory activities, including collecting and evaluating the existing tools, datasets and services relevant to executing all the steps of a research workflow. A significant aspect of this process is the semanticisation of selected data and metadata to promote interoperability among heterogeneous resources (data coming from archives, libraries, museums and/or produced by researchers). This process of resource alignment involves data cleaning, mapping and modelling, as well as defining standardised vocabularies and ontologies, in synergy with the deployment of tools and workflows that are specifically designed to implement the semantic transformation. The Pilots, presented as platforms or hubs, integrate domainspecific services, workflows and interfaces and are conceived as executable scientific workflows, combining different resources in specific computational chains. To create the executable workflows supporting scientific pilots, DARIAH-IT implemented • a set of core elements to manage the interaction between the selected services (i.e. the API manager) • a service-provision-oriented infrastructure that enables the actual execution of the services in a specific runtime environment (i.e. the AEON) By implementing this environment, DARIAH-IT aims to provide an innovative, sustainable and resilient research ecosystem for SSH research. Workflows for digital humanities In academic and industry settings, workflows represent organised sequences of tasks or operations. Within scientific research, the concept of scientific workflows specifically refers to formally defined computational processes that enable automation, reproducibility and the systematic execution of research steps (Atkinson et al. 2017). Thus, while “workflow” is a general concept, “scientific workflow” explicitly denotes computationally executable processes. By breaking complex research tasks into smaller, manageable steps, scientific workflows automate repetitive activities, enhance data management efficiency and foster reproducibility, transparency and replicability, all of which are particularly critical in data-intensive research (Gil et al. 2008; National Academies of Sciences, Engineering, and Medicine 2019; Concordia, Meghini, and Benedetti 2020). To address emerging needs in the implementation of complex workflows, platforms have been developed across disciplinary domains. Among others, Galaxy (Giardine et al. 2005) has been successfully used in bioinformatics, while Kepler (Altintas et al. 2004; Ludäscher et al. 2009) is a relevant attempt known for its generic applicability across scientific disciplines. Within the SSH, relevant platforms include the CLARIAH Media Suite (Melgar-Estrada et al. 2019) and the SSH Open Marketplace (Barbot et al. 2020), a discovery portal that gathers and contextualises resources for the SSH research communities. The SSH Open Marketplace was developed within the Social Sciences and Humanities Open Cloud (SSHOC) project (SSHOC 2022), a E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 5 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC European initiative aimed at building a cloud-based infrastructure for SSH research by integrating services, tools and training resources. The primary objective of the SSH Open Marketplace is to establish a collaborative space where users can access and share digital tools and resources, fostering greater transparency and cooperation in research. 7 Today, the SSH Open Marketplace drives digital transformation within the social sciences and humanities by providing access to a wide range of tools, data, services and resources tailored to researchers and scholars in these fields. It also responds to the growing need for dedicated workflow management tools. The SSH Open Marketplace describes a research workflow as a sequence of steps that can be performed on research data throughout its lifecycle. Workflows can be executed using a variety of tools, methods and resources connected to each step.8 Scientific workflows are employed by researchers as a means of defining automated, scalable and portable experiments: “A scientific workflow is a composition of interconnected and possibly heterogeneous scripts that are used in a scientific experiment” (Concordia, Meghini, and Benedetti 2020). As a good example of workflow management researchers may consider WfCommons 9 , a tool that serves as a comprehensive framework aimed at advancing research and development in scientific workflows. It offers tools, datasets and infrastructure to support the creation, simulation and comparison of scientific workflow instances (Coleman et al. 2022). These workflows streamline the research process by providing a structured approach to data management and analysis, allowing researchers to focus on their core scientific questions. The formal description of an experiment as a workflow can improve the replicability and reproducibility of experiments. Replicability concerns the consistency of results obtained using the same data, computational steps, methods, code and analysis conditions. Reproducibility, on the other hand, concerns the consistency of results between different studies that attempt to answer the same scientific question. Reproducibility requires the use of original data and codes, while replication requires the collection of new data and the use of similar methods. Publishing datasets together with scripts or workflows is common practice, although it may not be sufficient to ensure the reusability of data and reproducibility. 10 Workflows become increasingly complex as they advance research endeavours. Considering the inputs provided by the literature and using the SSH Open Marketplace as a benchmark, the AEON platform design phase was informed by existing knowledge and best practices. When defining the runtime environment for executing scientific workflows, it became necessary to identify the different levels of complexity that typically characterise such workflows. The levels identified are: 7. https://marketplace.sshopencloud.eu/ 8. https://marketplace.sshopencloud.eu/about/service 9. Developed as an open-source platform, WfCommons addresses the complexities involved in running intricate workflows on distributed computing environments, such as cloud and high-performance computing (HPC) systems. For researchers in STEM fields focused on workflow management, WfCommons provides a stable foundation for testing, refining and benchmarking workflow management solutions across a wide array of scientific applications. For more details about WfCommons, see the official website: https://wfcommons.org/. 10. For a comprehensive reflection on the discussion around replicability and reproducibility of experiments, see https://nap.nationalacademies.org/read/25303/chapter/6. E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 6 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC •Unitary workflow: simple tasks requiring a single action •Generic workflow: general data management operations •Complex workflow: multiple steps with specific requirements •Domain workflow: tailored to a specific research domain Unitary workflows are the most basic type, representing simple tasks or processes that can be completed in a single step: examples include data ingestion, cleaning or basic analysis. Generic workflows (as opposed to domain-based workflows) are more comprehensive, outlining common data management operations such as data ingestion, transformation and analysis. These workflows can be applied to various research projects, providing a foundational framework. Complex workflows involve multiple steps, each requiring specific competencies, resources and tools for effective execution. These workflows are often tailored to address complex research questions or projects. Domain workflows (as opposed to generic workflows) are highly specialised, focusing on a specific research domain (e.g, philology, arts, philosophy). They are designed to address specific needs or challenges within that domain, and they may combine multiple complex workflows to achieve the desired outcomes. Any workflow that is completely automated (i.e. not requiring human intervention to be completed) it is called a pipeline. In the field of Information Technology (IT), the concept of pipelines has become a fundamental tool for managing automated workflows. Pipelines refer to a series of interconnected processes through which data flows and information is transformed and refined in a systematic and gradual manner. This model has been extensively developed, documented and refined over the years, enabling greater efficiency, accuracy and scalability in managing complex operations. It therefore seems natural to draw inspiration from the IT world to introduce the concept of a pipeline in the SSH scientific domain, in order to have an instance of tools that can streamline workflows in research and data analysis. Moreover, as SSH increasingly involve the use of digital tools to process large amounts of textual, visual and multimedia data, pipelines can help automate repetitive tasks, improve data processing capabilities and foster collaborative research. By adopting pipeline strategies from IT, humanities scholars can benefit from established automation methods, thereby reducing manual effort while ensuring the accuracy of tasks such as text analysis, metadata extraction and digital archiving. The following section presents how DARIAH-IT implemented these principles in the AEON platform through a systematic workflow development process. E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 7 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC Workflow implementation Within the above context, DARIAH-IT worked on defining performing workflows in the H2IOSC project. Within this framework, DARIAH-IT aims to provide its users with a complete system for creating and managing workflows by upgrading AEON’s current service provision capabilities to match the needs of the national digital humanities research community.11 A systematic approach was applied to the development of a robust workflow that also encompasses servification, virtualisation, and remotisation 12 in research infrastructures. To achieve this, the following key stages were identified: i) assessment of existing tools and services to identify their suitability for transformation; ii) design and development of standardized interfaces and protocols for interoperability and integration; iii) implementation of virtualisation and cloud-based platforms to provide scalable and accessible services; iv) development of user-friendly interfaces and workflows to facilitate seamless interaction for researchers; v) rigorous testing and quality assurance to ensure the reliability and performance of the services; and vi) continuous monitoring and evaluation to identify areas for improvement and adaptation to evolving needs. By implementing these key features, the research infrastructures involved in the project can effectively transition to a more serviceoriented and accessible model, enhancing collaboration and innovation within the research community (i.e. scholars, researchers and practitioners in the SSH and CH domains) who interact with and benefit from shared resources and workflows. DARIAH-IT will make services findable through the H2IOSC Marketplace (and the cooperating projects) and offer actual service provision via the AEON platform. This section presents the set of activities undertaken to achieve these goals, namely the design and implementation of the AEON platform, as well as the evaluation and refactoring of existing services. To strengthen the infrastructure, DARIAH-IT is also concerned about the development of specific policies and guidelines to ensure effective interoperability and security for existing and newly created resources, which is becoming increasingly important due to recent cyber attacks on major cultural heritage institutions.13 The first step in implementing this vision was to define users and the actions they can undertake within the system. AEON supports four user roles, each with clearly defined permissions and responsibilities summarised in table 1. Workflow creation and management permissions begin with the Basic Users, who can create and manage personal workflows. They can also request that a personal workflow be added to the catalogue of services, 11. The AEON platform has not been made available to the general user. To know more about the service, visit the DARIAH-IT and H2IOSC websites. 12. Servification refers to the transformation of resources into standardised, reusable services accessible via APIs; virtualisation involves abstracting physical resources into virtual environments for flexible allocation; remotisation enables remote access and management of resources through network-based platforms. 13. An illustrative example is the recent cyber attack on the British Library, which raised concerns about the resilience of cultural heritage institutions. See: Financial Times, ’Cyber attack on British Library raises concerns over lack of UK resilience’, https://www.ft.com/content/ 642ee014-4768-4c65-b1ee-0d4f39a8a63d (accessed 15 November 2024). A comprehensive summary is also provided in ’British Library cyberattack’, Wikipedia, The Free Encyclopedia,https://en.wikipedia. org/w/index.php?title=British_Library_cyberattack&oldid=1255404045 (accessed November 15, 2024). E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 8 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC Role Permissions Functions Basic User •Access to the catalogue and public services •Create, publish and share personal workflows •Search and execute services and applications •Create and manage workflows •Publish and share workflows (under approval) •Participate in the community Contributor •All Basic User permissions •Submit new services or applications to the catalogue (subject to approval) •Update descriptions of existing services •Submit new services •Update documentation •Provide technical support for contributed services Curator •All Contributor permissions •Review and approve new services or applications •Manage categories and tags in the catalogue •Review and approve submissions •Organise catalogue content •Generate reports on service usage Administrator •Full system access •Manage users and roles •Configure system settings •Monitor performance and security •Oversee system operations •User management and role assignment •System security and performance monitoring •Log analysis and software updates Table 1:AEON user roles, permissions and functions. which will happen after validation and authorisation by administrator-level users. Administrators oversee all aspects of the platform, including the workflows created by other users, ensuring system security, performance and proper configuration. Different user roles typically represent different usage scenarios. A Basic User, for example, might be a researcher who searches the catalogue for relevant services, creates a workflow, executes it on specific data and adjusts it as necessary, or a museum curator who explores applications suitable for creating virtual exhibitions, customises the selected application with appropriate content, tests the result and publishes it for public access. A Contributor may be a developer who is responsible for developing new services, testing and documenting them, submitting them for approval and providing user support once the services are included in the catalogue. Administrators monitor the platform’s system performance, manages user access, reviews security logs, applies software updates and oversees integration with external systems. For all user roles, the workflow creation process begins with a clear description of the workflow’s scope and intended outcomes. Administrators, in addition, have the ability to modify any existing workflow via the workflow manager. If the purpose is to publish a descriptive workflow, a narrative account is sufficient, and the administrator may choose to save and publish the workflow on the platform. In such cases, the inclusion of services is not required. Conversely, to create an executable workflow, at least one service must be selected. When only one service is involved, it must be associated with a graphical user interface (GUI). If the service does not include a native GUI, a standard one is E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 9 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC g = Graph ( ) # T r i p l i f y data for _ , row in s e l e c t e d _ d f . i ter ro ws ( ) : s u b j e c t = URIRef ( f " { manifest [ ’ namespace ’ ] } / { row [ manifest [ ’ id_column ’ ] ] } " ) for column , p r e d i c a t e in manifest [ ’ schema ’ ] . items ( ) : g . add ( ( sub je ct , URIRef ( p r e d i c a t e ) , \ L i t e r a l ( row [ column ] ) ) ) # Output TTL f i l e g . s e r i a l i z e ( d e s t i n a t i o n = ’ / tmp / ou tp ut_s er vice _b . t t l ’ , \ format= ’ t t l ’ ) # ## s e r v i c e c : Load TTL i n t o V i r t u o s o / GraphDB ### def s e r v i c e _ c ( ∗ ∗ context ) : t t l _ f i l e = ’ / tmp / o ut pu t_ s er vi ce _b . t t l ’ # T h i s i s a p l a c e h o l d e r . The a c t u a l c ode would depend \ on your s p e c i f i c setup for Virtuoso / GraphDB # I t would t y p i c a l l y i n v o l v e c o n n e c t i n g to t he SPARQL \ endpoint and performing an INSERT or LOAD query with open ( t t l _ f i l e , ’ r ’ ) as f i l e : ttl_data = f i l e . read ( ) # Example SPARQL Update query f o r l o a d i n g i n t o \ Virtuoso / GraphDB sparql_update = """ ␣ ␣ ␣ ␣ INSERT ␣DATA␣ { ␣␣␣␣␣␣␣␣%s ␣␣␣␣} ␣ ␣ ␣ ␣ " " " % t t l _ d a t a # Assume c o n n e c t i o n t o V i r t u o s o / GraphDB SPARQL e n dp o in t # You would use a package l i k e SPARQLWrapper or a \ custom connection to load the TTL # s p a r q l _ c l i e n t . query ( s p a r q l _ u p d a t e ) print ( " Loaded ␣ TTL ␣ in t o ␣ SPARQL ␣ database . " ) # ## D e fi n e A i r f l o w t a s k s ### task_a = PythonOperator ( t a s k _ i d = ’ s e r v i c e _ a ’ , p yt h on_ ca lla bl e = service_a , provide_context=True , params ={ ’ i n p u t _ f i l e ’ : ’ / path / to / input . xml ’ } , # o r i n p u t . j s o n dag=dag , ) E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 16 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC task_b = PythonOperator( t a s k _ i d = ’ s e r v i c e _ b ’ , p yt ho n_c al la bl e = serv ice_ b , provide_context=True , dag=dag , ) task_c = PythonOperator( t a s k _ i d = ’ s e r v i c e _ c ’ , py th on_ cal la b le = s erv i ce_ c , provide_context=True , dag=dag , ) # Task d e p e n d e n c i e s task_a >> task_b >> t ask_c Example YAML Manifest (for Service A and B) manifest_service_a.yaml r e l e v a n t _ i n f o : ’ rec ord ’ # Path in XML where relevant # data i s st o re d fields : −id −title −d e s c r i p t i o n −date manifest_service_b.yaml relevant_columns : −id −title −d e s c r i p t i o n schema : id : ’ http : / / example . org / id ’ t i t l e : ’ http : / / purl . org / dc / elements / 1 . 1 / t i t l e ’ d e s c r i p t i o n : ’ http : / / p url . org / dc / elements / 1 . 1 / d e s c r i p t i o n ’ date : ’ http : / / purl . org / dc / elements / 1 . 1 / date ’ namespace : ’ http : / / example . org / r e s ource ’ id_column : ’ id ’ E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 17 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC References Altintas, I., C. Berkley, E. Jaeger, M. Jones, B. Ludascher, and S. Mock. 2004. “Kepler: an extensible system for design and execution of scientific workflows.” In Proceedings. 16th International Conference on Scientific and Statistical Database Management, 2004. 423–424. https://doi.org/10. 1109/SSDM.2004.1311241. Atkinson, Malcolm, Sandra Gesing, Johan Montagnat, and Ian Taylor. 2017. “Scientific workflows: Past, present and future.” Future Generation Computer Systems 75 (June): 216–227. https://doi.org/10.1016/j. future.2017.05.041. Barbot, Laure, Yoann Moranville, Stefan Buddenbohm, Klaus Illmayer, and Matej Ďurčo. 2020. MS42 Marketplace – Alpha Release. Zenodo. Milestone 42 of SSHOC project (alpha release of SSH Open Marketplace), June. https://doi.org/10.5281/ zenodo.4585700. Biemann, Chris, Gregory R. Crane, Christiane D. Fellbaum, and Alexander Mehler. 2014. “Computational Humanities – Bridging the Gap between Computer Science and Digital Humanities (Dagstuhl Seminar 14301).” Dagstuhl Reports (Dagstuhl, Germany) 4 (7): 80–111. issn: 2192-5283. https: //doi.org/10.4230/DAGREP.4.7.80. Coleman, Tainã, Henri Casanova, Loïc Pottier, Manav Kaushik, Ewa Deelman, and Rafael Ferreira da Silva. 2022. “WfCommons: A framework for enabling scientific workflow research and development.” Future Generation Computer Systems 128:16–27. issn: 0167-739X. https://doi.org/https: //doi.org/10.1016/j.future.2021.09.043. Concordia, Cesare, Carlo Meghini, and Filippo Benedetti. 2020. “Store Scientific Workflows Data in SSHOC Repository.” In Proceedings of the Workshop about Language Resources for the SSH Cloud, edited by Daan Broeder, Maria Eskevich, and Monica Monachini, 1–4. Marseille, France: European Language Resources Association, May. isbn: 979-10-95546-43-6. https : / / aclanthology.org/2020.lr4sshoc-1.1/. ESFRI. 2020. ESFRI White Paper: Making Science Happen. A New Ambition for Research Infrastructures in the European Research Area. Technical report. White Paper published 27 April 2020 by ESFRI. Brussels: ESFRI, April. https://www.esfri.eu/sites/default/ files/White_paper_ESFRI-final.pdf. European Union. 2013. Regulation (EU) No 1291/2013 of the European Parliament and of the Council of 11 December 2013 establishing Horizon 2020 – the Framework Programme for Research and Innovation (2014–2020) and repealing Decision No 1982/2006/EC. https://eur-lex.europa.eu/ legalcontent / EN/ TXT /?uri=CELEX : 32013R1291. Official Journal of the European Union, L347, 20.12.2013, pp. 104–173, December. Giardine, Belinda, Cathy Riemer, Ross C. Hardison, Richard Burhans, Laura Elnitski, Prachi Shah, Yi Zhang, et al. 2005. “Galaxy: a platform for interactive largescale genome analysis.” Published online 16 September 2005; printed October 2005; PMID 16169926; PMCID PMC1240089, Genome Research 15, no. 10 (October): 1451–1455. https://doi.org/10.1101/gr. 4086505. Gil, Yolanda, Ewa Deelman, Mark Ellisman, Thomas Fahringer, Geoffrey Fox, Dennis Gannon, Carole Goble, Miron Livny, Luc Moreau, and James Myers. 2008. “Examining the Challenges of Scientific Workflows.” Computer 40 (January): 24–32. https://doi.org/10.1109/MC.2007.421. Latour, Bruno. 1993. We Have Never Been Modern: Essai d’anthropologie symétrique. Translated by Catherine Porter. Translated by Catherine Porter; original French edition published 1991. Cambridge, MA, USA: Harvard University Press. isbn: 0674-94838-6. Latour, Bruno. 2005. Reassembling the Social: An Introduction to Actor - Network Theory. First edition published 2005; part of the “Clarendon Lectures in Management Studies”. Oxford & New York: Oxford University Press. isbn: 978-0199256044. E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 18 Supporting Executable Scientific Workflows in a Clustered Infrastructure: DARIAH-IT and H2IOSC Latour, Bruno. 2014. How Better to Register the Agency of Things: Ontology. Tanner Lecture, Yale University. Delivered 27 March 2014; accessed 20 June 2025, March. http: //www.bruno-latour.fr/node/563.html. Ludäscher, Bertram, Ilkay Altintas, Shawn Bowers, Julian Cummings, Terence Critchlow, Ewa Deelman, David De Roure, et al. 2009. “Scientific Process Automation and Workflow Management.” Chap. 13 in Scientific Data Management: Challenges, Technology, and Deployment, edited by Arie Shoshani and Doron Rotem, 467–508. Computational Science Series. Book chapter in the Computational Science Series. Boca Raton, FL: Chapman & HallCRC, December. https://doi.org/https://doi.org/ 10.1201/9781420069815. Melgar-Estrada, Liliana, Marijn Koolen, Kaspar Beelen, Hugo Huurdeman, Mari Wigham, Carlos Martinez Ortiz, Jaap Blom, and Roeland Ordelman. 2019. “The CLARIAH Media Suite: a Hybrid Approach to System Design in the Humanities.” In Proceedings of the 2019 Conference on Human Information Interaction and Retrieval (CHIIR ’19), 373–377. March. https://doi.org/10.1145/ 3295750.3298918. National Academies of Sciences, Engineering, and Medicine. 2019. Reproducibility and Replicability in Science. A consensus study report defining reproducibility and replicability, with recommendations to improve research rigour and transparency. Washington, DC: The National Academies Press. isbn: 978-0-309-48616-3. https://doi. org/10.17226/25303. Presidenza del Consiglio dei Ministri. 2021. Piano Nazionale di Ripresa e Resilienza (PNRR). https://www.governo.it/sites/ governo . it / files / PNRR . pdf. Versione trasmessa alla Commissione Europea il 30 aprile 2021, April. Snow, Charles P. 1961. The Two Cultures and the Scientific Revolution. The Rede Lecture. Cambridge: Cambridge University Press. Spadi, Alessia, Emiliano Degl’Innocenti, and Carmen Di Meo. 2024. “DARIAH.it: Data Integration Strategies and Solutions for Digital-Resources Management and Research in the Arts and Humanities.” Published in June 2024; Open Access under CC BY-NC-ND 4.0, Mimesis Journal 13 (2): 119–134. issn: 2279-7203. https://doi.org/ 10.13135/2389-6086/9920. SSHOC. 2022. SSHOC Legacy Booklet. Zenodo. Published on Zenodo covering SSHOC project outcomes. https://doi.org/10.5281/ zenodo.6394462. Wilkinson, Mark D., Michel Dumontier, IJsbrand J. Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, and et al. 2016. “The FAIR Guiding Principles for Scientific Data Management and Stewardship.” Published 15 March 2016; PMC 4792175, Scientific Data 3 (1): 160018. https://doi.org/10.1038/sdata.2016.18. E. Degl’Innocenti, F. Pinna, A. Spadi, and F. Spinelli 19