D4.2 - Training material, manual and documentation for RESILIENCE
Abstract
This document is jointly written by WP4 members to detail the achievements goaled in their work prompted by the third deliverable (number D4.2, preceded by D4.1.1 and 4.1.2).The experience and the feedback collected during the development and testing of DaMSym platform, and also during the preparation of D4.1.1 and D4.1.2 is systematised and integrated in this deliverable, in the form of a training package to be integrated with the RESILIENCE infrastructure. The package includes use-case samples (drawn from D4.1.2), as well as templates.
Full text
Ir00000014 - Itserr D4.2 - Training material, manual and documentation for RESILIENCE 1 Document reference: ITSERR-WP4-DAMSYM-D4.2 Version number: 01.00 Status: FINAL Last revision date: DD/MM/YYYY by: UNIPA Verification date: DD/MM/YYYY by: Board Approval date: DD/MM/YYYY by: MUR Subject: IR00000014 - ITSERR D4.2 - Training material, manual and documentation for RESILIENCE Filename: ITSERR_WP4_Deliverable4.2_FINAL.docx
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 This document is available in the ITSERR WP Management document repository at: Change history Version Numbe r Date Status Summary of main or important changes 00.01 03/02/2025 WORKING Working version 00.02 01/07/2025 DRAFT Contributions collected by WP4 members 00.03 14/07/2025 DRAFT First revision - Elia Scapini 00.04 19/07/2025 DRAFT Second revision - Marianna Napolitano, Elia Scapini 00.05 21/07/2025 DRAFT Submission to reviewer 01.00 11/12/2025 FINAL integration of the reviewer's observations and finalization of the document Distribution List Name Company Role ITSERR Members All ITSERR partners 2
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Table of Contents Document Overview...............................................................................................................5 Scope................................................................................................................................. 5 Objectives and content.......................................................................................................5 Structure.............................................................................................................................6 Team...................................................................................................................................6 Training package....................................................................................................................8 Latin and Ancient Greek...........................................................................................................8 Tracing ideas instead of words...........................................................................................8 What corpora are included?......................................................................................... 8 What will this guide cover?...........................................................................................8 How to use the tool 1: theory............................................................................................10 How to use the tool 2: use-case samples.........................................................................11 a) Guided Test Workflow.............................................................................................11 b) Free Exploration workflow......................................................................................12 Conclusion (Latin and Ancient Greek)..............................................................................15 Old Church Slavonic (OCS) and Church Slavonic.................................................................16 Introduction: The Context Behind the Development of the Tools.....................................16 User Manual 1..................................................................................................................18 1. Technical Description of the PUA Handling System (Tool for Data Processing – Text Normalisation Tools)........................................................................................... 18 Use-case sample..............................................................................................................21 Arabic.....................................................................................................................................26 Description of the tool: Arabic Semantic Retrieval........................................................... 26 User Manual.....................................................................................................................27 Use-case sample..............................................................................................................29 Sanskrit.................................................................................................................................. 31 Description of the tools.....................................................................................................31 What corpora are included?............................................................................................. 32 What will this guide cover?...............................................................................................32 User Manual (Sanskrit Search Engine)............................................................................33 How to use the tool 1: theory......................................................................................34 How to use the tool 2: use-case samples...................................................................34 User Manual (Sanskrit Semantic Retrieval)..................................................................... 35 How to use the tool 1: theory......................................................................................35 How to use the tool 2: use-case samples.........................................................................37 Below is an example of semantic retrieval based on the 1st verse of the Nicene-Costantinopolitan Creed in Sanskrit:..............................................................38 Conclusion (Sanskrit tools)...............................................................................................39 Conclusion............................................................................................................................40 3
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 4
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Document Overview Scope The global purpose of Work Package 4 (WP4) DaMSym (Data Mining: the Nicene-Constantinopolitan Symbol) is the semantic analysis of key terms found in our case study, namely the text of the Nicene-Constantinopolitan Creed in its various translations. The ITSERR project envisions that advanced knowledge for the history of religious studies will emerge from these studies, but also that new purposely developed cutting-edge IT tools will be made available to fellow scholars for extracting information from the ancient texts under consideration. Language models, particularly those based on Large Language Models (LLM), can be used to analyze the text of the Nicene-Constantinopolitan Creed in order to identify significant semantic patterns. These models will be optimized (e.g., fine-tuning operation) over a wide range of Christian texts for the different translation languages under consideration. To analyze the text of the Nicene-Constantinopolitan Creed for the purpose of identifying significant semantic patterns, rule-based or probabilistic architecture tools can be used. In particular, Large Language Models (LLMs) may eventually be finetuned to perform a wide range of semantic operations on the Christian texts for the various translation languages under consideration whose amount of data currently available allows it. With these models it will be possible to obtain advanced semantic representations to facilitate the identification of relationships, similarities and deep meanings for key terms in the Creed. Data mining will then be applied to this enriched corpus to extract detailed information about semantic relationships within the Creed, helping us to better understand the theological and historical significance of each term in the various translations and interpretations of the Nicene-Constantinopolitan Creed. Objectives and content This document is jointly written by WP4 members to detail the achievements goaled in their work prompted by the third deliverable (number D4.2, preceded by D4.1.1 and 4.1.2). This is the deliverable from which the title of the present document is derived: Whitepaper on the results using DaMSym on Nicene-Constantinopolitan Creed. In its original form this deliverable was composed in the following way: The experience and the feedback collected during the development and testing of DaMSym platform, and also during the preparation of D4.1.1 and D4.1.2 is systematised and integrated in this deliverable, in the form of a training package to be integrated with the RESILIENCE infrastructure. The package includes use-case samples (drawn from D4.1.2), as well as templates. To align the deliverable with the updated workflow and research path, D4.2 has been reorganized as follows: The experience and the feedback collected during the development and testing of DaMSym platform, and also during the preparation of D4.1.1 and D4.1.2 is systematised and integrated in this deliverable, in the form of a training package to be integrated with the RESILIENCE infrastructure. The package includes use-case samples (drawn from D4.1.2), as well as templates. 5
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 The global outcome expected from the quoted Deliverable 4.2 is therefore a training package that will be included in the marketplace of the project to allow scholars interested in the tools of DaMSym prompted by WP4 to be able to exploit them the best way. The submitting deadline was also updated to 21 July 2025. Structure Since WP4 takes into consideration several historical languages of translation of the Nicene-Constantinopolitan Creed and also in order to best comply with the requirements of Deliverable 4.2, we envision this document as a succession of four chapters each one dedicated to one or more tools for the same linguistic and technological domain. Thus, a first chapter for Latin and Ancient Greek is followed by a second for Church Slavonic, and yet another two chapters are dedicated to Arabic and Sanskrit. For each one of these chapters the reader will find three paragraphs: a general description of the tool, a theoretical user-manual and finally one or more use-cases to follow in the testing of the tool. Overall, across this user guide, query and workflow templates to learn how to use the tool are provided in the form of an introductory modular training package. Each module includes 1) a conceptual description; 2) a step-by-step workflow; 3) a guided test and a free exploration space for personal familiarization with the tool. Team Marianna NAPOLITANO Unimore Old Church Slavonic Usman NAWAZ Unipa – Liliana LO PRESTI Old Church Slavonic Ivana PANZECA Unipa Arabic Giovanni PUCCETTI ISTI-CNR Arabic/Old Church Slavonic/Greek/Latin Maria CASSESE ISTI-CNR Old Church Slavonic Igor SPANÒ Unipa Sanskrit Irfan ALÌ Unipa – Marco LA CASCIA Sanskrit Federico IEZZI Unimore Latin and Greek Elia SCAPINI Unimore Latin and Greek Costanza BIANCHI Unimore Coptic Fabio QUATTRINI Unimore - Silvia CASCIANELLI Coptic Lorenzo BIANCHI ISTI-CNR - Fabrizio SEBASTIANI Coptic The WP Leader is Fabrizio D’Avenia, while Costanza Bianchi (post-doc researcher - FSCIRE) and Marianna Napolitano (RTD-A -Unimore) are the Product owners. The team directly involves 4 Phd Candidates: Federico Iezzi (Unimore), Elia Scapini (Unimore), Usman Nawaz (UniPa) and Irfan Alì (UniPa). Ivanza Panzeca (UniPa) and Igor Spanò (UniPa) are the researchers (RTD-A) included in our work package. 6
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 As of Oct. 29, the team expanded for the Slavic section with the participation of the Centro Nazionale (ISTI-CNR, Pisa), represented by Maria Cassese, Giovanni Puccetti and Fabrizio Sebastiani. Activities for Latin and Greek languages are fully covered by CNR-ISTI, Pisa, but the preparatory work for this deliverable - specifically, until December 2024 for Greek and until June 2025 for Latin-was covered by WP6. Humanities research for Church Slavonic is supported by WP3 (Criterion). Regarding Coptic, due to the scarcity of data, the work is covered in a partial way and is supervised by an assignee from the University of Modena and Reggio Emilia. 7
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Training package Latin and Ancient Greek Tracing ideas instead of words DaMSym is a deep neural networks based tool that provides retrieval functionalities from a database of Ancient Greek and Latin text. It provides advanced text search capabilities that go far beyond traditional keyword or string-based retrieval systems. What makes DaMSym unique among state-of-the-art retrieval tools is its concept-oriented design: rather than retrieving exact words or lemmatized forms, DaMSym retrieves meaning. This fundamental shift, from verbatim string retrieval to correlated semantic concept retrieval, allows scholars to explore texts in a more intelligent, interpretive, and context-aware way. The purpose of this tool is to allow researchers to retrieve sentences that, despite lexical, grammatical and syntactical differences, share a very related meaning and, for this reason, may have the same relevance for the purposes of the scholar. At the core of DaMSym is a semantic search engine powered by sentence embedding models trained specifically on Ancient Greek and Latin corpora. Each sentence in the database is converted into a numerical representation, an embedding, that captures its contextual features in a multidimensional vector space. The user’s query, whether in Greek or Latin, is embedded in the same way, and the system returns those sentences whose semantic content is most closely related in the vector space to the query. This means DaMSym can surface passages that express a given idea without requiring any lexical overlap with the user’s query. For example, a Greek query about "divine generation" may retrieve Latin phrases from theological or conciliar sources that express the same doctrinal content, even if no surface-level terms are shared. DaMSym was designed with researchers in mind, especially those working on ancient text, classicists, but also patristics, conciliar texts, and early Christian literature. The tool can be used to: ● Discover parallels across time, languages, or traditions; ● Explore semantic variants in the transmission of speculative concepts; ● Identify intertextual connections and influence between authors; ● Conduct cross-lingual comparisons of meaning (e.g., between Greek and Latin formulations); ● Support philological and historical analysis. What corpora are included? DaMSym operates over curated corpora of Ancient Greek and Latin texts. The texts have been preprocessed, normalized, embedded using models fine-tuned for classical languages and indexed. The system supports retrieval in either language and can surface results in a cross-linguistic fashion. What will this guide cover? This tutorial will walk you through: 8
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 ● Getting familiar with the tool ● How to formulate a semantic query in Ancient Greek or Latin; ● How to obtain ranked retrieval results; ● How to integrate context in your query; ● How to use advanced options such as cross-language retrieval, filtering (e.g., by author, or work), sum of queries. DaMSym is a conceptual exploration engine, designed to help scholars go beyond surface forms and uncover the deeper connections of thought, doctrine, and language across the long arc of classical and religious history. Whether you're tracing a theological formula, analyzing symbolic expression, or investigating the development of an idea across centuries and cultures, DaMSym offers a powerful and flexible interface for semantic research in the ancient world. The ultimate goal of this training course is to help the user to: ● Develop skills for mastering the retrieval tool in its core functionalities ● Gain the essential knowledge to be fully familiar with the tool and to obtain the most from it ● Set out the proper attitude toward the deep learning based retrieval so to fine-tune the proper expectations from it Therefore, the intended audience of this tutorial primarily consists of researchers working with ancient sources in Latin and Ancient Greek. No advanced knowledge of computer science or informatics is required to benefit from the tutorial. However, a basic familiarity with deep learning and artificial intelligence systems may assist users in setting appropriate expectations, neither overly optimistic nor unduly modest, regarding the tool’s capabilities. This tool, together with the present tutorial, is especially designed for scholars who already make use of established resources for navigating ancient texts, such as Corpus Corporum and the Thesaurus Linguae Graecae. It is offered as a complementary aid for those who wish to enhance their research with a more flexible, semantic-oriented retrieval method, particularly for exploring texts and connections that may remain inaccessible within the constraints of rule-based or lemmatization-dependent systems. By the end of this tutorial, the reader should be able to obtain the best from the retrieval tool in terms of retrieval precision and efficiency. With that, the reader should also be able to formulate the query (or the queries) in the most appropriate way to navigate the corpus with a deep retrieval system. This self awareness in the navigation of the corpus should also ensure that the reader will be able to lower down or increase his/her expectations to the most realistic level. Readers will familiarize themselves with the tool in theory and in practice, taking advantage of both general instructions and actual steps to apply while testing the tool. By the end of this tutorial readers will have practiced themselves with all the features of the tool and will be able to autonomously exploit it the proper way in their work. The tutorial is composed by two parts that together should demand a training effort of around 20 minutes from the reader. The first part is a more theoretical one and allows the reader to familiarize with the core concepts, fields and terminology that compose the tool main functionalities. The second part is a practical vademecum that the user will follow to learn how to exploit the tool with proper use cases. The tutorial can be performed alone and does not require any previous knowledge 9
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Old Church Slavonic (OCS) and Church Slavonic Introduction: The Context Behind the Development of the Tools The guide presented in the following sections aims to address one of the main challenges encountered during the selection of sources for the creation of a dataset, a diachronic reference corpus that would bring together diverse sources in a unified and standardized format. In fact, Natural Language Processing (NLP) for Old Church Slavonic and Church Slavonic texts present specific challenges: these languages are under-resourced therefore the development of a comprehensive corpus capable of supporting linguistic analysis involves complex processes. In particular, two key challenges can be distinguished: (a) the data is scattered across multiple archives and online repositories that provide access to digitally transcribed texts; (b) many of the online texts rely on the use of Private Use Area (PUA) characters rather than standard Unicode. The decision to avoid Unicode encoding often reflects a deliberate attempt to restrict access to and reuse of the available texts. This issue however is not limited to Old Church Slavonic resources but affects a range of both ancient and modern languages. Our current research aims to address these limitations. The first part of the project focused on the creation of a dataset collecting texts that were already available online. This decision was primarily driven by the challenges involved in training a model for OCR and HTR on the digitized manuscripts currently accessible. Developing such a model was not considered a viable option due to the substantial volume of training data required to achieve an acceptable improvement curve in digitization accuracy. The correction and post-editing phases would have demanded extensive manual revision work. We then focused on selecting the most suitable source formats, prioritizing .txt and .html files for their ease of processing and integration into a unified format. The resulting dataset, discussed in Deliverable 2 and detailed in the forthcoming paper “DIACU: A Dataset for the DIAchronic Analysis of Church Slavonic” (Maria Cassese, Giovanni Puccetti, Marianna Napolitano, Andrea Esuli), to be presented at the next Slavic NLP (Vienna, 31 July 2025) and published in the ACL Anthology, organizes the texts by identifying, where available, the original title, the transliterated version, and the Latin title, along with metadata such as century, area of preservation and publication, historical and regional variants, and source information. The dataset includes a total of 652 texts spanning from the 10th to the 18th century, all in Cyrillic script. The texts are classified into four linguistic categories. The selection of texts was also influenced by the availability of appropriate fonts that allowed the materials to be read. As previously mentioned, a common issue affecting both Old Church Slavonic 16
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 and Church Slavonic sources is the frequent use of PUA characters. For further details on the use of PUA characters, please refer to Deliverable 4.1.2. A classifier was developed in conjunction with the dataset to perform supervised classification tasks of texts according to four historical periods from the IX to the XVIII century, corresponding to four linguistic variants: Old Church Slavonic (OCS): 9th – 11th century; Church Slavonic (CS): 12th – 17th century (with different revisions: Bulgarian, East Slavic, Serbian); New Church Slavonic (NCS): 18th century; Ruthenian, ruska mova (Rut): 15th – 18th century), each corresponding to a specific language variety. Each language includes several historical and regional variants, which are also taken into account in the classification process. In addition to the three officially recognized liturgical languages, Ruthenian has also been included. This choice is motivated by two interrelated factors: a) The Ruthenian corpus, derived from sources available through Universal Dependencies, is the most extensive within the dataset, comprising 344 documents; b) The overarching goal of this project is to trace semantic relations among words and phrases, and to investigate their diachronic evolution. Including texts from a different geographical area—particularly those used as a vernacular language—proves valuable in capturing these shifts and in understanding the influence of cultural factors on semantic variation. The source code of this tool is available and has already been included in Deliverable 4.1.2. It will be reported also at the end of this paragraph. Nonetheless, the user manual for this tool itself is not part of this deliverable, as its operation requires the presence of an IT professional, nevertheless all information to use the tool is made available together with the code itself. In this section, we provide a brief user guide to the tool developed in the context of the research and selection of sources to be included in the dataset, as well as in relation to the metadata definition process. The tool introduced here is a converter designed to generate texts encoded in Unicode, thus avoiding the use of unreadable PUA characters. The version presented in this short user guide was developed on the basis of PUA-to-Unicode mapping tables, specifically targeting fonts used on the Cyrillomethodiana web portal (https://histdict.uni-sofia.bg/). This is a preliminary version, which is expected to be extended in the future to incorporate additional mapping schemes. The guide is intended for students and researchers in the fields of Slavic philology, theology and history, as well as for specialists in Natural Language Processing (NLP). A working knowledge of at least one Slavic language is nevertheless required in order to verify the accuracy of the: ● results of the converter and to ensure that letters and diacritical marks are not misrepresented or incorrectly mapped to other symbols. ● connection between the chronological placement of a text and its linguistic and orthographic features. This may be particularly helpful when the exact date of a document is unknown. In such cases, the classifier can serve as a supportive method for validating existing hypotheses about the text’s dating based on its linguistic and grammatical characteristics. 17
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 User Manual 1 1. Technical Description of the PUA Handling System (Tool for Data Processing – Text Normalisation Tools) The PUA Handling System for Old Church Slavonic is a Windows-based utility developed in the framework of WP4, with specific reference to the task of text normalisation.Its primary aim is to convert textual data encoded using PUA (Private Use Area) characters into standard Unicode representations. This supports data harmonisation and machine-readability within the broader textual processing workflow of the project. As discussed in Deliverable D4.1.2, the use of non-standard fonts and encoding schemes, especially in the context of Church Slavonic digital sources, represents a major obstacle to automated analysis and integration into digital corpora. This tool addresses that challenge by providing a fast, scalable, and user-friendly solution for PUA-to-Unicode conversion. System Requirements and Deployment ● Operating System: Windows (installation required) ● Distribution Format: Installer package (.exe) ● Installation: The application must be installed locally in order to function correctly. It cannot be run as a portable executable. Installation Instructions: Download the installer file and double-click to begin the setup process. If prompted by Windows SmartScreen, click “More info” followed by “Run anyway” to proceed with the installation. Once installed, the application can be launched from the Start menu or desktop shortcut. The main interface window is titled “PUA Handling System for Old Church Slavonic”. Functional Overview The tool offers two key functionalities: Single Text Conversion This feature allows for the manual conversion of short PUA-encoded text segments: 1. Select mapping database (e.g., Database 1, Database 2), corresponding to a predefined PUA-to-Unicode scheme. 18
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 2. Paste input: the PUA-encoded text is entered in the top input box. 3. Run conversion: clicking Submit applies the mapping and displays the Unicode output in the Processed Text box. 4. Copy result: standard keyboard shortcuts (Ctrl+A, Ctrl+C) can be used to extract the converted text. 19
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 This mode is especially useful for testing mappings, performing quick validations, or processing small textual units manually. Batch Processing (Folder Conversion) For large-scale operations, the tool supports batch conversion of entire folders of PUA-encoded files: 1. The user initiates the process by clicking “Select folder and process”. 2. A first dialog window is used to choose the source folder with PUA text files (documents should be processed as .txt files). 3. A second dialog allows for the selection (or creation) of an output folder. 4. Once the conversion is complete, a confirmation pop-up appears. All processed files are saved in .txt format in the output folder. This function is tailored for the normalisation of the entire corpora and is aligned with the data preparation needs identified in WP4. Implementation Basis and Future Developments The tool relies on an initial set of PUA-to-Unicode mapping tables, developed from fonts used on the Cyrillomethodiana web portal. This implementation constitutes a first version, with the following extensions foreseen in the next development phases: ● Inclusion of additional mapping tables from diverse sources; ● Expansion of coverage to rare glyphs and diacritics; ● Cross-platform implementation (e.g., Linux, macOS); ● Development of a CLI (command-line interface) for integration in NLP pipelines and automated workflows. Conclusion and Contribution to WP4 Objectives The PUA Handling System is an important component of the toolkit foreseen in WP4. By enabling the standardised conversion of encoded texts, it facilitates the transition from heterogeneous, legacy data formats to interoperable and analysable text resources. It supports both manual and large-scale processing scenarios, and contributes directly to the project’s goal of creating robust digital infrastructure for (Old) Church Slavonic textual analysis. 20
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Use-case sample This use case illustrates the application of the PUA Handling System to textual materials extracted from the Cyrillomethodiana web portal. The primary aim is to demonstrate how texts encoded with non-standard fonts using Private Use Area (PUA) characters can be converted into Unicode-compliant format, enabling downstream processing and integration into the project’s dataset. 2. Context and Source Description The Cyrillomethodiana (https://histdict.uni-sofia.bg/) portal provides access to digitised liturgical and theological texts in (Old) Church Slavonic, many of which are rendered using custom fonts that encode characters in the PUA range of Unicode. These encodings are not machine-readable and require custom conversion logic to ensure standardisation. For this use case, a set of texts was manually downloaded in .txt format from the portal. Visual inspection of the files confirmed that the texts rely on a proprietary font with extensive PUA usage. 3. Mapping Strategy As no existing mapping tables were available, a custom PUA-to-Unicode mapping was created manually through the following process: ● Glyphs were visually compared using other available maps of PUA-Unicode conversion ● Each PUA code point was matched with its intended Unicode equivalent ( using character charts, reference grammars, and Church Slavonic orthographic conventions.A brief excerpt follows: 21
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Context The source text, Пандекты Антиоха – Pandectes Scripturae divinitus inspiratae sancti patris Antiochi, is a corpus of 130 chapters that circulated widely in the Byzantine Middle Ages. This is evidenced by its extensive Greek manuscript tradition and by the Slavic version, whose earliest known witness is the 11th-century manuscript Moscow, State Historical Museum, Voskresenskij 30P, as well as by numerous other contemporary sources. It saw its first printed edition in the Latin version, published in Paris in 1543 by Godefroy Tilmann (†1562). The version published on the HistDict portal is accessible through a web interface and can be copied manually. Like many digitised texts from similar projects, it may contain non-standard character encoding, including PUA (Private Use Area) characters, especially for Cyrillic glyphs and diacritics not supported in standard fonts. Conversion Procedure 1. Open the PUA Handling System. 2. Select the appropriate mapping (HistDictMap_v1). 3. Paste the copied text from the HistDict portal into the input field. 4. Click Submit to convert the text. 5. Copy the result from the output field. Original (PUA-encoded, copied from site): 22
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Results ● The text is successfully normalised and machine-readable. 23
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 For example, in this case, the main unrecognised characters were those shown below, which the system was able to identify thanks to the implemented mapping. The output changes when a text from a different website is used. In this instance, it was not possible to find a version of the text available on the Cyrillomethodiana portal that was also published on other platforms, thus preventing a direct comparison. However, it is evident, by taking a fragment from the Psalter available on the Azbuka.ru portal, that the PUA characters fail to be recognised. In this case, bibliographic information related to the edition is not available, due to staff-related constraints on the part of the website's administrators. However, future developments of the converter aim to incorporate additional mapping tables in order to address this gap more comprehensively. 24
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Classifier The dataset and the code of the experiments are available at https: //github.com/MariaCassese/DIACU 25
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 This means the Sanskrit Semantic Retrieval can surface passages that express a given idea without requiring any lexical overlap with the user’s query, and even if no surface-level terms are shared. The Sanskrit Semantic Retrieval is a conceptual exploration engine, designed to help scholars go beyond surface forms and uncover the deeper connections of thought, doctrine, and language across the long arc of textual history. Whether one is engaged in the analysis of the ritualistic density underlying formalized expressions, the interpretation of symbolic language, or the diachronic study of conceptual developments, Sanskrit Semantic Retrieval provides a robust and versatile framework for conducting semantic investigations into the intellectual traditions of South Asia expressed in Sanskrit literature. Furthermore, it is essential to emphasise that Sanskrit Semantic Retrieval, already tested and currently in the prototyping phase, not only allows scholars to retrieve semantically related passages from a corpus of Sanskrit texts but also allows the user to designate a target word within the input phrase, thus refining and focusing the search according to specific research needs. The Sanskrit tools were designed with researchers in mind, especially Indologists whose approach is philological and historical at the same time. The tools can be used to: ● Discover parallels across time, languages, different religious texts (Vedic, Hindu, Buddhist, etc.), and different textual genres; ● Explore semantic variants in the transmission of speculative concepts; ● Identify intertextual connections and influence between different works; ● Support philological and historical analysis. What corpora are included? The Sanskrit Semantic Retrieval and the Sanskrit Word Search Engine operate over a corpus created from the Digital Corpus of Sanskrit (DCS) repository, whose texts comprise a very large range of periods, from the Vedic period to the Modern epoch. We reprocessed the content of this dataset and removed all unwanted data from the dataset and all unnecessary information. What will this guide cover? This tutorial will walk you through: ● Getting familiar with the tools ● How to formulate a lexical query (Sanskrit Word Search Engine) or a semantic query (Sanskrit Semantic Retrieval) in Sanskrit; ● How to obtain ranked retrieval results (with the Sanskrit Semantic Retrieval); ● How to integrate context in your query; ● How to use advanced options such as filtering (e.g., multiple word search, by period or type of work, or by focused research). The ultimate goal of this training course is to help the user to: ● Develop skills for mastering the Sanskrit Semantic Retrieval and the Sanskrit Word Search Engine in their core functionalities ● Gain the essential knowledge to be fully familiar with the tools and to obtain the most from them 32
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 ● Set out the proper attitude toward the deep learning-based retrieval to fine-tune the proper expectations from it The present tutorials are primarily intended for scholars and researchers engaged in the study of Sanskrit texts. While no advanced background in computer science or informatics is required to benefit from the material, a general familiarity with the foundational principles of artificial intelligence—particularly deep learning—may assist users in developing appropriately calibrated expectations regarding the tool’s capabilities, avoiding both undue optimism and unwarranted skepticism. These tools, along with the present instructional guide, have been developed with the specific needs of philologists and textual scholars in mind—particularly those already working with established digital resources such as the Digital Corpus of Sanskrit (DCS), the Göttingen Register of Electronic Texts in Indian Languages (GRETIL), or the Thesaurus Indogermanischer Textund Sprachmaterialien (TITUS). The Sanskrit Semantic Retrieval interface is offered as a complementary, semantically-oriented tool to enhance research workflows, especially in those areas where rule-based or lemmatization-dependent systems may fall short—such as the exploration of implicit conceptual linkages, non-canonical terminology, or syntactically opaque constructions. Upon completion of the tutorial, users should be able to formulate effective queries tailored to their research objectives and the particular architecture of the corpus under examination. They will acquire the skills necessary to navigate both traditional lemmatization-dependent platforms and systems based on deep retrieval, with a refined understanding of each system’s affordances and limitations. This methodological self-awareness will enable scholars to modulate their expectations realistically and to evaluate retrieval outcomes with critical insight. The tutorials are organized into two integrated parts, requiring approximately 20 minutes of active engagement. The first component introduces the theoretical foundations and terminological framework necessary for understanding the core functionalities of the tools. The second part functions as a practical vademecum, guiding users step-by-step through representative use cases to facilitate hands-on familiarity with the system’s features. No prior technical training is assumed beyond the standard competencies expected of researchers already accustomed to navigating digital Sanskrit resources. Finally, as the tools are fully web-based, no local installation or software dependencies are required. Users may access the platform simply by logging into the ITSERR marketplace, thereby ensuring ease of adoption and broad accessibility within the research community. User Manual (Sanskrit Search Engine) This user manual provides step-by-step instructions for using the Sanskrit Search Engine tool and its interface. The tool is designed to retrieve words from a corpus and features an advanced filtering function not available in existing tools. This filter allows users to narrow their searches to either 33
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Vedic and ritual texts or non-Vedic, later-period texts. Additionally, the retrieved words are displayed along with their relative positions within the sentence, providing deeper understanding. This guide will take you through the setup, guided testing, understanding of the functionalities, and evaluation of the tool. How to use the tool 1: theory Step 1: Accessing the Tool ● Open the Interface: Navigate to the tool’s web-based interface using the provided link. ● Log In: If required, enter your credentials to access the tool. For testing purposes, use the temporary credentials provided. ● Select Corpus: Choose between Vedic and ritual texts or non-Vedic/all other texts from the dropdown menu. ● Proceed: Click on this button to proceed to search for words. Step 2: Querying the corpus for word(s) search ● How many words to search for: Enter the number of words to search for (e.g, 1,2,3, etc.,) then hit the proceed button. ● Enter words to search: Write the desired word(s) as indicated above. ● Press the “start search” button to see the results. Step 3: Navigate the results ● The results are listed at the bottom of the page and at the top of the page, you will also find the words searching for and words found/not found in the corpus. ● Choose one of the two options: Select Yes (search again), this will take you to the step 1 option ‘Select Corpus’ then follow the same process as discussed above and if select No(Exit) option then the tool will be closed. How to use the tool 2: use-case samples The guided test helps you familiarize yourself with the tool by following a test case. Here we show the layout of the tool, researchers may want to search in a corpus easily. 34
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Once familiar with the basic working of the tool, you can conduct independent searches based on your research goals for a varying number of words in the corpus. User Manual (Sanskrit Semantic Retrieval) How to use the tool 1: theory This user guide offers step-by-step instructions for using the transformer-based tool, Sanskrit Semantic Retrieval, and its interface, which is designed to analyze semantic textual similarity within ancient Sanskrit corpora. The tool utilizes multidimensional vector embeddings to support advanced semantic searches, enabling users to explore conceptual and meaning-based connections that go beyond exact wording matches. This guide will take you through the setup, guided testing, free exploration, understanding of the functionalities, and evaluation of the tool. 35
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Step 1: Accessing the Tool ● Open the Interface: Navigate to the tool’s web-based interface using the provided link. ● Log In: If required, enter your credentials to access the tool. For testing purposes, use the temporary credentials provided. Step 2: Querying the Sanskrit Corpus for Semantic Similarity Search ● Enter Sanskrit Query Sentence: Copy and paste the sentence from the corpus containing the semantic features you wish to search for into the designated input box labeled Sanskrit Query Sentence. ● Enter Focus Word: In the Focus Word box, you may enter a specific Sanskrit word that you want the search to emphasize. This helps refine the semantic search around a central concept or keyword. Choose Search Mode: ● Select With Focus Word if you want the search engine to prioritize results based on the given focus word. ● Select Without Focus Word to perform a broader semantic similarity search based only on the query sentence. ● Submit the Query: Once all desired inputs are provided, click the Submit button to initiate the search. The system will process the input using its semantic model and display the most relevant results in the Results pane on the right. Step 3: Navigate the results ● Review the Results: The right-hand panel labeled Results will display the top semantically related Sanskrit sentences retrieved from the corpus based on your query. Each result is presented along with the name of the source text, indicating the origin of the sentence within the corpus. This feature allows users to not only compare semantic similarity but also understand the textual context and source of each match. ● Flagging Output (Optional): If a result appears incorrect or irrelevant, you can click the Flag button below the results panel to report it for review or refinement. ● Clear the Fields: To reset the input fields and start a new query, click the Clear button. 36
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 How to use the tool 2: use-case samples Guided Test Workflow The guided test helps you familiarize yourself with the tool by following a test case. Here we provide a sample layout of the tool to illustrate how researchers can easily search the corpus. Sanskrit nainaṃ sarpo na pṛdākur hinasti na vṛściko na tiraścīnarājiḥ nainaṃ sarpo na pṛdākur hinasti na vṛściko na tiraścīnarājiḥ With Focus Word : vṛścik Without Focus Word : --- Copy the Sanskrit sentence from the above box and paste it in the query box and select with focus word or not focus word. Click the “Submit” button to run the search across the entire corpus. The tool will then retrieve all sentences that are semantically similar to your query, along with the source text. 37
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Below is an example of semantic retrieval based on the 1st verse of the Nicene-Costantinopolitan Creed in Sanskrit: We used as input for our research the Sanskrit translation of the Creed published in 1823 by the Anglican Reverend William Hodge Mill. Textual comparison, philological analysis, and examination of the historical and religious context enable us to draw some conclusions that have been inspired by the results of our tool, and which further deepen those findings. The first half of the first verse of the Creed in Sanskrit reads: devam ekam mahātātaṃ sraṣṭāram bhūminākayoḥ |. The first two results reported by our Semantic Retrieval are two verses taken from the Liṅgapurāṇa and the Sātvatatantra, respectively: 1) śuklasyāpy ammayaṃ śuklaṃ padaṃ ṣoḍaśaraśmivat || (LiṅgaP I, 61, 23); 2) jambūdvīpasamaḥ sraṣṭā pīyūṣotpattikāraṇam || SātvT VI, 60. A first observation is that these results do not concern the famous Vedic cosmogonies, but significantly the Purāṇic cosmogonies and theistic texts. Here, bhakti, that is the attitude of “participation” of the faithful towards the deity, has taken the place of both the more ancient Vedic religiosity. 38
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Conclusion (Sanskrit tools) This user manual provides a comprehensive guide to using the Sanskrit Search Engine and the transformer-based Sanskrit Semantic Retrieval tool for exploring words with locations and semantic similarity within ancient Sanskrit corpora. By following the outlined steps from query formulation to interpreting results you can effectively navigate the tool’s interface, perform meaningful searches, and analyze conceptual connections across texts. Your constructive feedback is vital to refining the tool and expanding its relevance for the study of Sanskrit texts. 39
Ir00000014 - Itserr Status: FINAL ITSERR-WP4-DAMSYM-D4.2 Version: 01.00 Conclusion The ITSERR WP4 Deliverable 4.2, "Training material, manual, and documentation for RESILIENCE," provides a comprehensive guide to some of the advanced tools developed under the DaMSym platform for semantic analysis of historical religious texts. This document serves as a resource for scholars working with ancient languages such as Latin, Ancient Greek, Old Church Slavonic, Arabic, and Sanskrit, offering detailed user manuals, practical use-case samples, and theoretical foundations for each tool. The tools introduced—DaMSym for Latin and Greek, the PUA Handling System for Old Church Slavonic, the Arabic Semantic Retrieval, and the Sanskrit Semantic Retrieval and Word Search Engine—represent significant advancements in the field of digital humanities. By leveraging cutting-edge technologies like deep neural networks and semantic embedding models, these tools enable researchers to move beyond traditional keyword-based searches and explore textual corpora at a conceptual level. This shift allows for the identification of semantic parallels, intertextual connections, and doctrinal variations across languages and historical periods, enriching the study of religious and philosophical texts. The training materials included in this deliverable are designed to empower scholars with the skills needed to harness these tools effectively, regardless of their technical background. Through guided workflows and free exploration exercises, users can familiarize themselves with the functionalities, refine their queries, and interpret results with confidence. The emphasis on practical application ensures that the tools can be seamlessly integrated into existing research methodologies. As the ITSERR - WP4 DaMSym project progresses, the tools and methodologies outlined here will continue to evolve, incorporating user feedback and expanding their capabilities. The ultimate goal is to foster a collaborative environment where scholars can access, utilize, and contribute to a growing digital infrastructure for the study of historical religious texts. This deliverable not only marks a milestone in the project’s development but also lays the groundwork for future innovations in semantic analysis and digital philology. In closing, Deliverable 4.2 underscores the transformative potential of combining humanities research with advanced computational tools. By bridging the gap between traditional scholarship and modern technology, the WP4 tool open the way for new discoveries and deeper insights into the rich tapestry of religious and cultural history. Other tools developed within the framework of the Work Package and additional training based on the improvements made to the tool during this period will be presented in Deliverable 4.3: The analysis of data resulting from the application of DaMSym tools to the various corpora is published, powered by DaMSym platform for browsing, consultation and user contribution, through RESILIENCE. On the one hand, this presents to the scholarly community a powerful prototype for producing new research acquisitions and exploring the connections between different religious texts; on the other hand, this displays the potentialities of DaMSym tools, setting a new standard for tools dedicated to Religious Studies and, potentially, also for other domains. 40