Representing Texts as LOD: a Systematic Literature Review
Abstract
The presentation was delivered at the XIII AIUCD National Conference (AIUCD 2024), held in Catania, and presents a systematic literature review on the representation of texts as Linked Open Data (LOD) in the Digital Humanities. It examines existing models, formats, and levels of granularity used to publish textual corpora on the Semantic Web, discussing current approaches and open issues. Based on the results of this work, a paper was published in the conference proceedings and a journal article in Umanistica Digitale. The study is carried out within the framework of the ItAnt – Languages and Cultures of Ancient Italy project.
Full text
Representing Texts as LOD: a Systematic Literature Review Michela Bandini & Valeria Quochi CNR-ILC Pisa
➔Significant growth in LOD interest in digital humanities over the last decade. ➔The majority of linguistic resources in the LLOD cloud are dictionaries, lexica, thesauri, terminologies, and controlled vocabularies. ➔The publishing of textual corpora as LOD is still debated and controversial. ➔Some recent projects have started to explore different approaches to converting into and representing textual corpora as LOD. ➔A systematic literature review to assess benefits and identify existing models to represent textual corpora as LOD. ➔Further goal is to find the most suitable model for ancient language inscriptions for our ongoing project purpose 2 Why a systematic literature review?
Project ItAnt: Languages and Cultures of Ancient Italy. Historical Linguistics and Digital Models ❖Main goal: creation of an ecosystem of language resources for ancient “fragmentary” languages (Oscan, Neo Faliscan, Venetic, Celtic). ❖Very small number of texts, lack of knowledge, and epigraphic evidence. ❖The publication of text corpora (in TEI/EpiDoc) on the Semantic Web is still at an early stage: still deciding what approach/model to follow. 3
4 Systematic Literature Review Methodology 4 1. Define terms and questions 2. Define criterias 3. Screening 4. Quality assessment 5. Data analysis
➔What are the most relevant works that have already attempted to transform/represent (text) corpora as LOD? ➔What are the models and formats already in use by the digital humanities community to represent corpora as LOD? ➔How can we classify projects according to the model used for representation or the granularity of the representation? 5 Terms and Questions of the Research
❖Sources: regular/advanced search on digital libraries (e.g. ACL Anthology, DBPL, Google Scholar). ❖Authors: authors strictly related to LOD works on corpora (e.g. C. Chiarcos). ❖Keywords/seed terms: 3 seed keywords combined with other terms, multi-term search, extra keywords (e.g. "POWLA", "NIF”). ❖Filtering criteria: date, language, number of citations, kind of publications (e.g. 2000 and up, ita or eng, conference papers, etc.). 6 Criteria of the Research
7 Screening for Inclusion Title reading Abstract, intro & conclusion reading Deep reading Filtering criteria: We excluded papers that… ❖provided not interesting works or not relevant topics ❖provided papers related to theoretical topics as RDF or XML formats. Filtering criteria: We excluded papers that… ❖illustrated general and theoretical topics; ❖provided a LLOD-compliant model or activity clearly not related to texts. Filtering criteria: We excluded papers that… ❖targeted corpus-derived data represented as lexicons, or CSV / TSV data; ❖mentioned textual data, but dealt with platforms, websites, or ontologies. Results: 219 relevant works Saved on a Zotero Library Results: 136 relevant works Results: 77 relevant works EXAMPLE excluded: Zinn, C., Hinrichs, M., & Hinrichs, E. (2022). Adapting GermaNet for the Semantic Web. In Proceedings of the 18th Conference on Natural Language Processing (KONVENS 2022) (pp. 41-47). EXAMPLE excluded: Tittel, S., & Chiarcos, C. (2018). Historical lexicography of Old French and linked open data: transforming the resources of the Dictionnaire étymologique de l'ancien français with OntoLex-Lemon. EXAMPLE excluded: Van Assem, M., et all (2004). A method for converting thesauri to RDF/OWL. In The Semantic Web–ISWC 2004: Third International Semantic Web Conference, Hiroshima, Japan, November 7-11, 2004. Proceedings 3 (pp. 17-31). Springer Berlin Heidelberg.
8 Quality Assessment & Data Analysis Full-article Reading Data Classification Categorization according to 2 criteria: 1. level of granularity in data representation; 2. model and formats used for the representation.
9 Granularity of the Data Representation Textual data represented at a document level Partial representation of textual data Granular representation of textual data ❖data represented as bibliographic entities or cultural objects; ❖no representation of the textual data contained. ❖only some extracted text parts are represented as RDF triples (e.i. named entities). ❖data represented taking into account detailed info (e.i. sentences, tokens, morphological information, etc.).