scieee AI-readable full text Open interactive document viewer

How do we integrate community data from Wikidata and the fuzzy-sl Wikibase into a cultural heritage knowledge graph?

Stefan, Daria; Thiery, Florian; Schenk, Fiona

Abstract

Community-curated Wikibase ecosystems, most notably Wikidata, FactGrid, and wikibase.cloud instances such as fuzzy-sl, have become significant sources of Cultural Heritage (CH) and archaeological data. In parallel, infrastructure initiatives (e.g., NFDI4Objects) are building CIDOC-aligned knowledge graphs that demand robust provenance, semantic interoperability, and reproducibility. This paper presents a semi-automated workflow for integrating Wikibase data into infrastructure-scale graphs, including entity selection, ontology design aligned with CIDOC CRM (with CRMarchaeo, CRMsci, and CRMdig as needed), scripted RDF transformation, open publication of code and data, versioned snapshot releases, and ingestion into the NFDI4Objects Knowledge Graph. The approach preserves Wikidata-style qualifiers and references while yielding an event- and provenance-centric representation. Two use cases demonstrate feasibility and limits. The Irish Holy Wells dataset showcases richly reified statements (e.g., use, sources) mapped to CIDOC events. The Campanian Ignimbrite findspots in fuzzy-sl focus on location-centric modelling and interdisciplinarity, requiring E53 Place as baseline with CRMarchaeo/CRMsci specialisation for stratigraphy, sampling, and analysis. We discuss challenges in ontology alignment, granularity, explicit treatment of uncertainty (“fuzzy/wobbly” data), and sustaining semi-automated pipelines amid evolving community schemas. We argue for identifier discipline, machine-actionable provenance, and FAIR Digital Objects for each release. The outcome is an interoperable, federated ecosystem, spanning triplestores, Wikibases, and FDOs, in which community knowledge bases become partners of infrastructure graphs.

Full text

How do we integrate community data from Wikidata and the fuzzy-sl Wikibase into a cultural heritage knowledge graph? Daria Stefan* 1,2 , Florian Thiery* 2 , Fiona Schenk 3,2 1 TU Wien – Vienna, Austria 2 Research Squirrel Engineers Network – Mainz/Vienna, Germany/Austria 3 Johannes Gutenberg Universität Mainz – Mainz, Germany * Corresponding author(s) Correspondence: [email protected] ; [email protected] 1 2 3 4 5 6 7 8 9 10 11 12 A BSTRACT Community-curated Wikibase ecosystems, most notably Wikidata, FactGrid, and wikibase.cloud instances such as fuzzy-sl, have become significant sources of Cultural Heritage (CH) and archaeological data. In parallel, infrastructure initiatives (e.g., NFDI4Objects) are building CIDOC-aligned knowledge graphs that demand robust provenance, semantic interoperability, and reproducibility. This paper presents a semi-automated workflow for integrating Wikibase data into infrastructure-scale graphs, including entity selection, ontology design aligned with CIDOC CRM (with CRMarchaeo, CRMsci, and CRMdig as needed), scripted RDF transformation, open publication of code and data, versioned snapshot releases, and ingestion into the NFDI4Objects Knowledge Graph. The approach preserves Wikidata-style qualifiers and references while yielding an eventand provenance-centric representation. Two use cases demonstrate feasibility and limits. The Irish Holy Wells dataset showcases richly reified statements (e.g., use, sources) mapped to CIDOC events. The Campanian Ignimbrite findspots in fuzzy-sl focus on location-centric modelling and interdisciplinarity, requiring E53 Place as baseline with CRMarchaeo/CRMsci specialisation for stratigraphy, sampling, and analysis. We discuss challenges in ontology alignment, granularity, explicit treatment of uncertainty (“fuzzy/wobbly” data), and sustaining semi-automated pipelines amid evolving community schemas. We argue for identifier discipline, machine-actionable provenance, and FAIR Digital Objects for each release. The outcome is an interoperable, federated ecosystem, spanning triplestores, Wikibases, and FDOs, in which community knowledge bases become partners of infrastructure graphs. Keywords: Wikibase, Wikidata, Linked Data, Ontology, Archaeology, Geosciences 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 Introduction Community-driven repositories and knowledge bases – most prominently Wikidata, FactGrid, and Wikibase instances on wikibase.cloud such as fuzzy-sl – constitute substantial reservoirs of cultural heritage (CH) and archaeology-related data [1], [2]. Curated by volunteers, citizen scientists, and independent researchers, these platforms frequently close gaps left by project-bound initiatives and enrich the broader research ecosystem with timely, fine-grained observations and links to heterogeneous sources. In parallel, large-scale infrastructure efforts (e.g., within the German National Research Data Infrastructure, NFDI [3]) are building knowledge-graph-based services for feeding the interdisciplinary federated Knowledge Graph Ecosystem [4], [5] (Fig. 1), grounded in established semantic frameworks (RDF/OWL) and domain ontologies such as CIDOC CRM and its extensions. A key challenge – and opportunity – lies in integrating community-driven resources with these infrastructure graphs in a way that is methodologically sound, FAIR, and sustainable. Figure 1 - A distributed Knowledge Graph Scheme bringing together Linked Open Data and FAIR Digital Objects approaches. Florian Thiery & Andreas Noback, CC BY 4.0, via Wikimedia Commons. Within this landscape, Germany’s NFDI4Objects consortium [6] acts as an interdisciplinary hub for archaeology and cultural heritage, spanning an exceptionally broad temporal and material scope. The NFDI4Objects Knowledge Graph 1 already interlinks diverse datasets 2 (e.g., Linked Open Ogham, African Red Slip Ware, Linked Open Samian Ware), employing CIDOC CRM and selected extensions (e.g., CRMarchaeo, CRMdig) to enable interoperable querying across collections, sites, and research outputs 3 . Leveraging Linked Open Data (LOD) practices [7], [8] here is not merely a technical choice: it underpins reproducible scholarship, traceable provenance, and 3 cf. https://doi.org/10.5281/zenodo.13946052 2 cf. https://graph.nfdi4objects.net/collection/ 1 cf. https://graph.nfdi4objects.net/ 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 cross-dataset discovery. The question, then, is how to interface curated, evolving community platforms – each with its own modelling idioms and governance – with a harmonised, CIDOC-aligned knowledge graph at infrastructure scale. Our proposed approach is a semi-automated, six-step workflow that (1) identifies CH-relevant entities in Wikibase instances, (2) designs or selects an ontology aligned to CIDOC CRM (supplemented where appropriate), (3) transforms the source data into RDF via scripted pipelines, (4) publishes code and data for citation and versioning, (5) releases regular dataset snapshots to ensure currency and persistence, and (6) ingests these releases into, e.g. the NFDI4Objects Knowledge Graph. This pipeline aims to respect the open-world assumption of LOD and preserve the rich context of statements and qualifiers. It references typical Wikidata-style modelling, while producing a coherent, eventand provenance-centric representation in CIDOC CRM. The Wikiverse provides complementary affordances in this setting [9], [10]. Wikidata offers multilingual breadth, mature property constraints, and widely adopted identifiers; FactGrid supports humanities-oriented projects with detailed historical context; and fuzzy-sl targets spatial uncertainty explicitly, introducing modelling patterns for vagueness and ambiguity in findspot descriptions. By triangulating across these strengths, we can derive a least common denominator of classes and properties that is both mappable to CIDOC CRM and serviceable for downstream analytical tasks. The objective is not to erase local nuances but to encode them faithfully, for example, by transforming qualifiers and reifications into event-centred CIDOC constructs with explicit timespans, actors, and sources. Two use cases anchor our contribution. First, the Irish Holy Wells dataset (Wikidata’s WikiProject Holy Wells 4 ) exemplifies richly reified statements that combine status, conservation, and sourcing practices across heterogeneous evidence (e.g., Ordnance Survey maps, local histories, Wikimedia Commons media). Here, the task is to preserve referential integrity (Q/P identifiers) and transform qualifiers into CIDOC CRM event patterns, thereby enabling temporal reasoning and provenance-aware queries. Second, the Campanian Ignimbrite (CI) findspots [11] from the fuzzy-sl Wikibase focus on spatial uncertainty and cross-domain alignment (archaeology–geosciences). In this case, fuzzy-sl categories and qualifiers can be mapped to CIDOC CRM (with CRMsci/CRMarchaeo as needed), while preserving uncertainty envelopes and observational provenance. Both cases stress machine-actionable provenance and interoperable uncertainty, which are prerequisites for meaningful aggregation across NFDI4Objects and beyond. Methodologically, three principles guide the integration: 1. Identifier discipline . We retain or reference authoritative URIs (e.g., Wikidata Q-codes) and, where appropriate, assert equivalences (e.g., owl:sameAs) to maintain global coherence. This minimises fragile string-matching and ensures stable joins across sources. 4 cf. https://www.wikidata.org/wiki/Wikidata:WikiProject_HolyWells 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 2. Event-centric modelling . Rather than collapsing complex, qualified statements into timeless triples, we materialise the implied activities (assessment, documentation, naming, celebration, sampling) as CIDOC CRM events with agents, timespans, and documentary evidence. This preserves semantics essential for historical reasoning and data quality assessment. 3. Uncertainty as first-class data . Spatial vagueness, competing attributions, and evolving states are modelled explicitly (e.g., via condition assessments and temporally scoped states; fuzzy spatial footprints), aligning with the open-world assumption and avoiding misleading false precision. In sum, we position community knowledge bases not as external “feeds” but as coequal knowledge partners whose specificity and dynamism enrich infrastructure-scale graphs. By combining scripted workflows, CIDOC-compliant patterns, and rigorous identifier strategies, our approach enhances the FAIRness and analytic utility of CH data while maintaining the flexibility for periodic, automated refreshes that keep infrastructure graphs in sync with ongoing community efforts. In this way, Wikibase instances can become integral parts of an interdisciplinary federated Knowledge Graph ecosystem that spans triplestores, Wikibases, and FAIR Digital Objects. Use Case: Irish Holy Wells from Wikidata The WikiProject HolyWells (Q126443484) provides a complex and richly structured dataset. It combines contributions from both academic researchers and citizen scientists, complements textual data [12] with richly annotated media from Wikimedia Commons, and integrates information from heterogeneous sources, both digital and analogue (Fig. 2 and 3). By using reifications, the project documents the evolution of conservation and use status, specifying the context in which each condition was recorded. These features are optimal for demonstrating the process of aligning a flexible model like Wikidata to an interoperable and consistent semantic framework like CIDOC CRM, while also exposing the practical challenges that arise when attempting to automate the transformation of such multifaceted data. 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 Figure 2 - Example: The Citizen Science Wikidata Project Holy Wells. CC0. Figure 3 - Freshford: St Lachtain's Well in Wikidata. CC0 and OSM Contributors. This project aims to develop a semi-automatic ontology design and alignment workflow that can be implemented entirely with Python libraries. The current workflow prototype begins by querying Wikidata and storing the data in relational tables within Postgres. A CIDOC CRM compliant ontology is designed in Protégé, along with mappings that associate the data stored in Postgres with its new form. Then the triples are materialised using the Ontop plugin, creating the final knowledge graph. While the extraction, transformation, and loading processes can be largely automated using Python libraries, designing a CIDOC CRM–compliant ontology remains a non-trivial task, as it requires careful analysis and alignment of the loosely structured information from Wikidata. In the following sections, we outline how the data was acquired, structured, and ultimately aligned with CIDOC CRM, as well as the specific challenges we encountered. 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 Querying Wikidata For the data exploration step, we set up a SPARQL endpoint and queried for all triples with an instance of Holy Well Semantic Concept (Q126443332) as their subject. Holy Well Semantic Concept is a subclass of Holy Well (Q1371047), and every entry in the Irish Holy Wells project is an instance of this class. To ensure both accuracy and interoperability, queries retrieve not only the human-readable labels for entities and properties but also their unique Wikidata identifiers: Q-IDs for entities and P-IDs for properties. Embedding these Q-IDs directly into the URIs of individuals modelled in Protégé while keeping the Wikidata prefix is our strategy for maintaining referential integrity. Human-readable labels, descriptions, and aliases are still incorporated as rdfs:label, skos:altLabel, and schema:description, allowing semantic clarity for human users and multilingual applications. The initial query retrieved 229 distinct subjects and 37 unique statements. Among the most frequently used properties are ‘described by source’ (P1343), ‘instance of’ (P31), ‘located in the administrative territorial entity’ (P131), ‘inventory number’ (P217), and ‘collection’ (P195). Wikidata’s data model is built around statements, qualifiers, and references [13]. Each fact is expressed as a statement, while qualifiers add contextual conditions and references point to supporting evidence, but are not directly connected to the subject. All statements were examined for their domain and range constraints to facilitate semantic alignment and to ensure accurate mapping into a CIDOC CRM-based ontology. For some entries, both status information and contextual information are missing altogether; for others, statuses are declared directly as properties, without further supporting information. A smaller subset contains reified statements, which allow qualifiers such as the state of use to be attached to a property, such as ‘instance of’ (P31) and for references to be added as well, mentioning the source of the claim. Both the classification of the entity and an assessment of its condition are linked to the same piece of evidence. From a CIDOC CRM perspective, this practice aligns with the event-centric modelling of knowledge, where the assertion of a property is itself situated in time, attributed to an actor, and supported by evidence. Queries over such statements aim to preserve their structure and context, retrieving not only the complete statement, but also the superclass of the source and all additional information contained on the source’s page, such as author and publication date. Wikidata incorporates and references external resources such as YouTube, OpenStreetMap, Irish Sites and Monuments Records through specific properties known as external identifiers. Those properties are all instances of subclasses of the class Unique Identifier (Q6545185). For each identifier, queries collect both the reference itself and its associated qualifiers, which may include retrieval date, publisher, or language, and in some cases, the classification of the qualifier values. For example, a linked ‘YouTube video’ (P1651) can be queried alongside the video’s ‘publication date’ (P577) 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 and ‘duration’ (P2047), showing how audiovisual content is anchored in Wikidata with descriptive and temporal attributes. A different strategy is required for Wikimedia Commons, since metadata such as authorship, credit, and licensing is not fully exposed through the Wikidata SPARQL endpoint. Instead, filenames retrieved from Wikidata must first be normalised, after which the Wikimedia API is used to fetch extended metadata. HTML fragments are then parsed to extract structured information, such as artist, credit name and links. While the technical querying mechanism differs, the goal remains the same—enriching identifiers with contextual metadata, like date and place of the creation of an image, its author and license policy. Alongside reified statements and external identifiers, some properties are expressed as direct statements without additional qualifiers. These include spatial and categorical attributes such as geographic coordinates (converted into both WKT and DMS formats for standardised representation), diocese affiliation, and medical condition. For these cases, a simplified query pattern retrieves the subject–predicate–object triple directly and supplements it with the ‘instance of’ (P31) classification of the object, ensuring that even simple statements remain embedded within a structured and interpretable data model. CIDOC CRM Ontology Mapping Translating the acquired information into CIDOC CRM requires the construction of a custom ontology with mappings. While the ultimate aim is to develop a semi-automated workflow, we currently produce them manually, providing a valuable “ground truth” for future experiments with automation. The procedure begins with identifying salient features of both ontologies: for Wikidata, these include entity Q-codes, labels, descriptions, ‘instance of’ (P31) values, and property domain and range restrictions, while for CIDOC CRM they consist of class descriptions and property domain and range constraints as defined in the documentation. The representative example of Columbkile’s Well (Q126456441) was modelled according to Table 1: 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 Table 1 – correspondences when modelling relationship instance of (P31) in wikidata to a CIDOC compliant ontology for the example entity Columbkille’s well. Columbkile’s Well Wikidata Ontology Well is Instance Of X X is Subclass Of Well is Instance Of Y Y is Subclass of Archaeological site (Q839954) location of discovery (Q1291195), human-geographic territorial entity (Q15642541), geographic location (Q2221906) Archaeological Site E27 Site, E53 Place Water source (Q10713454) body of water (Q15324), water resource (Q1049799) Water Source E27 Site, E53 Place Holy well (Q1371047) fountain(Q483453), structure of worship(Q1370598) spring (Q124714), water well(Q43483) Holy Well Fountain, Water Well (both are subclasses of E25 Human made Feature ) Holy well semantic concept (Q126443332) Holy well (Q1371047) Holy Well Semantic Concept Holy Well This step brings in another level of complexity: the relationships from Wikidata do not correspond one-to-one to the ones defined in CIDOC, since the latter is event-based, and most events, like documenting, assessing, naming, or celebrating, are not explicitly stated in Wikidata. The implied existence of these events has to be recognised by a human and placed in the correct context, modelling causality and temporality. Once the relevant classes have been identified, the relationships between their instances must also be translated. This can be automated through the use of mappings—queries that extract data from the relational database created after querying Wikidata and reshape it into the target ontology structure, thereby materialising individuals. Although Ontop automates the materialisation step itself, the mappings that define how the data is transformed (Figure 10) need to be written manually. The following cases illustrate this modelling process using Columbkille’s well as an example. For the ‘named after‘ property (P138), the process begins by retrieving the subject (Columbkille’s Well), predicate and object (Columba), together with the object’s Q-code (Q236326), class (human) and labels from Wikidata. A corresponding individual (Columba) is created in the ontology, as an instance of the correct CIDOC class, E21 Person . Columbkille’s Well is identified by its name, which is a new individual instance of E41 Appellation . This name ‘P67 refers to ’ Columba, the person. The feast day celebration of Columba was ‘ P17 motivated by ’ Columba and ‘ P4 has timespan ’ of the Feast Day of Columba, a new individual instance of E7 Activity . June 9 is a new individual instance of SP14 Time Expression and ‘Q16 defines time’ for the feast day without modifying the string returned by the wikidata query. 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 For archaeological sites, examples include Toplitsa Cave [15] (Q115), Franchthi Cave [16] (Q111; Fig. 11 and 12), and Crvena Stiljena Cave [17] (Q89). These locations are relevant because they combine evidence of human occupation with CI tephra, allowing for inferences about human – environment interactions, chronological anchors, and cultural responses to volcanic events. Within a CIDOC CRM perspective, these entities can remain modelled as E53 Place , but require linking to E27 Site (culturally defined locale) and, where excavation data exist, to CRMarchaeo’s A8 Stratigraphic Unit for layers containing volcanic ash. Events such as the deposition of tephra or its discovery can then be represented as E5 Event instances, enabling temporal and causal reasoning. For geological sites, entries such as DE3 Dehner Maar (Q85), Auel Maar AU3 (Q70), and Urluia Quarry [18] (Q73) document CI deposits identified through field surveys and geoscientific sampling [14]. Here, CIDOC CRM alone is insufficient. Integration with CRMsci is necessary to represent sampling activities (e.g., S4 Observation, S19 Encounter Event) and analytical outcomes, while CRMdig can support provenance of laboratory processes. This ensures that tephra samples are not merely attributes of places but are documented as entities generated by scientific procedures with traceable provenance and reproducibility. At present, the fuzzy-sl Wikibase captures these sites primarily as labelled places with coordinates and limited qualifiers. The properties and qualifiers are not yet mapped to CIDOC CRM, reflecting the still-experimental state of the modelling. Nonetheless, this dataset highlights a methodological challenge: how to align heterogeneous disciplinary perspectives on the “same” site. Archaeologists prioritise cultural layers and human interaction, whereas geoscientists focus on stratigraphic sequences and geochemical signatures. A federated graph must accommodate both perspectives without flattening them into an oversimplified schema. The interdisciplinary integration can be achieved by adopting an event-centric approach: ● The eruption itself is modelled as an E5 Event , producing deposits with a defined timespan. ● Geological observations (sampling, analysis) are modelled via CRMsci, linked to both the deposits and the places where they occur. ● Archaeological contexts are represented as E27 Sites , stratigraphic units, or material culture associations, connected to the same deposits but with additional cultural interpretation. By interlinking these elements, the knowledge graph supports queries that traverse disciplinary boundaries, e.g., “Which archaeological sites with CI deposits coincide with geologically dated findspots older than ~40.000 yr b2k?” or “Which stratigraphic contexts combine volcanic ash with artefactual assemblages?” Crucially, this case demonstrates the role of CIDOC CRM as an interoperability bridge. While E53 Place suffices as a baseline, interdisciplinary data integration requires layering additional ontologies. For CI findspots, this means that the same URI may function simultaneously as a E53 Place , an E27 Archaeological Site , and the locus of scientific observations 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 (CRMsci). Such polyhierarchical modelling aligns with the open-world assumption and preserves the capacity to expand as new data or disciplinary perspectives emerge. From the perspective of the NFDI4Objects Knowledge Graph, the CI case offers an opportunity rather than a limitation. Even though the implementation is not yet operational, the conceptual mapping illustrates how community-curated Wikibase data can flow into infrastructure-scale graphs. The benefit lies in establishing identifier discipline (stable URIs, equivalences to Wikidata Q-IDs), documenting provenance of scientific activities, and ensuring that uncertainty in spatial and chronological resolution is not collapsed but explicitly represented. In conclusion, the Campanian Ignimbrite use case underscores the potential of semi-automated workflows for federated knowledge graph construction. By integrating fuzzy-sl data into a CIDOC-aligned graph, we not only capture a pivotal Late Pleistocene volcanic event but also showcase how interdisciplinary collaboration – bridging geology and archaeology – can be encoded as Linked Open Data. This sets the stage for future implementation where lab analysis of CI sites becomes FAIR Digital Objects within the broader federated ecosystem of NFDI and beyond. Discussion & Conclusion & Outlook The two case studies highlight both the opportunities and the challenges of integrating community-curated Wikibase data into infrastructure-scale knowledge graphs. The Holy Wells illustrate the potential of richly reified cultural heritage statements, while the Campanian Ignimbrite (CI) findspots expose the difficulties of aligning archaeological and geoscientific perspectives. A key challenge lies in ontology alignment and modelling. Wikibase properties and qualifiers are often too flexible to be mapped directly into CIDOC CRM structures. While CIDOC CRM provides an event-centric framework, its extensions (CRMarchaeo, CRMsci, CRMdig) are required to represent stratigraphic reasoning, sampling procedures, and laboratory provenance. For the Holy Wells example, the first exploration step revealed a modelling inconsistency in Wikidata: instances of Holy Well Semantic Concept appear as the subjects of the ‘feast day’ (P841) property, despite it being restricted to humans, groups of humans, titles of Mary, legendary figures, attributes of God, Bible stories, or periscopes. Additionally, Holy Wells are not consistently named after or associated with Christian religious concepts, nor do they always have a fixed feast day. For example, the Well of the Rags (Q126647640) is named after a clootie tree (Q107257053) and has a feast day “Sunday after August 13”. This issue also ties in with the conservation and usage state qualifiers: according to the project description, if the feast day of a well is currently being celebrated, it should be marked as being in use (Q55654238) and as an instance of (P31) a religious site (Q105889895). 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 Although humans can intuitively understand these associations, modelling them in CIDOC CRM raises questions: If the celebration is to be modelled as an event, then additional contextual details are required: Is the celebration motivated by devotion to a specific religious figure? Does it take place at the well itself, or does it involve the well in any way? Has the celebration been documented, and if so, by whom and when? Was it historically tied to a particular period, and is it still practised today? These questions highlight which contextual information is not explicitly encoded and must therefore be added when modelling the data in CIDOC, taking the feast day from a simple recurring point in time to an ontologically appropriate entity that expresses the researchers’ intended semantic construct. Additionally, the project suggests a simplified modelling strategy that limits state qualifiers to two options (in use or abandoned), and conservation state qualifiers to three (preserved, in danger, or demolished/destroyed), even though Wikidata property constraints allow for a much wider range of values. This strategy implicitly reflects a closed world assumption in which the qualifiers are treated as mutually exclusive and exhaustive: if a well is not in use, it is assumed to be abandoned. LOD operates under an open-world assumption, where the absence of a qualifier does not imply its negation; the list of possible states is not seen as exhaustive, and states can co-occur. CIDOC CRM aligns with the open world assumption and supports fine-grained modelling of use and conservation states across time, while also drawing attention to the gaps that hinder automatic reasoning. Another challenge is posed by the classification step, where candidate correspondences are generated by comparing the Wikidata classes against CIDOC classes. This step proved challenging even for humans, as it requires a nuanced understanding of the fine-grained distinctions in CIDOC CRM (e.g. E22 Human-Made Object vs E25 Human-Made Feature or E31 Document vs E73 Information Object ), the implications of CIDOC’s inheritance hierarchy for property domains and ranges, as well as basic understanding of the modelled subjects and their individual particularities. Deciding on an appropriate granularity level for the classification step is a further non-trivial task: too generic, and valuable semantics are lost; too specific, and interoperability suffers. The decision-making process and heuristics employed to classify Columbkile’s Well (Q126456441, Table 1) were based mainly on “common sense” and background knowledge, which is notoriously flimsy and difficult to encode into a formal decision graph for computational use. If the process were to be automated, subsequent steps would involve computing similarity measures for all candidate correspondences, aggregating these results, and applying thresholds to filter out unreliable pairings. The mapping would then be refined through iterative cycles, incorporating structural indices, contextual neighbourhood information, map-discovery and map-repair techniques, until consistency is reached and all classes find a correspondence. This process would be applied not only to all 229 subjects in the Wikidata project but also to every entity appearing as an object in a related statement, qualifier, or reference. The scale and 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 heterogeneity of this task make it clear that one of the central challenges lies in developing an automated approach that is simultaneously scalable and reliable. The second challenge is interdisciplinarity. Archaeologists tend to focus on cultural layers, artefacts, and human activities, while geoscientists prioritise stratigraphy, geochemistry, and analytical workflows. Modelling a single location as both an archaeological site and a geological outcrop requires multi-perspective representations that can coexist without contradiction. The CI case shows that the same entity may need to function simultaneously as an E53 Place , an E27 Site , and the locus of a CRMsci observation. Such polyhierarchical modelling is conceptually demanding but essential if the federated graph is to support queries across disciplines. A further issue is uncertainty. Community data often contains vague coordinates, contested chronologies, or shifting attributions. Unless explicitly represented, these ambiguities risk being flattened into misleading precision. Handling “fuzzy” and “wobbly” data requires not only technical modelling patterns but also shared standards that balance usability with accuracy. Automation and sustainability present another layer of difficulty. Semi-automated pipelines (SPARQL extraction, RDF conversion, ETL) are indispensable for scalability, yet they depend on evolving community schemas. Quality assurance, persistent identifiers, and versioned releases are needed to ensure that infrastructure knowledge graphs remain stable even as community instances change dynamically. In sum, the discussion reveals a tension between community dynamism and infrastructure stability. Community Wikibases thrive on openness, rapid evolution, and volunteer contributions; infrastructures such as NFDI4Objects require persistence, citability, and reliability. Reconciling these modes demands careful governance, reproducible workflows, and machine-actionable provenance. Looking forward, the ontology used in the fuzzy-sl Wikibase will require refinement and closer alignment with CIDOC CRM, MaCHeCO, and the Object Core Metadata Profile. Semi-automated workflows for extraction and mapping must be stabilised, with each release packaged as a FAIR Digital Object. More importantly, the development of community standards for uncertainty and interdisciplinarity will be central to scaling beyond individual use cases. If these steps are taken, Wikibase instances such as Wikidata, FactGrid, and fuzzy-sl can evolve from experimental repositories into integral components of a federated, interdisciplinary knowledge graph ecosystem that bridges archaeology, cultural heritage, and the geosciences. Acknowledgements The authors would like to thank Stephen Stead for his advice as well as the CAA Germany, SIG Data Dragon and NFDI4Objects Community. The authors acknowledge the use of language assistance powered by artificial intelligence (ChatGPT, OpenAI) for stylistic editing and linguistic refinement. All content and arguments were authored and verified by the authors themselves. 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 Data, scripts, code, and supplementary information availability Distel, A.K.. et al. (2025). Wikidata:WikiProject HolyWells: https://www.wikidata.org/wiki/Wikidata:WikiProject_HolyWells ; Thiery, F. et al. (2025). fuzzy-sl Wikibase: https://fuzzy-sl.wikibase.cloud ; Thiery, F., & Schenk, F. (2023). Campanian Ignimbrite Geo Locations [DataSet] at https://github.com/Research-Squirrel-Engineers/campanian-ignimbrite-geo [19]; Stefan, D. (2025). Holy Wells to CIDOC [Software] at https://github.com/CopyKittyCode/Holy_Wells_to_CIDOC ; Conflict of interest disclosure The authors declare that they comply with the PCI rule of having no financial conflicts of interest in relation to the content of the article. Funding The authors declare that they have received no specific funding for this study. References [1] L. Rossenova, P. Duchesne, and I. Blümel, ‘Wikidata and Wikibase as complementary research data management services for cultural heritage data’, in Proceedings of the 3rd Wikidata Workshop 2022 co-located with the 21st International Semantic Web Conference (ISWC2022) , 2022. [Online]. Available: https://ceur-ws.org/Vol-3262/paper15.pdf [2] F. Thiery, A. W. Mees, and J. B. Kiesling, ‘Challenges in research community building: integrating Terra Sigillata (Samian) research into the Wikidata community’, AeC , vol. 34, no. 1, pp. 157–164, 2023, doi: 10.19282/ac.34.1.2023.17. [3] N. Hartl, E. Wössner, and Y. Sure-Vetter, ‘Nationale Forschungsdateninfrastruktur (NFDI)’, Informatik Spektrum , vol. 44, no. 5, pp. 370–373, Oct. 2021, doi: 10.1007/s00287-021-01392-6. [4] L. Rossenova et al. , ‘How are NFDI consortia using Knowledge Graphs? An overview of common functions and challenges by the Working Group “Knowledge Graphs”’, in Proceedings of the Conference on Research Data Infrastructure 2025 , Y. Sure-Vetter and G. Paul, Eds, Aachen: Squirrel Papers, Aug. 2025, p. 7(5), 𝒬5. doi: 10.5281/zenodo.16736077. [5] K. Fischer et al. , ‘Windows on Data: Federating Research Data with FAIR Digital Objects and Linked Open Data’, in Proceedings of the Conference on Research Data Infrastructure 2025 , Y. Sure-Vetter and P. Groth, Eds, Aachen: Squirrel Papers, Aug. 2025, p. 7(5), 𝒬3. doi: 10.5281/zenodo.16736221. [6] F. Thiery et al. , ‘Object-Related Research Data Workflows Within NFDI4Objects and Beyond’, in Proceedings of the Conference on Research Data Infrastructure , Y. Sure-Vetter and C. Goble, Eds, Hannover: TIB Open Publishing, Sept. 2023, pp. CoRDI2023-46. doi: 10.52825/cordi.v1i.326. [7] S. C. Schmidt, F. Thiery, and M. Trognitz, ‘Practices of Linked Open Data in Archaeology and Their Realisation in Wikidata’, Digital , vol. 2, no. 3, pp. 333–364, June 2022, doi: 10.3390/digital2030019. 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 [8] F. Thiery and P. Thiery, ‘Linked Open Ogham. How to publish and interlink various Ogham Data?’, Archeologia e Calcolatori , vol. 34, no. 1, pp. 105–114, 2023, doi: 10.19282/ac.34.1.2023.12. [9] F. Thiery, L. Rossenova, D. Mietchen, T. Homburg, and P. Thiery, ‘Distributed Research Data Knowledge Graphs - Challenges of federated queries using the Wikiverse and OpenStreetMap within the NFDI Knowledge Graph Ecosystem’, in Proceedings of the Conference on Research Data Infrastructure 2025 , Y. Sure-Vetter and P. Groth, Eds, Aachen: Squirrel Papers, Aug. 2025, p. 7(5), 𝒬2. doi: 10.5281/zenodo.16736047. [10] F. Thiery, L. Rossenova, and O. Simons, ‘Wikibase instances in the Cultural Heritage Domain: Examples from the German humanities NFDI consortia’, Squirrel Papers , vol. 6, no. 4, p. #4, Nov. 2024, doi: 10.5281/zenodo.14055699. [11] F. Thiery and F. Schenk, ‘Modelling of Uncertainty in Geo Sciences Sites’, Squirrel Papers , vol. 5, no. 1, p. #4, Dec. 2023, doi: 10.5281/zenodo.10255259. [12] P. Ó Dálaigh, ‘The Holy Wells of County Kilkenny - Volume 2’, Doctoral thesis, Mary Immaculate College, University of Limerick, 2018. Accessed: Jan. 11, 2025. [Online]. Available: https://dspace.mic.ul.ie/handle/10395/2584 [13] S. C. Schmidt, F. Thiery, and M. Trognitz, ‘Practices of Linked Open Data in Archaeology and Their Realisation in Wikidata’, Digital , vol. 2, no. 3, pp. 333–364, June 2022, doi: 10.3390/digital2030019. [14] F. Schenk, U. Hambach, S. Britzius, D. Veres, and F. Sirocko, ‘A Cryptotephra Layer in Sediments of an Infilled Maar Lake from the Eifel (Germany): First Evidence of Campanian Ignimbrite Ash Airfall in Central Europe’, Quaternary , vol. 7, no. 2, p. 17, Mar. 2024, doi: 10.3390/quat7020017. [15] T. Tsanova et al. , ‘Upper Palaeolithic layers and Campanian Ignimbrite/Y-5 tephra in Toplitsa cave, Northern Bulgaria’, Journal of Archaeological Science: Reports , vol. 37, p. 102912, June 2021, doi: 10.1016/j.jasrep.2021.102912. [16] F. G. Fedele, B. Giaccio, R. Isaia, and G. Orsi, ‘The Campanian Ignimbrite Eruption, Heinrich Event 4, and palaeolithic change in Europe: A high-resolution investigation’, in Geophysical Monograph Series , vol. 139, A. Robock and C. Oppenheimer, Eds, Washington, D. C.: American Geophysical Union, 2003, pp. 301–325. doi: 10.1029/139GM20. [17] M. W. Morley and J. C. Woodward, ‘The Campanian Ignimbrite (Y5) tephra at Crvena Stijena Rockshelter, Montenegro’, Quat. res. , vol. 75, no. 3, pp. 683–696, May 2011, doi: 10.1016/j.yqres.2011.02.005. [18] F. Thiery and F. Schenk, ‘How to locate the Campanian Ignimbrite site Urluia based on literature? How to provide and publish this data in a FAIR way?’, Squirrel Papers , vol. 5, no. 1, p. #5, 2023, doi: 10.5281/zenodo.10262720. [19] F. Thiery and F. Schenk, ‘Campanian Ignimbrite Geo Locations’, Squirrel Papers , vol. 5, no. 2, p. #2, 2023, doi: 10.5281/zenodo.10361309. 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610