Comparative Evaluation of GeoSPARQL Implementations
Abstract
Master Thesis
Full text
Comparative Evaluation of GeoSPARQL Implementations Master Thesis Name of the Study Programme Master Professional IT Business and Digitalization Faculty 3 from Samneang Seng Date: Berlin, 30.03.2025 1st Supervisor: Prof. Dr.-Ing. Thomas Schwotzer 2nd Supervisor: Dr. Simon Fokt
i Index 1. INTRODUCTION ............................................................................................................................. 1 1.1. PROBLEM STATEMENT ..................................................................................................................... 1 1.2. OBJECTIVES ................................................................................................................................... 2 1.2.1. GEOSPARQL IMPLEMENTATIONS EVALUATION WITH FUNCTIONALITY TESTING ............................................. 2 1.2.2. BENCHMARKING GEOSPARQL IMPLEMENTATIONS .................................................................................. 2 2. STATE OF THE ART ........................................................................................................................ 3 2.1. GEOSPARQL - THE ONTOLOGY ........................................................................................................ 3 2.1.1. CLASSES ............................................................................................................................................ 4 2.1.2. DATA TYPES ....................................................................................................................................... 5 2.1.3. QUERY TRANSFORMATION RULES .......................................................................................................... 5 2.1.4. DATA PROPERTIES ............................................................................................................................... 6 2.2. GEOSPARQL – THE QUERY LANGUAGE .............................................................................................. 7 2.2.1. NON-TOPOLOGICAL FUNCTIONS ........................................................................................................... 7 2.2.2. SIMPLE FEATURES ............................................................................................................................... 9 2.2.3. EGENHOFER ..................................................................................................................................... 10 2.2.4. RCC8 ............................................................................................................................................. 10 2.3. GEOSPARQL IMPLEMENTATIONS .................................................................................................... 12 2.4. APACHE JENA FUSEKI .................................................................................................................... 12 2.4.1. LICENSING ....................................................................................................................................... 13 2.4.2. INSTALLATION .................................................................................................................................. 13 2.5. ONTOTEXT GRAPHDB ................................................................................................................... 13 2.5.1. LICENSING ....................................................................................................................................... 14 2.5.2. INSTALLATION .................................................................................................................................. 14 2.6. GETTING STARTED WITH GEOSPARQL ............................................................................................. 15 2.6.1.1. Example 1 .................................................................................................................................. 15 2.6.1.2. Example 2 .................................................................................................................................. 16 2.6.1.3. Example 3 .................................................................................................................................. 16 2.6.1.4. Example 4 .................................................................................................................................. 17 3. METHODOLOGY .......................................................................................................................... 18 3.1. SELECTION OF GEOSPARQL IMPLEMENTATIONS ................................................................................. 18 3.2. DATASETS SELECTION .................................................................................................................... 19 3.3. DATASETS PREPARATION ............................................................................................................... 19 3.4. FUNCTIONALITY TESTING ON STATIC EXAMPLE QUERIES ........................................................................ 20 3.5. USE CASE DEVELOPMENT BASED ON GEOSPARQL FUNCTIONALITY ........................................................ 21 3.5.1. USE CASES SELECTION ....................................................................................................................... 22
ii 3.5.2. QUERIES DEVELOPMENT .................................................................................................................... 22 3.5.2.1. Use Case # 1 .............................................................................................................................. 24 3.5.2.2. Use Case # 2 .............................................................................................................................. 26 3.5.2.3. Use Case # 3 .............................................................................................................................. 27 3.5.2.4. Use Case # 4 .............................................................................................................................. 29 3.5.2.5. Use Case # 5 .............................................................................................................................. 30 3.5.2.6. Use Case # 6 .............................................................................................................................. 31 3.6. SCENARIOS DEVELOPMENT FOR GEOSPARQL FUNCTIONALITY TESTING ................................................... 33 3.7. BENCHMARKING .......................................................................................................................... 34 4. EXPERIMENTS ............................................................................................................................. 35 4.1. HARDWARE AND SOFTWARE ........................................................................................................... 35 4.2. DATA LOADING ............................................................................................................................ 35 4.2.1. APACHE JENA FUSEKI ......................................................................................................................... 35 4.2.2. GRAPHDB ....................................................................................................................................... 37 4.2.3. DATASETS VERIFICATION .................................................................................................................... 37 4.3. USE CASE IMPLEMENTATION IN GRAPHDB ........................................................................................ 38 4.4. USE CASE IMPLEMENTATION IN APACHE JENA FUSEKI .......................................................................... 40 4.5. MODIFIED BENCHMARKING APPROACH ............................................................................................ 43 4.6. SCENARIO-BASED QUERY TESTING AND EXECUTION ............................................................................. 45 4.7. EXPERIMENT WITH QLEVER ............................................................................................................ 47 5. RESULTS AND DISCUSSION .......................................................................................................... 51 5.1. FUNCTIONALITY TESTING ................................................................................................................ 51 5.2. BENCHMARKING .......................................................................................................................... 53 6. LIMITATIONS ............................................................................................................................... 55 7. CONCLUSION .............................................................................................................................. 56 ATTACHMENTS .................................................................................................................................. 57 LIST OF LITERATURE ........................................................................................................................... 66 STATUTORY DECLARATION ................................................................................................................ 67
iii List of Figures Figure 1: An overview of the Classes and Properties defined in GeoSPARQL ............................. 4 Figure 2: GeoSPARQL plugin turned on by default on GraphDB (Setup -> Plugins) ................ 14 Figure 3: Spatial objects representing geosparql-example.rdf ...................................................... 15 Figure 4: Query results of geof:rcc8ec from Cambridgesemantics on GraphDB ......................... 21 Figure 5: Query results of geof:rcc8ec after query modification on GraphDB ............................ 21 Figure 6: Details of osm:way 190942449 on GraphDB ............................................................... 24 Figure 7: Datasets manage page on Apache Jena Fuseki ............................................................. 36 Figure 8: Add-data option on Apache Jena Fuseki ....................................................................... 36 Figure 9: Results of query execution for Q1 on GraphDB ........................................................... 38 Figure 10: Direction from Galeria Alexanderplatz to Fernsehturm on OSM Maps ..................... 39 Figure 11: Results of query execution for Q4 on GraphDB ......................................................... 39 Figure 12: Direction from Ministry of Cat to Russian Market on OSM Maps ............................. 40 Figure 13: No results returned for Q1 query on Apache Jena Fuseki ........................................... 40 Figure 14: Results returned for Q1 query after function substitution ........................................... 42 Figure 15: Apache Jena Fuseki failed to parse JSON response for scenario 4 ............................. 45 Figure 16: Response successfully parsed in CSV format for scenario 4 ...................................... 46 Figure 17: Symdifferences of Eichepark and Neue Wuhle from scenario 4 rendered on geojsonto-wkt-converter.onrender.com ..................................................................................................... 46 Figure 18: No results returned on Apache Jena Fuseki for scenario 12 ....................................... 47 Figure 19: Jewish Hospital disconnected from Heinz-Galinski-Straße from scenario 12 visualized on geojson.io ................................................................................................................................. 47 Figure 20: Listing all German cities on Qlever ............................................................................ 50 Figure 21: Error message for geof:buffer on Qlever .................................................................... 50 Figure 22: Error message for geof:distance on Qlever ................................................................. 50 Figure 23: Graph showing GeoSPARQL Use Case Execution Time Comparison on both implementations ............................................................................................................................ 54
Introduction 1 1. Introduction The increasing use of geospatial data has necessitated the development of standardized methods for querying and managing geospatial data in a variety of domains including but not limited to, urban planning, environmental monitoring, transportation and especially geographic information systems (GIS). The GeoSPARQL standard, a specification for representing and querying geospatial data using and extending the SPARQL query language, is one of the most important developments in this area. GeoSPARQL provides a set of standardized functions and data types, designed to enable spatial queries of RDF data sets, making it an essential tool for working with geospatial information in the Semantic Web. As the geospatial data becomes more complex, the need for effective tools to manage and query this data is critical, especially for large datasets such as OpenStreetMap (OSM), which contain detailed information about geographic features around the world. OSM’s extensive collection of geographic and feature data makes it a valuable resource for evaluating and benchmarking GeoSPARQL implementations. Despite the increasing adoption of GeoSPARQL, there is a lack of comprehensive studies that evaluate the performance and functionality of different GeoSPARQL implementations. This thesis aims to address this knowledge by conducting a comprehensive evaluation of popular GeoSPARQL implementations, Apache Jena Fuseki and GraphDB, using OSM data as the ground of testing. 1.1. Problem Statement GeoSPARQL has become an increasingly utilized standard for the representation and retrieval of geospatial data within the RDF knowledge graphs. While numerous studies have been conducted to evaluate the GeoSPARQL implementations for standards compliance and query performance, these studies have mainly relied on synthetic datasets or simplified spatial data, designed and aimed to individually isolate spatial functions for testing purposes. However, the expanding adopting of volunteered geographic information (VGI) sources, such as OpenStreetMap (OSM), introduces a new non-trivial challenge for GeoSPARQL implementations. OSM datasets are known to be imperfectly organized, inherently noisy and highly diverse from one another depending on regions and span multiple spatial scales. They range from individual cities to entire continents including one of the most prominent projects, OpenHistoricalMap. The effectiveness of GeoSPARQL implementations in handling such real-world spatial data at varying scales has not been the subject of a comprehensive evaluation. This study aims to address this gap by assessing the functionality and performance of GeoSPARQL-compliant triple stores using real-world OSM datasets. The evaluation will be conducted at the city (Berlin), country (Cambodia), and continental (Antarctica) scales. The study’s primary objective is to comprehensively assess the compliance of these systems with the fundamental principles of GeoSPARQL, while also investigating their capacity to effectively navigate the unique challenges imposed by OSM data. These challenges include unique and non-standard available geometries, complex spatial relationships, and real-world tagging schemes.
Introduction 2 1.2. Objectives The primary objective of this research is to conduct comprehensive comparative evaluation of GeoSPARQL implementations with a focus on functionality testing and benchmarking. The specific objectives of the research are as follows: 1.2.1. GeoSPARQL Implementations Evaluation with Functionality Testing In order to achieve this objective, two steps must be taken. First, a comparison must be performed in GeoSPARQL support offered by the implementations. Second, real-world use cases must be built based on the supported features. The first step involves assessing the extent to which the implementations support the set of functions defined by the GeoSPARQL standards. These functions include, but are not limited to, spatial operators (e.g., sfWithin, sfIntersects), data properties (e.g., geo:asWKT), and data types (e.g., wktLiteral). The second step focuses on building use cases and developing GeoSPARQL queries based on those selected use cases using the actual OSM data. In Chapter 3, an investigative approach will be conducted on the methodology used to construct the use cases and queries. This investigation will address the fundamental question of how these components were integrated to facilitate the execution of functional testing procedures. 1.2.2. Benchmarking GeoSPARQL Implementations Benchmarking in this thesis focuses on measuring and comparing the performance of different GeoSPARQL implementations under varied query loads. Indicators such as query execution time, resource consumption (e.g., memory and CPU usage), and scalability with large datasets will be measured and analyzed. The benchmarking process will evaluate the speed and accuracy of results when executing the same set of queries on different implementations, Apache Jena Fuseki and GraphDB. This will help identify how well the implementations handle OSM data which is categorized as challenging due to its extensive nature and diverse types of geographic features (e.g., roads, parks, buildings, natural features) under limited computational resources, a local setup environment in this case during the experiments, detailed in Chapter 4.
State of the Art 3 2. State of the Art GeoSPARQL, an acronym for “A Geographic Query Language for RDF Data”, is a standard issued by the Open Geospatial Consortium (OGC) 1 . The first version of GeoSPARQL, designated as GeoSPARQL 1.0, was released in 2012. This was followed by the most recent update, GeoSPARQL 1.1, which was published in 2024 2 . The primary objective of this chapter is to introduce GeoSPARQL, and the additional concepts and properties that have been implemented in this latest version. At the time of writing this thesis, the future releases of GeoSPARQL are under development and in planning [1] with distinct updates to the current version as follows: • 1.1 (current release): Extensions that are fully compatible with GeoSPARQL 1.0; • 1.2: Contain necessary changes to GeoSPARQL 1.1 that needs to be accepted as an ISOcopublication 3 . • 1.3 (current draft) 4 : An actual update to GeoSPARQL 1.1 with new features support on fully-featured 3D and GeoCoding literals support. • 2.0: Future GeoSPARQL likely incompatible with GeoSPARQL 1.0. GeoSPARQL is actually two things: an ontology for geo-spatial data and a query language [2]. The Ontology provides a standardized vocabulary for representing spatial data in RDF, including concepts like geo:Feature, geo:Geometry, and spatial relationships. The query language extends SPARQL with geospatial functions to perform spatial queries, such as geof:distance(), and geof:sfIntersects(). 2.1. GeoSPARQL - The Ontology GeoSPARQL is primarily an ontology – a structured way to describe and organize geospatial data using RDF. GeoSPARQL follows a familiar way of handling spatial data. It defines two key concepts: • Feature – A real-world object (e.g., a building, an airport, or a school). • Geometry – The shape or location of that object, described using coordinates. GeoSPARQL is designed to support both qualitative spatial reasoning systems and quantitative spatial computation systems. Qualitative spatial reasoning systems (e.g., those based on Region Connection Calculus) [3] typically do not explicitly use geometries. Instead, they evaluate spatial relationships between features, such as whether one region touches or contains another. Quantitative spatial computation systems, on the other hand, work with explicit geometries and perform calculations based on coordinates, such as checking if a point is within a polygon or measuring distances. 1 https://www.ogc.org/publications/standard/geosparql/ 2 https://opengeospatial.github.io/ogc-geosparql/geosparql11/document.html 3 https://opengeospatial.github.io/ogc-geosparql/geosparql12/ 4 https://opengeospatial.github.io/ogc-geosparql/geosparql13/
State of the Art 4 The following listing shows all the ontology prefixes used in GeoSPARQL 1.1 ontology 5 . Listing: Prefixes used by GeoSPARQL 1.1 PREFIX : <http://www.opengis.net/ont/geosparql#> PREFIX dcterms: <http://purl.org/dc/terms/> PREFIX owl: <http://www.w3.org/2002/07/owl#> PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#> PREFIX sdo: <https://schema.org/> PREFIX vann: <http://purl.org/vocab/vann/> PREFIX skos: <http://www.w3.org/2004/02/skos/core#> PREFIX spec11: <http://www.opengis.net/spec/geosparql/1.1/specification.html#> PREFIX xsd: <http://www.w3.org/2001/XMLSchema#> 2.1.1. Classes The ontology starts with the super class geo:SpatialObject followed by two main subclasses, geo:Feature and geo:Geometry. The figure below 6 gives an overview of the Classes and Properties defined in GeoSPARQL. geo:hasGeometry and geo:hasDefaultGeometry connect (multiple) geo:Feature and geo:Geometry together. The following example 7 shows how it is structured. 5 https://opengeospatial.github.io/ogc-geosparql/geosparql11/geo.ttl 6 https://opengeospatial.github.io/ogc-geosparql/geosparql11/document.html#_core 7 ISPRS%Int.%J.%Geo-Inf.%2022,"11,"117""-"Page"8 Figure 1: An overview of the Classes and Properties defined in GeoSPARQL
State of the Art 5 Listing: Example of geo:hasGeometry and geo:hasDefaultGeometry PREFIX ex: <http://example.com/thing/> PREFIX geo: <http://www.opengis.net/ont/geosparql#> PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#> ex:brisbane a geo:Feature ; rdfs:label "Brisbane Federal Electorate" ; geo:hasGeometry [ a geo:Geometry ; geo:asWKT """POLYGON (( 153.099932 -27.445258, ... , 153.099932 -27.445258 ))"""^^geo:wktLiteral ; ] ; In GeoSPARQL 1.1, two new additional classes geo:FeatureCollection and geo:GeometryCollection were added into the release. Geometry model collections can be done by using these two with RDF. 2.1.2. Data Types The GeoSPARQL ontology specifies five data types which include: WKTLiteral, and GMLLiteral originally and geoJSONLiteral, kmlLiteral, and dggsLiteral following the new data properties addition in GeoSPARQL 1.1. As the names imply, WKTLiteral and GMLLiteral are used in encoding geometries for geo:asWKT and geo:asGML data properties following the other three data types and their corresponding data properties. Here is an example 8 of a geometry encoded as WKTLiteral. Listing: Example of encoding geometries to wktLiteral "http://www.opengis.net/def/crs/OGC/1.3/CRS84 POLYGON((21.5 18.5,23.5 18.5, 23.5 21,21.5 21,21.5 18.5))" ^^sf:wktLiteral. 2.1.3. Query Transformation Rules Query transformation rules add another layer of flexibility for SPARQL queries. While the common possible case is directly comparing specific geometry objects, it is also useful to 8 https://www.lirmm.fr/rod/slidesRoD04102018/RoD2018-tutorial.pdf
State of the Art 12 Listing: Example of using rcc8ntpp PREFIX geof: <http://www.opengis.net/def/function/geosparql/> SELECT (geof:rcc8ntpp(?x,?y) as ?is_ntpp) WHERE { VALUES (?x ?y) { ('POLYGON ((2 2, 5 2, 5 4, 2 4))' 'POLYGON ((1 1, 1 6, 4 1, 4 6))') ('POLYGON ((1 1, 1 4, 4 4, 4 1))' 'POLYGON ((1 1, 1 4, 4 4, 4 1))') } } 2.3. GeoSPARQL Implementations A GeoSPARQL implementation is primarily a SPARL engine that provides GeoSPARQL support. It enables the features such as: storing, querying, and providing reasoning over RDF data format. As previously discussed in section 2.2, GeoSPARQL extends SPARQL, the query language with added capabilities such as geometry representation, spatial relationships, and spatial functions. Since GeoSPARQL implementations are commonly and originally SPARQL engines, another feature is also found in the programs, SPARQL endpoint where every query is available for API calls. This specific feature will be used as a part of benchmarking process which will later be detailed in chapter 3. To summarize, key characteristics and features of a GeoSPARQL implementation are as follows: • Geospatial Data Storage – RDF triples could include geospatial information in WKT (Well-Known Text) or GML (Geography Markup Language) formats. • Spatial Querying – Allows spatial queries using SPARQL, such as distance computation between locations and points finding within a polygon. • Geospatial Reasoning – Supports topological relationships like intersects, within, contains, allowing inference over spatial data. 2.4. Apache Jena Fuseki Apache Jena, formerly known as Jena before Apache integrated it into The Apache Software Foundation in April 2012, is an open-source Semantic Web framework for Java. Similarly functioning as a SPARQL engine, it provides an API for data extraction from and to write RDF graphs. Jena provides support for various formats of serialization including 18 : a relational database, RDF/XML, Turtle, TriG, Notation 3, and JSON-LD. Apache Jena Fuseki, on the other hand is the Jena SPARQL server that provides interface that makes querying and managing triples and RDF easier and more convenient. 18 https://en.wikipedia.org/wiki/Apache_Jena
State of the Art 13 However, Apache Jena Fuseki does not provide GeoSPARQL support by default. jenafuseki-geosparql therefore, is made available for the purpose. It can be accessed as an embedded server using Maven as well as the binary format, which can be downloaded from the Maven central repository 19 . The latter option was actually used in this thesis for running the program. Since jena-fuseki-geosparql does not come with GUI, there is a certain way that it could be run using Apache Jena Fuseki as an UI wrapper. It is noted in the documentation that GUI is currently in development plan as future work, assisting users when querying the data. Next section shows exactly the way to do it and get the server up and running with the GUI. 2.4.1. Licensing As previously mentioned, Apache Jena Fuseki is an open-source project. Based on the documentation, it is claimed that the implementation follows 1.0 GeoSPARQL standard. The implementation is pure Java and does not require any set-up or configuration of any thirdparty relational databases and geospatial extension. Although being an open-source implementation, it is said to apply and comply with six Conformance Classes as described in the GeoSPARQL standard. Those six Conformance Classes include: Core, Topology Vocabulary, Geometry Extension, RDFS Entailment Extension, Query Rewrite Extension. Additionally, the three spatial relation families: Simple Feature, Egenhofer and RCC8 are also supported. 2.4.2. Installation The binary distribution of the Fuseki server can be obtained directly from the downloads site 20 and more options are available on the page 21 . Meanwhile, the self-contained JAR file of jena-fuseki-geosparql must also be downloaded 22 . Following the download of these files, the JAR file of jena-fuseki-geosparql must be placed in the /lib directory of the Apache Jena Fuseki directory. The following command starts the Apache Jena Fuseki with GeoSPARQL support, specifying the classpath -cp to include all the jar files in the /lib directory. Listing: Command for running Apache Jena Fuseki java -cp "fuseki-server.jar:lib/*" org.apache.jena.fuseki.cmd.FusekiCmd --config run/config.ttl --debug 2.5. Ontotext GraphDB Ontotext GraphDB, shortened as GraphDB later in the rest of the thesis content, developed and provided by Ontotext company, is a highly efficient, scalable and robust graph database 19 https://repo1.maven.org/maven2/org/apache/jena/jena-fuseki-geosparql/ 20 https://dlcdn.apache.org/jena/binaries/apache-jena-fuseki-5.3.0.zip 21 https://jena.apache.org/download/index.cgi 22 https://repo1.maven.org/maven2/org/apache/jena/jena-fuseki-geosparql/
State of the Art 14 with RDF and SPARQL support 23 . It implements the RDF4J 24 framework interfaces, the W3C SPARQL Protocol specification. Serialization support is also available in all RDF formats. Unlike Apache Jena Fuseki, GeoSPARQL feature on GraphDB could easily be turned on via the web-based administration portal and actually enabled by default. Section 2.5.2 shows the interface on how it could be turned on. In the next section, licensing of GraphDB is detailed and differences between the licenses will be listed. 2.5.1. Licensing GraphDB offers two kinds of licenses: GraphDB Free and GraphDB Enterprise. While both types give the same common features such as: • Fully compliant with RDF 1.1 and SPARQL 1.1, with RDF-Star and SPARQL-Star extensions. • Fully compliant reasoning for the standard rulesets RDFS, OWL 2 RL and QL. • 100% compatible with the RDF4J framework GraphDB Enterprise, however, provides additional features like Lucene connector for fulltext search, Solr connector for full-text search, Elasticsearch connector and Kafka connector for downstream synchronization. Workbench or web-based interface is also available for both types for managing repositories, data, user accounts and access roles. It is also important to note that with GraphDB Free which is actually used in the thesis for the experiments process, comes with single-core license only. The next section shows how to obtain the program and set it up. 2.5.2. Installation GraphDB can only obtained by making a request on the Ontotext page 25 . Once the request is made, users will receive an email containing the links to download the program. GraphDB is available in a variety of installation options, including desktop, standalone, and cloud setups on Azure and AWS. In this thesis, the standalone server version was selected due to the fact that certain command options are only available in this edition. The following command is used to start the engine with allocation of maximum heap size of 10GB. GeoSPARQL is turned on by default with GraphDB as shown in the following figure. Listing: Command for running GraphDB ➜ graphdb-10.8.3 ./bin/graphdb -Xmx10g Figure 2: GeoSPARQL plugin turned on by default on GraphDB (Setup -> Plugins) 23 https://graphdb.ontotext.com/documentation/10.8/index.html 24 https://rdf4j.org 25 https://www.ontotext.com/products/graphdb/
State of the Art 15 2.6. Getting Started with GeoSPARQL One of the prominent features provided by GraphDB documentation is a sample RDF document with a set of GeoSPARQL example queries, giving an overview and quick insight to users on how the program supports GeoSPARQL. Below is a list of four SELECT queries on geographic data. Sample dataset geosparql-example.rdf 26 can be downloaded and imported into a repository on GraphDB with GeoSPARQL plugin enabled as shown in the section 2.5.2. The RDF dataset defines the following figure for spatial objects. Listing: Prefixes used in the example queries PREFIX my: <http://example.org/ApplicationSchema#> PREFIX geo: <http://www.opengis.net/ont/geosparql#> PREFIX geof: <http://www.opengis.net/def/function/geosparql/> PREFIX uom: <http://www.opengis.net/def/uom/OGC/1.0/> 2.6.1.1. Example 1 Find all features that feature my:A contains, where spatial calculations are based on my:hasExactGeometry. Listing: Example 1 of GeoSPARQL query from GraphDB SELECT ?f WHERE { my:A my:hasExactGeometry ?aGeom . ?aGeom geo:asWKT ?aWKT . ?f my:hasExactGeometry ?fGeom . 26 https://graphdb.ontotext.com/documentation/10.8/_downloads/d10860b42c7c39e3d05ef397f5756ca1/geosparqlexample.rdf Figure 3: Spatial objects representing geosparql-example.rdf
State of the Art 16 ?fGeom geo:asWKT ?fWKT . FILTER (geof:sfContains(?aWKT, ?fWKT) && !sameTerm(?aGeom, ?fGeom)) } Listing: Result of Example 1 tested on GraphDB f http://example.org/ApplicationSchema#B http://example.org/ApplicationSchema#F 2.6.1.2. Example 2 Find all features that are within a transient bounding box geometry, where spatial calculations are based on my:hasPointGeometry. Listing: Example 2 of GeoSPARQL query from GraphDB SELECT ?f WHERE { ?f my:hasPointGeometry ?fGeom . ?fGeom geo:asWKT ?fWKT . FILTER (geof:sfWithin(?fWKT, ''' <http://www.opengis.net/def/crs/OGC/1.3/CRS84> Polygon ((-83.4 34.0, -83.1 34.0, -83.1 34.2, -83.4 34.2, -83.4 34.0)) '''^^geo:wktLiteral)) } Listing: Result of Example 2 tested on GraphDB f http://example.org/ApplicationSchema#D 2.6.1.3. Example 3 Find all features that touch the union of feature my:A and feature my:D, where computations are based on my:hasExactGeometry. Listing: Example 3 of GeoSPARQL query from GraphDB SELECT ?f WHERE { ?f my:hasExactGeometry ?fGeom . ?fGeom geo:asWKT ?fWKT . my:A my:hasExactGeometry ?aGeom .
State of the Art 17 ?aGeom geo:asWKT ?aWKT . my:D my:hasExactGeometry ?dGeom . ?dGeom geo:asWKT ?dWKT . FILTER (geof:sfTouches(?fWKT, geof:union(?aWKT, ?dWKT))) } Listing: Result of Example 3 tested on GraphDB f http://example.org/ApplicationSchema#C 2.6.1.4. Example 4 Find the 3 closest features to feature my:C, where computations are based on my:hasExactGeometry. Listing: Example 4 of GeoSPARQL query from GraphDB SELECT ?f WHERE { my:C my:hasExactGeometry ?cGeom . ?cGeom geo:asWKT ?cWKT . ?f my:hasExactGeometry ?fGeom . ?fGeom geo:asWKT ?fWKT . FILTER (?fGeom != ?cGeom) } ORDER BY ASC(geof:distance(?cWKT, ?fWKT, uom:metre)) LIMIT 3 Listing: Result of Example 4 tested on GraphDB f http://example.org/ApplicationSchema#A http://example.org/ApplicationSchema#E http://example.org/ApplicationSchema#D
Methodology 18 3. Methodology The thesis is focused on two primary objectives: functionality testing and benchmarking. To accomplish these objectives, the use of GeoSPARQL implementations, also known as “triple stores” 27 , is necessary. This chapter provides a detailed examination of the process of comparison and evaluation for GeoSPARQL implementations. Although many triple stores have adopted the standard and started providing support for GeoSPARQL in respective programs, the functionality, usability and performance may vary across the programs due to their own spatial indexing strategies and architectural differences. To date, a considerable number of papers and studies have been conducted to either assess and focus solely on individual implementation or to conduct benchmarking against the expected results of the GeoSPARQL standards using static simplified and synthetic datasets only. The objective of this therefore is to address the aforementioned gap by comparing and evaluating functionality and benchmarking using different use cases across various sized datasets on two selected implementations: Apache Jena Fuseki and GraphDB. An overview of the methodology carried out for achieving each objective is discussed in the following sections. 3.1. Selection of GeoSPARQL Implementations Both Apache Jena Fuseki and GraphDB are commonly used and chosen for RDF triplestores management and their GeoSPARQL support. In fact, a Journal Article “A GeoSPARQL Compliance Benchmark” showed that GeoSPARQL Jena Fuseki and GraphDB topped the results for complying with the requirements set out by GeoSPARQL standard, scoring 82.75% and 69.75% [5] respectively. While Apache Jena Fuseki is open-source, GraphDB also provides the freeware version along with their enterprise edition. Therefore, choosing these two implementations provides insights into trade-offs between free and paid solutions. Although they both support GeoSPARQL, their distinct indexing strategies may impact performance and scalabilities differently. To investigate the differences, functionality and benchmarking will be conducted to assess their support in handling queries and their claim in adhering to GeoSPARQL standards compliance. This choice of selection aims to provide insights into capabilities and limitations of the chosen implementations. 27 Triple stores: purpose-built databases for storing and retrieving triples using semantic queries - https://w.wiki/D7Yi
Methodology 19 3.2. Datasets Selection In order to conduct functionality testing and perform benchmarking, datasets containing proper geospatial data are required. Fortunately, there are a few sites that provide OpenStreetMap (OSM) datasets with RDF extracts; geofabrik 28 and osm2rdf 29 are selected in the thesis. The former provides daily updated dataset files in.osm.pbf and .shp.zip formats of sub-regions and countries followed by cities (where available) in every continent while the latter provides weekly updated zipped triples in Turtle format .ttl.bz2 files for entire planet OSM data as well as extracts for individual continents and countries. There are three datasets obtained for use in this thesis, including: • City-level: Berlin named berlin-latest-osm.pbf • Country-level: Cambodia named khm-osm.ttl.bz2 • Continent-level: Antarctica named ata-osm.ttl.bz2 Listing: Datasets obtained from geofabric and osm2rdf sites dataset type size source berlin-latest.osm.pbf city 88.1 MB geofabrik khm.osm.ttl.bz2 country 234 MB osm2rdf ata.osm.ttl.bz2 continent 86.7 MB osm2rdf 3.3. Datasets Preparation After obtaining the datasets, conversion and unzipping processes to RDF format files are necessary for the berlin-latest.osm.pbf and the other two datasets, respectively. Many tools are available to serve the conversion purpose. However, after researching and testing on so many programs, osm2rdf, [6] first introduced as osm2ttland provided by the University of Freiburg 30 was determined to be the easiest and most efficient to set up and process. The following command is used to build the Docker image following the cloning of the repository from osm2rdf’s github repository. Listing: Docker command for building osm2rdf docker build -t osm2rdf The following listing shows the command used for converting the “Berlin” dataset to RDF compatible for the GeoSPARQL implementations in .ttl format file. 28 https://download.geofabrik.de/ 29 https://osm2rdf.cs.uni-freiburg.de/ 30 https://github.com/ad-freiburg/osm2rdf
Methodology 20 Listing: Docker command for converting Berlin dataset to .ttl format docker run --rm -v `pwd`/input/:/input/ -v `pwd`/output/:/output/ -v `pwd`/scratch/:/scratch/ -it osm2rdf ./input/berlin-latest.osm.pbf -o /output/berlin-latest.osm.ttl -t /scratch/ Meanwhile, since khm and ata datasets are already available in zipped turtle format files, all was needed to be done is unzipping the files. Listing: Details of datasets after conversion and unzipping processes dataset type new size source berlin-latest.osm.pbf city 23.56 GB geofabrik khm.osm.ttl.bz2 country 2.56 GB osm2rdf ata.osm.ttl.bz2 continent 513.7 MB osm2rdf 3.4. Functionality Testing on Static Example Queries In addition to the functionality testing with examples provided exclusively by GraphDB in section 2.6 as an initial functionality testing phase, there are many other GeoSPARQL functions and spatial relations that shall also be tested in order to conduct the functionality testing between chosen implementations and make a comparison. However, instead of testing the functionality using the sample RDF document provided by either of the implementations, all sets of examples from Cambridgesemantics 31 will be used. This decision is motivated by the observation that, among the examples provided by GraphDB (see Section 2.6.1.4), example 4 shows a different result when executed on Apache Jena Fuseki. Listing: Result of Example 4 tested on Apache Jena Fuseki f http://example.org/ApplicationSchema#A http://example.org/ApplicationSchema#B http://example.org/ApplicationSchema#D Furthermore, the use of example queries in the testing of GeoSPARQL functions from Cambridgesemantics reflects the actual functionality observed in both implementations. For instance, geof:distance returned no results on Apache Jena Fuseki (also see Section 4.4 for further investigation), while GraphDB returned the expected results. The results of this GeoSPARQL functionality testing will be presented and discussed in Chapter 5. During the testing phase on GeoSPARQL functions, however, numerous samples provided by Cambridgesemantics failed due to syntax and incomplete query structure. Therefore, many of the examples were corrected to accommodate this failure. The following figures show an example of such failure and the result after making query adjustment and syntax correction. 31 https://docs.cambridgesemantics.com/graphlakehouse/v3.2/userdoc/geosparql-sparql.htm#GeoSPARQLFunctions
Methodology 21 Figure 4: Query results of geof:rcc8ec from Cambridgesemantics on GraphDB Figure 5: Query results of geof:rcc8ec after query modification on GraphDB 3.5. Use Case Development Based on GeoSPARQL Functionality In order to comprehensively and effectively evaluate the functionality of GeoSPARQL implementations, a set of six designed use cases has been selected corresponding to the datasets obtained and readily processed in section 3.2 and section 3.3 respectively. Two use cases have been designed and assigned for each dataset with different relevant real-world scenarios and questions. The six use cases that follow were developed in consideration of the functions that have demonstrated successful operation in the context of spatial relations. This development is founded on the initial functionality testing that was conducted utilizing the sample RDF document that was provided by GraphDB documented in section 2.5 followed by the documented capabilities in the previous section. They are aimed at testing various aspects of GeoSPARQL compliance, including spatial functions, query complexity and data retrieval accuracy. They also serve as practical scenarios that reflect real-world geospatial queries while ensuring that functionality is mainly tested and adhered to the GeoSPARQL standard.
Methodology 28 Query development 1. Retrieves provinces in Cambodia by defining osmkey:boundary "administrative" and osmkey:admin_level 4 as the pattern. 2. Extracts province centroids using geo:hasCentroid and gets their geometries with geo:asWKT . 3. Filters the results that have ?isoCode as KHprefixes only. 4. Gets attraction/historical sites using osmkey:historic and their spatial representations together with their geometries using geo:hasGeometry and geo:asWKT respectively. 5. Computes the distances between each province centroid to the sites’. 6. Filters the distances less than 50km. 7. Sets ?en for ?provinceName instead of ?name if either exists. 8. Aggregates the provinces using GROUP BY ?provinceName followed by ?isoCode . 9. Sorts the results by the count of historical sites of each province in descending order. Listing: Query for Use Case #3 SELECT ?provinceName ?isoCode (COUNT(?attraction) AS ?attractionCount) WHERE { # Get provinces ?province osmkey:boundary "administrative" ; osmkey:admin_level 4 ; osmkey:name ?name ; osmkey:name:en ?en ; osmkey:ISO3166-2 ?isoCode ; geo:hasCentroid ?provinceCentroid . ?provinceCentroid geo:asWKT ?provinceCentroidWKT . FILTER(REGEX(STR(?isoCode), "^KH-")) # Filters only Cambodian provinces # Get historical sites ?attraction osmkey:historic ?type ; geo:hasGeometry ?attrGeom . ?attrGeom geo:asWKT ?attractionWKT . # Compute distance from province centroid to attraction BIND(geof:distance(?provinceCentroidWKT, ?attractionWKT, uom:metre) AS ?distance) # Ensure the attraction is within a reasonable distance for counting FILTER(?distance < 50000) # English name PREFERRED if either exists
Methodology 29 OPTIONAL { BIND(COALESCE(?en, ?name) AS ?provinceName) } } GROUP BY ?provinceName ?isoCode ORDER BY DESC(?attractionCount) Expected results A list of provinces with their names, ISO codes, and the count of historical sites for each province. 3.5.2.4. Use Case # 4 Objective Finds the cafes that are within or intersect the Russian market by defining the bounding box to the market's coordinates and verifying that the cafes are actually within or intersect the bounding box. GeoSPARQL functions and properties • geo:wktLiteral – Serializes the geometries to WKT format. • geof:sfWithin – Checks if the coffee shops are within the bounding box of the market. • geof:sfIntersects – Checks if the coffee shops intersect with the bonding box defined for the market. Query development 1. Defines a bounding box of 750m for Russian Market using BIND and geo:wktLiteral . 2. Retrieves coffee shop and their OBB geometry using osmkey:amenity with osm2rdfgeom:obb . 3. Optionally gets the coffee names. 4. Converts the OBB geometries to WKT using geo:wktLiteral . 5. Filters the shops that are within the bounding box. 6. Filters the shops that intersect with the bounding box. Listing: Query for Use Case #4 # get coffee shops within or intersects with Russian Market Phnom Penh of 750m SELECT ?cafe ?cafeName ?coffeeShopPoint WHERE { # Define the 750m Bounding Box Around Russian Market BIND("POLYGON((104.9072 11.5338, 104.9072 11.5472, 104.9222 11.5472, 104.9222 11.5338, 104.9072 11.5338))"^^geo:wktLiteral AS ?ttpMarket)
Methodology 30 # manually define the coffee shop's point - La'r Cafe # BIND("POINT(104.9146614 11.5401775)"^^geo:wktLiteral AS ?coffeeShopPoint) # Get coffee shops ?cafe osmkey:amenity "cafe" ; osm2rdfgeom:obb ?obb . # Extract OBB geometry OPTIONAL { ?cafe osmkey:name ?cafeName} # Convert OBB to WKT BIND(REPLACE(STR(?obb), "POLYGON\\(\\(([^ ]+) ([^ ]+)\\)\\)", "POINT($1 $2)") AS ?pointWKT) BIND(STRDT(?pointWKT, geo:wktLiteral) AS ?coffeeShopPoint) # filter the shops within the bounding box FILTER(geof:sfWithin(?coffeeShopPoint, ?ttpMarket)) # filter the shops intersect with the bounding box FILTER(geof:sfIntersects(?coffeeShopPoint, ?ttpMarket)) } Expected results A list of coffee shops, their names, and their coordinates. 3.5.2.5. Use Case # 5 Objective Retrieves the research stations or institutes and their geometries by combining multiple conditions of searching for the stations. GeoSPARQL functions and properties • geo:hasGeometry – Gets the stations geometry representation. • geo:hasCentroid – Gets the station centroids. • geo:asWKT – Gets the station geometries in WKT format. Query development 1. Retrieves the stations by combining the patterns of triples with osmkey:research, with osmkey:amenity having "research_institute" as object and triples with osmkey:office .
Methodology 31 2. Retrieves station geometries and centroids in WKT formats using geo:hasGeometry, and geo:hasCentroid with geo:asWKT values. 3. Obtains station names as ?name and English names as ?name_en if exist. 4. Sets ?name as the value if ?name_en also exists for ?stationName. 5. Order by the station names in ascending order. Listing: Query for Use Case #5 SELECT ?station ?stationName ?stationCentroidWKT WHERE { { ?station osmkey:research ?type . } UNION { ?station osmkey:amenity "research_institute" . } UNION { ?station osmkey:office ?type . } ?station geo:hasGeometry ?stationGeom . ?stationGeom geo:asWKT ?stationWKT . ?station geo:hasCentroid ?stationCentroid . ?stationCentroid geo:asWKT ?stationCentroidWKT . OPTIONAL { ?station osmkey:name ?name} OPTIONAL { ?station osmkey:name:en ?name_en} OPTIONAL { BIND(COALESCE(?name, ?name_en) AS ?stationName) } } ORDER BY ?stationName Expected results A list of stations, their names and their centroids. 3.5.2.6. Use Case # 6 Objective Retrieves the research stations or institutes and their geometries and the nearest volcanoes by calculating the distances between the institute centroids and the volcanoes and filtering those that are less than 1 km.
Methodology 32 GeoSPARQL functions and properties • geof:distance – Calculates distance between the stations and the volcanoes. • geo:hasGeometry – Gets station geometries representation. • geo:asWKT – Gets station geometries stored in WKT format. • geo:hasCentroid – Gets station centroids. Query development 1. Retrieves the stations by combining the patterns of triples with osmkey:research, with osmkey:amenity having "research_institute" as object and triples with osmkey:office . 2. Retrieves station geometries and centroids in WKT formats using geo:hasGeometry, and geo:hasCentroid with geo:asWKT values. 3. Optionally gets station names with osmkey:name if they exist. 4. Retrieves volcanoes with osmkey:natural "volcano" and their geometries geo:hasGeometry with geo:asWKT. 5. Optionally gets volcano names and statues with osmkey:name and osmkey:volcano:status as ?volcanoName and ?status. 6. Calculates the distance between station and volcano centroids using geof:distance in metre unit and filters the stations that have volcanoes less than 1km away. 7. Order by volcano names in ascending order. Listing: Query for Use Case #6 SELECT distinct ?station ?stationName ?stationCentroidWKT ?volcano ?volcanoName ?status ?volcanoWKT (geof:distance(?stationCentroidWKT, ?volcanoWKT, uom:metre) AS ?distance) WHERE { # Get stations { ?station osmkey:research ?value . } UNION { ?station osmkey:amenity "research_institute" . } UNION { ?station osmkey:office ?value . } ?station geo:hasGeometry ?stationGeom . ?stationGeom geo:asWKT ?stationWKT .
Methodology 33 # Get station centroids ?station geo:hasCentroid ?stationCentroid . ?stationCentroid geo:asWKT ?stationCentroidWKT . OPTIONAL { ?station osmkey:name ?stationName } # Get volcanoes ?volcano osmkey:natural "volcano"; geo:hasGeometry ?geom . ?geom geo:asWKT ?volcanoWKT . OPTIONAL { ?volcano osmkey:name ?volcanoName } OPTIONAL { ?volcano osmkey:volcano:status ?status } FILTER(geof:distance(?stationCentroidWKT, ?volcanoWKT, uom:metre) < 1000) } ORDER BY ?volcanoName LIMIT 10 Expected results A list of stations, their names, their centroids together with the nearest volcanoes, the distances between them and the volcano statuses. 3.6. Scenarios Development for GeoSPARQL Functionality Testing In addition to the examples from Cambridgesemantics that were modified and used to test the functionality of GeoSPARQL functions on both Apache Jena Fuseki and GraphDB in section 3.4, scenarios in which different functions were utilized were also developed. These scenarios were used to further test the functionality of the implementations on one of actual datasets, the “Berlin” dataset. The experiment and the results of this functionality testing using the developed scenarios are presented in Chapter 4 and Chapter 5 respectively. The listings that follow present the scenarios in question, along with the relevant functions utilized. The queries developed in this section will be included in the Attachments section. Listing: Scenarios development list for Functionality Testing # Description Functions Used 1 Get intersections of parks and rivers geof:sfIntersects, geof:intersection 2 Get convexHull of parks and rivers and combine them geof:convexHull, geof:union 3 Find parks without overlapping with rivers geof:difference, geof:sfEquals
Methodology 34 4 Get symmetrical differences of parks and rivers geof:symDifference 5 Extract boundaries and envelopes geof:boundary, geof:envelope 6 Check coordinate systems geof:getSRID 7 Get roads touching but not crossing buildings geof:ehMeet 8 Get coffee shops inside malls geof:sfContains, geof:sfTouches, geof:sfCrosses 9 Find associated infrastructure that pass-through university campuses geof:relate 10 Get parks not related to lakes geof:sfDisjoint 11 Get parks that overlap with forests geof:sfOverlaps 12 Get roads that disconnect or from hospital but within 100m geof:rcc8dc, geof:rcc8eq 3.7. Benchmarking Meanwhile, “HOBBIT” 32 , a benchmarking platform is chosen to perform Benchmarking process for the implementations. The platform was introduced in the paper [7] “Software for the GeoSPARQL compliance benchmark” where the author introduced the program and emphasized its importance of usage. The objective of benchmarking is to assess and compare the performance of Apache Jena Fuseki and GraphDB. To achieve this, three components must be taken into consideration: query execution time, indexing time, and resource consumption monitoring. In order to ensure the fairness of the evaluation, use case queries will be executed in the same hardware setup environment using the same available computational resources. In addition, the results of the query execution are derived from the execution of the queries multiple of times for each use case on each implementation. The total execution time for each use case is subsequently calculated and averaged for analyzing the caching effects of each implementation. 32 https://hobbit-project.github.io/
Experiments 35 4. Experiments This chapter presents a series of experiments conducted on both GeoSPARQL implementations, encompassing the initial phase of data loading and the subsequent benchmarking process. It is important to note that certain changes were made during the course of these experiments, and these modifications will be addressed throughout the subsequent sections. Ultimately, an experiment with an unplanned program, called “Qlever” will also be included, to showcase its potential and prospective contributions in GeoSPARQL implementation and support. 4.1. Hardware and Software During the preparation and experimentation phases of the thesis, a local server environment was used, consistent throughout both stages. The hardware and software specifications of this environment are as follows: Listing: Details of hardware and software setup environment Chip Apple M1 Pro Software Apache Jena Fuseki GraphDB Memory 16 GB Version 5.2.0 10.8.3 OS macOS Sonoma 14.5 License Open Source Freeware Cores 8 Edition Standalone Standalone 4.2. Data Loading First and foremost, the datasets were loaded into the chosen implementations of choice so that the experiment could commence. Indexing is the most important and crucial process in loading the datasets to the GeoSPARQL implementations. Both implementations provide both command-line and interface options for loading the data. Additionally, each system has its own parameters used for tweaking to fit the setup environment, such as load, along with parallel mode activation and preload provided by GraphDB, etc. While both systems offer the interface option for serving the purpose, during the experiments, however, indicated that GraphDB failed to index the triples properly for the Berlin dataset, which is the largest at over 20GB, without producing any warning or error messages, despite the fact that the successfully imported message was shown on the screen. The duration and the storage taken for indexing each dataset on each implementation will be presented and discussed in the following chapter. 4.2.1. Apache Jena Fuseki In the case of Apache Jena Fuseki, the datasets were successfully loaded and indexed through the interface provided by the program itself when initiated with sufficient amount of heap size memory. Especially, for this particular experiment, customized amount of heap size was assigned when running the program before attempting to load the datasets; Berlin dataset, for instance, with a size greater than 20GB. The following command shows the command that
Experiments 36 allocate 10GB of maximum heap size with -Xmx10g option when running the program, while the following figures after, show the steps onto creating and loading the datasets on Apache Jena Fuseki via the interface after the program has started. Listing: Command for running Apache Jena Fuseki with 10GB of memory allocation java -Xmx10g -cp "fuseki-server.jar:lib/*" org.apache.jena.fuseki.cmd.FusekiCmd --config run/config.ttl --debug Figure 7: Datasets manage page on Apache Jena Fuseki On the interface, select manage and new dataset tab where name and type of dataset are expected. During the experiment, all datasets were created with Persistent (TDB2) mode to avoid getting the data erased after closing the program. Figure 8: Add-data option on Apache Jena Fuseki
Experiments 37 After creating the datasets, the data turtle files that were completely processed in Section 3.3, could be loaded with add data option in the existing datasets tab. 4.2.2. GraphDB In addition to the previously mentioned possibility of indexing failure during the data importing process as outlined earlier, the Berlin dataset was only able to be loaded and indexed successfully using the load option with parallel mode enabled after multiple trials and errors. The following commands were used to load all the datasets into GraphDB: Listing: Command for loading (indexing) Berlin dataset ➜ graphdb-10.8.3 ./bin/importrdf load -f -i berlin -m parallel /Users/macbookpro/Documents/ProITD/S3/Thesis/testing/osm2rdf/output/berlinlatest-3.osm.ttl Listing: Command for loading (indexing) Cambodia (khm) dataset ➜ graphdb-10.8.3 ./bin/importrdf preload -f -i khm /Users/macbookpro/Documents/ProITD/S3/Thesis/testing/osm2rdf/output/khm.osm.ttl Listing: Command for loading (indexing) Antarctica (ata) dataset ➜ graphdb-10.8.3 ./bin/importrdf load -f -i ata -m parallel /Users/macbookpro/Documents/ProITD/S3/Thesis/testing/osm2rdf/output/ata.osm.ttl 4.2.3. Datasets Verification To ensure that the datasets were correctly loaded and that geospatial data exist in the datasets and did not encounter any corruption during the importing process, the following two queries were used. The first query aims to find spatial entities (?s), retrieves their geometric representation (?geo) and ensure that geometry has a WKT representation (geo:asWKT) with limit of 100 and the second query is for retrieving the number of triples, subjects, predicates, and objects. Listing: Retrieves entities with their geometries PREFIX geof: <http://www.opengis.net/def/function/geosparql/> PREFIX geo: <http://www.opengis.net/ont/geosparql#> PREFIX osm2rdfgeom: <https://osm2rdf.cs.uni-freiburg.de/rdf/geom#> SELECT ?s ?geo WHERE { ?s geo:hasGeometry ?geo . ?geo geo:asWKT ?wkt .
Experiments 44 SPARQLWrapper and visualizing the result are enclosed in the Attachments page, the following command shows how a python environment was set up and installed the necessary mentioned libraries. Listing: Setting up virtual environment for the benchmarking process python3 -m venv venv # install a virtual environment called venv source venv/bin/activate # activate the virtual environment pip install SPARQLWrapper pandas matplotlib seaborn # install libraries Once the virtual environment was completely set up and necessary libraries were installed, there was one additional thing needed to be done, the file containing the GeoSPARQL implementation details, SPARQL endpoints precisely, in order to make requests for queries execution. Initially, only JSON file was prepared with the engine and datasets/repositories info on each implementation as shown in the following listing. Listing: Implementations config JSON file (lines truncated for readability) { "jena_fuseki": { "endpoint": "http://localhost:3030", "repos": { "berlin3": { "queries": { "q1_jena": "", "q2_jena": "" } }, "graphdb": { "endpoint": "http://localhost:7200/repositories", "repos": { "berlin": { "queries": { "q1_graphdb": "", "q2_graphdb": "" } However, after making numerous attempts and although not specified on the documentation of the SPARQLWrapper library itself, the requests were only successfully made through the implementations when queries were sent as string literals. As a result, another separate Python file was developed to contain all the queries for the six use cases with variations in the function and syntax for each implementation as documented in Section 4.3 and Section 4.4. The following listing presents the simple Python file containing the queries with the variables
Experiments 45 corresponding to those in the aforementioned config file. The results from this modified benchmarking process using the Python based script will be presented in the Chapter 5. Listing: Python file containing all the queries (lines truncated for readability) q1_jena = """ PREFIX osmkey: <https://www.openstreetmap.org/wiki/Key:> PREFIX spatialF: <http://jena.apache.org/function/spatial#> PREFIX geo: <http://www.opengis.net/ont/geosparql#> PREFIX uom: <http://www.opengis.net/def/uom/OGC/1.0/> # Top 10 places to visit starting with Galeria as the starting point SELECT ?attraction (COALESCE(?attrName, ?operator) AS ?name_operator) ?type ?distance """ 4.6. Scenario-Based Query Testing and Execution In section 3.6, a series of scenarios were introduced and developed to further assess the functionality of GeoSPARQL functions across both implementations. The majority of queries developed based on these scenarios were executed successfully on both platforms. However, two of functions, geof:rcc8eq and geof:rcc8dc, returned no results on Apache Jena Fuseki. This outcome is at contrary to the results obtained when these functions were utilized on the modified dataset from Cambridgesemantics. Additionally, in scenario 4, Apache Jena Fuseki failed to parse the JSON response onto the UI, likely due to validation constraints imposed on response length. The following figures illustrate some of the results obtained and verified through the use of WKT-GeoJSON converter 34 , which facilitates the conversion of WKT values of the obtained geometries to GeoJSON format and their further visualization on geojson.io 35 site. Figure 15: Apache Jena Fuseki failed to parse JSON response for scenario 4 34 https://geojson-to-wkt-converter.onrender.com/ 35 https://geojson.io/
Experiments 46 Figure 16: Response successfully parsed in CSV format for scenario 4 Figure 17: Symdifferences of Eichepark and Neue Wuhle from scenario 4 rendered on geojson-to-wktconverter.onrender.com
Experiments 47 Figure 18: No results returned on Apache Jena Fuseki for scenario 12 Figure 19: Jewish Hospital disconnected from Heinz-Galinski-Straße from scenario 12 visualized on geojson.io 4.7. Experiment with Qlever In Chapter 3, osm2rdf was introduced as a tool for converting the dataset into a compatible working format with the GeoSPARQL implementations. The University of Freiburg actually also introduces their own SPARQL engine called “Qlever” 36 . As stated on the documentation, it is also claimed to be extremely fast and able to efficiently index and query very large queries from knowledge graphs with over 100 billion triples on a single standard PC or server. Qlever also provides a demo site 37 where many knowledge graphs are hosted with given sample 36 https://github.com/ad-freiburg/qlever 37 https://qlever.cs.uni-freiburg.de/wikidata
Experiments 48 SPARQL queries for different use cases. After going through the demo site and seeing how fast, convenient and effective the engine is, during the experiments phase, Qlever was also tried out to see if it also supports GeoSPARQL since it is also on the list of notified contributors for GeoSPARQL updates on Github. The easiest way to install qlever is using pip, the package manager for Python. It is also important to note that docker and docker demon are required throughout the process. Listing: Command for installing qlever pip install qlever After the installation, several commands and options are available for use for setting up the data repositories, starting/stopping the server and accessing the UI, etc. The commands are as followed: Commands List qlever setup-config olympics # Get Qleverfile (config file) for this dataset qlever get-data # Download the dataset qlever index # Build index data structures for this dataset qlever start # Start a QLever server using that index qlever example-queries # Launch some example queries qlever ui # Launch the QLever UI There are several configuration templates available 38 which could be chosen from based on actual needs. In fact, it is even greater that the actual qlever project which is referred to as qlever-control, equivalent to pip install qlever could also be cloned for modification and further development with the following steps: Steps to cloning and developing qlever-control git clone https://github.com/ad-freiburg/qlever-control cd qlever-control pip install -e . vvz Qleverfile template was chosen for setting up Berlin dataset on the engine with the “qlever setup-config vvz ” command as the configuration parameters in the template file were the closest to the actual local machine environment. However, due to the limited computational resources, some changes were made and adjusted until the successful attempt at indexing the dataset. The following shows the final Qleverfile that successfully managed to index the dataset. Qleverfile for Berlin dataset (lines truncated for readability) 38 https://github.com/ad-freiburg/qlever-control/tree/main/src/qlever/Qleverfiles
Experiments 49 [data] NAME = berlin [index] INPUT_FILES = berlin-latest.osm.rdf.bz2 CAT_INPUT_FILES = bzcat ${INPUT_FILES} PARSER_BUFFER_SIZE = 50M STXXL_MEMORY = 2G SETTINGS_JSON = { "ascii-prefixes-only": false, "num-triples-per-batch": 50000 } [server] MEMORY_FOR_QUERIES = 12G [ui] UI_CONFIG = olympics UI_PORT = 8180 The most important section index where necessary adjustments were made and could be done depending on the server specifications. Addtionally, on UI_CONFIG line, only olympics could be used. Otherwise, the program was not running properly. It is unknown whether this was a bug. Once the Qleverfile is ready, qlever get-data was used to prepare for dataset retrieval which is located locally in this case, and export the settings properties to json file for indexing process. Next, indexing process started with qlever index command. During the indexing process which is usually the most time-consuming process in loading the RDF documents, an index text log file was also produced alongside the command line prompt which is helpful for debugging and monitoring the process. The following shows the result from the index log file. It is noticeable that the indexing process took approximately 20 minutes for Berlin dataset which is more than 20 GB in size which is the fastest among the chosen GeoSPARQL implementations. Index log file (lines truncated for readability) 2025-03-08 13:12:19.384 - INFO: [1mQLever IndexBuilder, compiled on Sat Feb 22 02:04:18 UTC 2025 using git hash 8fe064[22m 2025-03-08 13:12:19.392 - INFO: You specified "num-triples-per-batch = 50,000", choose a lower value if the index builder runs out of memory 2025-03-08 13:31:34.612 - INFO: Statistics for PSO: #relations = 5,999, #blocks = 12,014, #triples = 373,720,412 2025-03-08 13:31:35.406 - INFO: Index build completed Once the dataset was completely and successfully indexed, qlever start and qlever ui commands were used to start the server and launch the UI workbench. Once the UI started, testing with GeoSPARQL queries on Qlever could be done. The following figure shows the results of testing a sample query provided by Qlever for listing all German cities.
Experiments 50 Figure 20: Listing all German cities on Qlever However, when tested with the actual GeoSPARQL functions such as geof:buffer and geof:distance, errors were returned. The following figures show the errors when testing both functions. In the case of geof:buffer, the error message indicated that the function is not currently supported. This error would be very informative in the development and debugging processes if GeoSPARQL were supported. Figure 21: Error message for geof:buffer on Qlever On the other hand, the error provided when trying to use geof:distance is misleading and appears to be inaccurate, despite the syntax and query structure being correct. Figure 22: Error message for geof:distance on Qlever
Results and Discussion 51 5. Results and Discussion 5.1. Functionality Testing The results of the functionality testing on GeoSPARQL functions using static example queries with modifications (see section 3.4) indicate that GraphDB supported a greater number of functions than Apache Jena Fuseki when tested on a different dataset than the one provided by either of the implementation, as well as the one found in on compliance and benchmarking studies 39 . Some of the functions below from the list were used when building the queries for use case development (see section 3.5). It is particularly interesting to note that geof:distance does not appear to be supported in Apache Jena Fuseki, as indicated by the list, that aligns with the functionality of the function when executing the queries in the specified use cases (see section 4.4 for further details). Listing: Result of Functionality Testing using static example queries Type Function Jena Fuseki GraphDB Non-Topological geof:distance ☑ Non-Topological geof:buffer Non-Topological geof:convexHull ☑ ☑ Non-Topological geof:intersection ☑ ☑ Non-Topological geof:union ☑ ☑ Non-Topological geof:difference ☑ ☑ Non-Topological geof:symDifference ☑ ☑ Non-Topological geof:envelope ☑ ☑ Non-Topological geof:boundary ☑ ☑ Non-Topological geof:getSRID ☑ ☑ Non-Topological Geof:relate ☑ ☑ Simple Feature geof:sfEquals ☑ ☑ Simple Feature geof:sfDisjoint ☑ ☑ Simple Feature geof:sfIntersects ☑ ☑ Simple Feature geof:sfTouches ☑ ☑ Simple Feature geof:sfCrosses ☑ ☑ Simple Feature geof:sfWithin ☑ ☑ Simple Feature geof:sfContains ☑ ☑ Simple Feature geof:sfOverlaps ☑ ☑ Egenhofer geof:ehEquals ☑ ☑ Egenhofer geof:ehDisjoint ☑ Egenhofer geof:ehMeet ☑ ☑ Egenhofer geof:ehOverlap ☑ ☑ Egenhofer geof:ehCovers ☑ ☑ 39 https://github.com/OpenLinkSoftware/GeoSPARQLBenchmark/tree/master/src/main/resources/gsb_dataset
Results and Discussion 52 Egenhofer geof:ehCoveredBy ☑ ☑ Egenhofer geof:ehInside ☑ Egenhofer geof:ehContains RCC8 geof:rcc8eq ☑ ☑ RCC8 geof:rcc8dc ☑ ☑ RCC8 geof:rcc8ec ☑ RCC8 geof:rcc8po ☑ RCC8 geof:rcc8tpp ☑ RCC8 geof:rcc8tppi ☑ RCC8 geof:rcc8ntpp ☑ RCC8 geof:rcc8ntppi ☑ Meanwhile, the results from conducting functionality testing on the “Berlin” dataset where the experiments of execution on scenario-based queries were detailed in section 4.6 suggest that approximately, 65% of the functions out of the 35 functions in the previous listing were tested and resulted to match with results from testing on the static sample queries with the exception of two functions geof:rcc8eq and geof:rcc8dc used in scenario number 12. Listing: Result of Functionality Testing using static example queries Type Function Jena Fuseki GraphDB Non-Topological geof:distance ☑ Non-Topological geof:buffer Non-Topological geof:convexHull ☑ ☑ Non-Topological geof:intersection ☑ ☑ Non-Topological geof:union ☑ ☑ Non-Topological geof:difference ☑ ☑ Non-Topological geof:symDifference ☑ ☑ Non-Topological geof:envelope ☑ ☑ Non-Topological geof:boundary ☑ ☑ Non-Topological geof:getSRID ☑ ☑ Non-Topological geof:relate ☑ ☑ Simple Feature geof:sfEquals ☑ ☑ Simple Feature geof:sfDisjoint ☑ ☑ Simple Feature geof:sfIntersects ☑ ☑ Simple Feature geof:sfTouches ☑ ☑ Simple Feature geof:sfCrosses ☑ ☑ Simple Feature geof:sfWithin ☑ ☑ Simple Feature geof:sfContains ☑ ☑ Simple Feature geof:sfOverlaps ☑ ☑ Egenhofer geof:ehMeet ☑ ☑
Results and Discussion 53 RCC8 geof:rcc8eq ☑ RCC8 geof:rcc8dc ☑ 5.2. Benchmarking One of the components for the benchmarking process outlined in the objectives section in chapter 1 is indexing time, defined as the duration and space required for each dataset to be imported and loaded onto the repositories and datasets of both platforms. The following list illustrates that Apache Jena Fuseki generally takes less time to import and index small and mid-sized datasets. However, it took nearly twice as much time to index the “Berlin” dataset. Additionally, it requires more storage space for indexed directories. Listing: Result of benchmarking indexing time Dataset Indexing Duration Storage Taken Jena GraphDB Jena GraphDB ata 2mn 7mn21s 1.06 GB 792.9 MB khm 5mn 10mn 5.6 GB 3.31 GB berlin 4hr50m 2hr48m 49.83 GB 34.81 GB Furthermore, the results of the benchmarking on the use case queries execution using a Python-based script demonstrate that both implementations returned the same number of results for all the use cases, with the execution time varying for each case between the two. Apache Jena Fuseki outperformed GraphDB in five out of the six use cases, with the time reduction ranging from approximately 40% to more than 150% in milliseconds. Listing: Result of benchmarking the use cases execution Use Case Returned Results Execution Time (ms) Difference in % Jena GraphDB Jena GraphDB Q1 10 10 83.9 219.54 161.67% Q2 10 10 3723.55 5336.51 43.32% Q3 25 25 263 510.34 94.05% Q4 67 67 61.44 126.01 105.09% Q5 137 137 14.06 36.54 159.89% Q6 10 10 43.31 27.87 -35.65%
Attachments 60 Listing: Python script used for transforming and visualizing the benchmark results import pandas as pd import matplotlib.pyplot as plt import seaborn as sns def transform_results(csv): df = pd.read_csv(csv) # transform Implementation column values df["Implementation"] = df["Implementation"].map(lambda x: "Apache Jena Fuseki" if "3030" in x else "GraphDB") # transform use case column values df["Use Case"] = df["Use Case"].map(lambda x: x.replace("_jena", "") if "_jena" in x else x.replace("_graphdb", "")).astype('str') # transform execution time df["Execution Time (ms)"] = round(df["Execution Time (ms)"],2) df.to_csv(f"transformed_benchmarked_results_{csv}.csv") return df def visualize(df, file_name): df = df.sort_values(['Implementation', 'Use Case']) df["Use Case"] = df["Use Case"].astype(str) plt.figure(figsize=(10, 5)) sns.lineplot(data=df, x="Use Case", y="Execution Time (ms)", hue="Implementation", marker="o") plt.title("GeoSPARQL Use Case Execution Time Comparison") plt.xticks(rotation=45) for implementation in df["Implementation"].unique(): subset = df[df["Implementation"] == implementation] max_index = subset["Execution Time (ms)"].idxmax() max_row = subset.loc[max_index] peak_x = max_row["Use Case"] peak_y = max_row["Execution Time (ms)"] print(f"Implementation: {implementation}, Peak: {peak_x}, {peak_y}")
Attachments 61 plt.annotate( f"{peak_y:.2f} ms", (peak_x, peak_y), textcoords="offset points", xytext=(0, 10), ha='center', fontsize=10, color="black", weight="bold" ) plt.savefig(f"query_execution_comparison_{file_name}.png", dpi=350, bbox_inches='tight') plt.show() def pivot_results(df, file_name): new_df = df.pivot(index="Use Case", columns=["Implementation"], values=["Results", "Execution Time (ms)"]).sort_values(['Use Case']) new_df.to_csv(f"benchmarks_{file_name}.csv") file_name = "benchmarks_result_26032025804.csv" df = transform_results(file_name) pivot_results(df, file_name) visualize(df, file_name)
Attachments 62 Listing: Query for scenario 1 SELECT ?park ?parkName ?river ?riverName ?intersection WHERE { ?park osmkey:leisure "park" ; geo:hasGeometry ?parkGeo . ?parkGeo geo:asWKT ?parkWKT . OPTIONAL{?park osmkey:name ?parkName} ?river osmkey:waterway "river" ; geo:hasGeometry ?riverGeo . ?riverGeo geo:asWKT ?riverWKT . OPTIONAL{?river osmkey:name ?riverName} FILTER(geof:sfIntersects(?parkWKT, ?riverWKT)) BIND(geof:intersection(?parkWKT, ?riverWKT) AS ?intersection) } order by ?intersection Listing: Query for scenario 2 SELECT ?park ?river (geof:convexHull(?parkWKT) AS ?parkHull) (geof:convexHull(?riverWKT) AS ?riverHull) (geof:union(geof:convexHull(?parkWKT), geof:convexHull(?riverWKT)) AS ?combinedHull) WHERE { ?park osmkey:leisure "park" ; geo:hasGeometry ?parkGeo . ?parkGeo geo:asWKT ?parkWKT . ?river osmkey:waterway "river" ; geo:hasGeometry ?riverGeo . ?riverGeo geo:asWKT ?riverWKT . FILTER(geof:sfIntersects(?parkWKT, ?riverWKT)) Listing: Query for scenario 3 SELECT ?park ?parkName ?river ?riverName (geof:difference(?parkWKT, ?riverWKT) AS ?parkWithoutRiver) WHERE { ?park osmkey:leisure "park" ; geo:hasGeometry ?parkGeo . ?parkGeo geo:asWKT ?parkWKT . OPTIONAL {?park osmkey:name ?parkName} ?river osmkey:waterway "river" ;
Attachments 63 geo:hasGeometry ?riverGeo . ?riverGeo geo:asWKT ?riverWKT . OPTIONAL {?river osmkey:name ?riverName} FILTER(!STRSTARTS(STR(?parkWKT), "GEOMETRYCOLLECTION")) FILTER(!STRSTARTS(STR(?riverWKT), "GEOMETRYCOLLECTION")) FILTER(!geof:sfEquals(?parkWKT, ?riverWKT)) } Listing: Query for scenario 4 SELECT ?park ?parkName ?river ?riverName (geof:symDifference(?parkWKT, ?riverWKT) AS ?uniqueAreas) WHERE { ?park osmkey:leisure "park"; geo:hasGeometry/geo:asWKT ?parkWKT. OPTIONAL {?park osmkey:name ?parkName} ?river osmkey:waterway "river"; geo:hasGeometry/geo:asWKT ?riverWKT. OPTIONAL {?river osmkey:name ?riverName} FILTER(geof:sfIntersects(?parkWKT, ?riverWKT)) } LIMIT 10 Listing: Query for scenario 5 SELECT ?river (geof:boundary(?riverWKT) AS ?riverBoundary) (geof:envelope(?riverWKT) AS ?riverBoundingBox) WHERE { ?river osmkey:waterway "river" ; geo:hasGeometry [ geo:asWKT ?riverWKT ]. FILTER(!STRSTARTS(STR(?riverWKT), "GEOMETRYCOLLECTION")) } Listing: Query for scenario 6 SELECT ?river (geof:getSRID(?wkt) AS ?srid) WHERE { ?river osmkey:waterway "river"; geo:hasGeometry/geo:asWKT ?wkt. FILTER(geof:getSRID(?wkt) != "EPSG:4326"^^xsd:string) } LIMIT 5
Attachments 64 Listing: Query for scenario 7 SELECT ?road ?building WHERE { ?road osmkey:highway []; geo:hasGeometry/geo:asWKT ?rWKT. ?building osmkey:building []; geo:hasGeometry/geo:asWKT ?bWKT. FILTER(geof:ehMeet(?rWKT, ?bWKT)) } Listing: Query for scenario 8 SELECT ?mall ?mallName ?amenity ?amenityName (geof:sfContains(?mallWKT, ?amenityWKT) AS ?contains) (geof:sfTouches(?mallWKT, ?amenityWKT) AS ?touches) (geof:sfCrosses(?mallWKT, ?amenityWKT) AS ?crosses) ?mallWKT ?amenityWKT WHERE { ?mall osmkey:shop "mall" ; osmkey:name ?mallName ; geo:hasGeometry/geo:asWKT ?mallWKT . ?amenity osmkey:amenity "cafe" ; osmkey:name ?amenityName ; geo:hasGeometry/geo:asWKT ?amenityWKT . FILTER( geof:sfContains(?mallWKT, ?amenityWKT) || geof:sfTouches(?mallWKT, ?amenityWKT) || geof:sfCrosses(?mallWKT, ?amenityWKT) ) } LIMIT 10 Listing: Query for scenario 9 SELECT ?university ?universityName ?road ?roadName (geof:relate(?buildingWKT, ?roadWKT, "T*F**FFF*") AS ?relate) WHERE { ?university osmkey:amenity "university"; geo:hasGeometry/geo:asWKT ?buildingWKT. OPTIONAL {?university osmkey:name ?universityName} ?road osmkey:highway ?roadType; geo:hasGeometry/geo:asWKT ?roadWKT. OPTIONAL {?road osmkey:name ?roadName} FILTER(geof:sfIntersects(?buildingWKT, ?roadWKT)) } LIMIT 5
Attachments 65 Listing: Query for scenario 10 SELECT ?park ?parkName ?lake WHERE { ?park osmkey:leisure "park"; geo:hasGeometry/geo:asWKT ?parkWKT. OPTIONAL {?park osmkey:name ?parkName} ?lake osmkey:water "lake" ; geo:hasGeometry ?lakeGeo . ?lakeGeo geo:asWKT ?lakeWKT . FILTER(geof:sfDisjoint(?parkWKT, ?lakeWKT)) } LIMIT 5 Listing: Query for scenario 11 SELECT ?park ?forest WHERE { ?park osmkey:leisure "park" ; geo:hasGeometry ?parkGeo . ?parkGeo geo:asWKT ?parkWKT . ?forest osmkey:landuse "forest" ; geo:hasGeometry ?forestGeo . ?forestGeo geo:asWKT ?forestWKT . FILTER(geof:sfOverlaps(?parkWKT, ?forestWKT)) } LIMIT 5 Listing: Query for scenario 12 SELECT ?hospital ?residential (geof:rcc8dc(?hWKT, ?rWKT) AS ?is_disconnected) (geof:rcc8eq(?hWKT, ?rWKT) AS ?is_equal) WHERE { ?hospital osmkey:amenity "hospital"; geo:hasGeometry/geo:asWKT ?hWKT. ?residential osmkey:highway "residential"; geo:hasGeometry/geo:asWKT ?rWKT. FILTER(geof:rcc8dc(?hWKT, ?rWKT)) || FILTER(geof:rcc8eq(?hWKT, ?rWKT)) FILTER(geof:distance(?hWKT, ?rWKT, uom:metre) < 100) } LIMIT 5
List of Literature 66 List of Literature [1] N. J. Car and T. Homburg, “GeoSPARQL 1.1: Motivations, Details and Applications of the Decadal Update to the Most Important Geospatial LOD Standard,” ISPRS Int. J. Geo-Inf., vol. 11, no. 2, p. 117, Feb. 2022, doi: 10.3390/ijgi11020117. [2] Luís Moreira de Sousa, “Spatial Linked Data Infrastructures,” Oct. 05, 2024, Zenodo. doi: 10.5281/ZENODO.13892963. [3] A. G. Cohn, B. Bennett, J. Gooday, and N. M. Gotts, “Qualitative Spatial Representation and Reasoning with the Region Connection Calculus,” Geoinformatica, vol. 1, no. 3, pp. 275–316, 1997, doi: 10.1023/A:1009712514511. [4] R. Battle and D. Kolas, “Enabling the geospatial Semantic Web with Parliament and GeoSPARQL,” Semantic Web, vol. 3, no. 4, pp. 355–370, 2012, doi: 10.3233/SW-20120065. [5] M. Jovanovik, T. Homburg, and M. Spasić, “A GeoSPARQL Compliance Benchmark,” ISPRS Int. J. Geo-Inf., vol. 10, no. 7, p. 487, Jul. 2021, doi: 10.3390/ijgi10070487. [6] H. Bast, P. Brosi, J. Kalmbach, and A. Lehmann, “An Efficient RDF Converter and SPARQL Endpoint for the Complete OpenStreetMap Data,” in Proceedings of the 29th International Conference on Advances in Geographic Information Systems, Beijing China: ACM, Nov. 2021, pp. 536–539. doi: 10.1145/3474717.3484256. [7] M. Jovanovik, T. Homburg, and M. Spasić, “Software for the GeoSPARQL compliance benchmark,” Softw. Impacts, vol. 8, p. 100071, May 2021, doi: 10.1016/j.simpa.2021.100071.
Statutory Declaration 67 Statutory Declaration I herewith formally declare that I have written the submitted thesis independently in all parts. I did not use any outside support except for the quoted literature and other sources mentioned in the paper. I declare that the thesis has not been presented in the same or a similar form in any other examination. I clearly marked and separately listed all of the literature and all of the other sources which I employed when producing this academic work, either literally or in content. I am aware that the violation of this regulation will lead to failure of the thesis. Information on the use of AI-based tools I declare that I have not used any AI-based tools whose use has been explicitly excluded in writing by the examiner. I am aware that the use of texts or other content and products generated by AI-based tools does not guarantee their quality. I am fully responsible for the adoption of any machine-generated passages used by me and bear responsibility for any incorrect or distorted content generated by the AI, incorrect references, violations of data protection and copyright law or plagiarism. I also declare that my creative influence predominates in this work. Samneang Seng Student’s name Student’s signature 590688 30.03.2025 Matriculation number Berlin, date