scieee AI-readable full text Open interactive document viewer

KGSEC: A Modular Framework for Knowledge Graph Schema Extraction and Comparison

Skoufis, Petros; Skoutas, Dimitrios

Abstract

Finding the underlying schema in knowledge graphs is an imperative operation for various tasks, such as query formulation or exploration. This task becomes even harder, when data are incomplete, noisy or are collected via multiple sources with different schemata that are combined. Several algorithms for extracting an implicit schema from a given knowledge graph have been proposed in the literature. However, the lack of a common framework and evaluation metrics makes it difficult to combine them and compare the results. To fill this gap, we present a modular three-stage framework and we have developed a Python library and web application that performs schema extraction and allows users to visually assess and compare the results. The developed tool, called KGSEC, facilitates experimentation and increases interactivity. Given that the quality of a schema is largely subjective, depending on the user's needs and preferences, KGSEC can make it easier and faster for users to generate a schema that is better tailored to their task.

Full text

KGSEC: A Modular Framework for Knowledge Graph Schema Extraction and Comparison Petros Skoufis IMSI, “Athena” Research Center Athens, Greece [email protected] Dimitrios Skoutas IMSI, “Athena” Research Center Athens, Greece [email protected] Abstract—Finding the underlying schema in knowledge graphs is an imperative operation for various tasks, such as query formulation or exploration. This task becomes even harder, when data are incomplete, noisy or are collected via multiple sources with different schemata that are combined. Several algorithms for extracting an implicit schema from a given knowledge graph have been proposed in the literature. However, the lack of a common framework and evaluation metrics makes it difficult to combine them and compare the results. To fill this gap, we present a modular three-stage framework and we have developed a Python library and web application that performs schema extraction and allows users to visually assess and compare the results. The developed tool, called KGSEC, facilitates experimentation and increases interactivity. Given that the quality of a schema is largely subjective, depending on the user’s needs and preferences, KGSEC can make it easier and faster for users to generate a schema that is better tailored to their task. I. INTRODUCTION The amount of data produced in our era is exponentially growing, due to the emergence and wide use of new data sources. Knowledge Graphs (KG) provide a powerful and versatile means for combining and integrating diverse data from heterogeneous sources, including structured, unstructured and semi-structured ones. However, this flexibility is also an obstacle for usability. In particular, the lack of an explicit schema of the data contained in the KG impedes users from effectively evaluating their value and their adequacy for many data analysis tasks [1]. Thus, providing effective ways to automatically extract an implicit schema from a given KG is a critical task for democratizing the use of KGs and increasing their adoption. The task of extracting or discovering an implicit schema from a given KG has been studied in the literature over the past years, with proposals varying in terms of approach and tools used [1]. However, despite the fact that these approaches share some commonalities, there is a lack of a unifying framework, which makes it difficult to combine and compare results from different algorithms. This becomes even more problematic given the subjective nature of what constitutes a good schema. In particular, there is no consensus on how to evaluate the produced schemata. Certain works formulate the schema extraction task as a clustering problem, while using the explicit type declarations available in the KG as a ground truth and calculating metrics, such as F1-score on that [2]–[4]. However, this approach essentially limits the scope of schema extraction into the operation of reconstructing the explicit schema that is suggested by the included type declarations, which in practice is often incomplete or even inaccurate. Other works focus on the characteristics of the final schema, such as the number and size of classes and the depth of the resulting hierarchy [2], [5]. However, it is not straightforward to assess the quality of an extracted schema from such metrics without other alternative schemata for comparison. Additionally, while the proposed methods often share some commonalities, there is a lack of a unifying framework that would facilitate experimenting with different approaches and comparing the results. This lack of modularity and extensibility impedes experimentation with certain modules. There is also a lack of interoperability between the input and output of the various implementations that exist. These differences demand extra assumptions and conversions in order to conduct any comparisons across different works. To address these shortcomings, we propose a unifying, modular and extensible framework for KG schema extraction, and we have developed KGSEC1, a Python library and web application that supports and facilitates this task. KGSEC employs a three-stage processing pipeline. In each stage, users can choose and parameterize various options. The resulting schemata can then be evaluated from multiple perspectives, from objective ground-truth comparison and cluster quality metrics to subjective visualizations. The user of the application can exploit the aforementioned functionality to iteratively generate, explore and compare schemata in an effort to draw conclusions on the graph and its underlying structure and data. II. OVERVIEW AND MAIN COMPONENTS We present KGSEC, a modular framework and tool that is built upon a unifying view over several KG schema extraction works in the literature. KGSEC comprises three main modules, as shown in Figure 1: (a) characteristic set extraction; (b) schema generation, and (c) schema refinement. We present each module below. We also describe the user interface in Section III. A. Characteristic Set Extraction This module takes as input a KG and extracts all its characteristic sets. Intuitively, a characteristic set (CS) is a set 1https://github.com/stelar-eu/kg-schema-extraction Fig. 1. KGSEC architecture overview. of properties that appear together in one or more instances. E.g., the CS of an instance of type Person may contain the properties {name, age, address, phone, ...}. These CSs are used as input by the subsequent schema generation module, as described in Section II-B. CSs are commonly used by KG schema extraction methods in the literature [1]. In the most common and simplest cases, the CS of a non-literal entity ein a KG Gis defined as: CS(e) := {p| ∃o: (e, p, o)∈G}(1) i.e., the set of properties pfor which there exists a triple having eas its subject [6]. Notice that this only includes the outgoing edges of e. The definition can be extended to also include the incoming edges [5]: CS(e) := {p| ∃o: (e, p, o)∈Gor ∃s: (o, p, e)∈G}(2) Even further, the definition can be extended to also take into consideration predefined classes of the graph, through the rdf:type declarations [2], or alternatively the range of the properties [7]. For instance, such more fine-grained CSs can be helpful for splitting a class such as Employee into subclasses based on the type of the entity being the object of the property occupation. It is especially useful when attempting to extract a schema from KGs that have a very large number of classes but relatively few distinct properties. Our CS extraction module supports all above variants, which can be configured by the user. It also allows users to filter out certain properties that may not be relevant for schema extraction. Parsing the input KG is implemented using the Python library RDFLib2. B. Schema Generation This module takes as input the extracted CSs and produces an initial KG schema. The schema is a grouping of the graph’s entities into classes. Depending on the employed method, overlapping or hierarchical classes may be produced. Based on related works in the literature, KGSEC supports two main approaches for schema extraction, as described below. 2https://rdflib.readthedocs.io The first approach is clustering-based. Several works have employed clustering algorithms, including density-based clustering [8], Gaussian Mixture Models [2], and network community detection [9]. In KGSEC, for this purpose, we utilize the clustering module of the popular Python library scikit-learn3, which provides several well-known algorithms (e.g., K-means, DBSCAN, HDBSCAN, etc.) Our library can be used with any sci-kit algorithm that takes as an input a similarity matrix, as well as variants of established approaches [2]. It is also possible to select the desired similarity function (e.g., Jaccard, Cosine, etc.) to be applied when computing the similarity score between two CSs. In this case, each class in the produced schema corresponds to a cluster of similar CSs. The second approach is based on frequent itemsets. Instead of grouping together similar CSs, classes are extracted by discovering frequent patterns [10]. This is based on frequent itemset mining, where each CS is treated as an itemset. In KGSEC, frequent itemset mining is implemented using the Python library MLxtend4. In this case, each class in the produced schema corresponds to a set of properties that frequently co-occur in the input CSs. By assigning each CS to the most appropriate class, based on a scoring function, we manage to produce again a schema as a set of clusters of similar CSs, having this time a predefined description for the class, its associated frequent-itemset. C. Schema Refinement This module performs a post-processing step on the output of the clustering or frequent itemset mining functions. The goal is to produce a final schema that better satisfies certain desirable properties that may be specified by the user. This includes operations such as pruning classes that contain a low number of instances, enabling or disabling overlaps between classes and switching between hierarchical and flat class structures. III. USER INTERFACE KGSEC includes a Web-based graphical user interface that allows users to easily configure the schema extraction process and, most importantly, to assess and compare the quality and characteristics of schemata produced by alternative methods or different configurations. This is crucial, given that different methods can produce different schemas and there is no objective criterion to determine which schema is the “best”. Hence, generating and comparing multiple alternative schemas may be more helpful for understanding and exploring a given KG than relying only on a single schema. A. Schema Extraction Configuration Panel This panel allows users to configure and execute KGSEC through a graphical interface. They can select a pre-loaded dataset to experiment, or load their own RDF dataset. The 3https://scikit-learn.org/stable/modules/clustering.html 4https://rasbt.github.io/mlxtend/ Fig. 2. Schema extraction configuration panel. Fig. 3. Schema overview panel. users are then presented with a set of configurable options, associated with different stages of the schema extraction process, in order to curate their own schema extraction configuration. This is done via the configuration panel presented in Figure 2. The user has the option to select between different types of characteristic sets, schema extraction methods and similarity functions. Once the request is submitted, feedback regarding the validity of the request is shown on the screen, while the schema extraction process is carried out by another thread in the back-end of KGSEC. B. Schema Overview Panel This panel allows users to visually explore a generated schema and its characteristics. It provides several global metrics, such as the number of classes, the hierarchy depth, as well as typical metrics used to assess the quality of a resulting clustering, such as silhouette score. It also allows users to examine each class in more detail. For each class, they can visualize the number of instances and characteristic sets, depth, dominant types and properties and missingness ratio. Additionally, they can generate an index-like hierarchy file using the Ontospy5library, through the schema overview panel, presented in Figure 3. Finally, through a distance preserving two-dimensional projection, users can visualize 5http://lambdamusic.github.io/Ontospy/ Fig. 4. Schema comparison panel. class proximity. When clicking on a class in the scatter-plot, the corresponding row on the table presenting the classes is highlighted, to enable users to explore easier the schema components. The whole set of views presented on this module is such that can allow users to discover homogeneous classes not explicitly defined in the graph, or explicitly defined classes with potential to be further clustered. At this point, users also have the option to select a specific class of the schema to load and further perform schema extraction on it. This enables a progressive schema extraction approach that gradually drills down to specific classes that are of interest to the user. C. Schema Comparison Panel This panel allows users to compare either two schemata generated by KGSEC or a generated schema against a provided ground truth. The comparison view, shown in Figure 4, includes classic clustering comparison metrics, such as Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI), or F1-score, precision and recall in the case of comparison to a given ground-truth schema. Users can visualize the distribution of instances and CSs between the two schemata through a Sankey diagram and the class proximity through a distance preserving two-dimensional projection of classes of both schemata. Additionally, certain classes can be selected to view their details, enabling a more in-depth comparison. The toolkit and data provided by this page allow users to investigate the impact of different configurations, when generating the classes of a schema for a graph. IV. DEMONSTRATION SCENARIOS During the demonstration, the users will be able to select a KG, execute KGSEC to perform schema extraction using one or more methods, and visually explore and compare the results. We describe a typical usage scenario below. As an example, we use the nobelprize6dataset, which contains information regarding laureates and awards of the Nobel prize. The user in this demonstration scenario uses KGSEC to iteratively explore the dataset, extracting alternative schemas from this KG in a series of steps, as described below. Step 1: Initial schema extraction configuration. The user starts by selecting a dataset, choosing a schema extraction algorithm to be applied, and configuring its parameters. In this example, the user proceeds using a directed-property characteristic set extraction method, with DBSCAN as the schema extractor, parameterized with epsilon set to 0.6 and min samples set to 1, while the similarity function is set to Jaccard. Step 2: Schema overview and exploration. The schema is stored in the schema repository. By viewing the schema overview panel, shown in Figure 3, the user can browse all previously generated schemata and select one to explore. Viewing the produced schema the user observes that it contains only 7 classes. However, the biggest class in terms of size Cluster 0 has a very high missingness ratio of 0.84. This hints the user that instances described by diverse CSs belong to this cluster. Given the relative size of the class, which indicates importance for the dataset and its underlying diversity and sparsity, which indicates difficulty in understanding and handling, the user is hinted to reiterate over the schema extraction process, demanding this time a more detailed schema. Step 3: Iteration and schema comparison. In an effort to to explore the KG further by generating a more finegrained schema, the user submits a new request with the same parameters as before but now setting a lower value in the epsilon parameter, namely 0.4, which is expected to produce smaller but denser clusters. Analyzing the newly produced schema, through the schema overview panel, the user now validates the fact that it has increased detail, by observing the increase in the schema’s number of classes, from 7 to 18. Most importantly, the user is hinted from the class details table about the initial reason of the high missingness ratio of previous Cluster 0. In the new schema both Cluster 3 and Cluster 1, with respectively 6585 and 1160 instances each, have AWARDFILE as dominant type, but missingness ratio is only 0.13 for the first one and 0.62 for the second one. This implies that a high proportion of the instances of the initial cluster had very similar CSs, a fact which was hidden from the initial extracted schema, but can be proven valuable for the user’s analysis. To validate this assumption, the user navigates to the comparison panel, presented in Figure 4. By selecting both schemata, the user loads the views regarding their general 6https://www.nobelprize.org/about/linked-data-examples/ information and by clicking compare she loads the comparison field. As initially suspected, the Sankey diagram used for instance flow inspection between classes of the two datasets shows that both Cluster 1 and Cluster 3 of the new schema belong to Cluster 0 of the initial schema. Step 4: Experimentation with alternative schema extractors. While the user so far executed a fine-tuning process, KGSEC also provides the possibility to generate potentially radically different schemata, by choosing a different schema extraction method. In our scenario, the user can identify the need to try alternative approaches when pointing out that in the detailed schema there is great disparity between the number of instances of the classes. Trying to tackle this with an alternative schema, the user performs a frequent-itemset based schema extraction, whose implementation allows to set some thresholds of minimum global support for the initial clusters. Generating a schema with frequent-itemset approach and min global-support set to 0.2, the user observes through the schema overview panel that a relative disparity between the class sizes persists. However, the schema managed to capture a big AWARDFILE class with zero missingness ratio, which can be proven valuable for the user’s analysis. Additionally, data from the comparison panel indicate similarity with the previous schemata. ACKNOWLEDGEMENT This work was partially funded by the EU Horizon Europe project STELAR (101070122). REFERENCES [1] K. Kellou-Menouer, N. Kardoulakis, G. Troullinou, Z. Kedad, D. Plexousakis, and H. Kondylakis, “A survey on semantic schema discovery,” The VLDB Journal, vol. 31, no. 4, pp. 675–710, 2022. [2] A. Bonifati, S. Dumbrava, and N. Mir, “Hierarchical clustering for property graph schema discovery,” in EDBT, 2022, pp. 2:449–2:453. [3] R. Bouhamoum, Z. Kedad, and S. Lopes, “Incremental Schema Discovery at Scale for RDF Data,” in The Semantic Web, 2021, pp. 195–211. [4] M.-D. Pham, L. Passing, O. Erling, and P. Boncz, “Deriving an Emergent Relational Schema from RDF Data,” in WWW, 2015, pp. 864–874. [5] R. Bouhamoum, K. Kellou-Menouer, S. Lopes, and Z. Kedad, “Scaling Up Schema Discovery for RDF Datasets,” in IEEE 34th International Conference on Data Engineering Workshops (ICDEW), 2018, pp. 84–89. [6] G. Papastefanatos, M. Meimaris, and P. Vassiliadis, “Relational schema optimization for RDF-based knowledge graphs,” Information Systems, vol. 104, p. 101754, 2022. [7] S. K. Mohamed, “Unsupervised Hierarchical Grouping of Knowledge Graph Entities,” in Joint Proc. of the 1st Int. Workshop on Knowledge Graph Building and 1st Int. Workshop on Large Scale RDF Analytics co-located with ESWC, vol. 2489, 2019, pp. 65–74. [8] K. Kellou-Menouer and Z. Kedad, “Schema discovery in RDF data sources,” in ER, vol. 9381, 2015, pp. 481–495. [9] A. Lutov, S. Roshankish, M. Khayati, and P. Cudr´ e-Mauroux, “Statix - statistical type inference on linked data,” in IEEE BigData, 2018, pp. 2253–2262. [10] S. Issa, P. Paris, F. Hamdi, and S. S. Cherfi, “Revealing the conceptual schemas of RDF datasets,” in CAiSE, vol. 11483, 2019, pp. 312–327.