User-oriented exploration of semi-structured datasets
Abstract
This is the slides of my PhD defense.
Full text
User-oriented exploration of semi-structured datasets Nelly Barret Inria Saclay and Institut Polytechnique de Paris Supervised by Ioana Manolescu and Karen Bastien March 15, 2024 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 1 / 100
Outline 1Motivation: exploring semi-structured data 2Overview of our approach 3Abstra: first-sight overview of a dataset 4Pathways: efficiently finding interesting paths 5Systems developed 6Conclusion Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 2 / 100
Motivation: exploring semi-structured data Outline 1Motivation: exploring semi-structured data 2Overview of our approach 3Abstra: first-sight overview of a dataset 4Pathways: efficiently finding interesting paths 5Systems developed 6Conclusion Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 3 / 100
Motivation: exploring semi-structured data Data exploration by non-technical users (NTUs) Conflicts of Interest in the biomedical domain [ABB+21] w/ S. Horel Is this dataset useful for the investigations? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 4 / 100
Motivation: exploring semi-structured data Data exploration by non-technical users (NTUs) Conflicts of Interest in the biomedical domain [ABB+21] w/ S. Horel Is this dataset useful for the investigations? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 4 / 100
Motivation: exploring semi-structured data Data exploration by non-technical users (NTUs) Conflicts of Interest in the biomedical domain [ABB+21] w/ S. Horel Is this dataset useful for the investigations? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 4 / 100
Motivation: exploring semi-structured data Data exploration by non-technical users (NTUs) Conflicts of Interest in the biomedical domain [ABB+21] w/ S. Horel How are authors connected to biomedical companies? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 5 / 100
Motivation: exploring semi-structured data Semi-structured data exploration Several semi-structured data models: XML documents JSON documents RDF graphs Property graphs Semi-structured dataset exploration is hard: complex, irregular structure Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 6 / 100
Motivation: exploring semi-structured data Semi-structured data exploration Several semi-structured data models: XML documents JSON documents RDF graphs Property graphs Semi-structured dataset exploration is hard: complex, irregular structure Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 6 / 100
Overview of our approach Research contributions Abstra: data overviews [BMU22,BMU24] Lightweight Entity-Relationship diagrams Compact yet meaningful data overviews Ideal for first-sight dataset discovery PathWays: interesting Named Entity connections [BGLM23b,BGLM23a,BGLM24] Interesting entity paths in and across datasets Complete set of NE-to-NE interesting connections Ideal for exploring connections within and across datasets Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 9 / 100
Abstra: first-sight overview of a dataset Outline 1Motivation: exploring semi-structured data 2Overview of our approach 3Abstra: first-sight overview of a dataset 4Pathways: efficiently finding interesting paths 5Systems developed 6Conclusion Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 10 / 100
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 11 / 100
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 11 / 100
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 11 / 100
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! What about semi-structured data models (nesting)? Keep it simple and of controllable size Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 12 / 100
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! What about semi-structured data models (nesting)? Keep it simple and of controllable size Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 12 / 100
Abstra: first-sight overview of a dataset What does the dataset describe? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 13 / 100
Abstra: first-sight overview of a dataset The Abstra approach 1Integrate all data sources in a graph (ConnectionLens) [ABC+22] 2Summarize the graph 3Among summary nodes, identify entities and their attributes 4In the summary, identify relationships between the entities 5Propose a simple category to each entity (best-effort) Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 14 / 100
Abstra: first-sight overview of a dataset Background Background: from heterogeneous data to data graphs ConnectionLens [ABC+22]: 1Ingests any dataset into a directed graph Generic, flexible, fine granularity 2Extracts Named Entities (NEs) from all text nodes date , email address , People , Place , Organization , ... Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 15 / 100
Abstra: first-sight overview of a dataset Data graph summarization The summary (collection graph) G Collection node for each equivalence class paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 19 / 100
Abstra: first-sight overview of a dataset Data graph summarization The summary (collection graph) G Collection node for each equivalence class Collection edge Cs→Ctif a data edge exists paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 20 / 100
Abstra: first-sight overview of a dataset Data graph summarization The summary (collection graph) G Collection node for each equivalence class Collection edge Cs→Ctif a data edge exists Entity profile for each leaf collection node: reflects NEs in the leaves paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 21 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 22 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 22 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 22 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Requirements and algorithm We need an algorithm to identify entity roots and attributes for the E-R diagram For complex, potentially cyclic, collection graphs Greedy selection of few entities in G 1Assign a score to each collection node 2While less than Emax entity roots, or data coverage <covmin 1Elect the next highest-scored eligible collection node as an entity root 2Compute its boundary , i.e., attribute set 3Update the collection graph to reflect the selection of an entity 4Recompute the scores Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 23 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Requirements and algorithm We need an algorithm to identify entity roots and attributes for the E-R diagram For complex, potentially cyclic, collection graphs Greedy selection of few entities in G 1Assign a score to each collection node 2While less than Emax entity roots, or data coverage <covmin 1Elect the next highest-scored eligible collection node as an entity root 2Compute its boundary , i.e., attribute set 3Update the collection graph to reflect the selection of an entity 4Recompute the scores Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 23 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships How to score a collection node? Reflect the weight of this node and its structure in the dataset 1wdesck,wleafk: # descendants, leaf descendants, at depth k ×Not clear how to pick k Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 24 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships How to score a collection node? Reflect the weight of this node and its structure in the dataset 1wdesck,wleafk: # descendants, leaf descendants, at depth k ×Not clear how to pick k Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 24 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val The reverse collection graph GR Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 31 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.5 1 1 1 1 1 0.5 1 1 0.5 0.5 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith PR edge weights Collections distribute their score based solely on their connectivity Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 32 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.5 1 1 1 1 1 0.5 1 1 0.5 0.5 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith PR edge weights Collections distribute their score based solely on their connectivity Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 32 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships How to score a collection node? 1wdesck,wleafk: # descendants, leaf descendants, at depth k 2wDAG :dw bottom-up propagation on G(outside cycles) 3wPageRank : PageRank algorithm on G 4wdwPageRank : PageRank algorithm on Gwith dw-tuned PR edge weights XReflects both the topology and where actual data is Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 33 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith dw-tuned PR edge weights Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 34 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 35 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 36 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Propagates scores across the collection graph Works on cyclic collection graphs The score reflects the topology and where the data is A collection node distributes its weight Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 37 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships How to compute an entity boundary? Collections in Grepresenting attributes of this entity “Those that contribute to the entity’s weight” The boundary may go far (for deep-structure entities) Easy to define for wdesck,wleafk,wDAG . Example for wdesc2 paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Does not apply for PageRank-based scores Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 38 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships How to compute an entity boundary? Collections in Grepresenting attributes of this entity “Those that contribute to the entity’s weight” The boundary may go far (for deep-structure entities) Easy to define for wdesck,wleafk,wDAG . Example for wdesc2 paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Does not apply for PageRank-based scores Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 38 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships How to update the collection graph after selecting an entity? Reflect the allocation of data nodes and edges to one entity 1updateboolean Collection nodes and edges in the boundary of the entity Very efficient Sufficient for wdesck,wleafk,wDAG 2updateexact Graph nodes and edges Much more costly Required for wPageRank ,wdwPageRank Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 42 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Exact graph update Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 43 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Exact graph update paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 44 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Selected entities and their boundaries paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 45 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Finding relationships between entities Relationship: a path from an entity to another paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 paper →wB →author paper →pIn →conf author →hW →paper conf →inv →author Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 46 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Assign a semantic category to each entity Input: an entity E, categories K, semantic properties P K: Person, ScientificPaper, Event, Website, Mountain, ... P:{label:"address", domain:[Pers., Org.], range:[Place]}, ... Output: a category for E Algorithm: Compare: The common name of all nodes in the entity root (if it exists) with k∈ K (conf, paper, author) Its attribute names with p∈ P (affiliation, email, ...) Its entity profiles with p.range ∈ P (,,, ...) Each good match votes for one or few categories Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 47 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Assign a semantic category to each entity Input: an entity E, categories K, semantic properties P K: Person, ScientificPaper, Event, Website, Mountain, ... P:{label:"address", domain:[Pers., Org.], range:[Place]}, ... Output: a category for E Algorithm: Compare: The common name of all nodes in the entity root (if it exists) with k∈ K (conf, paper, author) Its attribute names with p∈ P (affiliation, email, ...) Its entity profiles with p.range ∈ P (,,, ...) Each good match votes for one or few categories Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 47 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Name Similar to Votes for paper ResearchPublication (0.85) ResearchPublication News (0.63) News paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 48 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Attribute Similar to Votes for abstract abstract (1.0) ResearchPublication summary (0.92) Book preface (0.47) title title (1.0) ResearchPublication honorific title (0.87) Movie Person year year publication (0.85 + )Event Book ResearchPublication, ... paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 49 / 100
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Attribute Similar to Votes for abstract abstract (1.0) ResearchPublication summary (0.92) Book preface (0.47) title title (1.0) ResearchPublication honorific title (0.87) Movie Person year year publication (0.85 + )Event Book ResearchPublication, ... paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 50 / 100
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 55 / 100
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 56 / 100
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 57 / 100
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 58 / 100
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Abstra selects frequent, coherent and semantically central entities Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 59 / 100
Abstra: first-sight overview of a dataset Experimental evaluation Experimental evaluation: scalability Our abstraction method scales up linearly in the data size Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 60 / 100
Abstra: first-sight overview of a dataset Related work Related work Data summarization Structural Quotient [GGM20,KC10,MS99] (the one we adopt to build G) Non-quotient [GW97] Pattern mining [ZLVK16] Statistical [HS12] Hybrid [RGSB17] Schema inference XML [CGS11] JSON [BCGS19] RDF [GLSW22] PG [LBH21] Data summarization and schema inference are tied to one data model Schemas are often not suited to NTUs Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 61 / 100
Abstra: first-sight overview of a dataset Related work A JSON schema from social network data using [BCGS19] Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 62 / 100
Pathways: efficiently finding interesting paths Outline 1Motivation: exploring semi-structured data 2Overview of our approach 3Abstra: first-sight overview of a dataset 4Pathways: efficiently finding interesting paths 5Systems developed 6Conclusion Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 63 / 100
Pathways: efficiently finding interesting paths Data is often used to find connections Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 64 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths connecting Person NEs () to Organization NEs () ←#val ←Name ←Author →Affiliation →#val → ←#val ←Name ←Author ←Authors ←Article →Journal →#val → ←#val ←COI ←Article →Journal →#val →←#val → Which paths are most interesting and deserve to be evaluated? Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 66 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths are unreliable: we face entity extraction errors E.g., “John Hopkins | {z } person University Hospital” False positives, or wrong entity type attribution, e.g., “THC | {z } org. ” Some paths are structurally weak: we face information dilution E.g., a paper has 50 authors Path interestingness : based on edge reliability and edge force Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 67 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths are unreliable: we face entity extraction errors E.g., “John Hopkins | {z } person University Hospital” False positives, or wrong entity type attribution, e.g., “THC | {z } org. ” Some paths are structurally weak: we face information dilution E.g., a paper has 50 authors Path interestingness : based on edge reliability and edge force Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 67 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths are unreliable: we face entity extraction errors E.g., “John Hopkins | {z } person University Hospital” False positives, or wrong entity type attribution, e.g., “THC | {z } org. ” Some paths are structurally weak: we face information dilution E.g., a paper has 50 authors Path interestingness : based on edge reliability and edge force Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 67 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 68 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 68 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 68 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 68 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths connecting Person NEs () to Organization NEs () 1.0 ←−− #val 1.0 ←−− Name 1.0 ←−− Author 1.0 −−→ Affiliation 1.0 −−→ #val 0.91 −−→ Reliable; strong 1.0 ←−− #val 1.0 ←−− Name 1.0 ←−− Author 0.02 ←−− Authors 1.0 ←−− Article 1.0 −−→ Journal 1.0 −−→ #val 0.41 −−→ Reliable; weak 0.09 ←−− #val 1.0 ←−− COI 1.0 ←−− Article 1.0 −−→ Journal 1.0 −−→ #val 0.05 −−→ 0.09 ←−− #val 0.04 −−→ Not reliable; strong Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 69 / 100
Pathways: efficiently finding interesting paths NE-to-NE path enumeration PathWays output: data paths as tables Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 70 / 100
Pathways: efficiently finding interesting paths Related work Related work Structured querying SQL, SPARQL, GQL [DFG+22] Assisted struct. querying Interactive queries [DAB16] Guided query writing [ERAAL18,KKBS10] NL2SQL [KSHL20] Keyword-based search Unidirectional [ABC+02,LOF+08] Bi-directional [ABC+22] Path search in struct. queries SPARQL extensions: [ASMH18,AMSH18, AMM23] For PGs: [DFG+22] Pathways users need no knowledge of the graph structure or values Less intimidating for NTUs Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 76 / 100
Systems developed Outline 1Motivation: exploring semi-structured data 2Overview of our approach 3Abstra: first-sight overview of a dataset 4Pathways: efficiently finding interesting paths 5Systems developed 6Conclusion Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 77 / 100
Systems developed Systems developed Abstra for data abstraction: https://team.inria.fr/cedar/projects/abstra/ 65 Java core classes and 10K LOC Demonstrated at CIKM 2022 [BMU22] (also BDA 2022) PathWays for NE-to-NE paths: https://team.inria.fr/cedar/projects/pathways/ 18 Java core classes and 4K LOC Demonstrated at ESWC 2023 [BGLM23b] (also BDA 2023) ConnectionStudio for NTU data exploration: https://connectionstudio.inria.fr/ 4K Java LOC and 21K JavaScript LOC (w/ T. Galizzi, S. Ebel, M. Mohanty) Demonstrated at CoopIS 2023 [BEG+23] (also BDA 2023) Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 78 / 100
Systems developed Systems developed Abstra for data abstraction: https://team.inria.fr/cedar/projects/abstra/ 65 Java core classes and 10K LOC Demonstrated at CIKM 2022 [BMU22] (also BDA 2022) PathWays for NE-to-NE paths: https://team.inria.fr/cedar/projects/pathways/ 18 Java core classes and 4K LOC Demonstrated at ESWC 2023 [BGLM23b] (also BDA 2023) ConnectionStudio for NTU data exploration: https://connectionstudio.inria.fr/ 4K Java LOC and 21K JavaScript LOC (w/ T. Galizzi, S. Ebel, M. Mohanty) Demonstrated at CoopIS 2023 [BEG+23] (also BDA 2023) Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 78 / 100
Systems developed ConnectionStudio software pile All deployed using Maven, hundreds of unit tests, etc. Help from T. Galizzi, M. Mohanty Several rounds of re-engineering (ML model memory consumption, etc.) ConnectionStudio Pathways Abstra ConnectionLens, incl. [AMM23]RDFQuotient (14K LOC) OntoSQL (85K LOC) Jena Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 79 / 100
Systems developed A comprehensive data exploration tool for NTUs ConnectionStudio: a data lake for ingesting, exploring and querying heterogeneous data 1Data abstractions as E-R diagrams (Abstra) 2NE-to-NE paths as tables (PathWays) 3“Gentle introduction” to the data lake (w/ journalist input) Demonstrated to journalists at DataJournos (40) and CFI (60) ConnectionStudio interesting for a first look at the data. Still maturing... Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 80 / 100
Systems developed A comprehensive data exploration tool for NTUs ConnectionStudio: a data lake for ingesting, exploring and querying heterogeneous data 1Data abstractions as E-R diagrams (Abstra) 2NE-to-NE paths as tables (PathWays) 3“Gentle introduction” to the data lake (w/ journalist input) Demonstrated to journalists at DataJournos (40) and CFI (60) ConnectionStudio interesting for a first look at the data. Still maturing... Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 80 / 100
Systems developed A comprehensive data exploration tool for NTUs ConnectionStudio: a data lake for ingesting, exploring and querying heterogeneous data 1Data abstractions as E-R diagrams (Abstra) 2NE-to-NE paths as tables (PathWays) 3“Gentle introduction” to the data lake (w/ journalist input) Demonstrated to journalists at DataJournos (40) and CFI (60) ConnectionStudio interesting for a first look at the data. Still maturing... Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 80 / 100
Conclusion Outline 1Motivation: exploring semi-structured data 2Overview of our approach 3Abstra: first-sight overview of a dataset 4Pathways: efficiently finding interesting paths 5Systems developed 6Conclusion Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 81 / 100
Conclusion Takeaways and next steps We introduced: 1A unified view over heterogeneous semi-structured data models 2Abstra: a dataset abstraction system for semi-structured data 3PathWays: an entity-focused exploration system 4ConnectionStudio: a comprehensive data lake exploration tool Next steps: Generate PG schemas from abstractions [BEMM24] Migrate data graphs into PG graphs Enrich extracted NEs with RDF knowledge bases Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 82 / 100
Conclusion References III Nelly Barret, Antoine Gauquier, Jia Jean Law, and Ioana Manolescu. PATHWAYS: entity-focused exploration of heterogeneous data graphs (demonstration). In ESWC, 2023. Nelly Barret, Antoine Gauquier, Jia Jean Law, and Ioana Manolescu. Exploring heterogeneous data graphs through their entity paths. Inf. Systems SUBM, 2024. Nelly Barret, Ioana Manolescu, and Prajna Upadhyay. ABSTRA: toward generic abstractions for data of any model (demonstration). In CIKM, 2022. Nelly Barret, Ioana Manolescu, and Prajna Upadhyay. Computing generic abstractions from application datasets. In EDBT, 2024. Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. Schemas for safe and efficient XML processing. In ICDE. IEEE Computer Society, 2011. Gonzalo Diaz, Marcelo Arenas, and Michael Benedikt. SPARQLByE: querying rdf data by example. Proceedings of the VLDB Endowment, 9(13):1533–1536, 2016. Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 88 / 100
Conclusion References IV Alin Deutsch, Nadime Francis, Alastair Green, Keith Hare, Bei Li, Leonid Libkin, Tobias Lindaaker, Victor Marsault, Wim Martens, Jan Michels, Filip Murlak, Stefan Plantikow, Petra Selmer, Oskar van Rest, Hannes Voigt, Domagoj Vrgoc, Mingxi Wu, and Fred Zemke. Graph pattern matching in GQL and SQL/PGQ. In SIGMOD ’22: International Conference on Management of Data, Philadelphia, PA, USA, June 12 - 17, 2022, pages 2246–2258, 2022. Ahmed El-Roby, Khaled Ammar, Ashraf Aboulnaga, and Jimmy Lin. Sapphire: querying rdf data made simple. arXiv preprint arXiv:1805.11728, 2018. Fran¸cois Goasdou´e, Pawel Guzewicz, and Ioana Manolescu. RDF graph summarization for first-sight structure discovery. The VLDB Journal, 29(5), April 2020. Benoˆıt Groz, Aur´elien Lemay, Slawek Staworko, and Piotr Wieczorek. Inference of shape graphs for graph databases. In ICDT, volume 220, 2022. Roy Goldman and Jennifer Widom. DataGuides: enabling query formulation and optimization in semistructured databases. In VLDB, 1997. Katja Hose and Ralf Schenkel. Towards benefit-based RDF source selection for SPARQL queries. In Proceedings of the 4th International Workshop on Semantic Web Information Management, pages 1–8, 2012. Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 89 / 100
Conclusion References V Shahan Khatchadourian and Mariano P Consens. ExpLOD: summary-based exploration of interlinking and RDF usage in the Linked Open Data Cloud. In Extended semantic web conference, pages 272–287. Springer, 2010. Nodira Khoussainova, YongChul Kwon, Magdalena Balazinska, and Dan Suciu. SnipSuggest: context-aware autocompletion for SQL. Proceedings of the VLDB Endowment, 4(1):22–33, 2010. Hyeonji Kim, Byeong-Hoon So, Wook-Shin Han, and Hongrae Lee. Natural language to SQL: Where are we today? Proceedings of the VLDB Endowment, 13(10):1737–1750, 2020. Hanˆa Lbath, Angela Bonifati, and Russ Harmer. Schema inference for property graphs. In EDBT, 2021. Guoliang Li, Beng Chin Ooi, Jianhua Feng, Jianyong Wang, and Lizhu Zhou. EASE: an effective 3-in-1 keyword search method for unstructured, semi-structured and structured data. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data, pages 903–914, 2008. Tova Milo and Dan Suciu. Index structures for path expressions. In International Conference on Database Theory, pages 277–295. Springer, 1999. Raghu Ramakhrishnan and Johannes Gehrke. Database Management Systems (3rd edition). McGraw-Hill, 2003. Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 90 / 100
Conclusion References VI Matteo Riondato, David Garc´ıa-Soriano, and Francesco Bonchi. Graph summarization with quality guarantees. Data mining and knowledge discovery, 31:314–349, 2017. Mussab Zneika, Claudio Lucchese, Dan Vodislav, and Dimitris Kotzinos. Summarizing linked data RDF graphs using approximate graph pattern mining. In 19th International Conference on Extending Database Technology, 2016. Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 91 / 100
Conclusion Data-acyclic flooding boundary mailbox email date #val content list item text #val The boundary is truncated due to cyclic collection edges Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 92 / 100
Conclusion Entity classification time The classification time is composed of: Loading the Word2Vec semantic model Constant, 4-8 seconds Comparing entity attributes with semantic properties Varies with the number of entities and their number of attributes May vary in a generated dataset of different sizes (different entity roots) Computing entity profiles Linear in the input size Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 93 / 100
Conclusion RDF quotient graph summarization [GGM20] Source clique: set of outgoing properties co-occuring together on at least one node Target clique: set of incoming properties co-occuring together on at least one node Properties “a”, “b”, “d” are in the same source clique Properties “a” and “e” are in the same target clique (c) Pawel Guzewic Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 94 / 100
Conclusion Strong summary [GGM20] Strong S summary: Two nodes are S equivalent iff they have both the same source and target cliques Source and target cliques for each node Strong summary (c) Pawel Guzewic Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 95 / 100
Conclusion Typed-strong summary [GGM20] Typed-strong TS summary: Two typed nodes are TS equivalent iff they have the same type set Two untyped nodes are TS equivalent iff they have both the same source and target cliques Source and target cliques for each node + an RDF type Typed-strong summary (c) Pawel Guzewic Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 96 / 100
Conclusion Disagreement between Flair and ChatGPT False Flair positives: Flair identifies “Av. Peter Henry Rolfs | {z } person 36570-900 Vicosa” Flair mislead by capitalization: Flair identifies “Claudin-7b | {z } person ” (but not ChatGPT) Different token allocation: “University of Alabama | {z } org. ”, “Birmingham | {z } loc. ” “University of Alabama, Birmingham | {z } loc. ” Missed non-English spelling/names: ChatGPT finds “Antonio Gonz´alez | {z } person ” ChatGPT finds “Yoshida, Sakyo-ku, Kyoto 606-8501, Japan | {z } loc. ” Nelly Barret (Inria) Semi-structured Data Exploration March 15, 2024 97 / 100