Extracting and Utilizing Structured Data from the Open Web Index
Abstract
Appeared in: Open Search Symposium 2025, 8-10 October 2025, CSC IT Center for Science, Helsinki, Finland.
Full text
EXTRACTING AND UTILIZING STRUCTURED DATA FROM THE OPEN WEB INDEX L. Caspari∗, M. Dinzinger, J. Mitrovic, M. Granitzer, University of Passau, Passau, Germany Abstract Structured data is a valuable source of information that can be found on many web pages and can be extracted efficiently during crawling. It is often encoded in the form of JavaScript Object Notation for Linked Data (JSON-LD) or Microdata using schema.org definitions for entities such as FAQ pages or addresses, allowing efficient parsing and extraction of data. The OpenWebSearch.EU (OWS) [Hendriksen et al . (2024)] project, which publicly releases crawled web data on a regular basis, is a useful source for fresh structured data, as published datasets contain specific columns for JSON-LD and Microdata encountered during crawling. In this paper, we present initial statistics on the occurrence of structured data in the OWS datasets, focusing on the presence of certain entities in schema.org, namely Frequently Asked Questions (FAQs), opening hours, phone numbers, and addresses. Additionally, we discuss two practical application scenarios of the extracted data. In our first use case, in line with our previous work [Dinzinger et al.(2025)], we demonstrate how FAQ data can be used to construct multilingual Q&A-style datasets, which can be used to train large language models (LLMs) for tasks like question answering or retrieval. In our second case, we show the potential of structured data to enrich map applications and improve user experience. These use cases exemplify the value of structured data and demonstrate the benefits of its systematic extraction and integration into real-world applications. INTRODUCTION Structured data provides an important source of information about webpages that can easily be parsed and ingested by downstream applications like search engines or map providers. It can be used to enrich the results shown to users, i.e. by displaying the address and opening hours of a shop. Structured data is specified by webmasters using schema.org definitions of entities like addresses, opening hours or FAQs. While it can be specified in a variety of formats, JSON-LD and Microdata have emerged as popular choices [Volpini et al . (2024)]. As defining information in these formats requires additional effort from webmasters, the data is generally of high quality and can easily be extracted during crawling. However, many downstream applications require the extracted data to be fresh, necessitating frequent revisits of webpages. An important source for fresh structured data is the OpenWebSearch project [Granitzer et al . (2024)], which releases crawled JSON-LD and Microdata as part of their regularly released datasets. Thus, the project presents an important resource for publicly available and fresh structured data. To better understand the preva- ∗[email protected] lence of JSON-LD and Microdata within the OWS index, we analyze datasets from five days in February 2025 with a total size of 1.6TB. We find that more than half (54.2%) of the crawled webpages use JSON-LD or Microdata. However, when attempting to extract specific schemas, the ratio drops significantly, i.e. phone numbers can only be found on 1.8% of webpages. Given these extracted schemas, we establish two use cases for structured data, namely leveraging FAQ pages to construct a Q&A dataset that can be used to train or evaluate LLMs on question answering or retrieval tasks, and extracting FAQs, opening hours, addresses and phone numbers to enhance map applications. Our use cases demonstrate the potential of structured data in downstream applications and the importance of having publicly available and fresh data to enable researchers and companies to build upon the wealth of information contained within. The code to reproduce our results is available on GitHub1. The remainder of this paper is structured as follows: After looking at related work, the methodology section gives an overview of the schema definitions we focus on and explains our extraction approach. The following section contains statistics about the occurrence of Microdata and JSON-LD as well as of specific schemas like addresses or phone numbers. We then detail our two use cases for the extracted data before discussing limitations and problems with data quality. Finally, we conclude our paper with a summary of our findings and future extensions of our work. RELATED WORK Using schema.org 2 classes to represent information about entities like restaurants, events or products has become increasingly prevalent since the introduction of schema specifications in 2011 [Brinkmann et al . (2023), Volpini et al . (2024)]. Established by Google, Bing, Yahoo and Yandex to offer webmasters a unified way of defining structured data, it provides information in machine-readable and widely accepted formats, with JSON-LD and Microdata being common ways in which structured data is made available. Microdata is an extension of the HTML5 specification 3 and is thus added directly to the HTML tags themselves, whereas JSON-LD is specified within one set of script tags 4 , making it easy to extract. While both formats have seen increasing adoption in the past years with JSONLD being present on 41% of webpages and Microdata on 26% in 2024 [Volpini et al . (2024)], their usage patterns differ. As analyzed by Volpini et al., JSON-LD is most 1https://github.com/padas-lab-de/owi-sdm 2https://schema.org/ 3https://html.spec.whatwg.org/multipage/microdata.html 4https://json-ld.org/ https://doi.org/10.5281/zenodo.17228339
commonly used for organization data, local businesses and product listings, whereas Microdata often specifies webpage structure or site navigation. Apart from the imbalance of schema usage between the different formats, adoption of schema annotations also varies depending on the domain with a much higher usage of entities like products or local businesses [Brinkmann et al . (2023)] than entities related to educational resources [Navarrete et al.(2019)]. The higher prevalence of structured data increases the value of extracting its content for downstream applications. Apart from using structured data to increase search visibility [Recalde et al . (2021)], it can also be extracted to provide training data for machine learning models [Peeters et al . (2020), Dinzinger et al . (2025)]. While the usage of schema-based annotations requires additional effort by webmasters and thus is generally of high quality, applications using structured data still need to filter out low quality samples. For instance, a high percentage of schema.org dataset annotations do not describe actual datasets [Alrashed et al . (2021)], drastically limiting the usability of this schema for dataset search. Similarly, certain properties of common entities like products, e.g. the product ID or category, are seldom filled [Brinkmann et al.(2023)]. METHODOLOGY The following paragraphs offer a general introduction into the specification of structured data and define the exact schema classes that are of interest. Subsequently, an overview of the extraction pipeline is provided along with the format in which extracted data is stored. Defining Structured Data with Schemas In our context, structured data is specified using entities defined by the schema.org type hierarchy. For each entity, e.g. an FAQ page, schema.org defines the properties and its type, which are specified in a key-value-based manner. While there are various ways of specifying structured schema data, this paper will focus on JSON-LD and Microdata, which are part of the datasets published by OWS. Due to our current use cases, we will specifically consider schemas for defining phone numbers 5 , addresses 6 , opening hours 7 and FAQ pages8. Figure 1 shows an excerpt of the aforementioned schemas in JSON-LD which were encountered when crawling the webpage of a Subway store located in Seattle. The structured data contains important information about the store which can be used to enrich downstream applications. Extracting and Merging Schemas While structured data is a valuable resource, it is not available for every webpage. Therefore, we first use owilix 9 to 5https://schema.org/telephone 6https://schema.org/address 7https://schema.org/openingHours 8https://schema.org/FAQPage 9https://opencode.it4i.eu/openwebsearcheu-public/ owi-cli { ... "telephone": "(425) 614-3256", "address": { "@type": "PostalAddress", "addressCountry": "US", "addressLocality": "Bellevue King", "addressRegion": "WA", "postalCode": "98007", "streetAddress": "1410 156th Ave NE" }, "openingHours": ["Mo 08:00-22:00", "Tu 08:00-22:00", "We 08:00-22:00", "Th 08:00-22:00", "Fr 08:00-22:00", "Sa 09:00-22:00", "Su 09:00-22:00"] "@type": "FAQPage", "mainEntity": [{ "@type": "Question", "name": "How can I place a Subway Catering order?", "acceptedAnswer": { "@type": "Answer", "text": "To place an order, visit us online at catering.subway.com or call your local restaurant." } ... }] ... } Figure 1: An excerpt of JSON-LD extracted from the page of a Subway store in Seattle. download OWS datasets from five different days in February 2025, which contain files in Parquet format 10 . Specifically, we use the datasets published on the 19th and 21st-24th of February. As the Parquet files include specific columns for Microdata and JSON-LD, we subsequently filter out all entries for which both columns are empty and only use columns that are of interest to us, reducing the initial size of 1.6TB by a factor of four. We then apply our extraction code on the filtered data to obtain FAQs, opening hours, addresses and phone numbers contained in the structured data, with the extracted information being saved to Parquet files. As the structure of the data is quite dependent on the schema, we store each in a separate Parquet file with the exception of phone numbers and addresses, which are merged together. The resulting files organized per day are available for download from our MinIO instance11. DATA EXPLORATION To get an initial idea about how often structured data appears in the OWS crawls, we analyze the filtered datasets from February 2025 and find that 54.2% of webpages contain microdata or JSON-LD. However, as shown in Table 1, this number quickly drops when looking at a specific schema. In fact, all schemas we are interested in occur on less than 2.1% of webpages. Taking a closer look at the individual schemas, we also observe that a significant number of them are malformed or contain invalid data when applying simple sanity checks. To ensure some basic quality of the extracted data, we ignore entries that only contain empty values. Furthermore, for phone numbers and opening hours, we ensure that the ex10https://parquet.apache.org/ 11https://console.share.innkube.fim.uni-passau.de/ browser/public/ows-extracted%2F https://doi.org/10.5281/zenodo.17228339
Table 1: Occurrence of specific schemata in OWS datasets. Schema Name # % Telephone 5,352,078 1.78 +Address 6,121,713 2.04 ®FAQPage 1,645,691 0.55 OpeningHours 2,210,530 0.73 ÛOpeningHoursSpecification 1,639,338 0.54 tracted string contains at least one digit. Figure 1 illustrates that these simple measures already lead to a large number of discarded entries, showing that many webmasters struggle with obliging to the schema format or insert empty or invalid values. A manual inspection of parts of the extracted data further revealed that while most entries contain sensible information, some webmasters used unhelpful default values, i.e. "question" and "answer" in extracted Q&A pairs. Another issue with data quality is posed by entries that only contain partial information, i.e. an address that only mentions the city, but not the street address of the entity. Further processing of the extracted data to ensure high quality thus poses an important but non-trivial task for our multilingual data. USE CASES In the following sections, we describe two real-world use cases of structured data. The first use case, the extraction of FAQ-style annotations to build a Q&A dataset, has already been implemented. The second use case of extracting structured data to enrich map applications is a work in progress in collaboration with Murena 12 , a company that provides deGoogled and privacy preserving smartphones and cloud services. FAQ Dataset FAQ pages represented in structured data provide an interesting resource for building Q&A datasets. Their natural separation into questions and answers makes it easy to leverage them for question answering tasks. Furthermore, as the FAQ page schema requires an answer to be specified as either accepted or suggested, the schema contains an implicit relevance signal which can be extracted to make the dataset usable for retrieval tasks. Our recent work [Dinzinger et al . (2025)], in which we built a large-scale multilingual retrieval dataset by extracting FAQ page schemas from data provided by the Web Data Commons (WDC) project 13 , clearly demonstrates the use of FAQ-style structured data. Furthermore, we show that multilingual FAQs can be used to build bilingual corpora for a large number of language combinations. Both WebFAQ retrieval 14 and WebFAQ bitext 15 12https://murena.com/ 13https://webdatacommons.org/ 14https://huggingface.co/datasets/PaDaS-Lab/ webfaq-retrieval 15https://huggingface.co/datasets/PaDaS-Lab/ webfaq-bitexts telephone address openingHours 0.0M 5.0M 10.0M 15.0M 20.0M 25.0M 30.0M 35.0M 40.0M found extracted Figure 2: The number of found and extracted entries per schema in millions. are available on HuggingFace and as part of the Massive Text Embedding Benchmark (MTEB) [Muennighoff et al . (2023)] python package. While the WDC dumps provide a large resource of natural Q&A data, they are updated only on a yearly basis, thus likely containing many stale FAQs. The regularly published OWS datasets can alleviate this problem by providing fresh data for crawled web pages. We therefore apply the procedure developed to generate WebFAQ on the OWS data, extracting around 9.95 million Q&A pairs across the five days. To build a multilingual retrieval corpus, we perform language classification on the extracted Q&A pairs using FastText [Joulin et al . (2016)]. Figure 3 shows the distribution of the 10 most common languages found in the extracted FAQ data. While English unsurprisingly occurs most often, we also extract a large number of Q&A pairs for other languages like German, Spanish or French. Similarly to WebFAQ, the FAQs extracted from OWS data are available as a collection of multilingual retrieval datasets on HuggingFace16. Enriching Map Applications Apart from the FAQPage schema serving as a valuable starting point for Q&A datasets, the schemas we have extracted can also serve as a valuable resource for (non- )commercial map applications. To this end, we are currently collaborating with Murena, in an effort to enhance the data provided by OpenStreetMap 17 . While OpenStreetMap provides useful information like addresses or opening hours for points of interest, driven by a community of human mappers that contribute the data, certain parts of this information like the opening hours of a shop might change too frequently to be kept up to date. This can lead to undesirable situations if users rely on incorrect data, e.g. if they choose to visit a shop just to find that the opening hours are outdated and the shop has already closed for the day. To alleviate this 16https://huggingface.co/datasets/PaDaS-Lab/ owi-faq-retrieval 17https://www.openstreetmap.org https://doi.org/10.5281/zenodo.17228339
polish 1.4% italian 2.0% portuguese 2.1% dutch 2.4% japanese 3.1% russian 3.7% french 4.7% spanish 5.3% german 7.0% other 15.4% english 52.8% Figure 3: Language distribution for the 10 most common languages on FAQ pages. problem, we aim to crawl and extract information available in structured data for specific URLs that Murena is interested in on a regular basis. As an initial test, we crawled 10,547 URLs representing points of interest in the area of Seattle and extracted phone numbers from 122 (1.2%) webpages, addresses from 326 (3.1%), FAQs from 290 (2.7%) and opening hours from 711 (6.7%). Although the absolute number of extracted schemas remains low, they can still contribute valuable and fresh information for a large number of locations. As an example, Figure 4 demonstrates how the data extracted from the JSON-LD partially shown in Figure 1 can be presented to users. The data was extracted from the webpage of a Subway store in Seattle and clearly contains information that would benefit a map application. LIMITATIONS While extracting information from structured data seems straightforward at first glance, working with real-world data has proven to be more challenging. One such challenge is posed by the schema definitions themselves. For instance, there are two different ways to specify opening hours, namely using the openingHours 18 schema that provides the information as a dictionary with a defined set of keys or as a simple text as shown in Figure 1. Another common issue are missing values for some fields, fields containing placeholder values or data not conforming to the specified schema. As our main focus lies on extracting the information, we address the first problem by implementing extractors specific for each schema type and storing the information in separate columns of the output Parquet files. Thus, we leave it to downstream applications to merge data from different schemas describing the same entity. While we apply simple sanity checks to the extracted data like checking if the 18https://schema.org/openingHours 1410 156th Ave NE, Bellevue King, WA 98007, US (425) 614-3256 Mo-Fr: 08:00-22:00 Sa-Su: 09:00-22:00 Subway Questions and Answers Question: How can I place a Subway catering order? Answer: To place an order, visit us online at catering.subway.com or call your local restaurant. View all questions and answers Figure 4: The data extracted for a Subway store in Seattle and how it could be presented to users. schema contains only empty strings or whether dates or phone numbers contain at least one digit, doing comprehensive filtering on a multilingual corpus is a non-trivial task. As such, we do not apply any complex filtering techniques on the extracted data to ensure its semantic validity. Additionally, we are unable to verify the correctness or freshness of the extracted data, i.e. if a phone number found on the page of a shop really belongs to it and whether the number is still up to date. Thus, downstream applications wishing to use the extracted information will likely have to implement additional filtering techniques on top of our data to ensure high quality. CONCLUSION Structured data has proven to be an easily extractable and valuable resource for various application scenarios. In this work, we focused on analyzing and extracting certain types of Microdata and JSON-LD from datasets provided by the OWS project. We found that while structured data is available on more than half of the crawled webpages, the occurrence of specific schemas like addresses or opening hours is much less common. Nevertheless, we demonstrate the usefulness of the extracted data in two application scenarios, first generating a question answering and retrieval dataset using the FAQPage schema, and then providing additional information like opening hours, addresses and phone numbers for points of interest, which can be used to enrich map applications. As we believe that providing information extracted from structured data is of general interest, we plan to integrate the extraction mechanism as a regular step in the OWS preprocessing pipeline and create a new collection index for extracted structured data. This would allow interested parties to download only the extracted data instead of the much https://doi.org/10.5281/zenodo.17228339
larger standard OWS datasets, as well as to update information on points of interest on a regular basis without having to set up their own extraction pipelines. We will also expand our work with Murena to crawl more points of interest and provide the extracted information as part of the new collection index. ACKNOWLEDGEMENTS This work has received funding from the European Union’s Horizon Europe research and innovation program under grant agreement No 101070014 (OpenWebSearch.EU, https://doi.org/10.3030/101070014). REFERENCES [Alrashed et al.(2021)] Tarfah Alrashed, Dimitris Paparas, Omar Benjelloun, Ying Sheng, and Natasha Noy. 2021. Dataset or Not? A Study on the Veracity of Semantic Markup for Dataset Pages. In The Semantic Web – ISWC 2021, Andreas Hotho, Eva Blomqvist, Stefan Dietze, Achille Fokoue, Ying Ding, Payam Barnaghi, Armin Haller, Mauro Dragoni, and Harith Alani (Eds.). Springer International Publishing, Cham, 338–356. [Brinkmann et al.(2023)] Alexander Brinkmann, Anna Primpeli, and Christian Bizer. 2023. The Web Data Commons Schema.org Data Set Series. In Companion Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (WWW ’23 Companion). Association for Computing Machinery, New York, NY, USA, 136–139. https://doi.org/10.1145/ 3543873.3587331 [Dinzinger et al.(2025)] Michael Dinzinger, Laura Caspari, Kanishka Ghosh Dastidar, Jelena Mitrović, and Michael Granitzer. 2025. WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval. arXiv:2502.20936 [cs.CL] https://arxiv.org/abs/2502.20936 [Granitzer et al.(2024)] Michael Granitzer, Stefan Voigt, Noor Afshan Fathima, Martin Golasowski, Christian Guetl, Tobias Hecking, Gijs Hendriksen, Djoerd Hiemstra, Jan Martinovič, Jelena Mitrović, Izidor Mlakar, Stavros Moiras, Alexander Nussbaumer, Per Öster, Martin Potthast, Marjana Senčar Srdič, Sharikadze Megi, Kateřina Slaninová, Benno Stein, Arjen P. de Vries, Vít Vondrák, Andreas Wagner, and Saber Zerhoudi. 2024. Impact and development of an Open Web Index for open web search. Journal of the Association for Information Science and Technology 75, 5 (2024), 512–520. https://doi.org/10.1002/asi.24818 [Hendriksen et al.(2024)] Gijs Hendriksen, Michael Dinzinger, Sheikh Mastura Farzana, Noor Afshan Fathima, Maik Fröbe, Sebastian Schmidt, Saber Zerhoudi, Michael Granitzer, Matthias Hagen, Djoerd Hiemstra, Martin Potthast, and Benno Stein. 2024. The Open Web Index. In Advances in Information Retrieval, Nazli Goharian, Nicola Tonellotto, Yulan He, Aldo Lipani, Graham McDonald, Craig Macdonald, and Iadh Ounis (Eds.). Springer Nature Switzerland, Cham, 130–143. [Joulin et al.(2016)] Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016. FastText.zip: Compressing text classification models. https://doi.org/10.48550/ARXIV.1612.03651 [Muennighoff et al.(2023)] Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2023. MTEB: Massive Text Embedding Benchmark. arXiv:2210.07316 [cs.CL] [Navarrete et al.(2019)] Rosa Navarrete, Lorena Recalde, Carlos Montenegro, and Sergio Luján-Mora. 2019. Analyzing Embedded Semantic with JSON-LD and Microdata for Educational Resources in Large Scale Web Datasets. In 2019 International Conference on Computational Science and Computational Intelligence (CSCI). 1133–1138. https: //doi.org/10.1109/CSCI49370.2019.00214 [Peeters et al.(2020)] Ralph Peeters, Anna Primpeli, Benedikt Wichtlhuber, and Christian Bizer. 2020. Using schema.org Annotations for Training and Maintaining Product Matchers. In Proceedings of the 10th International Conference on Web Intelligence, Mining and Semantics (Biarritz, France) (WIMS 2020). Association for Computing Machinery, New York, NY, USA, 195–204. https://doi.org/10.1145/ 3405962.3405964 [Recalde et al.(2021)] Lorena Recalde, Rosa Navarrete, and Fernando Pogo. 2021. Making Open Educational Resources Discoverable: A JSON-LD Generator for OER Semantic Annotation. In 2021 Eighth International Conference on eDemocracy & eGovernment (ICEDEG). 182–187. https: //doi.org/10.1109/ICEDEG52154.2021.9530872 [Volpini et al.(2024)] Andrea Volpini, Jarno van Driel, Ryan Levering, Nurullah Demir, and James Gallagher. 2024. Structured data. HTTP Archive, Chapter 3. https://doi.org/ 10.5281/zenodo.14065771 https://doi.org/10.5281/zenodo.17228339