Full text
Open Music Observatory Building an open data sharing space for the European music sector Daniel Antal, CFA Mester, Anna Márta 2025-11-22
Table of contents Open Music Observatory 7 DisclaimerofWarranties................................. 7 Glossary 9 Musicterms........................................ 9 Creatorsofmusicalworks ............................. 10 Datascienceterms.................................... 11 Dataprotectionterms .................................. 15 Data curation and collection terms . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 Statisticalterms ..................................... 16 Registers, authorities, standards and identifiers . . . . . . . . . . . . . . . . . . . . 17 Organisations....................................... 20 Otherabbreviations ................................... 21 Executive Summary 23 1 Introduction 26 2 Background & Concept 30 2.1 Why Europe Needs a Music Observatory . . . . . . . . . . . . . . . . . . . . . 31 2.2 Historical Precedent: CEEMID . . . . . . . . . . . . . . . . . . . . . . . . . . 31 2.3 Policy and Technological Evolution Enabling a New Observatory . . . . . . . 32 2.3.1 European Parliament and EU-level Mandates . . . . . . . . . . . . . . 33 2.3.2 Data (Sharing) Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 2.3.3 Preference for Open-Source and Open Standards in the EU . . . . . . 34 2.3.4 Alignment with Europeana and Cultural Heritage Infrastructures . . . 34 2.3.5 Alignment with the EU Open Data Portal and EU Open Data Strategy 35 2.3.6 European Interoperability Framework (EIF) . . . . . . . . . . . . . . . 36 2.3.7 EOSC and ECCCH: Open Science and Cultural-Heritage Clouds . . . 36 2.3.8 Summary .................................. 37 2.4 Why Open Music Europe Uses a Decentralised Dataspace Model . . . . . . . 37 2.4.1 Lessons from CEEMID . . . . . . . . . . . . . . . . . . . . . . . . . . 37 2.4.2 Requirements of the EU policy environment . . . . . . . . . . . . . . . 38 2.4.3 Requirements of the Grant Agreement . . . . . . . . . . . . . . . . . . 38 2.4.4 Technical rationale for decentralisation . . . . . . . . . . . . . . . . . . 39 2.4.5 Why decentralisation is essential for a European Music Observatory . 39 2.5 The Data-to-Policy Pipeline: How Open Music Europe Works . . . . . . . . . 40 2.5.1 1. Indicator and problem definition (WP1–WP3) . . . . . . . . . . . . 40 2.5.2 2. Data governance (WP1–WP3, WP6) . . . . . . . . . . . . . . . . . 40 2
2.5.3 3. Software for data collection (WP4) . . . . . . . . . . . . . . . . . . 41 2.5.4 4. Data acquisition (WP1, WP2, WP3) . . . . . . . . . . . . . . . . . 41 2.5.5 5. Processing, enrichment, and harmonisation (WP4, WP5) . . . . . . 41 2.5.6 6. Validation for analysis and dissemination (WP4, WP5) . . . . . . . 42 2.5.7 7. Analysis and modelling (WP1–WP3) . . . . . . . . . . . . . . . . . 42 2.5.8 8. Policy translation (WP5) . . . . . . . . . . . . . . . . . . . . . . . . 42 2.5.9 9. Dissemination and reuse (WP5) . . . . . . . . . . . . . . . . . . . . 43 2.5.10Summary .................................. 43 2.6 Stakeholder Engagement and Early Feedback . . . . . . . . . . . . . . . . . . 43 3 Core Services 45 3.1 Collect: Data Curation & Collection . . . . . . . . . . . . . . . . . . . . . . . 48 3.1.1 Microdata, Collections, Records . . . . . . . . . . . . . . . . . . . . . . 50 3.1.2 Primary data collection . . . . . . . . . . . . . . . . . . . . . . . . . . 51 3.1.3 Metadata .................................. 52 3.1.4 Statistical indicators and datasets . . . . . . . . . . . . . . . . . . . . 53 3.2 Repair........................................ 53 3.3 Process ....................................... 53 3.3.1 Processing & re-processing microdata . . . . . . . . . . . . . . . . . . 54 3.3.2 Documentation............................... 55 3.4 Disseminate..................................... 56 3.4.1 EU Open Data Portal . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 3.4.2 Europeana Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 3.4.3 European Collaborative Cloud for Cultural Heritage . . . . . . . . . . 58 3.4.4 European Open Science Cloud . . . . . . . . . . . . . . . . . . . . . . 59 3.5 Metadata ...................................... 61 3.5.1 Wikibase & Wikidata . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 3.5.2 Music Observatory Website . . . . . . . . . . . . . . . . . . . . . . . . 62 3.5.3 APIEndpoint................................ 62 4 Architecture 63 4.1 Why Wikibase Is the Right Foundation for the Open Music Observatory . . . 63 4.1.1 Proven in real-world scenarios highly similar to music . . . . . . . . . 64 4.1.2 Already aligned with Europe’s digital knowledge infrastructure . . . . 64 4.1.3 Demonstrated support for required OMO functionality . . . . . . . . . 65 4.1.4 Fits EU policy preference for open-source and trustworthy AI . . . . . 65 4.1.5 The most widely used graph-editing interface in the world . . . . . . . 65 4.1.6 A hybrid model that fits real institutional workflows . . . . . . . . . . 66 4.2 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline . . . 66 4.2.1 Wikibase supports each stage of the pipeline . . . . . . . . . . . . . . 66 4.3 Summary ...................................... 68 5 Data coordination 69 5.0.1 The European Interoperability Framework (EIF) . . . . . . . . . . . . 69 5.0.2 Extending the EIF to public and private service coordination . . . . . 71 5.0.3 Datasharingspace............................. 72 3
5.1 Ontologies and Vocabularies in the Open Music Observatory . . . . . . . . . 73 5.1.1 Lightweight Ontology Patterns . . . . . . . . . . . . . . . . . . . . . . 75 5.2 Multiple roles, multiple workflows to support . . . . . . . . . . . . . . . . . . 76 5.2.1 Polyhierarchy................................ 78 5.2.2 Formalisation................................ 80 5.3 Future-Proofing................................... 80 5.3.1 Future-proofing through graph architecture . . . . . . . . . . . . . . . 81 5.3.2 Stabilising definitions through internationally defined standard vocabularies.................................. 81 5.3.3 Future-proofing through translatability and multiple serialisations . . 82 5.3.4 Future services through institutional interoperability . . . . . . . . . . 83 5.3.5 A concrete example: ALOADED, Livonian folk music, and DDEX . . 83 6 Data Sources 86 6.1 How data sources were identified . . . . . . . . . . . . . . . . . . . . . . . . . 86 6.2 Administrative and register data . . . . . . . . . . . . . . . . . . . . . . . . . 87 6.3 Surveydata..................................... 88 6.4 Statistical and economic data . . . . . . . . . . . . . . . . . . . . . . . . . . . 89 6.5 Platform and streaming data . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 6.6 Harmonisation tools, metadata structures, and vocabularies . . . . . . . . . . 90 6.7 Role of data sources in the data-to-policy pipeline . . . . . . . . . . . . . . . 91 6.8 Integration with the Open Music Observatory . . . . . . . . . . . . . . . . . . 91 7 Data Collection 93 7.1 Overview of the Data-Collection Framework . . . . . . . . . . . . . . . . . . . 93 7.2 Administrative and Register Data . . . . . . . . . . . . . . . . . . . . . . . . . 94 7.3 SurveyData..................................... 94 7.4 Statistical and Economic Data . . . . . . . . . . . . . . . . . . . . . . . . . . 95 7.5 Platform and Streaming Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 7.6 Processing and Harmonisation . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 7.7 Software Components Developed in WP4 . . . . . . . . . . . . . . . . . . . . 96 7.7.1 Data-Ingestion Tools . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 7.7.2 Validation and Reconciliation Tools . . . . . . . . . . . . . . . . . . . 96 7.7.3 Harmonisation and Metadata Tools . . . . . . . . . . . . . . . . . . . 96 7.7.4 OMO Integration Tools . . . . . . . . . . . . . . . . . . . . . . . . . . 97 7.8 Position of Data Collection and Processing in the Pipeline . . . . . . . . . . . 97 7.9 Integration with the Open Music Observatory . . . . . . . . . . . . . . . . . . 97 8 Standardisation of Data & Terminology 98 8.1 Businessprocesses ................................. 98 8.2 Conceptual and information models . . . . . . . . . . . . . . . . . . . . . . . 99 8.3 Identification & Entity Linking . . . . . . . . . . . . . . . . . . . . . . . . . . 101 8.3.1 Registers & Authority Files . . . . . . . . . . . . . . . . . . . . . . . . 102 8.3.2 Open and persistent identifiers . . . . . . . . . . . . . . . . . . . . . . 103 8.3.3 Not open, music-industry specific identifiers . . . . . . . . . . . . . . . 104 8.3.4 Lyrics ....................................105 4
8.3.5 ISCC ....................................105 8.3.6 OMOIdentifiers ..............................106 9 Data Improvement & Innovation 108 9.1 Value-Added Data Services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108 9.1.1 DataSharing................................108 9.1.2 Fix-the-data ................................109 9.1.3 DataLinking ................................109 9.1.4 Registration services . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 9.2 UseCases......................................111 9.2.1 Data Health Services for Collective Management . . . . . . . . . . . . 112 9.2.2 Sustainability Reporting for Music Organisations . . . . . . . . . . . . 113 9.2.3 ListenLocal.................................115 9.2.4 Unlabel ...................................116 9.3 UseofAIsystems .................................117 10 Data Catalogue 120 10.1CollectionGuidelines................................122 10.2TopicalPillars ...................................123 10.2.1MusicEconomy...............................124 10.2.2MusicDiversity...............................126 10.2.3MusicSociety................................127 10.2.4Innovation..................................128 10.2.5Sustainability................................128 References 129 Appendices 135 Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network 135 ...............................................135 Open Music Data Exchange . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 SKCMDb: Slovak Comprehensive Music Database 137 The Slovak Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 Slovak Comprehensive Music Database (public) . . . . . . . . . . . . . . . . . . . . 138 Slovak Comprehensive Music Database (private) . . . . . . . . . . . . . . . . . . . 139 Microdata......................................139 Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 140 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140 LīvMDb: Livonian Music Database 142 The Livonian Metadata Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 Livonian Music Database (public) . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 Livonian Music Database (private) . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 Microdata......................................144 5
Statistical Data & Data Catalogue . . . . . . . . . . . . . . . . . . . . . . . . 145 Publications & Catalogue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145 6
Open Music Observatory ¾Partly updated This document was presented as a planning document in 2023. It has been partly updated till 18 November 2025. Some text is still reflecting the planning phase. You can access all versions on https://zenodo.org/records/16539570 Disclaimer of Warranties This project has received funding from the European Union’s Horizon Europe, research and innovation programme, under Grant Agreement No. 101095295. This document has been prepared by Open Music Europe (OpenMusE) project partners as an account of work carried out within the framework of this contract. Any dissemination of results must indicate that it reflects only the author’s view and that the Commission Agency is not responsible for any use that may be made of the information it contains. Neither Project Coordinator, nor any signatory party of Open Music Europe (OpenMusE) Project Consortium Agreement, nor any person acting on behalf of any of them: (a) makes any warranty or representation whatsoever, express or implied, (i). with respect to the use of any information, apparatus, method, process, or similar item disclosed in this document, including merchantability and fitness for a particular purpose, or (ii). that such use does not infringe on or interfere with privately owned rights, including any party’s intellectual property, or 7
(iii). that this document is suitable to any particular user’s circumstance; or (b) assumes responsibility for any damages or other liability whatsoever (including any consequential damages, even if Project Coordinator or any representative of a signatory party of the Open Music Europe (OpenMusE) Project Consortium Agreement, has been advised of the possibility of such damages) resulting from your selection or use of this document or any information, apparatus, method, process, or similar item disclosed in this document. For the version history of this document, please refer to our open repository, where the change history can be reviewed with timestamps for every single file used to create the report: https://github.com/dataobservatory-eu/open-music-observatory 8
Glossary Music terms audio recording: fixation of sounds (ISO 2019b) music video recording: fixation of sounds synchronized with pictures or moving pictures where (a) the fixed sounds are wholly or substantially a musical performance or (b) the recording is intended for viewing in association with a recording of a musical performance. This definition includes music videos and concert recordings, together with music-related interviews and documentaries, but does not extend to genera! audiovisual material, even if it includes music.(ISO 2019b) recording: result of a recording process independent of the type and number of audio or audiovisual carriers and technology used Note 1 to entry: The term “recording” applies to each recorded item which may be used as a separate unit regardless of whether it is issued as part of a larger recorded work (e.g. each separate track on an album of audio recordings). [SOURCE:ISO 3901:2001, definition 3.3] (ISO 2017b) work: distinct, abstract creation of the mind whose existence is revealed through one or more expressions (e.g. a performance) or manifestations (e.g. an object) (ISO 2022) musical work: composed of a combination of sounds, with or without accompanying text (ISO 2022) DSP or digital streaming platform: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. DSPs can provide access to music downloads, like Apple’s iTunes Store, or access to streaming music like Spotify, or even provide satellite-delivered content such as SiriusXM in the USA. rights management (organisations): the function of managing the rights on behalf of rights owners. It can be companies whose sole purpose is to ensure that content that has been licensed has delivered royalties that are identified and accounted for. The role can be taken by collective management organisations or by private companies on behalf of songwriters, composers, performers, music publishers, or record labels. duration: the elapsed playing time between the first and last recorded modulations of the recording. LP or Long Player: gramophone record usually on both sides comprising one or more sound recordings with a playing time of each side of normally round about 30 minutes and released and sold on its own (ISO 2017b) 9
anthology: document consisting of a collection of full documents or of extracts, usually of literary works (ISO 2017b) exhibition: curated display of objects on a clear concept and communicating a message [SOURCE:ISO 18461:2016, definition 2.4.6 modified] (ISO 2017b) curator: person responsible for overseeing a collection or exhibition (ISO 2017b) data curation: managed process, throughout the data lifecycle, by which data/data collections are cleansed, documented, standardized, formatted and interrelated (ISO 2017b) register: an official list or record of names or items; it aims to be a complete list of the objects in a specific group of objects or population, for example, all copyright-protected musical works in a country, or all legal person enterprises in another country; a document, usually a volume, in which data are entered in a formal manner by a statutory authority Note 1 to entry: In modern usage, usually a database. (ISO 2017b) registration: act of giving an entity a unique identifier on its entry into a system (ISO 2017b) set of rules, operations, and procedures for inclusion of an item in a registry (ISO 2023a) registrant: organization or person that has either registered an authentication protocol or registered the adoption of an authentication protocol [SOURCE: ISO/IEC 24727-6:2010, definition 3.4] (ISO 2017b); an entity wishing to assign an ISRC to an applicable recording (ISO 2019b); aparty that requests an ISNI from the Registration Authority (ISNI 3.2 (ISO 2012, p15)) party: natural person or legal person, whether or not incorporated, or a group of either (ISO 2012) aggregation: acquisition of sensitive information by collecting and correlating information of lesser sensitivity (ISO 2023b) Statistical terms administrative records: data generated by a non‐statistical source, usually a public body, the main aim of which is not the provision of statistics. code list: predefined list from which some statistical coded concepts take their values (ISO 2013) data pipeline: a method in which raw data is ingested from various data sources and then ported to data store. FAIR or FAIR Guiding Principles for scientific data management and stewardship: guidelines to improve the Findability, Accessibility, Interoperability, and Reuse of digital 16
assets, emphasising machine-actionability (i.e., the capacity of computational systems to find, access, interoperate, and reuse data with none or minimal human intervention.) indicator: the representation of statistical data for a specified time, place or any other relevant characteristic, corrected for at least one dimension (usually size) so as to allow for meaningful comparison. microdata: non‐aggregated observations or measurements of characteristics of individual units, without direct identifier. MVP or minimum viable product: a version of a work product with just enough features and requirements to satisfy early customers and/or provide feedback for future development [SOURCE:IEEE 2675-2021, 3.1] observation unit: an identifiable entity about which data can be obtained, it is also often called a statistical unit or data subject in case of a natural person. Open Policy Analysis Guidelines: a set of information management rules to make policy analysis more transparent. personal data: any information relating to an identified or identifiable natural person. pseudonymisation: processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information. survey: a systematic examination and record of a physical or social area and its features so as to construct a map, plan, or description. In social sciences it usually refers to a well-structured questionnaire and answers given to its items by a target population. statistics: quantitative and qualitative, aggregated and representative information characterising a collective phenomenon in a considered population. visualisations: schematic charts, drawings, photographs, and their collages will as still image files that help to explain the relationship between information carriers, data points, or processes. Registers, authorities, standards and identifiers IČO: The organisation identification number (IČO) is an identifier assigned to all types of legal entities, entrepreneurs and public authorities by the Statistical Office of the Slovak Republic. The Czech Republic’s organisation identifier is also called IČO. OpenCorporates: a public corporation database which sources data from national business registries. ISNI: an ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. VIAF: The Virtual International Authority File (VIAF) is an international service that consolidates multiple name authority files into a single database. Their primary goal is 17
to enhance the efficiency and usability of library authority files by linking and merging widely used authority records and making them accessible online. VIAF ID: The VIAF (Virtual International Authority File) combines multiple name authority files into a single OCLC-hosted name authority service. ISRC: The International Standard Recording Code (ISRC) is a standard identifying code that can be used to identify sound recordings and music video recordings so that each such recording can be referred to uniquely and unambiguously. ISWC: The purpose in creating an ISWC for musical works is to enable more efficient administration of rights to those works on a worldwide basis. The ISWC provides an efficient means of identifying musical works in computer databases and related documentation and for the exchange of information between rights societies, publishers, record companies and other interested parties on an international level. ISBN: the International Standard Book Number is an identification system for the publishing industry and its supply chains. ISMN: The International standard music number (ISMN) was developed by, and for, the music publishing sector as a separate system to complement the International standard book number (ISBN). The existence of the ISMN as a separate identifier system makes it possible to identify printed and notated music as a distinct category of publication within the global supply chain and to develop trade directories and similar services for the specialized market for music publications. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. DOI: The Digital Object Identifier is a standardised unique number given to many (but not all) articles, papers and books, by some publishers, to identify a particular publication. ORCID: the Open Researcher and Contributor ID is a unique, persistent identifier free of charge to researchers. URI: A Uniform Resource Identifier (URI) is a string of characters used to identify a resource on the internet. This resource can be either abstract or physical, such as a website, an email address, or a file. URIs are essential for enabling interactions with resources over a network using specific protocols. W3C: The World Wide Web Consortium (W3C) is an international community that develops standards for the World Wide Web. Their mission is to lead the Web to its full potential by creating technical specifications and guidelines that are designed to be open and royalty-free. These standards include HTML, CSS, and other web technologies, which ensure that web content is accessible across different browsers and devices. DDI: The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. Wikibase: Wikibase is a software system that help the collaborative management of knowledge in a central repository. It was originally developed for the management of Wikidata, 18
but it is available now for the creation of private, or public-private partnership knowledge graphs. It is developed by Wikimedia Deutschland. GSBPM: The Generic Statistical Business Process Model is a international standard model that “describes and defines the set of business processes needed to produce official statistics.” | GSIM:Generic Statistical Information Model: a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation standards such as SDMX or DDI. SDMX: Statistical Data and Metadata eXchange (SDMX), is an international initiative that aims at standardising and modernising (“industrialising”) the mechanisms and processes for the exchange of statistical data and metadata among international organisations and their member countries. ESRS: The European Sustainability Reporting Standards (ESRS) are a set of guidelines developed by the European Financial Reporting Advisory Group (EFRAG) to standardise sustainability reporting across the European Union. These standards are designed to align with the Corporate Sustainability Reporting Directive (CSRD), which mandates detailed corporate reporting on environmental, social, and governance (ESG) issues for many companies operating within the EU. CIDOC-CRM: The conceptual model of CIDOC, the standard conceptualisation of collection management systems in heritage organisations. RiC:Records in Context is a new conceptual model that replaces the four most important international archiving standards. DCTERMS or DCMI: the Dublin Core Metadata Terms is a vocabulary of metadata terms developed and maintained by the Dublin Core Metadata Initiative (DCMI). These terms are used to describe various aspects of digital resources, such as web pages, documents, and other online content. They provide a standardized way to assign metadata to resources, making them easier to discover, manage, and exchange. RDFS: the Resource Description Framework Schema is an extension of the Resource Description Framework (RDF) that provides a vocabulary for describing classes and properties of resources within an RDF graph. EDM: the Europeana Data Model is a framework for collecting, connecting, and enriching cultural heritage metadata. It’s designed to facilitate the sharing and reuse of cultural heritage information by providing a standardized way to represent and link data. Europeana: a digital platform provided by the European Union that aggregates digitized cultural heritage from institutions across Europe. ESCO: the European Skills, Competences, Qualifications and Occupations classification is is a multilingual classification system developed by the European Commission to standardize the description of skills, competences, and qualifications relevant to the European labor market and education. 19
NACE: the European Union’s standard classification of economic activities for statistical purposes. The abbreviation stands for Nomenclature statistique des Activités économiques dans la Communauté européenne. ISCO: the International Standard Classification of Occupations (ISCO) is the International Labour Organization’s standardized system for classifying and organizing occupations according to jobs’ tasks and duties ISIC: the International Standard Industrial Classification of All Economic Activities (ISIC) is a standard classification system developed by the UN Statistics Division (UNSD) to categorize economic activities. PROV-O: the Provenance ontology is a formal ontology developed by W3C to represent and interchange provenance information. MARC: MAchine-Readable Cataloging, is a standard digital format used by libraries to represent and exchange bibliographic information. DCAT: an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. Organisations AEPO-ARTIS: Organisation representing European artists-performers. Regroups most of the European CMO representing performers. ALOADED: is a company which distributes and exploits recordings. CISAC: The International Confederation of Societies of Authors and Composers is an international non-governmental, not-for-profit organisation that aims to protect the rights and promote the interests of creators worldwide. CNM (former CNV): the Centre National de la Musique is a public organisation managing a tax on concert tickets EFRAG: The European Financial Reporting Advisory Group is a private association established in 2001 with the encouragement of the European Commission to serve the public interest. EFRAG extended its mission in 2022 following the new role assigned to EFRAG in the CSRD, providing Technical Advice to the European Commission in the form of fully prepared draft EU Sustainability Reporting Standards and/or draft amendments to these Standards. EMO: The European Music Observatory (EMO) is envisioned as a hub for collecting and analysing data on the music sector across Europe. Its primary aim is to address the current gaps and inconsistencies in music data collection, which have been a significant challenge for the sector. GESAC: GESAC comprises together 32 European authors’ societies in music, audiovisual, visual arts, literature and drama. 20
GESIS: Leibniz Institute for the Social Sciences. IAML: International Association of Music Libraries, Archives and Documentation Centres | IAMIC: International Association of Music Centres, an international network of organisations that collectively and collaboratively provides information and promotes the music of their countries or regions. ICMP: the global trade body representing the music publishing industry worldwide. SCAPR: International association for the development of the practical cooperation between performers’ collective management organisations (CMOs) SOZA: SOZA (Slovenský ochranný zväz autorský pre práva k hudobným dielam, Slovak Performing and Mechanical Rights Society) is a legal entity, non-profit civic association of authors and publishers of musical works, association of natural persons and legal entities. Hudobné Centrum: Music Centre Slovakia is a music organisation with a mission to promote Slovak contemporaly music. Other abbreviations CEEMID: the Central European Music Industry Databases is a multi-country project that was a predecessor of Reprex’s Digital Music Observatory CSRD: The Corporate Sustainability Reporting Directive (CSRD) is European Union (EU) legislation, effective from 5 January 2023, that requires EU businesses—including qualifying EU subsidiaries of non-EU companies—to disclose their environmental and social impacts, and how their environmental, social and governance (ESG) actions affect their business. DSP: Digital service providers (DSPs), or Digital Streaming Platforms are companies or organisations that provide access to services online. EIF: The European Interoperability Framework (EIF) is a set of recommendations and guidelines that aims to facilitate communication and collaboration between public administrations, businesses, and citizens within the European Union and across national borders. ECCCH: The European Collaborative Cloud for Cultural Heritage is a European Union initiative for a digital infrastructure that will connect cultural heritage institutions and professionals across the EU. EOSC: The European Open Science Cloud (EOSC) aims to create a trusted, open, and multidisciplinary environment for researchers and innovators in Europe. PPP: A Public-Private Partnership (PPP) is a collaborative arrangement between government entities and private sector companies aimed at financing, designing, implementing, and operating projects or services traditionally provided by the public sector. RDM: Research Data Management refers to the suite of practices, policies, and processes used to handle data throughout the lifecycle of a research project. 21
Our glossary is harmonised with relevant music-sector specific standards (referred to in Chapter 8) and with the ISO Information technology — Vocabulary (ISO 2023b); Information technology — Cloud computing — Taxonomy based data handling for cloud services (ISO 2020); Information technology — Cloud computing — Interoperability and portability (ISO 2017a) and the Information and documentation — Foundation and vocabulary (ISO 2017b) and Information technology — Metadata registries (MDR) — 1. Framework (ISO 2023a) 22
Executive Summary Our ambition with the development of the Open Music Observatory is to provide the technological basis and a practical roadmap for creating a European Music Observatory in a bottom-up, decentralised way. Instead of waiting for a grand, central agreement on what should a European music observatory be collecting and who should control it, we suggest a pragmatic approach: allow any data owners and collectors who satisfy certain quality and cooperation rules to add their data to an Open Music Observatory; when it reaches a sufficient maturity for use in Europe, then decide if its maintenance requires a new institutional form or not. Creating the Open Music Observatory is a cornerstone task of the OpenMusE project. This task is running till the end of the project (31 December 2025) with the collection, processing, and dissemination of more data and providing innovative, new data services in line with our exploitation pathways. This report is an accompanying document for the creation of Open Music Observatory as a digital infrastructure on the World Wide Web. In the OpenMusE project, the development of the Open Music Observatory is coupled with a clear contractual expectation: the project must populate the Observatory’s four thematic pillars—Music Economy, Music Diversity, Music, Society & Sustainability, and Innovation & Future Trends —with initial, well-documented data and knowledge. This population process follows the project’s data-to-policy pipeline (see Chapter 1): each work package defines its indicators, establishes data governance and legal bases, collects or accesses relevant administrative, survey, statistical, and platform data, processes and harmonises them through WP4 tools, and finally activates them in reproducible analytical workflows in WP5. The result is that the Observatory is not only a technical prototype but a functional, databearing infrastructure: the first integrated demonstration of how Europe’s music data can be curated, linked, analysed, and made reusable across public, private, and civic actors. The Open Music Observatory is a digital service provider for the music industry that follows the European Interoperability Framework (EIF) definition for such services with a unique governance model. The governance model and the digital service infrastructure represent a unique innovation that considers many good examples from the European Union and other industries. An observatory has traditionally been a permanent location for observing terrestrial, marine, or celestial events. In the past 30 years, it has also been used for long-term digital data collection programs for markets, social sciences, and humanities. Our milestone requires the start of this observatory after a lengthy and intensive planning and prototyping phase. It can be seen as a modern reimagination of the data observatory model, or the observatory 2.0. We created a new observatory model that fully aligns with the European Interoperability 23
Framework but extends the governance of the digital services beyond public bodies, and allows the creation of a public-private partnership to manage the observatory. The European Interoperability Framework aims to create a four-layered approach to build digital research, marketing, rights management, collection management services for the sector. These layers are introduced in separate chapters of these documentation. 1. The technological alignment is introduced in Chapter 4; we decided to choose the technology of the world’s largest open knowledge graph, Wikidata, which already coordinates countless digital services in Europe’s cultural sector, and provides training and guardrails for many AI applications. 2. The semantic alignment is is introduced in Chapter 5. This chapter focuses on the semantic and organisational aspects of interoperability. 3. The organisational alignment means bridging actual data-driven and computer supported workflows with data semantics (what does a song’s title mean for a librarian, an ethnomusicologist, a collective rights management agency), and how they can work with various translated, alternative, historical, mistyped, and preferred titles in distributing royalties, loaning printed sheets, describing musical traditions. 4. The legal alignment creates a policy that lays out the rights, prohibitions and necessary permission processes to connect and use the data together. We give a concrete example in Section 5.3. We were informed and influenced by the creation of Europeana (which started out from a similar collaborative project, a cultural heritage oriented data sharing space) and the Commission’s new plans to extend their digital services into the European Collaborative Cloud for Cultural Heritage (ECCCH). We aimed for full interoperability with Europeana and we Reprex successfully sent data and concluded a Data Exchange Agreement. We were also aiming for interoperability with the ECCCH, which only published the first version of its Heritage Digital Twin conceptual model and ontology; we were the first to test them with music data. but we also bring a new element into their thinking. While they are mainly aggregating the work of public sector memory institutions, we are building a governance model that allows a more successful cooperation among the private sector and the public music sector. By the end of 2025, we aim to create an “observatory 3.0”, which already hosts many intelligent data improvement technologies and fuels innovative applications/services in line with our project’s exploitation pathways. These services are at different maturity levels, but they could not be brought to a testable MVP without building out the minimal digital infrastructure and governance model at this milestone. 24
ĹNote This document is licensed under the CC BY 4.0 LEGAL CODE Attribution 4.0 International license. You must refer to the document with the DOI 10.5281/zenodo.11385044. Canonical Licence URL:https://creativecommons.org/licenses/by/4.0/ Other formats:Plain Text;RDF/XMLlSee the deed 25
collective management societies. Over time, it grew to include more than 60 stakeholders in 12 European countries. Its purpose was to fill the most pressing evidence gaps by combining: • voluntary data integration among partners, • open-data reprocessing, and • co-financed data collection. This work is documented in (Antal 2020a). CEEMID operated according to principles that would later become central to the European Union’s data (sharing) space strategy (formalised only years later). Its decentralised organisational model, distributed data stewardship, and emphasis on transparent, reusable methods demonstrated that a modern observatory in the digital era does not need to be a centralised institution. Instead, it can function as a federated ecosystem connecting statistical offices, cultural institutions, CMOs, and private actors. Long before the EU formalised its dataspace strategy, CEEMID also aligned its workflows with emerging statistical-system standards such as GSIM,DDI, and SDMX, anticipating later European requirements for interoperable, machine-readable statistical metadata. This early adoption provided a methodological bridge between cultural-sector data, administrative registers, and official statistics, and formed a direct precursor to the metadata foundations of the Open Music Observatory. The Feasibility Study for the European Music Observatory explicitly recognised CEEMID as a potential building block for a new observatory model. Our proposal therefore sought to transform CEEMID’s prototype—referred to in the study as the Digital Music Observatory— into a scientifically robust and methodologically coherent system that could scale across Europe. This required grounding the work in state-of-the-art statistical science, data science, and computer science, and ensuring alignment with European interoperability and datagovernance frameworks. The prototype work that preceded Open Music Europe was shaped through two innovation environments: the Yes!Delft AI+Blockchain Lab, where product–market fit and technical feasibility were tested, and the JUMP Music Market Accelerator, where the first integrated prototype of a Digital Music Observatory was developed. These early iterations validated not only stakeholder demand but also the feasibility of a decentralised, standards-based architecture, and they informed the methodological and technical design choices taken forward in this project. 2.3 Policy and Technological Evolution Enabling a New Observatory Since the publication of the EMO Feasibility Study, the European Union has introduced a series of policy and infrastructure initiatives that strengthen the case for a decentralised, in32
teroperable, and federated European Music Observatory. These developments span European Parliament mandates, Commission-funded research, cultural-heritage clouds, opendata regulation, and the EU’s overarching data-space strategy. Together, they establish the policy and technological foundations on which the Open Music Observatory is built. Our policy alignment is discussed in more detail in - Music Metadata Mainstreaming and EU Law -A Green Paper on AI, Data Governance, and Metadata–Policies for Europe’s Music Ecosystem2 2.3.1 European Parliament and EU-level Mandates The European Parliament, in its resolutions on the future of the music sector, explicitly called for: • the establishment of a European Music Observatory, • improved evidence for competitiveness, diversity, and fair remuneration, and • stronger coordination of public, private, and community data sources. These mandates update and reinforce both the Music Moves Europe framework and the findings of the EMO Feasibility Study. They frame the Observatory as an instrument that must serve industry, civic, and public actors through interoperable, reusable, crossborder data services. The EU Music Ecosystem Study (2024) deepened this diagnosis, pointing to fragmentation across metadata, rights information, cultural statistics, and market data. It concluded that the sector requires a technical and governance model capable of linking these domains, rather than separate, siloed initiatives. The architecture of our dataspace responds directly to these recommendations3. 2.3.2 Data (Sharing) Spaces The EU’s adoption of data (sharing) spaces provides the organisational and legal model for an Observatory that is not a centralised institution but a federated ecosystem. Curry defines dataspaces as: “an emerging approach to data management… Data is integrated on an ‘asneeded’ basis, with the labour-intensive aspects of data integration postponed until they are required.” (Curry 2020) The Design Principles for Data Spaces position paper further describes them as: 2See (Senftleben et al. 2024); and (Antal 2025d), summarised in the internal document (Open Music Europe Consortium 2025). 3(Music Moves Europe 2024; European Parliament 2024) 33
“a federated data ecosystem within a certain application domain and based on shared policies and rules.” (Nagel and Lycklama 2021, p7) These principles are fully consistent with CEEMID’s decentralised model and form the conceptual basis for the Open Music Dataspace (see Chapter 5). Observatories created in the 1990s and early 2000s were built around centralised databases and slow-moving data-collection cycles. Since then, the rapid expansion of agentic AI in data collection, the widespread digitisation of live and recorded music, and the proliferation of large-scale, real-time data sources have made such centralised architectures obsolete. Modern evidence ecosystems require automated ingestion, continuous semantic enrichment, cross-domain reconciliation, and transparent provenance — all of which presuppose a federated, decentralised model rather than a single institutional database. The European Audiovisual Observatory (EAO), the European Market Observatory for Fisheries and Aquaculture Products (EUMOFA), and the European Observatory on Infringements of Intellectual Property Rights (EUIPO) provide valuable models of long-standing EU observatories. However, each operates within a centralised data-submission and aggregation framework appropriate to their legal mandates and sectoral data structures. The Feasibility Study acknowledged that the music sector lacks comparable legal obligations and contains far more fragmented, cross-domain, multilingual, and institutionally diverse datasets. Therefore, while these observatories offer important governance precedents, their centralised architectures cannot be replicated in the music ecosystem — strengthening the case for a federated dataspace model. 2.3.3 Preference for Open-Source and Open Standards in the EU Across the EU’s data and digital-transition strategies, there is a consistent preference for: •open-source software, •open standards, •open licensing, and •transparent, reproducible workflows. This aligns directly with the Observatory’s use of open-source R and Python pipelines, Wikibase for semantic interoperability, and FAIR-compliant metadata. 2.3.4 Alignment with Europeana and Cultural Heritage Infrastructures Europeana demonstrates how Europe manages distributed cultural-haritage collections at scale using: 34
• persistent identifiers, • multilingual metadata, • open licences (e.g. CC BY), • shared semantic standards (EDM, IIIF, rightsstatements.org), and • decentralised stewardship by libraries, archives, and museums. The Open Music Observatory follows the same principles. It uses: •semantic technologies, •PID-based cross-domain linking, and •open, reusable data models. This ensures interoperability with cultural-heritage collections, performing-arts archives, and national memory institutions, and aligns the music domain with the emerging European Collaborative Cloud for Cultural Heritage (ECCCH). 2.3.5 Alignment with the EU Open Data Portal and EU Open Data Strategy The EU Open Data Portal (data.europa.eu) establishes a common framework for: • open licences (e.g. CC BY 4.0), • machine-readable formats, • harmonised metadata (DCAT-AP), • and publication of public-sector information. The Open Music Observatory is designed so that: • public datasets can be harvested directly by the EU Open Data Portal, • indicators and derived datasets comply with open-data rules, and • metadata follow DCAT-AP and DataCite to support long-term reuse. This alignment ensures that the Observatory meets both Horizon Europe open-science requirements and broader EU open-data policy objectives. 35
2.3.6 European Interoperability Framework (EIF) The European Interoperability Framework (EIF) provides a four-layer model—legal, organisational, semantic, technical—for connecting: • public authorities, • cultural institutions, • rights-management organisations, • national statistical offices, and • private intermediaries. These are precisely the actors whose data must interoperate to support a European Music Observatory. By adopting the EIF, the Observatory can link diverse datasets into coherent, reusable services without centralising them. 2.3.7 EOSC and ECCCH: Open Science and Cultural-Heritage Clouds The European Open Science Cloud (EOSC) and the European Collaborative Cloud for Cultural Heritage (ECCCH) promote: • FAIR data, • open science workflows, • reproducible analysis, • transparent provenance, and • decentralised storage and processing. These principles inform the Observatory’s architecture through the use of: • open-source analytical pipelines, • SDMX and DataCite metadata, • persistent identifiers, and • federated linking across domains and institutions. 36
2.3.8 Summary Together, these EU policy instruments—the Parliament’s mandate, the EU Music Ecosystem Study, data-space strategy, Europeana, the EU Open Data Portal, the EIF, EOSC, and ECCCH—provide a unified rationale for an Observatory that is federated, decentralised, data-driven, and interoperable by design. They define the policy and technological environment in which the Open Music Observatory must operate and directly shape its architecture. The Chapter 4explains why we chose an architecture that is built around Wikibase and Wikiadta. 2.4 Why Open Music Europe Uses a Decentralised Dataspace Model The Open Music Observatory adopts a decentralised, federated dataspace model because this is the only architecture that meets the needs identified by the EMO Feasibility Study, the EU Music Ecosystem Study, and the European Parliament’s resolutions, while also complying with the newer EU frameworks for interoperability, data governance, and cultural-heritage infrastructures. A centralised database model, common in observatories built in the 1990s or early 2000s, is no longer feasible or desirable for the music sector. 2.4.1 Lessons from CEEMID The CEEMID collaboration demonstrated that most music-sector data—repertoire, rights, cultural-heritage descriptions, business metadata, and statistical evidence—originate from many different institutions, each with its own mandates, legal obligations, and technical systems. Centralising such data is: • legally constrained (e.g. GDPR, contractual confidentiality), • institutionally unrealistic (distributed ownership and stewardship), and • technically inefficient (rapidly evolving local systems). CEEMID showed that these data can nonetheless be made interoperable through: • shared identifiers and authority files, • open metadata standards, • reproducible R-based pipelines, and • rule-based, voluntary data sharing. These are the foundational principles of a data (sharing) space, which the EU has since elevated to a core strategic component of its digital-policy agenda. 37
2.4.2 Requirements of the EU policy environment As outlined in Section C, the EU now expects cultural and creative sectors to adopt: •federated data architectures, •FAIR and open data practices, •transparent governance models, •semantic interoperability, and •alignment with Europeana, EOSC, ECCCH, and data.europa.eu. This expectation reflects the broader transformation of European data governance, where sectors are encouraged to organise around data spaces rather than central repositories. A decentralised model also supports cultural and data sovereignty by allowing institutions to maintain control over their collections and data-processing rules. 2.4.3 Requirements of the Grant Agreement The Open Music Europe Grant Agreement defines the project explicitly as: “an open, scalable data-to-policy pipeline for European music ecosystems” and mandates the creation of: “a highly automated, decentralised intelligence hub that aggregates open data and creates dynamic, live policy documents.” To fulfil these contractual obligations, the Observatory must: • connect heterogeneous data sources without centralising them, • refresh indicators automatically as upstream data changes, • maintain legally sound provenance across many institutions, • support multilingual, cross-border metadata, and • integrate statistical, cultural-heritage, and industry systems. These requirements can only be met in a federated dataspace, not in a single, centralised database. 38
2.4.4 Technical rationale for decentralisation The dataspace model makes it possible to: • keep sensitive or personal data (e.g. rights, royalties) within the institution that controls them, • link sources through semantic federation (Wikibase/Wikidata), • enable distributed curation by librarians, archivists, CMOs, and researchers, • integrate permanent identifier (PID) systems across domains (ISNI, VIAF, ROR, company registers), • use open standards (SDMX, DDI, DataCite, DCAT-AP), and • scale to new partners, genres, languages, and Member States. This structure mirrors the actual distribution of data in the music sector and the technical direction of the EU’s digital transition. 2.4.5 Why decentralisation is essential for a European Music Observatory For the European music ecosystem, decentralisation enables: • lower administrative and compliance burdens, • institutional autonomy and data sovereignty, • cross-border comparability without forced data transfer, • communityand expert-driven metadata improvement, • GDPR-compliant handling of personal data, and • sustainable expansion of the Observatory. A decentralised dataspace is therefore not an architectural choice but a necessary governance model for an Observatory that spans cultural heritage, rights management, statistical registers, community archives, and private-sector metadata across the EU. The Open Music Observatory is consequently designed as a federated, rule-based dataspace: an ecosystem where public, private, and civic stakeholders contribute knowledge, maintain authority records, and generate indicators while preserving full control over their own data. 39
2.5 The Data-to-Policy Pipeline: How Open Music Europe Works The Open Music Europe action is contractually defined as “an open, scalable data-topolicy pipeline for European music ecosystems” (see Grant Agreement). This is not a slogan: it is the methodological core of the project and the organising principle of all work packages (WP1–WP5). The pipeline connects indicator design, data governance, data acquisition, semantic modelling, statistical analysis, and policy translation into a single reproducible workflow. This chapter introduces the logic of that pipeline and explains how it shapes the design of the Open Music Observatory. 2.5.1 1. Indicator and problem definition (WP1–WP3) Each thematic work package begins by identifying policy-relevant gaps and defining the indicators needed to address them. Deliverables D1.1, D2.1, and D3.1 specify: • the conceptual frameworks guiding each domain (economy, diversity, society), • the data requirements for measuring them, and • the procedures for ensuring comparability across countries and years. These definitions also appear in the Open Music Europe Data Management Plan (D6.3), which provides human-readable summaries and machine-readable metadata for all indicators. 2.5.2 2. Data governance (WP1–WP3, WP6) Before data can be collected or integrated, partners agree on: • sources, access rights, and sampling frames; • metadata standards (SDMX, DDI, DataCite); • ethical safeguards and GDPR-compliant procedures; • controlled vocabularies, authority files, and persistent identifiers. These agreements are formalised in D1.2, D2.2, D3.2, and the Data Management Plan (D6.3). They ensure compliance with FAIR, OPA, and EU data-governance principles. 40
2.5.3 3. Software for data collection (WP4) WP4 develops the open-source tools used to gather and ingest administrative data, survey data, platform usage data, and CMO records. These tools form the operational backbone of the pipeline. They include: • survey-data management scripts, • connectors for royalty and licensing accounts, • streaming API integration modules, • metadata templates for ingestion and harmonisation. All tools adhere to the reproducibility and interoperability requirements defined in Annex 1 and the DMP. 2.5.4 4. Data acquisition (WP1, WP2, WP3) Data are collected from: • collective management organisations (CMOs), • ministries and statistical offices, • cultural-heritage institutions, • surveys (enterprise and personal), • streaming-service APIs. Each domain follows its own protocol (e.g. WP1 T1.2 sampling frames; WP3 T3.1 participation and wellbeing indicators). Data collected are documented in the DMP and crossreferenced with OPA-compliant folders. 2.5.5 5. Processing, enrichment, and harmonisation (WP4, WP5) Raw inputs are processed using REPREX’s R-based openmusic-pipeline: • cleaning and pseudonymisation, • metadata harmonisation, • cross-linking with authority records, 41
standards. The Data Documentation Initiatve (DDI), will ensure that we will remain compatible with official statistical microdata and metadata services and other social sciences archives, like GESIS, the official data archive of all European Commission-mandated survey research dating back over 50 years (Vardigan, Heus, and Thomas 2008). The application of SDMX ensures that our microdata and statistically processed data will be interoperable with official statistics of the UN, OECD, Eurostat, and national statistical services (Stahl and Staab 2018). Our data improvements, go beyond improvements of statistical quality and application of GSBPM; we aim to fix and improve music industry datasets for rights management or digital curation. The data enrichment and improvement are innovative solutions that are not part of the services of an open data portal or an observatory. We aim to offer these value added services to create new value and therefore motivation for music industry data owners to work with the observatory. •“Fix-the-data” means improving the data quality by finding or imputing missing values or finding and replacing erroneous data entries. In terms of metadata, adding further machine-actionable information to already existing datasets can improve their usability. •“Data linking”, data fusion, or data matching means correctly joining data from different datasets (data sources.) We ensure that data coming from sources can be meaningfully joined together; the variables have consistent meanings, the codebooks applied are harmonised; the timeframe or geographical frame is consistent. •Aggregation services: we turn your music-related datasets into statistical products or data publications. We clean, validate, and structure it to a format that it can be placed on the EU Open Data Portal, Europeana, or Wikibase for integration with Wikidata/Wikipedia. •Confidential data sharing: our data sharing space can be used for confidential data sharing and cross-pollination (for example, looking up missing ISWC/ISRC identifiers or misspelt names in each other’s datasets) without making the data public. These planned services will be discussed in Chapter 9. 3.1 Collect: Data Curation & Collection Data curation is the organisation and integration of data collected from various sources. It involves annotation, publication and presentation of the data so that the value of the data is maintained over time, and the data remains available for reuse and preservation. Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO feasibility study intuitively defines data gaps without an apparent reference to a data or conceptual model, but it recognises and stresses the need for terminological harmonisation. 48
ĹNote Four types of data-collection principles have been identified as essential both by various branches of the music sector and also by policymakers at European, national and local levels: • The data-collection service provided by a European Music Observatory should help in mapping, understanding and analysing the main characteristics, trends and idiosyncrasies of the music sector in Europe; • The data collected should be neutral and available to decision-makers, music sector operators, and the public; • The data itself should cover the activities of the music sector across the entire European Union, be comparable between Member States, and rely on identified and stable indicators; • The data collection methods should be transparent and provide a strong degree of scientific accuracy. (European Commission et al. 2020, p28) In short, we collect data about music, as defined in the cultural statistics of any European Economic Area and EU candidate statistical office or by a representative European or international music organisation. In more detail, we a systematic data collection program requires a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. ĹNote Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. We work with a concept of a musical work, not with the entire work. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. Again, in an information system we obviously work with a concept of an author, and instances of authors represented by their unique data. The EMO feasibility study catalogues 45 data gaps that a future European music observatory should fill. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. We need agreed concepts of the composer,sound recording,work, to answer such questions. 49
The initial data collection guidelines of the Open Music Observatory are derived from the EMO Feasibility study. We see them as a starting point for further discussion with the Observatory Stakeholder Network. We introduce them with our data catalogue in Section 10.1. These guidelines are supported by our first conceptualisation, which is built on some widely used conceptualisations of creative works and statistics. This is the topic of Chapter 8. 3.1.1 Microdata, Collections, Records We treat “microdata” as a collection of structured data. Aregister is a document [in modern usage, usually a database], in which data are entered in a formal manner by a statutory authority (ISO 2017b). In statistical data collection ian official list or record of names or items; it aims to be a complete list of the objects in a specific group of objects or population, for example, all copyright-protected musical works in a country, or all legal person enterprises in another country. Acollection is a group of objects, for example, musical works, sound recordings, printed scores, music enterprises, musician biographies, gathered together for some intellectual, artistic, or curatorial purpose. This is how radio playlists and charts, festival line-ups, local content guideline monitoring works; music labels and publisher select and musical works and their recordings or scores to place into commercial circulation. Such collections form the basis of census or sample surveys for statistical data collection. The documentation of collections relies on the work of registers. For example, music publishers can claim their revenues based on ISWC and ISMN identifiers provided to them by the collective management organisations that register works, or national libraries or other organisations that identify printed sheets. The maintenance of registers requires ongoing investment, and therefore registrars like the ISRC Authority or CISAC (the manager of the ISWC register) often restrict access to their data, or do not exchange data. In an increasingly globalised, automated music ecosystem where the number of identifiable works, recordings, scores, and related claims is growing exponentially, this situation puts the entire industry at a disadvantage, for example, against tech platforms. The Open Music Observatory is experimenting with innovative ways how registers can work together in some aspects of metadata standardisation, improvement and exchange in a way that keeps their core product intact and exclusive to them See: (Antal and Mester 2025). The Open Music Observatory works with metadata in a way that helps managing and improving registers, and it helps to create data about collections with authoritative data from registers. ĹNote Our first large database is the Slovak Comprehensive Music Database. Our aim is to publish a constantly refreshed database of every music composed or recorded in the territory of the current Slovak Republic, or composed and recorded by people from 50
Slovakia, or sung in the Slovak language. This database is partly based on registers, and partly on curated holdings of Slovak stakeholders. ⊠Our collections are always available on https://reprexbase.eu/skcmdb/�. Further details in Section 3.5.1. ⊠Whenever our collections fit in the collection and publication guidelines of Europeana, we make the collections available there, too. □We are investigating the possibility of synchronising our collections to the European Collaborative Cloud for Cultural Heritage. The Data Documentation Initiative is originating for the world of social sciences data archives and more and more in use in statistical organisations for the documentation of microdata. The DDI plays a particularly important role in the creation of statistical surveys, particularly using questionnaires and question banks. The new Records in Context has replaced the international standards on archives in 2023. Its central concept is the record, which is a document according to DDI; a collection is a set of records. Our standardisation of microdata is explained in more detail in Chapter 8. 3.1.2 Primary data collection The Open Music Observatory is supporting high-quality primary data collection, and itself is carrying out such collection activities. The indicators derived from the processing of survey questionnaires will be comparable if the same concepts of interest (for example, concert visiting frequencies) are measured via the same questions and answering instructions. 51
Aconcert is a standard concept of a live performance of music. How many times in the previous [12 months] have you been to a concert? is a standard question accompanied by standardised answer options and processing in the Cultural Access and Participation surveys following the ICET model. Using standardised concepts and question banks, including question and instruction labels with standardised translations, is a cornerstone of ex-ante survey harmonisation. This process is a prerequisite for retrospective survey harmonisation and the subsequent creation of comparable statistical indicators, underscoring the importance of uniformity in data collection. ⊠We provide API and download access to harmonised, multi-language question banks. This allows music stakeholders to use the same question formulations and translations for comparability with European statistical and policy research programs. ⊠We provide tutorials to retroharmonize, a background open-source software of Reprex, which is an R library to retrospectively harmonise data from different survey programs that had asked the same questions. ⊠The Open Music Europe project will carry out some harmonised surveys to show and improve the methodology of harmonised data collection within the music sector of Europe. This data will be available as metadata (questionbank), as microdata (individual answers), and as processed statistical data. 3.1.3 Metadata The most common—and perhaps least useful—definition of metadata is that it is “data about data.” As catchy as this definition is, however, it is entirely ambiguous. First of all, what is data? And second, what does “about” mean? (Pomerantz 2015a, p19) The new ISO standard on Information technology — Metadata registries (MDR) defines metadata as data that define and describe other data. As Pomerantz eloquently argues, this is a definition that is not very helpful. We use his more functional (but not contradictory) definition. “Data is only potential information, raw and unprocessed, prior to anyone actually being informed by it. […] Data must be understood not as an abstract concept but as objects that are potentially informative. […] Metadata Is a Statement about a Potentially Informative Object.” (Pomerantz 2015a, p26) Following the metadata definition of “a statement about a potentially informative object,” we believe that any high-quality data can be used as metadata in certain circumstances. ĹNote Data or metadata? The data of birth can be seen as a metadata for disambiguation among authors with the exact same name in a copyright register. It can be seen as data for a curator of a 52
young author prize, or a music sociologist. Either way, the date of birth should be precise, and encoded in a way that makes it portable and interoperable. From a data management point of view, we do not distinguish between data and metadata. Of course, we acknowledge the fact that some types of data will always remain under the hood and will only serve the proper functioning of an information system. The music industry’s famous “metadata problems” usually arise when a music enterprise or institution wants to use metadata information from an authoritative source that is somehow corrupted. The Open Music Observatory can help with these metadata problems by disseminating proper, open authoritative data (as registers or collection) or by providing data improvement services that fix the metadata problems of a user. 3.1.4 Statistical indicators and datasets ÁWarning We will place our first statistical datasets to the EU Open Data Portal this week (pending their approvals) and will provide a screenshot and access conditions here. 3.2 Repair Throughout the project we realised that data and metadata repair is perhaps a more urgent challenge then data processing. While our team and our stakeholders gradually embraced the concept of a data sharing space, i.e., the idea that instead of starting new data collections from scratch it is more economic and useful utilise existing data, given the high level of music industry digitalisation and that almost all transactions leave a digital trail behind, we also realised that the music sector is “drowning in numbers”; it handles more data in various obsolete, undocumented, unstructured or ad hoc forms than it can utilise. In these cases, usually the data is already available somewhere, but in a format that prevents the data to be used to its potential. Most of our efforts therefore were concentrated on data and metadata repair instead of new collection. Metadata repair and data processing is usually hard-to-distinguish tasks that comprise of similar or same steps. They are various validation, normalisation procedures that allow that make the data informative. 3.3 Process We use the theory of metadata by Jeffrey Pomerantz, who defines Metadata as “a statement about a potentially informative object.” A dataset without such statements is not 53
findable, accessible, interoperable, and very hard to reuse. Pomerantz distinguishes among descriptive, administrative, structural, preservation, and use metadata. The Generic Statistical Information Model (GSIM) is a common abstract representation of data objects manipulated in official statistical production and elaborated as an overarching model for implementation metadata standards such as SDMX or DDI. GSIM since its inception aims to bridge two important standards, SDMX and DDI. The Statistical Data and Metadata Exchange has been developed for decades and it is an ISO standard; it is more geared towards the aims of data sharing and preservation in RDM. DDI on the other hand is more focused on the documentation and quality control of primary data collection, or the reuse of often messy data sources, and supports the processes that make the data available for research. As DDI provides information about a much wider range of objects and processes, we are even more selective when we turn to this standard than SDMX; however, we cannot disregard DDI for microdata. 3.3.1 Processing & re-processing microdata ÁWarning We will place here an example that goes to the EU Open Data portal The EU Open Data Portal uses the following namespace definitions; these definitions refer to machine readable, explicit definitions (ontologies) of the way our datasets must be understood by a software agent. To demistify the process, we provide here an example of the metadata that we need to compile from the various steps of the data production pipeline. @prefix rdf:<http://www.w3.org/1999/02/22-rdf-syntax-ns#> . @prefix foaf:<http://xmlns.com/foaf/0.1/>. @prefix rdfs:<http://www.w3.org/2000/01/rdf-schema#> . @prefix xsd:<http://www.w3.org/2001/XMLSchema#> . @prefix owl:<http://www.w3.org/2002/07/owl#> . @prefix adms:<http://www.w3.org/ns/adms#> . @prefix dcat:<http://www.w3.org/ns/dcat#> . First we must translate the metadata of our datasets to any of the standard serialisations (file formats) of the World Wide Web Consortium’s Resource Description Framework definition, which allows the connection of data across the open internet. At the time of writing this report, the EU Open Data Portal was changing its backend, and for testing purposes, we worked with a dataset from the background of the Open Music Europe project (which had been earlier published by Reprex on Zenodo under the title *The turnover of the ration broadcasting industry in Europe*. ) The dataset itself cannot be downloaded from a data catalogue. It is an abstract intellectual work, similar to musical work or a literary work. A musical work is accessible in printed sheets or recordings, and a dataset in a distributed data file. 54
<https://doi.org/10.5281/zenodo.5652118> <a>"dcat:Dataset" ; <dcat:distribution><https://zenodo.org/records/5652118/files/codebook_trb.csv>,"https://zenodo.org/records/5652118/files/codebook_trb.csv" ; <dct:creator><https://orcid.org/0000-0001-7513-6760>; <dct:description>"\"The turnover of the ration broadcasting industry in Europe.\"@en" ; <dct:identifier><https://doi.org/10.5281/zenodo.5652118>; <dct:issued>"2022-06-03T00:00:00Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> ; <dct:modified>"2022-06-04T00:00:00Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> ; <dct:publisher><https://isni.org/isni/000000050973936X>; <dct:title>"A rádió szektor forgalma Európában\"@hu","\"Turnover of the Radio Broadcasting Industry in Europe\"@en" ; <edp:originalLanguage><rdf:resource><http://publications.europa.eu/resource/authority/language/ENG>. We can provide further provenance information about the dataset; in production, we will provide information on software agents (tools) used, researchers, data managers and curators and their organisations involved. As a bare minimum, we provide machine-readable information about the technical publisher of the dataset, Reprex B.V: <https://isni.org/isni/000000050973936X> <a>"foaf:Agent" . And then we point the user the downloadable files (distributions) of the dataset with the rights statements and licenses. We use the Creative Commons CC BY 4.0 license, similar to Eurostat on the EU Open Data Portal, and we state that the dataset is open for the public. <https://zenodo.org/records/5652118/files/codebook_trb.csv> <a>"dcat:Distribution" ; <dcat:accessURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:byteSize>"41672" ; <dcat:downloadURL><https://zenodo.org/records/5652118/files/codebook_trb.csv>; <dcat:mediaType>"text/csv" ; <dct:license><http://publications.europa.eu/resource/authority/licence/CC_BY_4_0>; <dct:rights><http://publications.europa.eu/resource/authority/access-right/PUBLIC>; <owl:sameAs><https://zenodo.org/records/5652118>. 3.3.2 Documentation ÁWarning We will provide the link and screenshot of the documentation for each file that goes public. 55
3.4 Disseminate 3.4.1 EU Open Data Portal The portal is a central point of access to European open data from international, European Union, national, regional, local and geodata portals. It consolidates the former EU Open Data Portal and the European Data Portal. The portal is intended to: 56
1. give access and foster the reuse of European open data among citizens, business and organisations. 2. promote and support the release of more and better-quality metadata and data by the EU’s institutions, businesses, agencies and other bodies, and European countries, enhancing the transparency of European administrations. 3. educate citizens and organisations about the opportunities that arise from the availability of open data. It is funded by the EU and managed operationally by the Publications Office of the European Union in cooperation with the Directorate-General for Communications Networks, Content and Technology of the European Commission, responsible for EU open data policy. We publish our data primarily on the EU open data portal for statistically processed datasets (datasets that contain the generalised characteristics of many data subjects without personal data that could identify them). 3.4.2 Europeana Integration We promised interoperability with Europeana, and after a lengthy process Reprex concluded aData Exchange Agreement with Europeana Sound (the music and sound aggregation center of Europeana hosted by the British Library’s Sound and Vision area) and sent pilot data from Latvia. 57
Wikibase is adopted because it has already proven its suitability in domains facing the same structural issues as the music ecosystem: fragmented identifiers, inconsistent authority control, multilingual metadata, cross-domain vocabularies, and parallel institutional workflows. 4.1.1 Proven in real-world scenarios highly similar to music Wikibase is widely used by libraries, archives, museums, national cultural bodies, and openscience infrastructures. These institutions face the same challenges OMO addresses: • reconciliation of people, works, events, organisations, and places; • multilingual labels and aliases; • authority file alignment (ISNI, VIAF, ORCID, GND, BNF, corporate registers); • provenance tracking and version history; • SPARQL-based validation and constraint checking. The GLAM-Wiki ecosystem, national knowledge graphs, and EU-funded linked-data projects have collectively demonstrated that Wikibase is an effective intermediary between: • authoritative PID systems; • domain ontologies (CIDOC-CRM, RiC-O, DDI, DCAT); • community-curated knowledge models. This track record gives OMO a mature, future-proof, and standards-aligned foundation. 4.1.2 Already aligned with Europe’s digital knowledge infrastructure Wikibase aligns with existing institutional practice across Europe. Many major knowledge centres and initiatives already use Wikibase/Wikidata: • national libraries and archives, • national cultural-heritage aggregators, • research infrastructures, • public-sector linked-data programmes, • the EU Knowledge Graph initiative. Adopting Wikibase ensures that the Open Music Observatory fits directly into the European interoperability ecosystem (see Section 2.3). Footnote: See Wikibase as an Infrastructure for Knowledge Graphs: the EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021) and (2020 2020). 64
4.1.3 Demonstrated support for required OMO functionality Everything OMO needs has already been demonstrated in production Wikibase environments: •authority control for creators, ensembles, organisations, venues; •multilingual and multiscript modelling for names, places, works; •cross-domain entity linking (work–recording–performance–rights–heritage); •event-based and entity-based models; •SPARQL validation, schema constraints, and automated reconciliation. This means OMO does not invent an untested paradigm: the consortium integrates proven practices from: • national registries, • performing-arts knowledge graphs, • the Slovak pilot and Finno-Ugric metadata federations developed inside the project. 4.1.4 Fits EU policy preference for open-source and trustworthy AI Wikibase is open-source, auditable, and non-proprietary. It aligns with: • the EU’s preference for open-source digital public infrastructure, • FAIR and CARE principles, • trustworthy AI requirements (provenance, transparency, versioning), • cross-border interoperability mandates, • decentralised data governance models. This makes it compatible with Europeana, EOSC/ECCCH, DCAT-AP, and the European Interoperability Framework. It is used in EU organisations, too. 1 4.1.5 The most widely used graph-editing interface in the world Tens of thousands of data stewards, librarians, researchers, and citizen-scientists already know how to edit Wikibase/Wikidata. This provides OMO with: • an immediate user base, • a ready-made contributor community, • institutional familiarity across Europe, • workflows already adopted in GLAM and research sectors. No alternative open-source system has remotely this level of adoption. 1On official adoption: EU Knowledge Graph (Diefenbach, De Wilde, and Alipio 2021); SEMIC guidelines (SEMIC Support Centre 2023). 65
4.1.6 A hybrid model that fits real institutional workflows Wikibase uniquely accommodates: • spreadsheet-based workflows (Excel, CSV), • relational database exports, • statistical microdata reference linking, • complex semantic modelling, • API-based ingestion, • R and Python pipelines. It is a practical compromise between triple stores, document databases, and relational systems — perfect for a music ecosystem where many partners still rely on basic tools. It has proven to be useful in music services2, and more generally on smalland large scale European knowledge institutions (national libraries, libraries, archives, museums3.) 4.2 How Wikibase Fits into the Open Music Europe Data-to-Policy Pipeline The data-to-policy pipeline defined in the Grant Agreement and documented in the Background chapter (see Section 2.5) provides the methodological backbone of OME. Wikibase is the component that makes this pipeline operational. 4.2.1 Wikibase supports each stage of the pipeline 4.2.1.1 Data collection • imports from Excel, CSV, SQL, APIs, and legacy systems; • immediate linkage to persistent identifiers; • entity reconciliation as part of ingestion. 4.2.1.2 Validation and reconciliation • authority-control workflows for people, works, organisations, and places; • constraint-based quality checks; • SPARQL-driven validation; • alignment with external authority files. 2On Belgian pilots: MetaBelgica (Stallmann et al. 2023) and Flemish performing arts enrichment (Magnus and Van D’huynslager 2021). 3See for example: On Wikidata/Wikibase in heritage: (Bianchini, Bargioni, and Pellizzari di San Girolamo 2021; Sardo and Bianchini 2022). 66
4.2.1.3 Harmonisation and enrichment • multilingual labels and roles; • event-based and relationship-based modelling; • addition of contextual metadata by different institutions; • integration of domain vocabularies. 4.2.1.4 Activation for analysis • SPARQL endpoints for programmatic access; • JSON-LD, RDF dumps, and REST APIs; • R and Python pipelines use stable URIs for reproducibility. 4.2.1.5 Indicator construction • cross-domain indicators linking economic, cultural, rights, and heritage data; • entity-level referencing ensures indicators are traceable and verifiable. 4.2.1.6 Interpretation and contextualisation • experts review, annotate, and correct metadata through a human-readable interface; • provenance guarantees transparency. 4.2.1.7 Policy translation and observatory outputs • live, federated knowledge base powering the OMO front end; • entity profiles, metadata dashboards, and linked methodological documentation; • direct links from indicators to source entities and datasets. 4.2.1.8 Feedback loop • corrections flow back into the shared graph; • updated authority files update all downstream indicators; • the observatory improves over time. 67
4.3 Summary Wikibase is adopted because it best fulfils the policy requirements, semantic needs, technical constraints, and governance expectations described in the Background chapter: • It is proven in domains identical to ours. • It is aligned with the EU’s open-source, dataspace, and interoperability agenda. • It is widely adopted by the very institutions we must interoperate with. • It supports the entire Open Music Europe data-to-policy pipeline. • It allows the Observatory to operate as a decentralised, federated, evolving knowledge infrastructure. This architecture is therefore not an optional design preference but the only viable model for delivering the European Music Observatory described in the Grant Agreement. 68
5 Data coordination Cultural and music data in Europe is generated and maintained by a highly diverse set of organisations. National libraries, archives, museums, collective rights management organisations, academic research projects, streaming platforms, local cultural associations, and commercial distributors all describe the same creators, the same works, and the same recordings—but they do so for different purposes, using different conceptual models, and under different legal and organisational constraints. These systems have grown organically over decades, reflecting local traditions, professional cultures, and technical possibilities. As a result, the same piece of music may be catalogued differently in an archive, credited differently in a rights database, indexed differently in a library catalogue, and published differently on a digital music service. This fragmentation is not a sign of failure; it is the natural consequence of cultural stewardship being distributed across many institutions with distinct missions. However, fragmentation becomes a problem when people attempt to connect, reuse, or reinterpret information across domains. Without coordination, data cannot be trusted across systems: names do not match; contributor roles are incompatible; identifiers are missing or ambiguous; and important aspects of minority or community-based cultural heritage may be lost or become invisible. The need for coordination arises from this structural diversity. Institutions should not be forced to adopt a single metadata standard or a single software system, but they must be able to communicate meaningfully. Coordination provides the minimum shared semantic and technical foundations that enable such communication. Coordination also supports equity and cultural diversity. Smaller institutions—community archives, ethnographic collections, minority cultural groups, and local music organisations—often lack the resources to operate modern digital infrastructures. Without coordinated frameworks, their data cannot easily participate in larger European infrastructures or data spaces, which in turn reinforces cultural imbalances. Coordination therefore serves not only technical efficiency but also cultural policy objectives related to inclusion, multilingualism, regional diversity, and the representation of endangered traditions. 5.0.1 The European Interoperability Framework (EIF) The European Interoperability Framework (EIF) provides guidance for how heterogeneous public-sector systems can work together without sacrificing their autonomy. It identifies four 69
layers of interoperability—legal, organisational, semantic, and technical—and emphasises that meaningful interoperability requires all four to be addressed1. The EIF’s underlying philosophy is that institutions will always differ in their practices, constraints, and goals. Therefore, interoperability cannot be achieved by imposing uniformity. Instead, it must enable institutions to maintain their internal systems while still participating in a shared ecosystem. Although the EIF was not developed specifically for the cultural or music domains, but generally to public digital services, its principles apply directly, because most music services are digital, and they are often public in nature, too. Figure 5.1: The role of data interoperability is not to exchange or centralise data, but to create digital services in libraries, rights management, distribution, streaming platforms that work together well. For the service providers they should work with less data processing cost, and for the user they should give a higher value or better user experience. Cultural data infrastructures face the same challenges the EIF describes: • different legislation (e.g., copyright, orphan works, deposit laws), • different organisational workflows (cataloguing, digitisation, rights clearance), 1The EIF defines layered interoperability (legal, organisational, semantic, technical) (European Commission 2017). The European Strategy for Data frames subsidiarity as compatible with federation (European Commission 2020). BDVA and the Federation Working Group emphasise that interoperability frameworks are needed to operationalise federation (BDVA/DAIRO 2023; BDVA/DAIRO Federation Working Group 2023). 70
• and different semantic traditions (archival provenance, library authority control, rights metadata, music industry credits). The EIF also emphasises reusability, transparency, and proportionality—all central issues for cultural data governance, particularly when dealing with sensitive material, intangible heritage, and community-generated knowledge. Most importantly, the EIF recognises that interoperability is not a technical property alone. It is a sociotechnical negotiation across institutions, requiring trust, shared agreements, and mechanisms for maintaining alignment over time. This recognition makes the EIF a strong conceptual foundation for coordinating cultural and music data, even when many actors fall outside the classical public-sector domain. 5.0.2 Extending the EIF to public and private service coordination Cultural data ecosystems are unique because they span both public and private infrastructures. A library’s bibliographic record for a musical work interacts with rights information managed by collective management organisations, which in turn interacts with ISWC and ISRC registries maintained by industry actors. A museum’s digitised ethnographic collection may be reused by researchers, by minority communities seeking cultural revival, or by streaming services presenting curated heritage playlists. Public institutions depend on commercial metadata pipelines for up-to-date identifiers, contributor roles, and distribution metadata; commercial actors depend on public institutions for authoritative information, contextual enrichment, and long-term preservation. To coordinate across this mixed ecosystem, the EIF must be extended beyond its original scope. Public–private coordination requires shared semantics, shared identifiers, and shared rules for linking, even when underlying workflows remain distinct. For example, DDEX roles must be interpretable by library and archival systems, just as archival provenance must be interpretable by music distributors when republishing historical recordings. Similarly, community-generated metadata (such as language revitalisation initiatives or fieldwork collections) must be able to participate in rights and distribution pipelines without being forced into commercial templates that erase cultural meaning. Extending the EIF in this way does not imply that private systems become part of public digital government infrastructures. It means that coordination mechanisms—identifier mapping, semantic mediation, minimal ontological alignment, provenance tracking, rights documentation—must operate across both sectors. This allows public and private services to interoperate responsibly while preserving their organisational autonomy and legal boundaries. The result is a more inclusive, resilient ecosystem, where cultural data can move across domains without losing meaning or trustworthiness. 71
5.0.3 Data sharing space A data sharing space is a distributed, federated infrastructure in which institutions agree on a minimal set of rules for linking and interpreting each other’s data. Unlike a centralised database, a data sharing space does not require institutions to upload all their data into a single platform or adopt a single schema. Instead, it provides a semantic mediation layer—shared identifiers, alignment patterns, equivalence mappings, and reference ontologies—that allows heterogeneous systems to interoperate. Figure 5.2: A data sharing space is an ideal coordination mechanism for EIF: it ensures that data-driven digital services of libraries, streaming platforms, distributors or collective management organisations can exchange data and data definitions efficiently. See the Annex for more information on the Slovak Music Dataspace behind https://hudobnadatabaza.sk/en/. Data sharing spaces are particularly suitable for cultural and music ecosystems, where no single actor can control the entire landscape. They allow archives to link to music distributors, libraries to link to rights registries, and community collections to link to research datasets, without forcing any of these institutions to abandon their internal models. A data sharing space can support service integration in practical, operational ways. For example, archival metadata describing a traditional song can be transformed into DDEX metadata for legal distribution; a library’s authority record can enrich a rights organisation’s contributor database; a researcher’s MIR dataset can link to digitised collections; and a community can add linguistically and culturally meaningful information without overwriting institutional records. 72
Because data sharing spaces respect institutional autonomy, they can grow incrementally. Institutions can join at different levels of readiness, linking only what they are comfortable sharing. Over time, shared identifiers and alignment patterns accumulate, improving the quality and connectivity of the ecosystem as a whole. The data sharing space becomes the place where coordinated services emerge: discovery interfaces spanning multiple institutions, rights-aware access systems, cross-domain research environments, and community-led enrichment workflows. In this way, a data sharing space serves as the structural foundation for a coordinated, equitable, and future-proof European cultural and music data infrastructure. 5.1 Ontologies and Vocabularies in the Open Music Observatory The Open Music Observatory (OMO) does not aim to create new ontologies. Instead, it integrates and reuses existing, well-established data models from the open data and cultural heritage ecosystems. In practice, this means that our data spaces and pilot projects (e.g. SKCMDb, Finno-Ugric Data Space, TextileBase) combine metadata drawn from: •Open Government and Open Science: DCAT-AP — the European data portal metadata model. •Cultural Heritage and GLAM: EDM (Europeana Data Model), CIDOC CRM, and Records in Contexts (RiC). •Emerging Cultural Infrastructure: High-Definition & Collaborative Cultural Data Infrastructure (HDTO). •Common Metadata Principles: Dublin Core Terms (DCTERMS) — for general metadata interoperability. •Music Industry Standards: ISWC for works, ISRC for recordings, DDEX for metadata exchange and release categorisation. Our goal is not to add another layer of complexity, but to connect these existing models so that cultural heritage data, rights management systems, and research metadata can interoperate seamlessly. ĹOpen Music Observatory Ontology Approach • The Observatory uses a Wikibase-based ontology layer, derived from the Wikibase Ontology. This defines classes and properties that are compatible with the Wikidata ecosystem. 73
5.2.2 Formalisation Advocates of data spaces sometimes describe them as “just connecting databases on an asneeded or as-permitted basis,” as if connection were a loose, informal network of API calls. But this is misleading. From a legal point of view, a database is either connected or not connected: if a system can dereference identifiers, look up metadata, or reuse statements, then the obligation to respect licenses, consent frameworks, provenance, and organisational constraints is triggered regardless of how lightly the connection is described. The same precision applies on the semantic level. Our data sharing space does not avoid formal modelling. We represent our conceptual model in the Wikibase ontology, and this can be—and in practice must be—compiled into RDF and OWL axioms so that machines can interpret the classes, properties, role patterns, qualifiers, and constraints. We use property semantics, subclass hierarchies, equivalence mappings, reified statements, and alignment patterns that are fully compatible with OWL 2 DL or OWL 2 RL, depending on the use case. Thus, “semantic mediation without semantic conformity” does not mean informality or the absence of a schema. It means something conceptually subtler: The data-sharing space maintains a formally specified ontology—but one that does not require all participating institutions to conform to a single, domainunifying conceptualisation. The formalism exists at the mediation layer, not as a global schema that all domains must adopt. A data space cannot function without a shared URI space, identity management, explicit class/property declarations, equivalence/near-equivalence mappings, domain/range expectations, inference patterns (even if limited), or consistent referential semantics. We just understand that we are making heavy trade offs for interoperability of systems, instead of serving one type of system’s internal workflows. 5.3 Future-Proofing Future-proofing in a Data Sharing Space means designing systems so that the knowledge we curate today will still make sense — technically, legally, and conceptually — ten, twenty, or fifty years from now. Because our audience spans librarians, archivists, music-industry professionals, rights managers, researchers, and IT staff with very different technical backgrounds, the future-proofing strategy of the Open Music Observatory must be both technically rigorous and described in familiar terms. To make this clearer, we describe future-proofing at four levels: the data model, the technology, the semantics, and the organisational workflows. 80
5.3.1 Future-proofing through graph architecture Many library and industry IT systems are still based on relational databases (MySQL, Oracle, SQL Server) or on simple tables (Excel, Access). Increasingly, people have heard the word “NoSQL,” even if they haven’t used such systems directly, and associate it with flexibility and modernity. A knowledge graph is a type of NoSQL system — but it is more than that. In a graph database: • the schema is not a separate document living in an IT department’s folder; • the schema is encoded in the data itself; • every relationship (composer of, recorded at, part of, published by) is explicitly stored alongside the entities. This means: • the structure of the data can evolve without breaking old records; • old data remains meaningful even after conceptual models change; • future systems can read the RDF graph and rebuild the entire database without needing to know the original software. Sadly, we often hear about cases when a developers passed away, or retired, and the latest database schema only existed in their heads. Graph databases connect the schema definition to every “table cell” forever. Often, our metadata repair work is re-discovering the not formalised schema of a legacy system. For underfunded library IT environments, this is crucial. A “TTL dump” or “JSON-LD export” is not just a backup — it is the guarantee of a future database, often with a path to transitioning to a cheaper open-source loan or archive management software with the explicit help of documenting the data first. In a music label, such future proofing of Excel repertoires helps to automate catalogue transfer to a more affordable or better quality distributor. 5.3.2 Stabilising definitions through internationally defined standard vocabularies When libraries, archives, or rights organisations use an SQL database, the meaning of a field often lives only inside that software. For example: • “CreatorName” inside an old PHP/MySQL-based webshop might mean a composer, or a performer, or someone who once clicked the upload button. • An “Author” field in a legacy library system may mix lyricists, arrangers, field collectors, annotators, and editors. 81
But DCTERMS, RDFS/OWL, and DDEX give clear, internationally defined meanings to these concepts. This means: • once your data is mapped to these standard vocabularies, • future systems — even ones not invented yet — will know how to interpret it. This is why DDEX is so powerful: even if a label used a 20-year-old MySQL database, once mapped to DDEX, the meaning becomes future-proof. The same applies to: • ISWC (work identifiers), • ISRC (recordings), • VIAF/ISNI (persons), • RiC-O (archives), • MARC relator codes (libraries). Even if the technology changes, the semantics remain stable. 5.3.3 Future-proofing through translatability and multiple serialisations In legacy systems, a database backup is often useless outside its native software. For example: • An Access .mdb file from 2008. • A PHP webshop dumping csv files in an ad-hoc format. • A FileMaker Pro database from a defunct project. A graph database solves this because RDF data can be exported in many serialisations: • Turtle (.ttl) • JSON-LD • RDF/XML • N-Triples • JSON Each of these files is both human-readable and machine-readable, and any future graph system can rebuild the knowledge base from them — exactly the same way software can be rebuilt from source code. If you take a look at the description of the Structural Business Statistics (OpenMuse) dataset (Q600), it can be read as om:Q600.ttl or om:Q600.json For cultural heritage preservation, this is vital: If we preserve the RDF, we preserve the meaning — not just the data. This is why OMO’s RDF exports act as both: • a functional database backup, and 82
• a preservation format for future researchers, developers, and institutions. 5.3.4 Future services through institutional interoperability Future-proofing is not only technical. It is also about making sure that libraries, archives, rights organisations, and small labels can move into the future without needing to throw away their old systems. When a library aligns its catalogue with the SKCMDB (Slovak: hudobnadatabaza.sk) federated module of the Open Music Observatory: • the library gains a modern semantic layer, • the obsolete system gains a migration path, • future library software can import the mapped RDF as its starting point. This has already happened in Hungary and Slovakia: metadata that previously lived only inside ageing local catalogues has now been “semantic-lifted” into future-proof form. The same applies to: • music labels using old PHP/MySQL webshops, • rights organisations with 1990s-era SQL systems, • archives with hand-maintained Excel inventories, • community collections with no database at all. Once the data is mapped into the data sharing space, it becomes portable. 5.3.5 A concrete example: ALOADED, Livonian folk music, and DDEX The ALOADED proof-of-concept demonstrates future-proofing in action, because it shows how archival cultural-heritage metadata, contemporary music-industry workflows, and emerging data-space standards can all meet in a single semantic pipeline. 83
Figure 5.3: We explained this in a broader context in our Open Access Music Dataspaces – Open Music Observatory on LineCheck 2025. DOI: 10.5281/zenodo.17669739 Our workflow proceeded in the following steps: • We began with archival metadata about Livonian and Latvian folk music, including the tšitšōrlinki traditional dance song Q4169, with its associated field notes, handwritten scores, and contextual ethnographic information. • This material included a corresponding archival field recording Q4166, originally catalogued in a traditional archival system that follows a very different descriptive paradigm. • We then expressed this material using DDEX concepts Q4745, creating mappings between archival descriptions and DDEX’s contribution and release structures. This transformation enabled several outcomes: • the track could appear legally on Spotify https://open.spotify.com/track/184iSrExWt2K9z7S4FvJYd • the musical score and the archival recording became findable on Garamantas.lv and Europeana • creators and contributors could be credited transparently and correctly • rights metadata could be clearly documented • and the entire pipeline resulted in a FAIR, future-proof transformation of cultural heritage into a living, usable service We describe this workflow in more detail in A Feasibility Study With Two Datasets on the Application of the Early Stage HDTO Ontology in Practice (Antal 2025c). There, we 84
explore how early HDTO (Heritage Digital Twin Ontology) structures can support the future Europeana Data Space and the European Collaborative Cloud for Cultural Heritage, including creating digital-twin representations of cultural objects that: • have a non-commercial mechanical licence for scientific MIR processing • and a commercial streaming-safe version (e.g., Spotify) without download or AItraining permissions This demonstrates how one song can travel through different rights regimes and different infrastructures — all mediated through a coherent semantic layer. This is not only a proof-of-concept for technical interoperability. It also illustrates how a shared observatory infrastructure can enable the circulation of culturally vital but economically marginal repertoires. The Livonian ethnolinguistic community is critically endangered. Making their vocal music available in modern channels is essential for revitalisation, language learning, and community visibility. Without the shared services of the Open Music Observatory, distributing such high-cultural-value but low-market-value repertoires would be practically impossible. A listener can now: 1. legally stream a piece of folk music (Spotify, Apple Music Classical, Deezer); 2. then borrow or download the score from a library or from Garamantas.lv; 3. a researcher can analyse the non-commercially licensed version using MIR tools; 4. and anyone can explore the archival recording, contextual notes, and documentation connected to it. This is future-proof music librarianship — linking listening, learning, research, and cultural memory. One of our future development tasks is to model rights, permissions, obligations, and prohibitions directly inside the Observatory knowledge graph. We outline this in Wikibase as a Data Sharing Space: Connecting Rights, Communities, and GLAM through Federated Infrastructures (Wikidata Conference 2025) (Antal 2025e). 85
6 Data Sources This chapter describes the data sources used in Open Music Europe and explains how they connect to the data-to-policy pipeline introduced in Section 2.5 and to the governance framework defined in the Data Management Plan. It is written in the same methodological structure as the Background and Architecture chapters: it explains where the data come from, how they were identified, how they will be governed, and how they will eventually be imported into the Open Music Observatory (OMO). The Data Management Plan (DMP) submitted at the beginning of the project has not been updated during the subsequent project cycles. As a result, many data sources that were collected, harmonised, or accessed during the project do not yet appear in the DMP’s Section “Data Summaries”,Section “Metadata and Standards”, or Section “Access, Retention, and Reuse”. For this reason, this chapter contains placeholders for the DMP manager indicating where updates are required before ingestion into the OMO can begin. Once the DMP’s Section Data Summaries and Section Metadata and Documentation are updated, the data listed in this chapter will be integrated into the Observatory and the ingestion workflow will resume. The vocabularies, classifications, licensing conditions, and provenance declarations listed here must be added to the DMP’s Section Standards and Vocabularies before ingestion. The purpose of this chapter is therefore twofold: 1) to describe the actual data sources used for Open Music Europe; 2) to prepare a DMP-aligned structure for importing them into the OMO once the DMP is updated. 6.1 How data sources were identified The identification of data sources followed the project’s data-to-policy pipeline described in Section 2.5. Work packages 1, 2, and 3 defined their indicator sets (D1.1, D2.1, D3.1), which in turn determined the necessary data inputs. These sources were documented in the internal methodological folders (see internal.qmd) and evaluated for: • accessibility and legal basis 86
• licensing and reuse conditions • data quality and completeness • compatibility with FAIR and GDPR principles • alignment with statistical and metadata standards (SDMX, DDI, GSIM) The resulting set of data sources falls into four functional groups: • administrative and register data • survey data • statistical and economic data • platform and streaming data These four categories reflect the cross-domain nature of the indicators: the music ecosystem cannot be understood solely through economic or cultural data, rights metadata alone, or platform data alone. The project therefore worked with a combined evidence base. When the DMP’s Data Summaries are updated, each source described here must be added as its own summary with variables, formats, provenance, and access conditions. 6.2 Administrative and register data Administrative sources form the backbone of WP1 and WP3, and they were essential for: • constructing sampling frames (WP1 T1.2) • validating enterprise and personal surveys • linking cultural, economic, and rights-related activities • generating indicators combining economic and cultural dimensions These sources include: • royalty and licensing records from CMOs • grant registers, project registers, and cultural funding databases • company registries and economic-activity registers 87
• non-profit and cultural-organisation registers • venue and festival registers • personal/professional registers where processing is permitted under GDPR Art. 6(1)(e) and Art. 89 In the Slovak pilot, these sources included historical microdata from the national KULT survey register, unpublished cultural-statistical microdata, and data held by Hudobné centrum and the Ministry of Culture. The DMP currently contains only partial references to administrative sources. To bring it into alignment, the DMP manager must update: •DMP Section: Data Summaries → add each administrative dataset with variables, formats, provenance, controller/processor roles •DMP Section: Legal and Ethical Compliance → include GDPR bases for each administrative source •DMP Section: Standards and Vocabularies → include controlled vocabularies used in registers (e.g. activity classifications, grant categories) These administrative data summaries will be added to the DMP before ingestion begins. 6.3 Survey data Survey data in the project covered: • enterprise surveys of MSMEs in the music sector • personal surveys on cultural participation, wellbeing, and music use • experimental diversity and circulation modules • harmonised components aligned with SDMX/DDI/GSIM Survey planning was delayed relative to the Grant Agreement timeline, but eventually harmonised with national cultural-statistics structures, including the Slovak KULT survey framework. For full DMP alignment, the following must be added: •DMP Section: Data Summaries → questionnaire versions, sample frames, variable lists 88
•DMP Section: Ethics and Consent → informed-consent procedures, pseudonymisation workflow •DMP Section: Standards and Vocabularies → codebooks, classifications, harmonisation mappings Survey sources listed here will only be ingested once their Data Summaries and ethical documentation are registered in the DMP. 6.4 Statistical and economic data These datasets include: • national accounts • satellite cultural and creative industry accounts • business demography • labour-force and employment data • trade and export statistics • cultural participation statistics • price indices and cost-structure data In the Slovak pilot, the availability of cultural satellite accounting provided a sophisticated basis for integrating cultural sectors into national economic indicators. These sources support valuation (WP1), diversity indicators (WP2), and societal-impact indicators (WP3). DMP alignment requires: -DMP Section: Data Summaries → add each statistical dataset with source agency, years, access rights - DMP Section: Metadata and Documentation → ensure that Eurostat and NSO metadata are included -DMP Section: Access and Reuse Conditions → specify restrictions where NSO confidentiality applies Statistical datasets will be imported once referenced formally in the DMP’s metadata and access sections. 89
• cross-linking with authority files and identifiers • metadata enrichment • structural harmonisation (SDMX, DataCite, DDI) • conversion to formats suitable for ingestion into the OMO The openmusic-pipeline (WP4) implemented these processes in R, ensuring reproducibility and alignment with FAIR and OPA principles. The updated DMP must list controlled vocabularies, classifications, and harmonisation rules used. 7.7 Software Components Developed in WP4 WP4 created a suite of open-source tools that implement the pipeline and prepare data for the OMO. These are described here briefly, with technical detail deferred to annexes. 7.7.1 Data-Ingestion Tools These tools handle imports from Excel, CSV, SQL exports, APIs, and legacy systems. They include: - connectors for CMO and/or ministry datasets - survey-import scripts - API wrappers for streaming platforms 7.7.2 Validation and Reconciliation Tools These tools perform semantic alignment and quality checks: - authority-control reconciliation (ISNI, VIAF, ORCID, corporate registries) - SPARQL-based constraint checks - duplicate detection and entity merging tools 7.7.3 Harmonisation and Metadata Tools These implement the project’s semantic rules: • SDMX structure builders • DDI variable metadata generators • DataCite dataset metadata templates • vocabulary management utilities 96
7.7.4 OMO Integration Tools These tools prepare processed data for the Observatory: • Wikibase ingestion scripts • URI stabilisation and PID mapping utilities • JSON-LD and RDF exporters • R and Python client libraries for the OMO API The DMP must reference the software components used for metadata production and processing, as required by Horizon Europe guidelines. 7.8 Position of Data Collection and Processing in the Pipeline Data collection and processing bridge the conceptual work of WP1–WP3 and the semantic and technical infrastructure described in Chapter 4. They provide: • validated raw inputs • harmonised metadata • cross-domain entity linking • analysis-ready datasets These steps ensure the construction of traceable, reproducible indicators and policy outputs. 7.9 Integration with the Open Music Observatory Once the DMP is updated and the Data Summaries approved, all data collections and processing workflows described here will be integrated into the OMO: • datasets will be linked to persistent identifiers • metadata will be converted to humanand machine-readable forms • provenance will be documented in accordance with OPA and FAIR • ingestion routines will run automatically using WP4 tools Ingestion will begin once the DMP contains the complete, updated Data Summaries and metadata structures for all datasets. 97
8 Standardisation of Data & Terminology Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO Feasibility Study intuitively defines data gaps without an apparent reference to a data or conceptual model. Because standardisation is one of the key services of the envisioned European music observatory, we gave a lot of consideration to the standards to be applied, and the terminology negotiation process among the observatory’s stakeholders. ÁNot updated This section was created at an early planning stage, and had not yet been updated. This is no longer applicable, and should not be read or quoted. In information science, a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. How can we define a data gap in such circumstances, and how can we fill it? 8.1 Business processes Since the Open Music Observatory is primarily a data dissemination hub, the definition of our services (Chapter 3) apply elements of the Generic Statistical Business Process Model (GSBPM), an international standard that describes and defines the set of business processes needed to produce official statistics. The GSBMP is accompanied by the General Statistical Information Model, which builds on the Data Documentation Initiative (DDI) and the Statistical Data and Metadata eXchange (SDMX) (Pellegrino and Grofils 2013). 98
The DDI and SDMX are the foundations of working with social sciences archives, statistical microdata, and processed statistical data. Their key elements are described in the Resource Description Framework of the World Wide Web and can be used in Linked Data. Some elements of DDI are described with RDF: The DDI-RDF Discovery Vocabulary is a draft specification of the DDI Alliance. (Hartmann et al. 2024). Whenever possible, we rely in our observatory with this annotation; if that is not yet possible, we follow the DDI Lifecycle (3.3) Documentation (Data Documentation Initiative 2020). 8.2 Conceptual and information models Data can only be understood with the broader concepts of information and knowledge, because data in itself is unprocessed, raw knowledge, that cannot be understood. The EMO feasibility Study intuitively defines data gaps without an apparent reference to a data or conceptual model. In information science, a conceptualisation is an abstract, simplified view of some selected part of the world, containing the objects, concepts, and other entities that are presumed of interest for some particular purpose and the relationships between them. Usually, when we record information about a musical work, we do not make a copy of the entire work but record some identifying properties of the work, for example, the name of its author and the name (i.e., the title), its unique ISWC identifier, and the data or registration. Composers as human beings are represented by their names, IP Names or ISNI identifiers, and date of birth and death. A data gap can only be formally defined and filled with some reference to conceptual models of the world. A typical data problem plaguing the music sector is the amount of computer and human work needed to connect musical works and their recorded fixation, and eventually, the composers, producers, and performers linked to these objects for royalty payment. How can we define a data gap in such circumstances, and how can we fill it? Numerous knowledge institutions store information about musical works, as well as natural persons (humans) who composed or performed these works and contributed to their recorded fixation. If we want to inquire about composers, we must know that a composer is always a human (animals or software agents with AI algorithms cannot be entitled to composer copyrights.) We also must know that a musical work is an abstract creation, manifesting as a notation (physical or digital sheets, MIDI files) or recording (analogue or digital-physical object, or a file.) If we want to validate the composer’s information connected to a recording of a particular musical work, we must access databases containing information about humans concerning some identifying properties of works or recordings. We imagine a future European Music Observatory that is not a specialised knowledge institution and is not a library, archive, museum, or statistical agency. Instead, it should be able to consolidate knowledge from all such institutions and find ways to bring together data from private enterprises and data collection programs to fill the information gaps of the European music sector stakeholders. 99
Our services use the Wikidata Data Model as a data coordination and reconciliation model (Wikimedia Foundation n.d.). In this regard, we follow many successful EU and memberstate, (Alexiev et al. 2020; Diefenbach, De Wilde, and Alipio 2021; Rossenova, Duchesne, and Blümel 2022; Faraj and Micsik 2023) or music projects (Siler 2022). We particularly want to mention the excellent work of the University of Helsinki in creating WB CIDOC, a simple business process and data mapping between the Wikidata Data Model and the more complex CIDOC CRM used by extensive collection management systems (Kesäniemi, Koho, and Hyvönen 2022). The StatDCAT-AP and the more general DCAT-AP definition of the EU Open Data Portal provide a bridge among library metadata systems, such as DCMI Metadata Terms (Dublin Core) for libraries, the World Wide Web DCAT standard for publishing datasets, and some core terms of the Statistical Data and Metadata eXchange. Figure 8.1: Our most important reference is the DCAT-AP 3.0 specification, and its extension to statistical data by the EU Open Data Portal. The Europeana Data Model (EDM) similarly provides a more straightforward connection tool among various library, museological or musical collections; it mainly builds on Dublin Core and offers equivalent classes for the more complex CIDOC CRM (Europeana 2017). We see no problem in connecting the EDM towards RiC. The CIDOC Conceptual Reference Model (CRM) provides an extensible ontology for concepts and information in cultural heritage and museum documentation (Bekiari et al. 2024). 100
Last, we mention some novel standards and standard candidates related to documents, microdata, and metadata documentation, such as music survey questionnaires. The Records In Context (RiC) 1.0 CRM and ontology were adopted in November 2023 to replace four international archival standards with backward compatibility. The DDI-Discovery vocabulary is an evolving standard that aims to describe important DDI terms with the World Wide Web standard Resource Description Framework. To keep our systems future-proof, we adopt elements of RiC and DDI-Discovery to document our question bank and codebooks (International Council on Archives Expert Group on Archival Description 2023; Hartmann et al. 2024). ĹNote A future European Music Observatory could help with coordinating European research activities in the music sector. An EMO could also develop tools to establish cooperation between various data collection bodies. The Observatory should, therefore, also be involved in setting standards and developing common EU wide definitions that are crucial for consistency. (European Commission et al. 2020, p80) Since the adaptation of the European Interoperability Framework and similar FAIR measures in open science, such terminological standardisation has taken place in the definition of formal ontologies, i.e., knowledge bases that software applications can use, too. The music observatory should have competent knowledge engineers and ontologists and should be involved in the discussions of sector-agnostic ontology, for example, on the possible improvements of CIDOC or EDM, for a better representation of music. There is also a need for the development of more usable and more widely accepted musicsector ontologies. In T5.1, we have reviewed the Polifonia Ontology Network and the Music Ontology, but we believe both have shortcomings for a full adaptation. 8.3 Identification & Entity Linking Entity linking, also referred to as named-entity linking (NEL), named-entity disambiguation (NED), named-entity recognition and disambiguation (NERD) or named-entity normalisation (NEN) is the task of assigning a unique identity to entities (such as famous individuals, locations, or companies) mentioned in a digital resource, such as a file. ĎTip The MusicBrainz free music database contains records of 20 artists named Paris (artists)�, and 15 locations using the same name Paris (locations)�, which all may enter a data-driven service as artists who must be credited for attribution or royalties, and as a place of an event, release, or publication. Connecting the word Paris to the correct person, group or location is the task of entity linking. 101
Since the inception of the world wide web, data flows across organisations and countries, and the use of local identifiers is not a good solution. International organisations of music, heritage management, science, and national organisations are increasingly shifting to the use of persistent identifiers (or permanent Identifier or handle). ĎTip Apersistent identifier (or permanent Identifier or handle), is one that never changes, so that your bookmarks and links don’t break when a website or a database or an API service gets updated. In 2024, there will be no European or international standard procedure for using PIDs, but several EU member states (Austria, Czechia, Germany, Netherlands) and other countries will have already adopted national PID strategies. Because Reprex is the current technical registrar of the Open Music Observatory, we losely follow the Dutch national strategy (Cruz and Tatum 2021) and the ID allocation practice of the Nationaal Archief, but this means no bias towards data partners in the Netherlands. The Dutch PID strategy does not use mandatory practices; it only recommends practices, and offers a thought-through consistent policy of using global identifiers that are not country-specific. The structure and management of global identifiers strongly correlates with the grade of achievable automation and the potential for innovation within and across different sectors of the media industries. Because of the prevailing problems of named entity linking, we are planning value added services to resolve named-entity recognition and disambiguation (NERD.) For this purpose, we are planning the use of AI (see Section 9.3). 8.3.1 Registers & Authority Files Registers record every data subject belonging to a category or class: every music publisher operating in a jurisdiction, music composer with copyright claims, or statistical dataset published. Registers are essential in identifying persons and objects (“things” in information science.) Authority files play a similar role in collections management: they provide identification information about persons or objects and tools for disambiguation. Authority files, for example, give the preferred name title for persons and musical works when available in different name or title formats, and they provide a language-independent, machine-readable identifier pointing to the correct name title. For two or more authors or performers with the same name, these identifiers help reference the proper person (or object.) Registers are valuable and indispensable for many digital workflows. They serve as the foundation of various processes, such as statistical sampling (determining who should receive a questionnaire) or copyright management (deciding who should receive the royalty payment). Their absence or inefficiency can significantly hamper these operations. 102
Unfortunately, the music industry has long missed access to reliable, open registers. The reasons for this are beyond the scope of this report, but we highlight that the underlying reasons for closed and not interoperable registers are deeply rooted in the conflicts of interests among different sub-sectors of music and are unlikely to be solved in a short time. Therefore, music enterprises, researchers, professionals, and curators will need identification services and identity brokerage services for a long time. Creating and maintaining high-quality registers require significant professional and financial commitments, and they can form a vital service of a future European Music Observatory. Currently, we are experimenting with three service levels in the Open Music Observatory. • We create our own transparent and interoperable identifiers within the OMO for persons and their groups (ensembles, bands, orchestras, associations…), legal persons (music businesses, collective rights management agencies, …), events (recording, composing, performing events, festivals, conferences, …), musical works and their manifestation (books, works, recordings, sheets.) • We create integrity brokerage services and middle-term identification via Wikibase and Wikidata. Our identifiers are connected to middle-term Wikidata and Wikibase QIDs, which also serve as graph nodes to registry, library, collections, and industry-specific identifiers. • We are piloting data improvement services that can find erroneous identifiers or add correct identifiers to various datasets. 8.3.2 Open and persistent identifiers In line with the practice of the Netherlands, we prefer the use of the following identifiers: ISNI: preferred persistent identifier for names of people and groups. The use of ISNI is also preferred by Apple Music, Spotify, and as a pilot it was introduced by Teosto, the Finnish national collective management society; it is being considered in many use cases for adoption in all CISAC societies. ISNI is the ISO certified global standard number for identifying the millions of contributors to creative works and those active in their distribution. (Camp, Lieber, and IFLA 2022) For legal persons, we are discussing the terms to use the OpenCorporates ID, because many organisations at this point do not have an ISNI. ORCiD: preferred persistent identifiers for music researchers and scholars. This is in line with the Horizon Europe and the European Open Science Cloud recommendations; ORCiD itself only adds functionality to ISNI; i.e. each ORCiD ID is at the same time registered as an ISNI. VIAF: VIAF is the shared authority file of national libraries. It offers more services than ISNI and includes an ISNI for the author. DOI: we use the Digital Object Identifier for publicly released documents. ISBN: We issue ISBN identifiers for long-form publications of our partners. (ISO 2017c) 103
8.3.3 Not open, music-industry specific identifiers Book and music sheet publishing uses the ISBN and ISWN, professional and magazines and scholarly music journals use the ISSN, and the music rights management uses ISRC and ISWC. These standards usually resolve an identifier to some network location where metadata or the object itself can be found. There are many advantages and disadvantages of this model. For example, the ISWC identification of musical works is the backbone of copyright management, and it is a closed and consistent system developed over many decades by the member organisations of CISAC. The downside of this closed system is that the metadata about the works identified by ISWC is strictly available only to CISAC member societies. While CISAC offers an API for the individual lookup of ISWC for one example of a musical work, currently it does not allow bulk access to the registered data. We have already started a discussion with some music industry registers about connecting the Open Music Observatory to their systems. We are planning to present our proposals on the CISAC Good Governance seminar to be held in December 2024. Musical works ISWC: nternational Standard Musical Work Code is a unique identifier for musical works. It is adopted as international standard ISO 15707 (ISO 2022). OpenCollectons ID: Our ID for music works (only if we publish data about them.) Sound recordings ISRC: The International Standard Recording Code (ISRC) is the international identification system for sound recordings and music video recordings. (ISO 2019b; International ISRC Registration Authority 2021) OpenCollectons ID: Our ID for sound recordings (only if we publish data about them.) Music sheets ISWN: The International Standard Music Number currently identifies published music sheets (ISO 2022). ISBN-13: Before the introduction of ISWN, published sheets were identified by the ISBN book identifier. ISBN-10: The older format of the ISBN book identifier, which predates both the ISWN and the 13-digit ISBN used to identify music sheets. ISCC: The International Standard Content Code (ISCC) is an identifier for numerous types of digital assets. This is our preferred identifier for not published sheets. (ISO 2017c) For unpublished works, our preference is the use of the brand-new ISO-standard ISCC because it was designed precisely for the use case we were looking for. It is free to generate, generated from digital content (or its digital copy), and can connect various local or lesserused identifiers. Datasets DOI: DOIs are assigned to each distribution of a dataset. As datasets are often continuously filled, these datasets will have periodic versions with versioned DOIs (from Zenodo.) OpenCollectons ID: Our ID for unversioned (continous) datasets, pointing to the latest available version of the data. 104
Codebooks URI: Whenever possible, we use standard codebooks of SDMX or Eurostat, and provide a URI to the codebook, and provide dereferencing to the codebook definition. OpenCollectons ID: Our ID for our codebooks, regardless if they are same as the SDMX/Eurostat standards, or we create a non-standard coding for a novel dataset. Questionbank URI: Whenever possible, we use standard questionnaires, and provide a URI to the codebook, and provide dereferencing to the DDI questionnaire item definitions. OpenCollectons ID: Our ID for questionbank items. 8.3.4 Lyrics In many genres, lyrics are very important parts of a musical work, and there is a growing demand and need to provide or analyse the lyrics of the work. For example, in our X, we want to create location-aware music services and encourage the public performance of music made in Bratislava or music somehow specific to Bratislava within the public places or radio stations of Bratislava. One possible semantic connection to this environment is that a song is about Bratislava (Berlin, Paris, or Germany.) Access to the lyrics part of the music is not straightforward, mainly because the lyrics may be arranged from a literary work. We see lyrics identification and semantic analysis as the next immediate step to our location-aware application, for which we are looking for good industry solutions. In many cases, we will likely need to rely on the ISCC code as a temporary identifier for lyrics databases that were not available in a licensed format earlier. 8.3.5 ISCC The Open Music Observatory will start to implement the newest ISO-standard open identifier, the ISCC-CODE. ISCC is inverting the principle of a centralised register. It generates the ISCC code from the digital content object itself, therefore no third-party lookup is needed for finding the identifier of the object. ISCC registration becomes necessary when an ISCC code needs to be globally unique, publicly discoverable, resolvable, owned or authenticated. While these features inevitably require some kind of registry, not all of them require a centralised institutional registry. The ISCC specifies the necessary protocols to implement the aforementioned features in a decentralised, federated environment and across multiple public blockchains. Given a registered ISCC code, an application can unambiguously determine on what blockchain (if any), by which account, and at what time an ISCC has been registered. Registered ISCC codes refer to an authoritative public blockchain network. This indicator is part of the ISCC Code itself, such that codes registered on different networks cannot collide. This guarantees uniqueness of ISCC codes across multiple blockchains. Ownership of ISCC codes (not the identified content) is granted to the signatory of the first transaction for a given ISCC code on the corresponding blockchain. 105
Figure 9.1: app 9.2.1 Data Health Services for Collective Management Entity linking and data linking are among the biggest technical problems in rights management. Because music authors, producers, and performers have three royalty streams and do not share an interoperable registry, the connection of musical works (compositions, ideally identified by an ISWC code), their sound recording manifestations (identified on all digital services with and ISRC code), and the various identifiers of performers require costly manual and technical identification. There are numerous projects underway in the music industry to resolve this problem going forward. In the United Kingdom, PRS’s Nexus programme� is developing a solution with the provisioning of preliminary ISWC registration to keep the recording and composition connected from the birth of a new recording. The Open Music Europe project, on the other hand, is pioneering a different route for already existing sound recordings, with the linking of public sector catalogues of heritage and library collections with rights management information; particularly with relying on the VIAF shared authority files. SOZA and Reprex are expected to present their MVP on the CISAC Good Governance seminar in December 2025. Modern registers typically assign a unique identifier, known as a URI, to their data subjects (our registered objects). A ‘Cool URI’, which resembles a URL, offers a practical advantage. When used as a URL, it generates a human-readable HTML file about the registered person or object. This can be particularly useful when processed by a graph application, as it 112
provides crucial information about this person or object in a machine-readable (XML, JSON, TTL, or NQUAD) file. For example, the VIAF identifier number 89006617 can be placed into the http://viaf.org/ viaf/89006617 URL, which provides as access to the cataloging information of works created by, or written about the great etnomusicologist and modern composer, Béla Bartók. Modern platforms, such as Spotify, use similar identifiers. For example, the Spotify Artist ID 2fIUlieTjLTaNQUIKHX5B8 resolves to Celeste Buckingham’s available recordings on the platform via the URL https://open.spotify.com/artist/2fIUlieTjLTaNQUIKHX5B8. The problem is that music creators are often present on more than 200 digital platforms, each of which has its identifier policy and requires the repeated import of the artists’, works’, and recordings’ data. To consistently report such metadata is costly and complex, even for major labels and publishers with a dedicated IT system. No wonder we saw before our project in our own Feasibility study that more than 50% of artist data needed fixing on digital platforms. Relying on many local identifiers on otherwise interconnected computer systems will always create a costly and error-prone data exchange. Unfortunately, the music industry has never agreed to use genuinely open, high-quality registers. These changes were made during the period of our project. For example, large platforms like Apple, Spotify, and some collective rights management organisations started using the ISO-standard name identifier (ISNI) to avoid the high prevalence of multiple same-name persons and musical groups. This transition is yet to begin, and it is incomplete, so the music sector will likely need to invest large IT resources into entity resolution in the next decade. 9.2.2 Sustainability Reporting for Music Organisations The Music Innovation Hub and Reprex will develop a CSRD-compliant sustainability reporting tool in 2024-2025. The reporting tool aims to provide an accurate and affordable ESG reporting facility that follows the European ESRS standards for music enterprises that create their financial reports according to the simplified reporting rules allowed by member states for microenterprises. 113
More than 95% of European music enterprises (in some member states, this reaches 100%) apply simplified financial reporting. For such companies, there are no CSRD-compliant ESG reporting tools. We identify the reason for this market failure as follows: □The CSRD Directive imposes the responsibility of connected financial sustainability reporting on large and public companies and applies it to their entire value chain. The music industry lacks such large enterprises that would have taken a piloting role or played a pivotal role in establishing the standards. □Music enterprises and their trade associations do not act proactively because they believe they must follow the data provision instructions of the directly affected B2B buyers, financiers, or corporate sponsors. □The standardisation body EFRAG has de-prioritised the cultural and creative industries in setting industry-specific standards favouring sectors with a much higher adverse environmental impact. □While small music businesses do not feel a compliance push, as they are not directly responsible for applying the ESRS, they also miss out on the opportunities provided by green financing and insurance. ⊠MiH and Reprex will pilot a service suitable for microenterprises, reducing compliance costs from 1500 euros to 500 euros per entity. ⊠This new application will rely on the Open Music Observatory’s Music Economy and Sustainability pillars and will derive its benchmarks, science-based targets and coefficients, and input-output tables. The MVP of this service was developed with a MusicAIRE microgrant, and it is the project’s background. A scale-up will be demonstrated with the use the Open Music Observatory’s open data API. 114
9.2.3 Listen Local Figure 9.2: The Feasibility Study On Promoting Slovak Music in Slovakia And Abroad is an important background of our project. In 2020, with a microgrant from the Slovak Arts Council, we created a Feasibility Study and a demo application called Listen Local (Antal 2020b). The study examined why the Spotify algorithm struggled to recommend Slovak music within Slovakia for Slovak people. We also created a demo application that modified the user’s Spotify recommendations to voluntarily comply with the local content guidelines applicable to local radio stations. The user could also listen to a lower or higher percentage of regional works. Our critical finding was the very pool data coverage and quality of the Slovak repertoire, which is mainly sent to distribution without the professional assistance of a commercial music label. Self-releasing artists and micro labels do not have the necessary metadata know-how, IT and data specialists to prepare their new releases for algorithmic curation by recommender engines of digital streaming platforms, radio stations, or large festivals. 115
Figure 9.3: Our conceptual demo application was able to make recommendations on voluntarily meeting the local content guidelines, but it was only supported by a relatively small Slovak Demo Music Database, and could only work with Spotify, which has the most transparent and open API of all streaming providers licensed to the territory of the Slovak Republic. We aim to develop applications to create a local content-aware public performance music stream. ⊠HearDis! aims to integrate such location-aware metadata into its background music playlisting service. ⊠We are planning Listen Local applications for radio stations to voluntarily review their current playlists for compliance with local content regulations and, if they fail to reach the statutory local content quotas, to recommend suitable recordings to their playlists. ⊠The OMO will disseminate the necessary data for these new services. □The data is not yet available, as the creation of the Slovak Comprehensive Music Database is a task of its own that will be ready by November 2025 in WP2. 9.2.4 Unlabel Unlabel is a planned service aimed at self-releasing artists and micro labels that need a functional data/IT department. Therefore, they are at a disadvantage compared to significant independent and major releases because they usually need to meet the high documentation standards necessary for a successful digital distribution strategy and engagement with algorithmic curation of streaming-, radio-, or festival playlists. Self-releasing artists and micro labels bring ill-documented new content to digital distributors like ALOADED. Digital distributors must maintain an arm’s length standard for all 116
labels, small or large, independent or major. ALOADED or other distributors cannot crossfinance the data problems of self-releasing and micro-label artists from the client revenues of more prominent labels. We identify the problem as a market failure and a technical failure: □In some developing markets, insufficient royalty revenues do not allow the presence or professionalisation of record labels with an IT and data management function because the payment of IT or data specialists or to keep external suppliers at least on a retainer cannot be financed from the label-artist revenue split. □Manual metadata provision without metadata specialists and tools leads to inferior data quality. Our feasibility study has shown that more than 50% of the releases have data shortcomings, and 17% have poor data representations that make algorithmic recommendations for these releases impossible. This creates a vicious circle because poor data quality translates into low visibility, low usage of such repertoire, and, therefore, low income. The cost of data improvement has no sustainable financial basis. ⊠In 2025, ALOADED, Reprex, Slovak Music Center and SOZA will conceptualise and plan a new public-private business model that aims at those rightsholders who do not have a technically proper label representation as a substitute for non-available market services. Our planned “Unlabel” service will provide documentation and metadata improvement services for self-releasing artists. This service, similar to current white-label services, will strictly address market failures and not compete with label services. We aim to provide a necessary level of data consolidation and improvement so that these artists can have equal opportunities in digital distribution services. The service will be connected to the Slovak Music Dataspace and its Slovak Comprehensive Music Database. We will provide a PPP business model for the onboarding and proper documentation of self-releasing artists on a large scale and the efficient, API-based provision of their digital distributor. Aloaded will provide the distribution services, Reprex will provide the data services, and SOZA and the Slovak Music Center will work out the details of minimal customer service for such labels. 9.3 Use of AI systems For the entity linking, related to our planned value added Section 9.2.1, we are planning to use in the future AI algorithms, particularly inference engines. The main goal of the system is to help matching correctly named entities, particularly rightsholders, musical works and recordings. The system is not yet in place. An adequate description will be provided for overview and will be brought to the attention of the Ethics Advisor during the upcoming meeting of the Ethics Board. 117
We cannot provide a full risk assessment because the service is not planned in detail yet. However, our preliminary risk assessment suggests low levels of risk, partly, because we plan to deploy AI in music/culture, which as a domain not seen as a high-risk area by the European regulation, and partly, because our system will not autonomous, will retain human-in-control, and will not influence the decisions or anyhow engage with end-users. We were conscious of the potential risk involved, and both the control structure and the data governance were planned over the course of 10 months. □Is the AI system designed to interact, guide or take decisions by human end-users that affect humans or society? No. The system will only help qualified persons in rights management to faster and more efficiently preview potentially unlinked entities. □Could the AI system affect human autonomy by interfering with the end-user’s decision-making process in any other unintended and undesirable way? No. The system in no way is considered as an end-user system. ⊠Please determine whether the AI system (choose as many as appropriate) overseen by a human: Is overseen by a Human-in-Command. ⊠Have the humans (human-in-the-loop, human-on-the-loop, human-in-command) been given specific training on how to exercise oversight? Yes. The system is not making autonomous decisions. ⊠Is your AI system being trained, or was it developed, by using or processing personal data (including special categories of personal data)? Yes. ⊠Did you put in place any of the following measures some of which are mandatory under the General Data Protection Regulation (GDPR), or a non-European equivalent? Yes. ⊠Data Protection Impact Assessment (DPIA) Yes. ⊠Designate a Data Protection Officer (DPO)24 and include them at an early state in the development, procurement or use phase of the AI system? Yes. ⊠Oversight mechanisms for data processing (including limiting access to qualified personnel, mechanisms for logging data access and making modifications)? Yes. □Measures to achieve privacy-by-design and default (e.g. encryption, pseudonymisation, aggregation, anonymisation)? Not applicable for NERD. The aim of the application is to detect errors in name attribution and to protect the moral and economic rights of the (named) rightsholders. ⊠Did you implement the right to withdraw consent, the right to object and the right to be forgotten into the development of the AI system? Yes. ⊠Did you consider the privacy and data protection implications of data collected, generated or processed over the course of the AI system’s life cycle? Yes. ⊠Did you consider the privacy and data protection implications of the AI system’s non-personal training-data or other processed non-personal data? Yes. 118
We do not consider that the system has wider risks or negative impacts. The algorithm is designed to cure sources of data biases that result in a late or missed payment for some rightsholders. 119
10 Data Catalogue The Open Music Observatory curates, maintains, and disseminates a data catalogue with the resources within the data catalogue: individual datasets and their series and API endpoints where the data can be queried in a custom format. In creating our data infrastructure, we considered the specifications of our dissemination nodes, which provide our data with a wide range of interoperability and easy access: the EU Open Data Portal, Europeana, Wikibase Cloud and Wikidata. From a thematic point of view, we relied on the definition of the EMO feasibility study, and created topical pillars (Section 10.2). The data curators of the Observatory Stakeholder Network (see Annex) and for the duration of the Open Music Europe project, the work packages (WP1-4 represent each “pillar”) can define and provide datasets or data series according to their topical collection guidelines (Section 10.1). Not updated ÁWarning This section was created at an early planning stage, and had not yet been updated. Unfortunately, due to the problems of the WP6 Data Management Plan task we cannot yet show how we will fill up the pillars of the observatory. A data catalogue formally is a metadata dataset: a dataset on information about our available datasets and their downloadable or queriable distributions. It follows the global World Wide Web DCAT standard. DCAT is an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. This document defines the schema and provides examples for its use (Albertoni et al. 2020). It is a global standard, which was further extended and specified for the release of statistical datasets (StatDCAT-AP) and for the needs of the EU Open Data Portal (DCAT-AP). (Sofou and Dragan 2019; Fragkou 2023) These extensions provide further metadata and organisations standards, but essentially they do not change the definition of the global standards. A data catalogue (dcat:Catalog) represents a catalogue, which is itself a dataset in which each individual item is a metadata record describing some resource: a description of a dataset, a data service, or other type of resource. dcat:Dataset represents a collection of data, published or curated by a single agent or identifiable community. We currently support two types of datasets: statistical datasets that conform to the datacube definition of SDMX, or collection datasets for microdata, which contain non-aggregated, structured data representing some unity criteria, for example, music works and recordings that have been present in the official radio charts of a given country. 120
The _dataset_, similar to a musical or literary work, is an abstract concept which can be used, downloaded, and stored in its manifestation. For a musical work, a manifestation may be a sound recording or music sheet; for a dataset, it is a distribution. A URI identifies a dataset; the URI does not allow the downloading of the dataset, because it refers to the abstract idea of the dataset; the URL for downloading the dataset belongs to the individual distributions. dcat:Distribution represents an accessible form of a dataset, such as a downloadable file. When the same dataset is distributed in different file formats (for example, CSV and SPSS files), each distribution is listed in the catalogue separately with a separate download link. Each distribution has its own URL where the dataset can be downloaded. Figure 10.1: The music.dataobservatory.eu/tag/music-economy/ URL lists the downloadable datasets on the Open Music Observatory website. They can be found on EU Open Data Portal, too. In the first days after launching our new service, around 1-10 June 2024, the datasets may be missing from the EU Open Data Portal, which is changing in these days its complete backend, and may have some backlog in accepting our datasets. dcat:DataService represents a collection of operations accessible through an interface (API) that provides access to one or more datasets or data processing functions. Our datasets are accessible on different platforms with their own datasets, and our internal data-sharing space also has its API. As data is added to the different platforms (EU Open Data Portal for statistical and microdata datasets, Europeana for collections dataset, Wikibase Cloud for further microdata, metadata and collections, and Reprexbase for confidential microdata and collections), we are updating the catalogue with the DataService entries. dcat:DatasetSeries is a dataset that represents a collection of datasets that are published separately but share some characteristics that group them; for example, a (play)list of sound recordings that were present in the weekly charts or the annual budget of an institution. A time series dataset is usually not defined as a data series, but the new time observations are added to an updated distribution of the time series dataset. Stakeholders who provide data to the Open Music Observatory can commit to making a data series; however, we only define a data series when we have at least two items available from the series. dcat:CatalogRecord represents a metadata record in the catalogue, primarily concerning the registration information, such as who added the record and when. 121
10.2.4 Innovation The definition of the Innovation pillar in the EMO feasibility study is more a topic to be covered than a data need description. This pillar is less data-driven in that it will rely mostly on research conducted on topics relating to changes in the market place, new business models, disruptive technologies, etc. A European Music Observatory will have the latitude to pick certain topics based on priorities and input from sectoral stakeholders. An EMO should consider setting up an “innovation experts’ advisory committee,” constituted of respected professionals in their field who are known for their forward thinking views, to help identify key themes to be studied. (European Commission et al. 2020, p37) We will initiate an informal music innovation expert’s roundtable to discuss potential data needs in this pillar. 10.2.5 Sustainability In the EMO feasibility study the definition of sustainability was mentioned among the innovation topics. Because of the triple transition, introducing the Corporate Social Responsibility Directive and the European Sustainability Reporting Standards have increased the interest and need in sustainability data; we decided to create a separate topical pillar for environmental and social sustainability, or governance indicators (ESG.) We will publish datasets that will be used in the value-added service described in Section 9.2.2. 128
References 2020, SEMIC. 2020. “Providing Sustainable Data Services Through Wikibase and Wikidata.” https://joinup.ec.europa.eu/sites/default/files/custom-page/attachment/202011/Parallel-track-4_B-Fischer_J-Thill_A-Angjeli%20final%20ppt.pdf. Albertoni, Riccardo, David Browning, Simon Cox, Alejandra Gonzalez Beltran, Andrea Perego, and Peter Winstanley, eds. 2020. “Data Catalog Vocabulary (DCAT) - Version 2.” W3C. https://www.w3.org/TR/2020/REC-vocab-dcat-2-20200204/. Alexiev, Vladimir, Plamen Tarkalanov, Nikola Georgiev, and Lilia Pavlova. 2020. “Bulgarian Icons in Wikidata and EDM.” Digital Presentation and Preservation of Cultural and Scientific Heritage 10: 45–63. https://doi.org/10.55630/dipp.2020.10.2. Anders, Wallgren, and Wallgren Britt. 2007. Register-Based Statistics. Administrative Data for Statistical Purposes. 1st ed. Chichester:United Kingdom: John Wiley & Sons Ltd. Antal, Daniel. 2020a. “Central And Eastern European Music Industry Report 2020.” CEEMID, Consolidated Independent. https://doi.org/10.13140/RG.2.2.21450.31686. ———. 2020b. “Feasibility Study on Promoting Slovak Music in Slovakia & Abroad.” https://doi.org/10.5281/zenodo.6427514. ———. 2023. “Pilot Program for Novel Music Industry Statistical Indicators in the Slovak Republic.” Zenodo. https://doi.org/10.5281/zenodo.8399254. ———. 2024a. “A szlovák adatkicserélési tér magyarországi föderációjának lehetőségei.” In Az oktatás, a kutatás és a közgyűjtemények digitális transzformációja felsőfokon : NETWORKSHOP 2024 : 33. Országos Informatikai Konferencia : 2024. április 3–5. Eszterházy Károly Katolikus Egyetem, Eger, edited by József Tick, Károly Kokas, and András Holl, 192–98. Budapest: HUNGARNET Egyesület. https://doi.org/10.31915/ NWS.2024.25. ———. 2024b. “Building a Music Data Sharing Space with Wikibase.” Open Music Observatory. https://doi.org/10.5281/zenodo.17078911. ———. 2024c. “Trustworthy AI and Data-Sharing Spaces for the Slovak Music Centre. Poster Presentation at the IAMIC Conference 2024 on November 21, 2024, at Music Austria, Vienna.” Open Music Observatory. https://doi.org/10.5281/zenodo.16540605. ———. 2025a. “SKCMDb. Interoperability of Music Libraries and Archives with Public and Private Music Services. Presentation at the IAML 2025 Conference in Salzburg, Austria Held on the 7th of July 2025.” Open Music Observatory. https://doi.org/10. 5281/zenodo.16634558. ———. 2025b. “Slovak Music Data Sharing Space. Poster Presentation at the IAML 2025 Conference in Salzburg, Austria Held on the 7th of July 2025.” Open Music Observatory. https://doi.org/10.5281/zenodo.15814286. ———. 2025c. “A Feasibility Study with Two Datasets on the Application of the Early Stage HDTO Ontology in Practice.” https://doi.org/10.5281/zenodo.17492178. 129
———. 2025d. “A Green Paper on AI, Data Governance, and Metadata Policies for Europe’s Music Ecosystem.” Open Music Observatory. https://doi.org/10.5281/zenodo. 17244314. ———. 2025e. “Wikibase as a Data Sharing Space: Connecting Rights, Communities, and GLAM Through Federated Infrastructures. Presentation on Wikidata Conf 2025.” Open Music Observatory. https://doi.org/10.5281/zenodo.17496740. ———. 2025f. “Open Music Ontology.” Open Music Observatory. https://doi.org/10.5281/ zenodo.17541357. Antal, Daniel, Mester Anna Márta, Ieva Pigozne, and Mihaly Nagy. 2025. “Dataset of the Multilingual Gazetteer of the Settlements on the Livonian Coast of Northern Kurzeme.” Finno-Ugric Data Sharing Space. https://doi.org/10.5281/zenodo.17545241. Antal, Daniel, Mária Kmety Barteková, and Katarína Remeňová. 2023. “Economy of music in Europe: Novel data collection methods and indicators.” Zenodo. https://doi.org/10. 5281/zenodo.8334648. Antal, Daniel, and Anna Márta Mester. 2025. Open Music Registers.https://doi.org/10. 5281/zenodo.14767717. Antal, Daniel, Ieva Pigozne, and Anna Márta Mester. 2025. “Remapping the Livonian Coast: A Multilingual Gazetteer of the Settlements of Northern Kurzeme.” Finno-Ugric Data Sharing Space. https://doi.org/10.5281/zenodo.15668712. Artisjus, HDS, SOZA, and Candole Partners. 2014. “Measuring and Reporting Regional Economic Value Added, National Income and Employment by the Music Industry in a Creative Industries Perspective. Memorandum of Understanding to Create a Regional Music Database to Support Professional National Reporting, Economic Valuation and a Regional Music Study.” BDVA/DAIRO. 2023. “Data Sharing Spaces and Interoperability: BDVA Discussion Paper.” BDVA/DAIRO. https://www.bdva.eu. BDVA/DAIRO Federation Working Group. 2023. “Federated Data Spaces: Position Paper.” BDVA/DAIRO. https://www.bdva.eu. Bekiari, Chryssoula, George Bruseke, Erin Canning, Martin Doerr, Philippe Michon, Christian-Emil Ore, Stephen Stead, and Velios Athanasios, eds. 2024. “Definition of the CIDOC Conceptual Reference Model.” CIDOC CRM Special Interest Group. https://www.cidoc-crm.org/sites/default/files/cidoc_crm_version_7.2.4.pdf. Bianchini, Carlo, Stefano Bargioni, and Camillo Carlo Pellizzari di San Girolamo. 2021. “Beyond VIAF Wikidata as a Complementary Tool for Authority Control in Libraries.” Information Technology and Libraries 40 (2). https://doi.org/10.6017/ital.v40i2.12959. Camp, Ann Van, Sven Lieber, and IFLA. 2022. “ISNI, a Top Tool for Quality Enhancement, Smooth Data Flows and Efficient Internal Processes.” Dublin: Ireland: International Federation of Library Associations; Institutions (IFLA). https://repository.ifla. org/handle/123456789/2008. Cruz, Maria, and Clifford Tatum. 2021. “NWO Persistent Identifier Strategy.” Zenodo. https://doi.org/10.5281/zenodo.4674513. Curry, Edward. 2020. “Dataspaces: Fundamentals, Principles, and Techniques.” In RealTime Linked Dataspaces: Enabling Data Ecosystems for Intelligent Systems, 45–62. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-296650_3. Data Documentation Initiative. 2020. “DDI Lifecycle (3.3) Documentation.” https://ddi130
lifecycle-documentation.readthedocs.io/en/latest/index.html. Diefenbach, Dennis, Max De Wilde, and Samantha Alipio. 2021. “Wikibase as an Infrastructure for Knowledge Graphs: The EU Knowledge Graph.” In The Semantic Web – ISWC 2021, 12922:626–42. Lecture Notes in Computer Science. Cham: Springer. https://doi.org/10.1007/978-3-030-88361-4_37. Ernštreits, Valts. 2020. “Livonian Place Names: Documentation, Problems, and Opportunities.” Eesti Ja Soome-Ugri Keeleteaduse Ajakiri. Journal of Estonian and Finno-Ugric Linguistics 11 (1): 213–33. https://doi.org/10.12697/jeful.2020.11.1.09. European Commission. 2017. “Commission Implementing Decision (EU) 2017/1348 of 25 July 2017 on the Interoperability Framework for European Public Services (European Interoperability Framework).” Official Journal of the European Union.https://eurlex.europa.eu/eli/dec_impl/2017/1348/oj. ———. 2020. “A European Strategy for Data.” European Commission. https://eurlex.europa.eu/legal-content/EN/TXT/?uri=celex:52020DC0066. ———. 2021a. “Brochure for Music Moves Europe Preparatory Action 2019.” European Commission. https://ec.europa.eu/culture/sites/default/files/library/mme_2019_ brochure_final-web.pdf. ———. 2021b. “Music Moves Europe - First Dialogue Meeting. Final Report.” European Commission. https://ec.europa.eu/culture/sites/default/files/library/mme-conferencereport-web.pdf. European Commission, Directorate-General for Education, Youth, Sport and Culture, M Clarke, P Vroonhof, J Snijders, A Le Gall, B Jacquemet, et al. 2020. Feasibility Study for the Establishment of a European Music Observatory : Final Report. Publications Office of the European Union. https://doi.org/10.2766/9691. European Parliament. 2024. “European Parliament Resolution of 17 January 2024 on Cultural Diversity and the Conditions for Authors in the European Music Streaming Market (2023/2054(INI)).” European Parliament. https://www.europarl.europa.eu/doceo/ document/TA-9-2024-0020_EN.pdf. Europeana. 2017. “Definition of the Europeana Data Model V5.2.8.” Europeana. https://pro.europeana.eu/files/Europeana_Professional/Share_your_data/Technical_ requirements/EDM_Documentation//EDM_Definition_v5.2.8_102017.pdf. Faraj, Ghazal, and András Micsik. 2023. “Enriching Wikidata with Cultural Heritage Data from the COURAGE Project.” In, 407–18. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-36599-8_37. Fragkou, Pavlina. 2023. “DCAT-AP 3.0.” Edited by Makx Dekkers, Pavlina Fragkou, Natasa Sofou, and Bert Van Nuffelen. https://semiceu.github.io/DCAT-AP/releases/3. 0.0/. Hartmann, Thomas, Sarven Capadisli, Franck Cotton, Richard Cyganiak, Arofan Gregory, Benedikt Kämpgen, Olof Olsson, Heiko Paulheim, Joachim Wackerow, and Benjamin Zapilko. 2024. “DDI-RDF Discovery Vocabulary. A Vocabulary for Publishing Metadata about Data Sets (Research and Survey Data) into the Web of Linked Data.” Edited by Thomas Hartmann, Richard Cyganiak, Joachim Wackerow, and Benjamin Zapilko. W3C. https://rdf-vocabulary.ddialliance.org/discovery.html. International Council on Archives Expert Group on Archival Description. 2023. “Records in Contexts–Conceptual Model. Version 1.0.” International Council on Archives. https: //www.ica.org/app/uploads/2023/12/RiC-CM-1.0.pdf. 131
International ISRC Registration Authority. 2021. “International Standard Recording Code (ISRC) Handbook. 4th Edition.” International ISRC Registration Authority. https: //www.ifpi.org/wp-content/uploads/2021/02/ISRC_Handbook.pdf. ISO. 2012. “International Standard Musical Work Code (ISNI). ISO 27729:2012.” International Organization for Standardization. https://www.iso.org/standard/44292.html. ———. 2013. “ISO 17369:2013(en) Statistical Data and Metadata Exchange (SDMX).” London:United Kingdom: International Organization for Standardization. https://www. iso.org/obp/ui/en/#iso:std:iso:17369:ed-1:v1:en. ———. 2017a. “ISO/IEC 19941:2017(en), Information Technology — Cloud Computing — Interoperability and Portability.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso-iec:19941:ed-1:v1:en. ———. 2017b. “ISO/IEC 5127:2017(en), Information and Documentation — Foundation and Vocabulary.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:iso:5127:ed-2:v1:en. ———. 2017c. “ISO 2108:2017 (En), Information and Documentation — International Standard Book Number (ISBN).” International Organization for Standardization. https: //www.iso.org/standard/65483.html. ———. 2019a. “ISO/IEC 20546:2019 Information Technology — Big Data — Overview and Vocabulary.” London:United Kingdom: International Standards Organisation. https: //www.iso.org/obp/ui/en/#iso:std:iso-iec:20546:ed-1:v1:en. ———. 2019b. “International Standard Recording Code (ISRC). ISO 3901:2019.” International Organization for Standardization. https://www.iso.org/standard/64817.html. ———. 2020. “ISO/IEC 22624:2020(en), Information Technology — Cloud Computing — Taxonomy Based Data Handling for Cloud Services.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:isoiec:22624:ed-1:v1:en. ———. 2022. “International Standard Musical Work Code (ISWC). ISO 15707:2022.” International Organization for Standardization. https://www.iso.org/standard/83125. html. ———. 2023a. “ISO/IEC 11179-1:2023(en), Information Technology — Metadata Registries (MDR) — Part 1: Framework.” London:United Kingdom: International Organization for Standardization. https://www.iso.org/obp/ui/en/#iso:std:iso-iec:11179:-1: ed-4:v1:en. ———. 2023b. “ISO/IEC 2382:2015(en), Information Technology — Vocabulary.” London:United Kingdom: International Standards Organisation. https://www.iso.org/obp/ ui/en/#iso:std:iso-iec:2382:ed-1:v2:en. Kesäniemi, Joonas, Mikko Koho, and Eero Hyvönen. 2022. “Using Wikibase for Managing Cultural Heritage Linked Open Data Based on CIDOC CRM.” In New Trends in Database and Information Systems, edited by Silvia Chiusano, Tania Cerquitelli, Robert Wrembel, Kjetil Nørvåg, Barbara Catania, Genoveva Vargas-Solar, and Ester Zumpano, 542–49. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-03115743-1_49. Magnus, Bart, and Olivier Van D’huynslager. 2021. “Podiumkunstendata Op Wikidata: De Stap Naar Echte Linked Open Data.” Kunstenpunt / Flanders Arts Institute. February 11, 2021. https://www.kunsten.be/nu-in-de-kunsten/podiumkunstendata-op-wikidatade-stap-naar-echte-linked-open-data/. 132
Mikš, Tomáš. 2025. “OpenMusE: Towards a Sustainable Licensing Market for AI Use of Protected Works.” Vilnius, Lithuania: Slovak Performing; Mechanical Rights Society (SOZA); OpenMusE Consortium; Presentation at the CISAC European Committee Meeting. https://www.openmuse.eu/wp-content/uploads/2025/05/20250427_CISAC_ EC_2025_OpenMusE_final.pdf. Music Moves Europe. 2024. “Music Ecosystem 2025: Study on the Music Ecosystem.” Publications Office of the European Union. Luxembourg: European Commission, Directorate-General for Education, Youth, Sport; Culture. https: //doi.org/10.2766/95340. Nagel, Lars, and Douwe Lycklama, eds. 2021. “Design Principles for Data Spaces. Position Paper. Version 1.0.” Open DEI. https://doi.org/10.5281/zenodo.5244997. Open Music Europe Consortium. 2025. “Policy Brief: An Open, Scalable Data-to-Policy Pipeline for European Music Ecosystems.” EU Horizon Europe Deliverable D5.7. Open Music Europe Consortium. https://openmuse.eu/. Pellegrino, Marco, and Denis Grofils. 2013. “DDI-SDMX Integration and Implementation. Working Paper.” United Nations Economic Commission for Europe. https://unece.org/ fileadmin/DAM/stats/documents/ece/ces/ge.40/2013/WP5.pdf. Pomerantz, Jeffrey. 2015a. “Definitions.” In Metadata, 19–64. Cambridge, MA, USA: The MIT Press. http://www.jstor.org.proxy.uba.uva.nl/stable/j.ctt1pv8904.6. ———. 2015b. Metadata. The MIT Press Essential Knowledge Series. Cambridge, MA, USA: MIT Press. Rossenova, Lozana, Paul Duchesne, and Ina Blümel. 2022. “Wikidata and Wikibase as Complementary Research Data Management Services for Cultural Heritage Data.” In CEUR Workshop Proceedings.https://serwiss.bib.hs-hannover.de/frontdoor/deliver/ index/docId/2573/file/rossenova_etal2022-wikidata_research_data_mgmt.pdf. Sardo, Lucia, and Carlo Bianchini. 2022. “Wikidata: A New Perspective Towards Universal Bibliographic Control.” JLIS.it : Italian Journal of Library and Information Science 13 (1): 291–311. https://doi.org/10.4403/jlis.it-12725. SEMIC Support Centre. 2023. “Wikidata and Wikibase — SEMIC Support Centre.” https://interoperable-europe.ec.europa.eu/collection/semic-support-centre/wikidataand-wikibase. Senftleben, Martin, Thomas Margoni, Joost Poort, Kacper Szkalej, and Etienne Valk. 2024. “Policy Brief 1: Music Metadata Mainstreaming and EU Law.” EU Horizon Europe Deliverable D5.6. OpenMusE Consortium. https://www.openmuse.eu/. Siler, M. 2022. “Beyond the Fountain: Mapping a New Entry Point to the Society of Independent Artists.” Art Documentation 41 (2): 219–41. https://doi.org/10.1086/ 722172. Sofou, Natasa, and Adina Dragan. 2019. “StatDCAT-AP – DCAT Application Profile for Description of Statistical Datasets. Version 1.0.1.” European Commission. https://joinup.ec.europa.eu/collection/semantic-interoperability-communitysemic/solution/statdcat-application-profile-data-portals-europe/release/101. Stahl, Reinhold, and Patricia Staab. 2018. Measuring the Data Universe: Data Integration Using Statistical Data and Metadata Exchange. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-319-76989-9. Stallmann, Claudia, Koen Deneckere, Ruben Verborgh, et al. 2023. “MetaBelgica Project: A Linked Data Infrastructure Between Federal Scientific Institutes in Belgium.” In Pro133
ceedings of the 19th Extended Semantic Web Conference (ESWC 2023). Cham: Springer International Publishing. https://doi.org/10.1007/978-3-031-33455-9_24. UNECE. 2014. “Generic Statistical Information Model. GSIM V2.0 Documents. UNECE Statswiki.” 2014. https://statswiki.unece.org/display/gsim/GSIM+v2.0+documents. ———. 2019. “Generic Statistical Business Process Model. GSBPM V5.2 Documents.” UNECE Statswiki. January 2019. https://statswiki.unece.org/display/GSBPM/GSBPM+ v5.1. Vardigan, Mary, Pascal Heus, and Wendy Thomas. 2008. “Data Documentation Initiative: Toward a Standard for the Social Sciences.” International Journal of Digital Curation 3 (1): 107–13. Wickett, Karen M., Antoine Isaac, Katrina S. Fenlon, Martin Doerr, Carlo Meghini, Carole L. Palmer, and Jacob Jett. 2013. “Modeling Cultural Collections for Digital Aggregation and Exchange Environments.” CIRSS Technical Report 201310-1, October. https://hdl. handle.net/2142/45860. Wikimedia Foundation. n.d. “Wikibase Data Model.” Wikimedia Foundation. Accessed May 26, 2024. https://www.mediawiki.org/wiki/Wikibase/DataModel. 134
Annex 1 - Stakeholder profile data sheet for the Observatory Stakeholder Network Name of the stakeholder: □Corporate (institutional) name: This is mandatory for legal persons and groups; also, please provide a contact person. □given name: mandatory for natural persons □family name: mandatory for natural persons Legal status of the stakeholder: ⊠private person: your name will be made public among the members; we also ask to provide a persistent ID (ISNI, ORCiD or VIAF.) ⊠Legal person: your name and legal person ID will be made public among the members, but not the contact person. Please provide ISNI, or OpenCorporates ID. □other association or group without legal personality: your name and persistent ID will be made public among the members, but not that of the contact person. Please provide ISNI identifier. Logo or icon of the stakeholder: □not mandatory, but if provided, we will make it public Online contact details of the stakeholder: □Official website: we will make it public if provided □LinkedIn page: we will make it public if provided □Facebook page: we will make it public if provided □YouTube channel: we will make it public if provided □Instagram account: we will make it public if provided Permanent residence or seat of the stakeholder: ⊠Only the country and municipality will be made public □Please provide full postal address Official email address of the stakeholder: 135
□only used for invitations to stakeholder meetings, giving and revoking data handling; never made public. Contact telephone number of the stakeholder: □is only used to clarify potential problems and consents; it is never made public and is not used unless necessary. For the intake, we will also ask for a few lines of statement of interest in the Observatory Stakeholder Network and topic interests for data if they apply. We will make this information public, too. We are also happy to create an intake interview and publish it on our website to allow the members of the Observatory Stakeholder Network to get familiar with each others ideas and interests. ĺImportant We will provide in the final deliverable a link for the official intake to the stakeholder network. I would like to invite SOZA, Hudobné centrum, Aloaded, and HearDis! as first partners. Open Music Data Exchange ĺImportant This advisory body will not be open for invitations. The representative pan-European stakeholders can nominate here their technical providers to consult the technical aspects of data exchange. 136
SKCMDb: Slovak Comprehensive Music Database The Slovak Comprehensive Music Database (SKCMDb) is a national initiative aimed at making Slovak music more accessible, discoverable, and usable across libraries, archives, streaming services, and rights management organisations. It connects scores, recordings, and metadata using open standards and collaborative governance. As a functional module of the Open Music Observatory, the SKCMDb also serves as a testbed for developing shared data services, addressing the conceptual models, workflows, and governance rules required to link diverse music stakeholders. ‘ The SKCMDb is supported by a data-sharing space consisting of both shared and private databases. The data-sharing space currently comprises the following initial components: •Slovak Metadata Database: A database that facilitates connections between various Slovak stakeholders’ systems. •SKCMDb (public): A public database containing microdata on musical works, their recordings and scores, biographical and institutional information, statistical datasets, and a catalogue of publications. 137
Livonian Music Database (private) The private components of the LīvMDb are a staging area for data that has unclear provenance or legal status. Unlike in the case of LīvMDb, we hold minimal business confidential data (related to the royalty accounts of music that we published), but some data may have, for example, unclear GDPR status. Microdata Microdata consists of information before it is aggregated into statistical datasets or formal publications. Metadata can also be considered microdata: while it is never aggregated, it plays a critical role in describing the provenance, semantics, and usability of aggregated data. •Collections: Structured sets of similar items created through curatorial activities, where inclusion is based on discretionary selection to serve end users (e.g., a library’s holdings or a curated playlist). Our work is centerred around the work of Hõimulõimed, a Finno-Ugric NGO, which curated many Finno-Ugric language collections, including the collection of Livonian-language songs available on Spotify. •Registers: Authoritative lists created through administrative processes with defined rules, aiming to capture all known items in a category. Our work buids on the register of Livonian placenames, because they offer the most straighforward curatorial help to find new Livonian (folk) music1. •Metadata: Relevant elements from the Livonian Metadata Database that support the use of collections or registers. The LīvMDb’s collections and register datasets are organized as a document database. This database stores structured data in RDF format describing musical works, sound recordings, printed and manuscript scores, as well as biographical information about music professionals and their organizations. Each music-related object or agent (person, corporate body, or organization) is represented as a microdata dataset. These datasets share common definitions via conceptual models and data structures, enabling automatic aggregation. All datasets are available with RDF annotation and can be exported in all standard RDF serializations. Microdata is intended for institutional and professional use, not for the general public. It is annotated with standardized metadata suitable for applications such as music library cataloguing, distribution platforms, and rights management systems. 1Livonian place names: documentation, problems, and opportunities (Ernštreits 2020) and our gazetteer dataset: (Antal, Pigozne, and Mester 2025). 144
ĹNote Example The village of Mazirbe (Livonian: Irē, German: Klein-Irben, Russian: �������) is the central location of the Livonian culture, which hosts the Livonian Community House. Folk songs collected and recordings made in Mazirbe may have a provenance of Irē, Mazirbe, Klein-Irben (or its Finnish and Estonian versions), or various Cyrillic transliterations. Mainitaing a clear metadata dataset on the geographical sources of Livonian music is necessary. •Graphical view: Navigate and contribute to detailed entries on musical works, sound recordings, and related assets. � Explore on Wikibase •Semantic view: Export structured data in XML, JSON-LD, Turtle, or NTriples formats for reuse in research or digital projects. � Example Turtle file or download XML We made available the dataset in standard RDFXML, JSON-LD, and TTL serialisations accompanies by a data paper explaining its use (Antal et al. 2025; Antal, Pigozne, and Mester 2025) Statistical Data & Data Catalogue Our statistical data and catalogue consist of datasets aggregated using statistical methodologies. These comply with the SDMX standard and the W3C Data Cube vocabulary, making them compatible with spreadsheet software, statistical packages, and data science workflows in R, Python, or similar environments. Datasets are offered in multiple formats. In addition to RDF serializations, we provide standard CSV files and, when required, Excel or SPSS formats. We also publish data papers and related documentation that describe dataset usability and highlight key insights. Publications & Catalogue The LīvMDb’s most important publications are sound recordings, which are made available as a archival recordings (not ready to be communicated to the public on commercial platforms) and publicly available sound recordings. We also provide access to some printed and and hand-written scores. Microdata and statistical datasets are treated as publications and are listed both in the general catalogue and a machine-readable data catalogue. The LīvMDb also includes methodological and musicological publications, as well as data papers explaining the use of datasets. The catalogue is designed for interoperability with libraries, archives, museums, and similar institutions. 145