scieee AI-readable full text Open interactive document viewer

Making Corpus Creation in Early Music Rewarding and Effective: Finding the Optimum Between Standardisation and Autonomy

Wiering, Frans; Bergwall, Erik; van Berchum, Marnix; Goebl, Werner; Van Kranenburg, Peter; Lewis, David; Plaksin, Anna; Rodríguez-García, Esperanza; Smith, David J.; Visscher, Mirjam; Weigl, David M.

Abstract

Several early music projects, such as the Stanford Josquin Project, have demonstrated the potential for attaining valuable new musicological insights using a corpus-based approach. However, the available musical corpora tend to be relatively small and exhibit considerable variation in encoding practices. Aspiring corpus researchers are confronted with a lack of suitable data, which needs to be addressed before they can embark on their proper research. The EarlyMuse Short Term Scientific Mission CORSICA has surveyed the current state of corpus creation and digital editing in early music. Based on this information, it has developed a vision for the future of corpus building in this field, which aims to speed up the production of digital encodings while respecting the autonomy of the encoders and acknowledging their efforts. This is important because much high-quality encoding is carried out outside the field of professional musicology, and engaging citizen scientists could help address the current shortage of research data. The CORSICA team‘s vision is informed not only by a study of the available data, standards and technologies, but also by Human-Computer Interaction, placing human goals and values before the creation of technology and work processes. The core of the vision is that successful corpus creation must be an inclusive endeavour in terms of both technology and human participation. The report concludes with an implementation plan outlining the initial steps required to realise the vision.

Full text

A New Ecosystem of Early Music Studies 1 COST ACTION 21161 MAKING CORPUS CREATION IN EARLY MUSIC REWARDING AND EFFECTIVE FINDING THE OPTIMUM BETWEEN STANDARDISATION AND AUTONOMY EDITED BY Frans Wiering (Utrecht University, Netherlands) WITH THE CONTRIBUTIONS OF Erik Bergwall (Uppsala University, Sweden), Marnix van Berchum (Huygens Institute, Netherlands), Werner Goebl (University of Music and Performing Arts Vienna, Austria), Peter van Kranenburg (Utrecht University, Netherlands), David Lewis (Goldsmiths, University of London, UK), Anna Plaksin (Paderborn University, Germany), Esperanza Rodríguez-García (Universidad Complutense de Madrid, Spain), David J. Smith (Northumbria University, UK), Mirjam Visscher (Utrecht University, Netherlands), and David Weigl (University of Music and Performing Arts Vienna, Austria) HOW TO CITE THIS TEXT Frans Wiering (ed.), Making Corpus Creation in Early Music Rewarding and Effective: Finding the Optimum Between Standardisation and Autonomy (Utrecht University, 2025), 73 pp., https://doi.org/10.5281/zenodo.17543932. 2 COST ACTION 21161 Abstract Several early music projects, such as the Stanford Josquin Project, have demonstrated the potential for attaining valuable new musicological insights using a corpus-based approach. However, the available musical corpora tend to be relatively small and exhibit considerable variation in encoding practices. Aspiring corpus researchers are confronted with a lack of suitable data, which needs to be addressed before they can embark on their proper research. The EarlyMuse Short Term Scientific Mission CORSICA has surveyed the current state of corpus creation and digital editing in early music. Based on this information, it has developed a vision for the future of corpus building in this field, which aims to speed up the production of digital encodings while respecting the autonomy of the encoders and acknowledging their efforts. This is important because much highquality encoding is carried out outside the field of professional musicology, and engaging citizen scientists could help address the current shortage of research data. The CORSICA team‘s vision is informed not only by a study of the available data, standards and technologies, but also by Human-Computer Interaction, placing human goals and values before the creation of technology and work processes. The core of the vision is that successful corpus creation must be an inclusive endeavour in terms of both technology and human participation. The report concludes with an implementation plan outlining the initial steps required to realise the vision. 3 COST ACTION 21161 Table of contents 1 INTRODUCTION ........................................................................................................................... 4 2 PROBLEM DESCRIPTION ............................................................................................................... 5 3 CORPORA AND TECHNOLOGIES .................................................................................................... 7 3.1 AN OVERVIEW OF EARLY MUSIC CORPORA AND ENCODING SYSTEMS ........................................................... 7 3.2 OPTICAL MUSIC RECOGNITION: HOPES AND REALITIES ........................................................................... 10 3.3 METADATA AND INTEROPERABILITY ................................................................................................... 14 3.4 PROVENANCE ................................................................................................................................ 17 4 LOOKING AT CORPORA FROM A SCHOLARLY VIEWPOINT .......................................................... 21 4.1 SYSTEMATIC CORPORA CREATION: WHERE DO WE COME FROM AND WHERE ARE WE NOW? .......................... 21 4.2 AIMS AND CHALLENGES OF AN ANALYST .............................................................................................. 25 4.3 CRITICAL EDITING OF MUSIC AND THE MONUMENTAL/COLLECTED EDITION IN THE DIGITAL AGE ..................... 28 4.4 HOW TO SELECT A REPRESENTATIVE CORPUS FOR HISTORICAL STUDY ........................................................ 36 5 BEYOND THE SCHOLARLY COMMUNITY ..................................................................................... 40 5.1 CPDL CONTRIBUTORS: A CASE STUDY ................................................................................................. 40 5.2 PARTICIPATION AND COLLABORATION IN CORPUS CREATION ................................................................... 42 6 EARLY MUSIC ENCODING: A PACT ANALYSIS .............................................................................. 45 6.1 BRIEF INTRODUCTION TO PACT ........................................................................................................ 45 6.2 PEOPLE ........................................................................................................................................ 46 6.3 ACTIVITIES .................................................................................................................................... 47 6.4 CONTEXTS .................................................................................................................................... 48 6.5 TECHNOLOGIES .............................................................................................................................. 49 6.6 CONCLUSION ................................................................................................................................ 50 7 NARROWING THE GULF BETWEEN PEOPLE AND SYSTEMS IN MUSIC ENCODING ........................ 51 7.1 INTERVIEWS .................................................................................................................................. 51 7.2 PERSONAS .................................................................................................................................... 52 7.3 ROLES .......................................................................................................................................... 53 8 VISION OF THE FUTURE OF CORPUS BUILDING ........................................................................... 56 9 IMPLEMENTATION OF THE VISION ............................................................................................. 59 REFERENCES ....................................................................................................................................... 61 APPENDICES ....................................................................................................................................... 65 APPENDIX 1: CORSICA PROPOSAL ................................................................................................................. 65 APPENDIX 2: EARLY MUSIC CORPORA .............................................................................................................. 69 4 COST ACTION 21161 1 Introduction This report presents a vision for the future of corpus creation and digital editing within the early music domain. It was developed in the EarlyMuse Short-Term Scientific Mission CORSICA (an acronym for Creation Of eaRly muSIc CorporA). 1 In a nutshell, the problem addressed by this report is that there isn’t enough data available for systematic computational analysis of early music. Since the number of potential participants in encoding projects is limited, we need to devise strategies that make data creation more effective, while at the same time the participants keep their sense of autonomy and feel respected for their contribution. Chapter 2 explains this problem further. The following chapters summarise the current state of early music corpus creation and digital editing from three perspectives: technology (Chapter 3), music research (Chapter 4) and community (Chapter 5). Chapter 6 organises the main insights from the state of the art by means of a PACT analysis, a framework for exploring a design situation from a human-centred perspective. Our human-centred approach is further elaborated in Chapter 7 by looking into the motivations, skills and work practices of professional and amateur corpus creators. This section also proposes a high-level model for collaborative projects. Based on the collected insights, the CORSICA team created a vision for the future of corpus creation (Chapter 8). As a starting point for the implementation of this vision, we present several recommendations to the EarlyMuse community as well as the wider community of researchers, practitioners and enthusiasts of early music (Chapter 9). Throughout the report we highlight two areas in particular: • various dimensions that describe the variety that can be observed in the early music encoding domain; • human aspects of corpus creation such as uses, requirements, motivations, expertise and communities. These relate to our view that diversity of approaches, motivations and technologies are an integral aspect of the music encoding domain and that it is more productive and rewarding to embrace these rather than strive for unification. This report was written by the STSM participants and researchers from the Music Information Processing group of Utrecht University. Frans Wiering acted as the general editor of the document. Authors are of the chapters and sections are named below the title. Unnamed text was written by Frans Wiering. The content of chapters 6 onwards represents the ideas and inputs of the entire group, not just the person who wrote the text. 1 The proposal for this STSM is reproduced in Appendix 1. There is also a blog post about the STSM meeting in Utrecht in May 2024: https://earlymuse.eu/publications/blog/creation-ofearly-music-corpora-corsica-activities-and-a-workshop-about-building-collections-of-digitaleditions/ (accessed 31 October 2025). 5 COST ACTION 21161 2 Problem description In a nutshell, the main problem CORSICA addresses is: we would like to conduct corpus research on early music but there aren’t enough suitable encodings. This description needs some unpacking, though. First, by early music we understand music composed before c. 1700 and secondly, we focus on polyphonic music. Despite the great variety of musics that fall within this description we believe that these share enough commonalities with respect to editing and encoding to separate them from monophonic and later music to be treated together. The term ‘corpus’ is often used for large, bounded collections of texts or compositions, as in Corpus mensurabilis musicae, a series of scholarly editions covering a large cross-section of late-medieval and Renaissance music. 2 A more precise description stems from linguistics: corpora are balanced, often stratified collections of authentic, ‘real world’, text of speech or writing that aim to represent a given linguistic variety. Today, corpora are generally machine-readable data collections. 3 Here we already encounter a first dimension in the encoding domain, between creation of digital editions on the one hand, and corpus creation for analytical purposes on the other. What connects the two is the availability of a machinereadable, manipulable encoding of the music that in the former scenario is rendered (usually) as a PDF, and in the latter is fed into analytical software such as music21. 4 For corpus research the availability of encodings is obviously a crucial matter. Whereas textual corpora have been created at a large scale and are relatively homogeneous, musical corpora are comparatively small and display a considerable variety in encoding practices. They were often created during research projects with a specific focus, which then impacted encoding decisions. For pre-1700 polyphonic music, the situation is especially complicated: few corpora contain more than several hundreds of items; different types of music notation such as mensural notation and tablature pose specific encoding problems; and transcription to modern notation can be supported in multiple ways. Some of the consequences of this situation are: • available corpora are often not a representative selection from the known repertoire, making bias in the corpus almost inevitable (see Section 4.4); 2 See https://www.corpusmusicae.com/cmm.htm (accessed 1 August 2025). 3 ‘Corpus Linguistics,’ Wikimedia Foundation, https://en.wikipedia.org/wiki/Corpus_linguistics (accessed 23 October 2025). The definition from this article is based on C.F. Meyer, English Corpus Linguistics: An Introduction, 2nd ed. (Cambridge University Press, 2023), Section 1.1, ‘Defining a Corpus.’ 4 Music21 is ‘a Python-based toolkit for computer-aided musicology’, see https://www.music21.org/music21docs/ (accessed 1 August 2025). 6 COST ACTION 21161 • quality and interoperability issues increase when using multiple corpora (see Section 3.3); • choosing and mastering technologies involves a substantial intellectual challenge, at the expense of scarce research resources; • off-the-beaten-track research requires an excessive amount of encoding effort, while studying that which has already been studied by others is easy; • PhD students and early-career researchers do not have sufficient resources to embark on careers as digital musicologists. On the other hand, there are some important positive developments: • the emergence of MEI (Music Encoding Initiative 5 ) as a versatile musicological encoding system with a growing user community (see Sections 3.1 and 4.1); • improvements in Optical Music Recognition for early music (see Section 3.2); • the general practice of musicologists to use music notation software to create their transcriptions; • the widespread practice of online sharing of editions created by both professional and amateur musicians, and by citizen scientists possessing significant musicological expertise (see Sections 5.1 and 5.2); • digital turn in musicology, accelerated by pandemic. 6 The solution for the lack of encodings has often been sought in standardisation and training. These may indeed boost productivity, especially in the context of funded, large-scale projects such as Polish Musical Heritage. 7 But elsewhere such endeavours have met with limited success for at least four interconnected reasons: complexity and limitations of tools; required time investment for learning; loss of already created work; and most importantly loss of autonomy. Since small-scale, autonomous projects are the norm for both music researchers and citizen scientists, it seems logical to take their autonomy, motivation and variety of practices as a starting point for designing a new ecology for early music corpus creation. 8 This means first understanding why and how encoders (which to large extent means creators of digital editions) do their work. Next comes the question if, how, and under what conditions encoders would be willing to share their work, and how their contributions should be recognised. The last question is how their work could be effectively coordinated and supported with suitable tooling and processes. This document will present the first iteration in the elaboration of these ideas. 5 https://music-encoding.org/ (accessed 1 August 2025). 6 F. Wiering and C. Inskip, ‘The Impact of the Pandemic on Musicologists’ Use of Technology,’ Digital Humanities Quarterly 19, no. 2 (2025), https://dhq.digitalhumanities.org/vol/19/2/000786/000786.html (accessed 23 October 2025). 7 https://polish.musicsources.pl/en (accessed 1 August 2025). 8 It is standard practice in Human Computer Interaction to investigate human aspects of system design first and start designing processes and technology only after those are sufficiently understood. See e.g. D. Benyon, Designing User Experience (Pearson UK, 2019). 7 COST ACTION 21161 3 Corpora and technologies The aim of this chapter is to give an overview of the current state of corpus creation from a primarily technological viewpoint: the availability of data and metadata, the software and the encoding formats used for data production. Section 3.1 presents an overview of the available early music encodings and their formats. There is considerable variation in technical approaches, making conversion between formats an issue. Section 3.2 describes Optical Music Recognition, a collection of technologies that potentially could reduce the need for tedious data entry work. Encodings are useless when we do not know what precisely they represent: metadata (Section 3.3) ideally records such information extensively, in a way that is shared between different resources. Section 3.4 describes how provenance information could be recorded by means of Data-Envelopes. 3.1 An overview of early music corpora and encoding systems Music encoding and corpus building have been part of computational musicology since the 1960s and 1970s. The legendary Princeton Josquin Project aimed to encode the composer’s complete output in IML (Intermediary Music Language). 9 Numerous other projects produced larger or smaller collections using a variety of encoding systems. 10 Unfortunately, through a near-perfect storm involving paradigm changes in both musicology and computing as well as personal circumstances and career choices of key researchers, only minimal traces of the encoding efforts from these early projects have survived. The main exception are the monophonic musical incipits of the RISM Catalog: 11 the work of encoding these in PAEC (Plaine and Easie Code) started in the 1960s and continues today. The next stage in music encoding was shaped by the personal computer revolution and the emerging Internet. The MuseData and Humdrum corpora that go back to the 1980s are still available for research and KernScores (the current name of the main Humdrum corpus) is still growing. This is also when music enthusiasts started to create MIDI collections that later developed into resources like the Classical MIDI Files. Today, encodings of early music can be found on numerous websites. Those that are known to the STSM participants are listed in Appendix 2, with a brief description of their content, estimated number of items, encoding system(s), urls, and some other 9 A. Mendel, ‘Some Preliminary Attempts at Computer-Assisted Style Analysis in Music,’ Computers and the Humanities 4, no. 2 (1969): 41–52, https://www.jstor.org/stable/30199321. 10 E. Selfridge-Field, ed., Beyond MIDI: The Handbook of Musical Codes (The MIT Press, 1998). 11 See https://rism.online/ and https://opac.rism.info/ (accessed 24 October 2025). 8 COST ACTION 21161 name description documentation ABC Encoding system designed to notate music in plain text format, often used for encoding tunes but also for encoding complete scores https://abcnotation.com/wiki/abc:standard CMME XML-based representation for mensural notation, supporting transcription to CMN https://cmme.org/data/music/cmme.xsd (XML Schema) Humdrum Family of music representations in tabular format, such as **kern for CMN, **fret for tablatures and **mens for mensural notation https://www.humdrum.org/rep/index.html; https://doc.verovio.humdrum.org/humdrum/mens/ for **mens IML Encoding language for representing CMN on punch cards, developed for the Princeton Josquin Project Robison (1967) Lilypond LaTeX-like typesetting language for music notation: early music notations and tablature are also supported https://lilypond.org/doc/Documentation/notation/ MEI XML-based representation for encoding music notation documents, customisable for a variety of notation types https://music-encoding.org/guidelines/ MIDI Binary format for storing music information as messages for an electronic instrument. Despite its limitations, it is also used for data exchange between music printing programs and as storage format https://midi.org/midi-1-0-detailed-specification (version 1.0, 1996); https://midimusic.github.io/tech/midispec.html (version 1.1, 1999) MuseData Text-based encoding system for the logical content of scores in CMN https://www.ccarh.org/publications/books/beyondmidi/online/musedata/ MusicXML XML format for sharing sheet music between applications, with support for mensural notation and tablature https://www.w3.org/2021/06/musicxml40/ (version 4.0, 2021) PAEC System for representing music notation with typewriter symbols, mostly used for encoding incipits https://www.iaml.info/plaine-easie-code Tab Encoding system for typesetting lute music https://www.cs.dartmouth.edu/~wbc/lute/AboutTab.html Tabcode ASCII format for French and Italian lute tablature https://igor.gold.ac.uk/isms/ecolm/?page=TabCode Table 1: Music encoding formats mentioned in this report. 9 COST ACTION 21161 information. A comprehensive overview of music corpora research is lacking and very likely some useful collections of early music encodings have escaped our notice. Most scholarly corpora, such as the Stanford-based Josquin Research Project (JRP), the Electronic Medieval Music Score Archive Project (EMMSAP) and Electronic Linked Annotated Unified Tablature Edition (E-LAUTE) have a rather specific focus and contain less than a hundred to a few thousand encodings. The largest corpora stem from individual or collective projects by citizen scientists who edit vocal and lute music. Helping the community by sharing music editions seems to be an important motivation for contributing. Reusable encodings are generally a byproduct of their effort and may not be made available to others. There is a wide range of editorial approaches, and the quality of the work is rather uneven. Nevertheless, the best and most active contributors to Choral Public Domain Library (CPDL) and similar online collections have produced substantial numbers of editions and encodings that offer an interesting potential for musicological research. For the scholarly perspective on several of these corpora see Section 4.1). Numerous encoding formats are employed in the corpora (see Table 1 for a survey). MuseScore, one of the most popular software tools for music editing, supports the open encoding format MusicXML. Other formats are regularly used as well, especially in music research, such as MEI (Music Encoding Initiative; with extensive support for critical editions and active development of markup for non-standard notation forms), Humdrum, Lilypond, CMME (Computerized Mensural Music Editing; for mensural notation) and various systems for encoding lute tablature. Many projects support MIDI as a format for playback and data exchange. As a storage format it has significant shortcomings, for example the inability to encode lyrics. The basic task of entering music into the computer can be time-consuming, errorprone and (depending on one’s motivation) tedious, even when encoding only a limited set of properties. Good tooling (Graphical User Interface design in particular) can obviously make a difference. Optical Music Recognition (OMR) may offer a solution in some cases (see Section 3.2 for an in-depth discussion), but even the best OMR output currently needs substantial error correction and post-processing. Different communities of encoders can be identified, each with its own mix of aims, tools, encoding format and publication practices (e.g. MEI, CPDL, Lilypond, Tab). While it is unlikely (for both intellectual and social reasons) that these communities will eventually merge, it makes sense to study how and to what extent one community’s output can be another community’s input. Many resources offer encodings in multiple formats that are automatically derived from the storage format. Tools for conversion of encodings include: • LuteConv, 12 a command-line tool for conversion between lute tablature encoding formats, also available as a Web interface and API for ease of use; 13 12 https://bitbucket.org/bayleaf/luteconv/ (accessed 24 October 2025). 13 https://luteconv.mdw.ac.at/ (accessed 24 October 2025). 16 COST ACTION 21161 and this composition is not one of his more famous works, we assume that the results of our case study are indicative for the situation concerning metadata about Renaissance music. We queried a number of data sources to find other sources and editions for this composition. The results can be summarized as follows. • Répertoire International des Sources Musicales (RISM) ○ The composition is in the RISM, but not this particular source (Codex Lerma) • Digital Image Archive of Medieval Music (DIAMM) ○ The Codex Lerma is in the DIAMM database ○ The list of works in Codex Lerma is incomplete ○ There is a dead link to the scan at Utrecht University Library • Wikidata 45 ○ The composer is present ○ The composition is not present • Muziekweb 46 ○ The composer is present ○ Three instances of a recording are present • MusicBrainz 47 ○ The composer is present ○ The list of works by Lassus is incomplete • Choral Public Domain Library (CPDL) 48 ○ Five editions of this composition are present ○ Only one includes a reference to a source • International Music Score Library Project (IMSLP) / Petrucci Music Library 49 ○ The composer is present ○ The list of works is incomplete ○ There are 3 editions and 4 arrangements of this composition. ○ There is a dead link to CPDL caused by a spelling variation (‘amerai’ instead of ‘aimerai’). During this exercise, one particular disappointment was that there does not seem to exist a machine-readable data source that gives all compositions for such a prominent composer as Lassus, although there exists a Lasso Verzeichnis (LV) in which our composition is number 39 (as reported in IMSLP). 50 45 https://www.wikidata.org/ (accessed 31 October 2025). 46 https://www.muziekweb.nl/ (accessed 31 October 2025). 47 https://musicbrainz.org/ (accessed 31 October 2025). 48 https://www.cpdl.org/ (accessed 31 October 2025). 49 https://imslp.org/ (accessed 31 October 2025). 50 H. Leuchtmann, Orlando di Lasso. Seine Werke in zeitgenössischen Drucken 1555–1687 (Bärenreiter, 2001), 3 vols. 17 COST ACTION 21161 These results show that even for a prominent composer such as Lassus it is not possible to automatically retrieve the requested overviews from the various data sources, despite the promises of Linked Open Data. The level of integration at machine-readable level is disappointing. 3.4 Provenance Marnix van Berchum Provenance information provides the ‘history’ of the content one is using at the moment, tracing all intermediate steps of its state. Basically, it models the changes of the state of the content, answering the questions on the agents involved in the changes: Who?, How?, What?, When?, Why?, and Where?. Provenance information gives the context for (re)use and enables the reproduction of research results. Two types of provenance can be distinguished: provenance of the content and technical provenance. Figure 2 shows, as example, a simplified view of where provenance information ‘lives’ in the context of early music editions. At the level of a corpus provenance information should include information on how the corpus came into being and which selections were made by whom (for more activities which can be captured as provenance, see Section 6.3). Figure 2. Sources of provenance information. Three stages of the edition are discerned: the Source, a Transcription, and the (final) Edition. At each stage relevant elements should be captured and documented: 1. Source: Which source or sources is/are used (siglum, description, identifier, etc.)? Is the physical item or a (digital) reproduction used? If a reproduction: how is the source reproduced, with what technology, settings, etc.? Information on the source may extend to detailed analysis of the codicological and paleographical status of a source. 2. Transcription: What is the format of the transcription (paper, MEI, **kern, etc.)? Who made the transcription, etc. (i.e. metadata of the transcription)? 3. Edition: What is the format of the transcription (paper, MEI, **kern, etc.)? Who made the edition, etc. (i.e. metadata of the edition)? 18 COST ACTION 21161 Other provenance information maps the (technical) processes, when moving from one stage to the other. It includes answers to the Who?, How?, What?, When?, Why?, and Where? questions: 1. How is the transcription made: manually or (semi)automatically with OMR? Which software is used, with which version? Which choices are made in the transcription (e.g. note values, error correction, etc.)? Why and on what basis are transcription conventions used? 2. What conversions are made between the transcription and the edition (e.g. from MEI to PDF)? Which software is used, with which version? Are more editorial choices made, and if so, by whom, on what basis etc.? Figure 3. Modules available in the Huygens Data-Envelopes, taken from Luthra and Eskevich. 51 Examples of implementation of provenance information, at the larger scale of a corpus, can be found at the Huygens Institute for History and Culture of the Netherlands. 52 In case of provenance of content, addressing ‘the particular needs and challenges of the cultural heritage field’ with regard to the heterogeneous 51 M. Luthra and M. Eskevich, ‘Data-Envelopes for Cultural Heritage: Going beyond Datasheets,’ in Proceedings of the Workshop on Legal and Ethical Issues in Human Language Technologies @ LREC-COLING 2024, eds. I. Siegert and K. Choukri (ELRA and ICCL, 2024): 52–65, https://aclanthology.org/2024.legal-1.9 (accessed 25 October 2025). 52 https://www.huygens.knaw.nl/en/ (accessed 25 October 2025). 19 COST ACTION 21161 nature (complex, diverse, uncertain, incomplete etc.) of humanities data, the data management department of this research institute developed the so-called ‘DataEnvelopes’. 53 This ‘documentation framework’ captures all intricacies of the data ‘while combining machine-readability and user-friendliness’. Furthermore, the dataenvelopes take into account the legal and ethical sensitivities of the data. The combination of information is aimed at the potential (re)use of the data, by both human and machine agents. It is ‘structured into modular sections, each designed to encapsulate different facets of the dataset in a systematic manner’, see Figure 3. Figure 4. Flow diagram of provenance trails generated in the REPUBLIC project. For the capture of technical provenance information an implementation was devised in the REPUBLIC project. 54 The goal of the project was to ‘digitally unlock all resolutions (decisions) that the States General [of the Netherlands] took during their existence as an autonomous political entity (1576–1796).’ These resolutions can be found in a scanned version of the originals in the National Archives and are made available in an online search interface (relating to the scope of this paper, it can be considered as a fully searchable digital corpus of resolutions). 55 All data processing steps were tracked, making ‘it possible to trace down to the smallest detail the decisions we have made in the entire process from scanning the resolution books to 53 See the summarising article on this work: M. Luthra and M. Eskevich, ‘Data-Envelopes for Cultural Heritage.’ All following quotes taken from this article. 54 See https://goetgevonden.nl/ (accessed 25 October 2025). 55 The online interface is available at https://app.goetgevonden.nl (accessed 25 October 2025). 20 COST ACTION 21161 displaying them in the web interface’. 56 Figures 4 and 5 show a schematic (‘flow’) and detailed view (‘data trail’) of the provenance trails generated in the project. In the latter view the different questions (Who?, How?, What? etc.) can be clearly seen, including links to external resources, for the individual involved (ORCID) and the script used. Figure 5. Detailed view of a provenance trail generated in the REPUBLIC project. Developed by Team Structured Data / Digital Infrastructure department of the KNAW Humanities Cluster, https://di.huc.knaw.nl/home-en.html. 56 See https://goetgevonden.nl/en/the-method/provenance/ (accessed 25 October 2025). The software developed by the shared Digital Infrastructure department / Team Structured Data is available at https://github.com/knaw-huc/provenance (accessed 25 October 2025). 21 COST ACTION 21161 4 Looking at corpora from a scholarly viewpoint This chapter describes the dealings of scholars with corpora. Section 4.1 on systematic corpus creation resumes the treatment of corpora from Section 3.1 through an in-depth discussion of some important scholarly corpora of early music from the perspective of the researcher who wishes to use them. Section 4.2 continues this perspective with an analysis of analytical tasks and the challenges that an analyst might face in the execution of those tasks. A different but related perspective emerges in Section 4.3, namely that of critical editing in the digital age. Finally, Section 4.4 presents the perspective of the researcher who, in the absence of a complete encoding of everything, still wants to assemble a representative selection of composition for their research. 4.1 Systematic corpora creation: where do we come from and where are we now? Esperanza Rodríguez-García The impulse to assemble comprehensive music corpora is deeply embedded in musicology. From its institutional beginnings, the field has prioritised the creation of cohesive collections of music. Notably, the emergence of musicology as a scientific discipline in the mid-nineteenth century was closely linked to the publication of monumental editions, such as the Bach-Gesellschaft Ausgabe (from 1851), the Händel-Gesellschaft Edition (from 1858), and the Opera omnia Ioannis Petraloysii Praenestini (from 1862), among the earliest of their kind. Over a century later, the rise of computational musicology has echoed this same impetus, especially as ‘big data’ methods gained traction in the 2000s. 57 This trend has only accelerated with the popularisation of digital editing tools, online open repositories, and increasingly sophisticated analytical technologies. Collaborative efforts between historical musicologists and information technologists have resulted in more accessible digital tools, fostering broader scholarly engagement with digital corpora. My direct experience with such corpora stems from my ‘Coordinator of Sacred Repertories’ role at the Archive of Iberian Polyphony. 58 This small and still-incomplete repository, which contains 130 works of Iberian origin, was developed as part of the project The Anatomy of Late 15thand Early 16th-Century Iberian Polyphonic Music (2016–2019), directed by João Pedro d’Alvarenga at Universidade Nova de Lisboa, Portugal. The Archive offers critical editions in PDF format accompanied by extensive metadata, though it does not currently include encoded music. 57 S. Tuppen et al., ‘Library Catalogue Records as a Research Resource: Introducing “a Big Data History of Music,”‘ Fontes Artis Musicae 63 (2016): 67–88, https://doi.org/10.1353/fam.2016.0011. 58 https://iberianpolyphony.fcsh.unl.pt/ (accessed 25 October 2025). 22 COST ACTION 21161 Nonetheless, it provided the foundation for a dataset of Iberian and Franco-Flemish motets ca. 1500, created to explore stylistic traits within these repertories and to investigate the authorship of two anonymous motets. In collaboration with Cory McKay (Marianopolis College), we analysed this dataset using jSymbolic for featurebased comparisons. 59 The resulting corpus consists of 175 motets encoded as MIDI files. 60 These were compiled by reusing edited materials from the Archive and supplementing them with works drawn from other repositories. The Iberian motets were modelled on editions produced by the Archive. The Franco-Flemish items were adapted from the Josquin Research Project. Newly produced scores were explicitly created to fill gaps in the repertoire. Given the heterogeneous nature of the sources, we followed a workflow aligned with the methodologies and standards outlined by Julie Cumming, Cory McKay, et al. in their seminal 2018 article, which systematises best practices for building interoperable symbolic corpora of pre-1600 Western music. 61 Along the encoding process, the crucial principle was to keep consistency. Whatever their origin, all the items were encoded as MusicXML files using Sibelius. They were manually adjusted and revised to create the master files and then exported as MIDI. Finally, they were double-checked with jSymbolic to detect any possible inconsistency. As a historical musicologist focusing on analysis, my assessment will concentrate on scholarly corpora that provide encoded outputs suitable for computational analysis. I will highlight the most notable examples without attempting to be exhaustive but rather aiming to underscore trends and challenges in the field. The earliest scholarly initiatives appeared in the first decade of the 21st century. Similar to the formation of traditional corpora of critical editions, they were discrete, uncoordinated and responded to various objectives. One such effort was the Tomas Luis de Victoria database, a personal project by Nancho Álvarez, mainly intended for performance. 62 While it does not qualify as a scholarly repository, it merits mention here as a representative case that illustrates the limitations of non-academic corpora for scholarly purposes. Started around 2002 and significantly updated in 2009, the project initially focused on the works of Tomás Luis de Victoria, eventually expanding to include 59 E. Rodríguez-García and C. McKay, ‘Composer Attribution of Renaissance Motets: A Case Study Using Statistical Features and Machine Learning,’ in The Anatomy of Iberian Polyphony around 1500, eds. E. Rodríguez-García and J. P. d’Alvarenga (Reichenberger 2021): 401–38. 60 E. Rodríguez-García, & C. McKay, ‘Composer Attribution of Renaissance Motets (Iberian Polyphony around 1500): MIDIs and Extracted Features,’ dataset, https://doi.org/10.5281/zenodo.4027957. 61 J. Cumming et al., ‘Methodologies for Creating Symbolic Corpora of Western Music Before 1600,’ in Proceedings of the19th International Society for Music Information Retrieval Conference, eds. E. Gómez et al. (2018): 348–354, https://archives.ismir.net/ismir2018/paper/000026.pdf (accessed 26 October 2025). 62 https://victoria.uma.es/ (accessed 26 October 2025). 23 COST ACTION 21161 compositions by Morales, Guerrero, Vásquez, and other Spanish Renaissance composers. An impressive undertaking, it currently offers over 1,500 encoded works. However, the editions are not critical, and the encodings lack standardised metadata essential for robust scholarly analysis. Furthermore, its reliance on Lilypond and MIDI encoding formats presents challenges for interoperability and reusability. Another early example is the Computerised Mensural Music Editing Project (CMME), created in 2006 by Theodor Dumitrescu and Marnix van Berchum. 63 Although it only contained 59 encoded pieces, 64 the project was groundbreaking in its approach to handling the complexities of mensural notation. However, its tailored encoding format (.cmme) limits compatibility with other formats. 65 The disjointed landscape of digital music notation in the early 2000s – described by Hankinson et al. (2011) as ‘highly fragmented’ – began to be more systematically addressed in 2010. 66 The Josquin Research Project, launched in 2010 and led by Jessie Rodin and Craig Sapp, represents a turning point in standardising procedures. 67 Designed with interoperability and reusability criteria in mind, it houses 902 works by Josquin and his contemporaries, offering a broad range of formats. In addition to the standard encoding trio of MusicXML, MEI and MIDI, the database includes Humdrum, MuseData, NoteArray, and JSON Piano Roll. It also features MP3 and PDF formats with/without editorial accidentals alongside an embedded analytical tool. It has helped establish a widely accepted model for what academically oriented music corpora should provide. Similarly, Tasso in Music, directed by Emiliano Ricciardi and Craig Sapp, launched in 2015 (and still ongoing), contains a collection of 778 music settings of Tasso’s poetry. 68 The project also offers a variety of encoding formats and capabilities. Another lesser-known dataset, the Symbolically Encoded Il Lauro Secco (SEILS) directed by Emilia Parada-Cabaleiro, is worth mentioning. 69 It comprises 30 pieces in both early and modern notation, amounting to 150 codified scores. Its range of formats is even more expansive, catering to both analysis and OMR applications: LilyPond, MusicXML, MIDI, Finale, **kern, **mens, MEI, agnostic, semantic, and PDF formats in both white mensural and modern notation. However, its hosting on GitHub – while an accessible platform – makes it somewhat more challenging to discover than a standard webpage, hindering findability. 63 http://www.cmme.org/ (accessed 26 October 2025). 64 And around 200 pieces in the Github Repository https://github.com/tdumitrescu/cmmemusic (accessed 26 October 2025). 65 But see Fiala et al., ‘A New XML Conversion Process.’ 66 A. Hankinson et al., ‘The Music Encoding Initiative as a Document-Encoding Framework,’ in Proceedings of the 12th International Society for Music Information Retrieval Conference, eds. A. Klapuri and C. Leider (2011): 293–298, https://archives.ismir.net/ismir2011/paper/000011.pdf (accessed 26 October 2025). 67 https://josquin.stanford.edu/ (accessed 26 October 2025). 68 https://www.tassomusic.org/ (accessed 26 October 2025). 69 https://github.com/SEILSdataset/SEILSdataset (accessed 26 October 2025). 24 COST ACTION 21161 The 1520s Project, directed by Ben Ory and Craig Sapp, is another step ahead. 70 Started in 2019, it contains 400 music items from ca. 1510–40, encoded in MEI, MusicXML, MIDI, and Humdrum. The project features an exquisite and thorough treatment of the metadata and a sophisticated array of analytical tools. Launched in 2023, E-LAUTE is the newest project in the field and one of the most ambitious. It aims to produce a comprehensive digital edition of lute tablatures written in German cypher. Supported by research agencies in Austria, Germany, and Switzerland, the project brings together a substantial interdisciplinary team under the leadership of Kateryna Schöning, Irene Holzer, and Martin Kirnbauer. The current focus is on tablatures dating from 1450 to 1550, with plans to extend coverage to works up to 1600 in subsequent phases, ultimately reaching a corpus of roughly 2,000 pages. The project demonstrates carefully considered choices in its use of technological standards (MEI and Web technologies), its balance between scholarly and performance-oriented perspectives, and its commitment to public engagement—most notably through opportunities for collaborative editing. 71 This progress represents a decisive move toward greater standardisation and accessibility in creating corpora, addressing the earlier fragmentation challenges. The importance of supporting as broad a range of formats as possible has been systematically articulated by Julie Cumming, Cory McKay, et al. in the aforementioned article. Despite this general trend, a significant group of projects developed around the Centre d’études supérieures de la Renaissance and Haverford College – directed by Philippe Vendrix and Richard Freedman – have taken a more focused approach, offering only MEI and PDF files. These include The Lost Voices (2012), 72 CRIM (2018), 73 and Gesualdo Online (2019). 74 The first two are particularly strong on analysis, and the latter exemplifies innovative dynamic editions. Well-suited to their specific aims, the restricted choice of formats can pose challenges for interoperability, as converting MEI into other symbolic formats remains technically challenging. Ideally, a digital corpus should consist of a representative selection of works curated for a specific scholarly objective, systematically edited and encoded per the established musicological conventions. It should provide PDF scores alongside symbolic files in as many formats as feasible, enriched with comprehensive metadata. This level of openness and structure not only enhances scholarly utility but 70 https://1520s-project.org/about/ (accessed 26 October 2025). 71 https://e-laute.info/ (accessed 26 October 2025). 72 http://digitalduchemin.org/ (accessed 26 October 2025). 73 https://crimproject.org/ (accessed 26 October 2025). 74 https://ricercar.gesualdo-online.cesr.univ-tours.fr/ (accessed 26 October 2025). 25 COST ACTION 21161 also aligns with the FAIR principles: 75 making data Findable, Accessible, Interoperable, and Reusable, thereby ensuring its long-term relevance and usability. 4.2 Aims and challenges of an analyst Anna Plaksin A key task for the analyst of a corpus is the mapping of the rationale present in the creation of a corpus with the goal of its interrogation. This is especially true when preexisting corpora of digital early music resources are reused and reaggregated. Data collections are gathered based on different goals, encompassing the spectrum of research-driven and curation-driven aims. 76 The role of a corpus in research can be understood as a model of the world we would like to gain knowledge of. According to Herbert Stachowiak’s definition, a model is always a representation of something. It generally captures not all the attributes of the original it represents, and it always represents the original not per se but for a specific context or under the restriction of certain operations. 77 Relying on this definition, this means that a corpus is always a representation of something in a different medium. It is not the original or a copy, it is a reduction. And it always aims at a special context – it only represents the original for a certain (limited) use. That means the re-use of a corpus in a different context is not guaranteed, and it is necessary to check whether the corpus and its assumptions fit the research question. Because both the creation rationales as well as the activities encompass nearly infinite possibilities. Activities that involve the interrogation of early music corpora may be grouped into three categories: 1. Search / Retrieval: Finding resources by certain criteria. E.g. finding teaching resources by composer, genre, or provenance or sheet music for performance for a certain setting, vocal range or difficulty. These scenarios are mainly focused on metadata searches and aim at finding specific resources within a large collection. 2. Content-based analytical operations. E.g. melodic search, automated music analysis or training an algorithm for automatic composition. These activities need machine-readable representations of music. They do work fine with straightforward representations or modern transcriptions, e.g. as MIDI files, MusicXML or **kern. 75 https://www.go-fair.org/fair-principles/ (accessed 26 October 2025). 76 See J. Flanders and F. Jannidis, eds., The Shape of Data in Digital Humanities: Modeling Texts and Text-Based Resources (Routledge, 2019), 86–89. 77 See H. Stachowiak, Allgemeine Modelltheorie (Springer,1973), 131ff. 32 COST ACTION 21161 expense of the others; the second is to conflate the texts. However, in reality we are dealing with a spectrum with diplomatic transcription at one end and critical edition Figure 8. Beginning of critical commentary for Peter Philips, Dolorosa Pavan. © Copyright 1999 by The Musica Britannica Trust and Stainer & Bell Ltd. Reproduced from Peter Philips: Complete Keyboard Music, edited by David J. Smith, Musica Britannica Vol. 75 (p.196) by permission of Stainer & Bell Ltd, London, www.stainer.co.uk. All rights reserved. at the other. The conflation of sources is the norm for Musica Britannica volumes of keyboard music, and indeed for some seventeenth-century instrumental collections. 91 At the other end of the spectrum there are source-based editions such as the one of Fitzwilliam Virginal Book by Jon Baxendale and Francis Knights which, while scholarly in terms of its introductory material, is not critical – indeed, the editors 91 For example, the Le Strange manuscripts, especially GB-Lbl, Add. MSS 39550–4, record every difference between his versions of music for viol consort and those in manuscripts owned by his friends, regardless of whether it affects how the music sounds. See Peter Philips and Richard Dering: Consort Music, ed. David J. Smith, Musica Britannica Vol. 101 (Stainer & Bell, 2016), xxiv–xxviii, 186–187. 33 COST ACTION 21161 describe it as a ‘pseudo-facsimile’, 92 and it preserves many graphical features of the manuscript, including the beaming of the shorter note-values and idiosyncratic notation of semibreves as two tied minims. The editors make minimal changes to the text, correcting obvious errors, and make suggestions concerning, for example, accidentals, but they do not refer to other sources of pieces contained in the Fitzwilliam Virginal Book: given the scale of the manuscript, this is understandable. The focus here is on the source, rather than on the composer. Most editions sit somewhere in the middle of the spectrum: a source edition – as opposed to transcription – would correct obvious mistakes by reference to concordant sources; a composer edition usually has a main, principal, base text. Players have a use for both kinds of edition: here, for example, someone making a recording of Fitzwilliam Virginal book would need a source-based edition, but a concert of keyboard music by Peter Philips might require a composer-based one. An ideal edition, then, would be able to cater for both. However, traditional, paperbased publications cannot provide editions of all the sources alongside a critical edition combining them because of the constraints of space. This is where the dynamic, digital edition of the future comes in: the two editorial approaches can be reconciled, with sourceand composer-editions coexisting in the same virtual space. A critical edition needs to record all the information in one form or another, and the co-existence of source editions and composer editions makes this easier to achieve. During the STSM in Utrecht, we thought a great deal in terms of ‘layers’. In the context of scholarly, critical music editions, encoding would involve a minimum of three levels, each building on the previous one: • Diplomatic Source Transcription = Transparency: all sources should be encoded to produce diplomatic source transcriptions so that the user is able to refer back to the ‘raw material’. • Critical Source Editions: critical editions of the sources where the main text is changed only where absolutely necessary and by reference to the concordant sources. • Critical Composer Editions: the focus is on the composer so, following the example of seventeenth-century sources such as the Le Strange consort manuscripts, a base text forms the foundation of an edition where there is greater flexibility in combining readings from different sources. In the digital domain, such a multifaceted edition could follow the model of existing collected editions: at a click of a mouse, the textual commentary could appear but using music notation rather than text. Notes could be coloured on screen to alert the user to disagreement between sources. However, it would also be feasible to display texts in parallel in a manner not dissimilar to Figure 6, or to use IIIF 93 to create links 92 The Fitzwilliam Virginal Book, eds. J. Baxendale and F. Knights (Lyrebird Music, 2020) vol. 1, xxxi. 93 International Image Interoperability Framework, https://iiif.io/ (accessed 26 October 2025). 34 COST ACTION 21161 between a bar in the edition and the corresponding place on a digitized image of the source. 94 The latter is more easily envisaged for source editions, and has been implemented in some of the Polish scores: see, for example, Giovanni Matteo Asola’s Ave Sanctissima Maria where double-clicking on a note brings up an image of the corresponding system in its source. The same piece may be used to illustrate how digital editions can facilitate the co-existence of original and modern notation: on the click of a mouse the display may be changed from original clefs, note-values (including stemming), part-names, spelling and orthography to a modern score of modern clefs, with barlines added, modern part-names, spelling and orthography. Different users of sheet keyboard music have different priorities, and there is disagreement over the extent to which graphical aspects of the notation (such as the beaming of quavers and semiquavers) have implications for performance practice. A digital edition would therefore be able to take the user from a version with original notation (number of staff lines, clefs, beaming, placement of accidentals) to a modern one. There could be similar options for downloading PDFs for printing or use on a tablet. For a digital edition along the lines suggested above to be possible, there is a need for corpus creation. All texts for each piece need to be encoded and ideally the entire contents of each manuscript so that the edition becomes a research tool to explore scribal interaction with the music. A digital edition such as is being proposed here shifts the focus away from composers and on to scribes and users of manuscripts in the seventeenth century. On the face of it, IIIF links to images might suggest that not everything needs to be encoded. However, a player who believes that beaming of smaller note values is indicative of phrasing would want it preserved in a source edition. System-breaks can be vital to understanding placement of accidentals and can provide valuable evidence of inter-source relationships. Furthermore, anything that remains within a digitised image of a source without being encoded becomes ‘invisible’ to a machine and therefore makes the data less useful when analysing the music. One significant advantage of a scholarly digital edition is that it can be more than the visual presentation of musical notation: the data required for the creation of a multidimensional, dynamic edition as envisaged above can be manipulated for a range of other purposes. Whereas a research project involving a particular research question may involve making compromises in what to include, leaving out any elements that will not be needed in addressing it, a digital edition needs to remain as open and flexible as possible since it needs to meet the needs of a range of users. As part of the Utrecht STSM, a number of interviews were conducted (see Section 7.1), and one participant commented that ‘you don’t know what it is that you need to know until you need to know it’. A monumental edition need not have a specific goal in mind, so needs to involve 94 https://polishscores.org/?id=16xx:900 (accessed 26 October 2025). 35 COST ACTION 21161 encoding as much as possible in order to make it useful to the widest range of audiences. Manually encoding music from scratch is labour-intensive. However, the use of existing digital material speeds things up immeasurably. Any existing transcription or edition of a piece can serve as the basis: OMR can be used if necessary to create a Music XML file that can then be tidied up in a notation software package before being exported as XML and converted to MEI. The Music XML file can be edited multiple times to generate the various layers needed for a dynamic digital edition. However, it is still a mammoth effort, especially if the encoding is to include the degree of detail suggested. There is a trade-off to be made between speed and quantity on the one hand and inclusivity on the other. Some projects have involved the beneficiaries – the user communities – in the creation of a resource, but this raises issues of Intellectual Property and copyright when working with established series, not to mention quality assurance. During the STSM, Mark Gotham gave an illuminating presentation on encoding German Lieder and string quartets. Although the repertoire is later, we benefited from his experience in crowd-sourcing the encoding of nineteenth-century editions which seemed to have been highly successful. By using such editions as sources, he avoided issues of IP and copyright. However, Julie Cuming and Cory McKay reflected on their projects, including ELVIS and SIMSSA. Their experience of crowdsourced encodings of music which were intended primarily for analysis differed markedly from that of Mark mainly because the sources used as their basis varied: some encodings employed reduced note values whereas others did not. They were sceptical about using crowdsourcing for the creation of encoded corpora of music and preferred smaller but accurate and consistent datasets to larger ones. The contrast between the two perspectives illustrated the trade-off between quality and scale. A project involving the creation of a new series of critical editions might well involve users in their creation. Ironically, the lack of lute music in monumental editions such as Musica Britannica has led to the creation of digital resources by enthusiasts which can now provide the source transcriptions needed for critical editions of, for example, lute music by Dowland. 95 However, the issues of IP and copyright are that much greater for a well-established monumental edition, especially when new editions are made from existing material. In the case of existing series, there is scope to involve user communities in the design of a future digital critical edition. It may be unrealistic to expect a user necessarily to pay for an edition if they have contributed to its creation. This is perhaps where there needs to be a distinction between data and edition: an encoded manuscript is a digital representation of our 95 T. Crawford, ‘A Digital Corpus for Exploring the Lute Music of John Dowland (1563–1626),’ Journal of New Music Research 53, no. 3–4 (2025): 1–11, https://doi.org/10.1080/09298215.2025.2486139. 36 COST ACTION 21161 shared heritage so perhaps should not be monetized, whereas an edition arising from it is invested with the intellectual property of whoever created it. Clearly, were crowdsourcing to be used for encoding texts in the creation of a critical edition for an established series, there would need to be absolute transparency surrounding the arrangement. There is an obvious need to be able to combine data for editions of English keyboard music made by several individual editors or editorial teams, especially when one is working on a composer, and another on an individual source. For scholarly, critical editions of keyboard music to flourish in the digital sphere, there needs to be a degree of standardisation, consistency and transparency. In the longer term, however, it is possible to envisage combining datasets of this keyboard repertoire with those of other kinds of music – maybe traditional, folk music? – to explore broader questions of dissemination and transmission. 4.4 How to select a representative corpus for historical study Mirjam Visscher and Frans Wiering The first problem one encounters when doing corpus-based music research is the creation of the corpus itself, for the composition of the corpus is a critical factor in the validity of the research outcomes. A suitable corpus minimally meets the following criteria: • it fits the research question: it must contain the features that are relevant to the research; • it is representative of the music under investigation, in terms of spread over different strata such as genre, country and composers. The latter problem is the topic of this section. It seems that most of the corpus research (insofar as it doesn’t study a closed corpus, such as the complete output of a composer) uses an approach of convenience sampling, making the most of the available data sets and/or the (usually limited) amount of time for data creation. We are not aware of a solid method to select compositions for inclusion in a musical corpus. However, in computational linguistics a comprehensive, straightforward method has been defined by Biber, that can be transferred to the music domain. 96 Biber’s approach has three important steps: • define the boundaries of the population; • define the organisation within the population; • set a sampling frame. We will illustrate this approach with a short example. For this we define the boundaries of the population as all western music surviving in manuscript or print dated between 1500 and 1700. The organisation within the population refers to 96 D. Biber, ‘Representativeness in Corpus Design,’ Literary and linguistic computing 8, no. 4 (1993): 243–257, https://doi.org/10.1093/llc/8.4.243. 37 COST ACTION 21161 properties known as strata such as the composer, genre, secular or sacred nature, instrumentation, country, language and date. The choice of strata relates to the research question: a representative corpus reflects the distribution of compositions over the selected strata of the population. Figure 9. Conceptual population breakdown of all music there ever was (dark blue) into the partially overlapping categories known (yellow), known & lost (red) and underrepresented (green). The visible dark blue area represents the unknown repertoire, of which we do not know its size. Figure 9 shows that the population has some problematic aspects. First, it is already a selection from all western music there ever was in the 16th and 17th centuries. We know very substantial amounts of music were lost. In some cases, we know that compositions and sources existed in the past, but there is an unknown amount of music that has vanished without a trace. There is no reason to assume that this was a purely random process in which every piece stood the same chance of surviving, so we can be sure this has influenced the distributions of strata. There is some specialised research into the estimation of numbers of items lost without a trace, 97 but we are not aware of methods for doing so at a large scale. 98 Information on the known losses is scattered over the literature: it would be helpful if a centralised inventory of these could be created. A second problematic aspect is the occurrence 97 M.S. Cuthbert, ‘Tipping the Iceberg: Missing Italian Polyphony from the Age of Schism,’ Musica Disciplina 54 (2009): 39–74, https://www.jstor.org/stable/25750547; M. Kestemont, et al., ‘Forgotten Books: The Application of Unseen Species Models to the Survival of Culture,’ Science 375, issue 6582 (2022): 765–769, https://www.doi.org/10.1126/science.abl7655. 98 Recently, however, it was estimated how many Gregorian chants have disappeared, based on the chant distributions in the CANTUS database (https://cantusdatabase.org/, accessed 30 October 2025): J. Hajič, jr. and F.C. Moss, ‘Knowing When to Stop: Insights from Ecology for Building Catalogues, Collections, and Corpora,’ in DLfM ‘25: Proceedings of the 12th International Conference on Digital Libraries for Musicology, ed. E. De Luca (Association for Computing Machinery, 2024): 90–94, https://doi.org/10.1145/3748336.3748347. 38 COST ACTION 21161 of anonymous compositions. This is not a negligible category: in RISM, 18% of the items are anonymous and in CPDL 10% of the compositions are anonymous. Anonymous works are often hard to date, and their geographical origin may also be less easy to identify. The distribution of anonymous composers is not uniform geographically; we know for instance that the majority of 17th century Scottish compositions are by an anonymous composer. 99 They thus pose a serious problem to creating a representative selection. Third, not taking underrepresented groups into account may introduce additional bias. There are catalogues of music by underrepresented groups, such as the ones in Table 2. Such lists aren’t the final answer, however, since the attention for the underrepresented may lead to overrepresentation and other underrepresented groups may be missing. Being explicit about such problems, even if they cannot be solved, is the best strategy. category dataset name url Iberian polyphony Archive of Iberian Polyphony https://iberianpolyphony.fcsh.unl.pt/ Female composers BIG LIST of Women Composers https://donne-uk.org/the-big-list/ Spanish and New World Polyphony Books of Hispanic Polyphony https://hispanicpolyphony.eu Poland, 16th-20th century Polish digital scores https://polishscores.org/ Spanish Renaissance Tomás Luis de Victoria https://www.uma.es/victoria/index.html Table 2. Sample catalogues of underrepresented groups. The third and last step in Biber’s methodology is creating a sample frame. Such a frame consists of one or more lists with attributes and their distribution over the population. Table 3 gives two simple examples of sample frames creating from RISM. In practice sample frames will be hierarchical (for example genre and country). After the sampling frame has been created, the actual corpus can be selected by randomly choosing compositions that fill the ‘slots’ in the frame. How many need to be selected depends on the minimum number of pieces per slot. In the end, we have a sample that is substantially smaller than the population but nevertheless is representative of it by the chosen criteria. 99 D. Coney, ‘The David Melville Bassus Partbook and Scottish Music Culture at the Turn of the 17th Century,’ paper presented at the International Medieval and Renaissance Music Conference, Durham, U.K, July 2025. 39 COST ACTION 21161 composer % genre % Anonymous 18.1 Sacred song 12.9 Palestrina 0.7 Opera 7.1 Lassus 0.4 Lied 4.4 Lully 0.3 Motet 4.0 Purcell 0.2 Song 3.7 Marenzio 0.1 Keyboard piece 3.6 Byrd 0.1 Mass 3.4 Victoria 0.1 Sonata 3.1 Gallus 0.1 Aria 3.1 Frescobaldi 0.1 Cantata 3.0 Table 3. Sampling frames for composers and genres based on the RISM records (top 10). 40 COST ACTION 21161 5 Beyond the scholarly community As has been repeatedly suggested in the previous chapters, the involvement of contributors from outside the scholarly community to corpus creation (in the widest sense) is quite important. Where possible the connection between music researchers, professional musicians and citizen scientists deserves to be fostered. Section 5.1 illustrates this by means of an analysis of the contributions to CPDL, a major online resource of choral music. In Section 5.2 digital editing is studied as a collaborative process where contributors are motivated to perform certain tasks: these tasks should be designed in such a way that they reduce technical complexity and allow focus on the musical task at hand. 5.1 CPDL contributors: a case study Mirjam Visscher and Frans Wiering Early 2024, we investigated the productivity of the contributors by scraping the contents of the site. At the time 60,868 editions of works by 4,753 composers were available. As our focus is on corpus building and reusability of encodings, we looked into the prevailing file format for sharing encodings, MusicXML. Our dataset contains 48,847 MusicXML files by 1,609 contributors, accompanied by metadata. Figure 10. Top 25 CPDL contributors (anonymised) and the years during which they were active. The contributors account for 69% of the total number of 48,874 MusicXML files contributed to CPDL. First, we looked into activity and productivity of the top 25 contributors. Their efforts account for 33,836 files or 69% of the total number of MusicXML files (see Figure 10). 41 COST ACTION 21161 Figure 11. Productivity metrics over the years. Figure 11 shows how over time new contributors joined the CPDL platform and that the long-term trend for the number of pieces added per year is going upward while the average productivity of the contributions seems stable. From a business economics perspective, these metrics are a sign of healthy continuity. Figure 12. Composition years (smoothed) of the contributions by the top 10 CPDL contributors (anonymised). 48 COST ACTION 21161 The processing of the data depends on the kind of desired output, with two main categories, analytical results or music. The former entails the use of specialised music software such as music21, software for machine learning or data analysis such as WEKA, or generic program languages such as Python. Music as output implies the ability to perform tasks such as conversion to CMN, transposition, choosing types of editorial detail, adaptation to performance circumstances, and sonification. Both user and maker perspectives spawn an extended workflow of tasks that differ in complexity and expertise needed. In principle, it seems possible to split these up in subtasks that can be carried out by workers with the right combination of qualities. However, even the simplest tasks are likely to require accuracy, concentration, and understanding of the content. Therefore, it is crucial to design user interfaces in such a way that they optimally support the task at hand and do not distract or hinder it, in other words, with full attention to usability aspects that are sadly often ignored. 6.4 Contexts Three types of context are distinguished in the PACT model: physical environment, social context, and organisational context. Considering the first, most of the work probably takes place in offices, libraries and home studies using a desktop computer. As an alternative setting one might think of a mobile device, on which micro tasks can be executed in a collaborative scenario, for example when travelling by public transport or while waiting in a queue. This would require a completely different approach to task and interface design. In terms of social context, most workers, even if they do their editing and encoding in isolation, are part of a community of practice (for example Lilypond) that shares the same tools and publication platforms and where they can ask each other for help. When designing a new corpus creation project, it is imperative not only to focus on tools and tasks, but to actively design and foster the social environment in which the work takes place as well. If workers don’t feel connected, supported and valued, they will go elsewhere. Organisational context brings in several interrelated factors that influence the longterm success of a project: • organisational embedding: projects are generally short-term, but are they endorsed and supported by organisations with a longer lifespan? • storage: where are the data hosted, how are they backed up? • technical sustainability: will data and interface for accessing them remain available and if necessary updated after the project? • intellectual sustainability: are sustainable formats used and is there documentation that records provenance of the data and encoding decisions, so that future users understand the potential and limitations of the data? • financial sustainability: what does it cost to maintain the environment and how is income generated to do so? 49 COST ACTION 21161 Other organisational factors include: • intellectual property: can data be lawfully shared, and is everyone’s share in creating them adequately recognised? • relation to other initiatives, such as digitisation projects of libraries and archives, cataloguing endeavours (RISM, CANTUS, DIAMM), and division of labour with related projects; • interoperability with the Linked Open Data world, for example by means of shared identifiers for various entities like persons and compositions; • observation of FAIR principles. The subtext here is that, although crowdsourcing and employing citizen scientists seem like cheap strategies to assemble a lot of data, corpus creation remains an expensive activity in a world where funding for cultural heritage and humanities research is scarce. When available, funding needs to be spent wisely and effectively, and when not available, strong and at the same time realistic cases need to be made to funding agencies to convince them that music corpus creation is a worthy destination for their money. 6.5 Technologies In the PACT framework, Technologies are divided in four categories, input, output, communication and content. We will skip input and communication, 106 important though they may be, since at this rather generic level of PACT analysis, little can be said about these that is specific to our problem. This is different for output, as this category describes the variety of products that corpus creation may aim at. Briefly summarised, the following dimensions have been identified: • data oriented to edition oriented encoding; • selective to near-complete encoding, with the qualification that complete encoding is unattainable and a well-articulated decision must be made about where to stop; • encodings and editions range from source-oriented to composer-oriented approaches; • whether symbols (e.g. ligatures) or semantics (e.g. pitches and durations) are encoded; • the notation output can be static, allow basic manipulation such as reformatting or selection of voices, or can be dynamic, affording notation-altering manipulations such as transcription and transposition; • whether the intended medium for the end product is paper or digital; • reusability of the encodings, ranging from none to basic features to complete encodings. 106 Input is mainly about input devices. For our purpose one could think of scanners and (music) keyboards as non-standard input devices. Communication addresses matters such as network connections, protocols, and transmission speed. 50 COST ACTION 21161 This range of output options implies a corresponding range of technical decisions in the design of the musical content. Such decisions take the form of computational models, abstract structures that capture aspects of the music in a formal logical structure. As explained in 4.2, a model is a reduction of the original for a certain purpose, and it is important to understand what is included in the model as well as what isn’t. Aspects of corpus creation that relate to content and modelling are: • encoding system(s) used in the system; conversion • quality and accuracy of the encoding • notation form of the source (mensural, tablature, CMN) • choice of features to encode • inclusion of digital images (potentially linking of image and encoding via IIIF) • metadata format describing source and encoding process Metadata design in particular seems an underdeveloped area: the choice seems to be in practice to either use a minimalistic approach and supply little beyond composer and title, or work with a verbose format such as the TEI/MEI header. There is a need for developing a middle ground, where metadata profiles are defined for specific repertoires but that can still be shared at a higher level. 6.6 Conclusion This PACT analysis can have multiple purposes. It can serve as a framework to describe current projects, identify opportunities for improving them, compare them and find chances for collaboration or data exchange. It can also be seen as a set of questions and attention points for designers of future music corpus creation environments that they can use as a starting point for user research and requirements gathering. An important generic observation about all PACT dimensions is that accurate documentation, not just of the specific encoding task but of the entire project, is vital, otherwise the user of the project output is left in the dark with respect to factors that fundamentally shape the encodings. Such documentation should include: • musical documentation, describing the aim and content of the project, as well as criteria for inclusion/exclusion of compositions; • technical documentation about the system design and implementation; • maker documentation, explaining the creation workflow in the system; • user documentation, explaining how the encodings can be manipulated: it should also discuss design choices and limitations that may influence the interpretation of the outcomes; • metadata documentation, explaining how the encoding should be formally described. In all documentation it is vital not just to present all the options (which sadly is what most user manuals document) but to give task oriented, practical advice that matches the problem the encoder is struggling with. 51 COST ACTION 21161 7 Narrowing the gulf between people and systems in music encoding In our activities, we focused on the people involved in the creation and use of corpora in three ways, interviews (Section 7.1), the creation of personas (Section 7.2), and the analysis of roles in collaborative corpus creation (Section 7.3). 7.1 Interviews We conducted nine interviews, six with musicologists and three with citizen scientists, to learn more about encoders’ motivations and experiences. We present a short overview of their responses. Multiple respondents mention that discovery is an important motivation for their work, as is sharing their discoveries with others. Another motivation is altruism: doing something for other musicologists or performers, or giving back to the community in exchange for the use of other people’s materials. Other motivations include obsessive collecting, editing as a way of gaining deep knowledge of the music, inaccessibility of expensive scholarly editions, and leaving something behind for future generations. All respondents see PDF as the primary format for distribution, though several mention sharing encodings; one of them does not want to share encodings for fear of others interfering with their work. There are some interesting thoughts though on the added value of digital editions, such as support for multiple completions of the same work (from an interviewee working with polyphony fragments), better visibility of sources in the edition, different ‘viewpoints’ on the edition for different user groups, searching for intertextual relationships, support for annotation, and finally quantitative analysis. Respondents generally expect acknowledgment of their work and appreciate feedback on their scores. No one expects material reward, but some would have a problem with their work being re-used in a commercial setting and have licenced their work accordingly. Some respondents plan to share their work at a later stage and wonder about the right platform to do so. One quite interesting suggestion is to create a centralised place (irreverently described as a ‘dumpster’) where researchers could deposit their encodings for others to use, after having finished researching them. Quality is a major concern for all. Everyone works directly from primary sources (we know that many encoders work from modern editions, but not our respondents) and aims to create accurate scores. Better access to sources is desirable, especially to underrepresented materials such as organ tablatures. Generally, they apply guidelines or principles, formally or informally. Such principles can be very concise (‘reversibility, traceability, transparency’) and sometimes evolve over time. Sometimes formal guidelines are followed. Critical commentaries range from none 52 COST ACTION 21161 to very extensive texts. One respondent suggests a peer approval process for contributions to CPDL-like resources. Our respondents come from several communities of practice (e.g. Humdrum, MEI, Lilypond, MuseScore, Finale, digital lutenists) and seem to be generally happy with the functionalities that their platform offer, minor complaints excepted. At times they mention limitations of other platforms. Technical help seems to be mostly available, and the citizen scientists have easy access to musicological expertise. It is not uncommon for digitised sources and unpublished transcriptions to be shared with peers. One participant is involved in a large-scale funded project and works with multiple colleagues. Others work individually and rather interact with their communities on an ad-hoc basis. 7.2 Personas Anna Plaksin and Frans Wiering We also used our experiences to create personas - fictional representations of potential users - each with a short profile including biographic background, experiences, motivations, and frustrations. This should help us to create a shared understanding of potential communities and a shared vision for the future of corpus building. In teams of two, we created five different personas with varying backgrounds. We envisioned people between their early 20s and their mid-70s, some of them professional musicians, others music enthusiasts with varying aspirations. Envisioned experiences, motivations, and frustrations were already astonishingly diverse, ranging from seeking material to play from, community, visibility of their own work, and recognition (for a compact survey see Table 4). However, the effect of creating these personas did not lie so much in their details but in their impact on our discussions. Motivations and frustrations of potential users increasingly guided our decisions and led us to take the consequences of design choices into account. It contrasted our earlier mapping of the many modalities of creating or using digital corpora of early music by giving us a perspective from within this system, showing potential considerations for choosing a certain direction instead of others. Moreover, it presented us with the diversity of aims and needs that are placed on technology and must be balanced in its use. Simply put: the creation of music content can be a rewarding activity for people from diverse backgrounds and with diverse technological expertise. Therefore, encodings systems and related technologies need to provide different modes of engagement e.g., addressing different levels of interest or commitment. This refers back to the very core of usercentred design. As Don Norman pointed out already in 1986: 107 humans use technology to achieve their goals. In their use of technology, they face two gulfs: the 107 D. A. Norman, ‘Cognitive Engineering,’ in User Centered System Design: New Perspectives on Human-Computer Interaction, eds. S.W. Draper and D.A. Norman (Erlbaum, 1986): 31–61. 53 COST ACTION 21161 Gulf of Execution when figuring out how something operates, and the Gulf of Evaluation when evaluating if an action has brought them closer to their goal. The design of technological systems is always concerned with narrowing this gulf, either by bringing people closer to technology or by bringing technology closer to people. name, age background motivation frustrations David (c. 60) choral conductor, has little interest in technology find repertoire that makes singers happy, help others in same situation existing editions always unsuitable Dirk (73) retired physics teacher, plays recorder and composes, unhappy marriage publish transcriptions and compositions online, wants recognition for his efforts web site doesn’t get much traffic, maintenance is tedious Elizabeth (late 60s) retired geography teacher, single, amateur singer learn more about the music she sings, find new music wary of technology, really needs sense of community Sebastian (mid 40s) tech savvy, plays violin in amateur orchestra save money by creating his own scores expensive orchestral materials Tanya (23) music student, basic MuseScore user publish transcriptions, write her own music, working with peers considers herself a noncomputer person, worried things go wrong Table 4. Personas created during the workshop. With regard to standardisation and data quality, this leads to a significant question in the role of technological systems for corpus creation: do we want to bring people closer to the technology by teaching them e.g., what ‘good data’ should look like? Or do we want to bring technology closer to the people who are using it? As the shift in our discussions based on the personas showed: aiming for the second option shifts the role of technology towards a system that helps negotiating people’s aims and needs, for example by providing conversion procedures or multi-faceted access points. With such an approach there would be less need for them to modify their technical skills and preferences, while their intrinsic motivation to create codes and editions would remain unaffected. 7.3 Roles We also discussed the collaborative process of music encoding in one of our sessions in connection with the ideas presented in Section 5.2. The aim was to create a generic workflow with tasks and roles that could be used as a template for 54 COST ACTION 21161 concrete encoding projects. Assuming a rather high level of organisation in the project, we initially we produced a quite elaborate diagram, which was subsequently reduced to the simple model of Figure 15 to fit a broader range of approaches. Figure 15. High-level workflow of music encoding projects. The model contains four main tasks, and multiple subtasks: • instigation: setting the goals of the project; selecting compositions; creating a template for encodings; defining the project’s success criteria; • gathering: collection the materials to be used: metadata; digitised and physical sources; digital and physical editions; reusable encodings; • editing: creating the encodings by means of OMR, manual transcription, encoding, and/or code conversion; integrating the raw transciption(s) into an edited encoding; correction and validation of the edited encoding; • publishing: making the validated encodings available to the users. This could be just a linear process, but more realistically it would involve feedback steps from publication to the three other tasks: • to editing: integrate feedback from the users in the encodings; • to gathering: enrich the encoding with newly-selected source materials, or reuse the encodings as base text for transcriptions of other versions of the same work, particularly close concordances; • instigation: check the result against the goals and success criteria, potentially instigating a next cycle through the workflow. In addition to the tasks, a number of roles is defined: • instigator of the project: performs the tasks necessary to start off the project (see above); • transcriber of sources: creates the raw encodings of the selected sources; • computational specialist: in charge of OMR, conversion and other computational processes; • editor: creates and edited encoding (or digital edition) from the raw encodings and the edition template; publishes edition after validation; 55 COST ACTION 21161 • validator: takes care of proof reading and quality checking; • edition user: uses the materials for their own purposes and gives feedback to the project. Leaving edition users aside, each role asks for a different set and/or level of expertise. For example, for transcribers a good working knowledge of music notation and a basic understanding of editing would suffice, while the editors need to possess both at a higher level. It depends on the ambition, scale, scope, embedding, funding and lifespan of a project if multiple (or even all) roles are relegated to a single person, or if each role is divided between multiple persons. In the case of transcribers and editors the latter seems obvious but large, structured projects are likely to have teams of instigators, computational specialists and validators. 56 COST ACTION 21161 8 Vision for the future of corpus building Increasingly, musicians and music researchers engaged in early music do their work in the digital environment. Music is searched, transcribed and analysed by means of specialised software and services, editions are shared online, and research results are published in digital journals, conference proceedings, and books. Yet if we overview the digital landscape of early music, the dominant impression is one of relative fragmentation and scarcity of resources. This is especially true of what is arguably the fundamental precondition for most digital work in the field: the availability of music notation in digital, computer-processible form. Currently one cannot assume that for a given research problem in early music sufficient musical data containing the necessary features are available that can be analysed to solve it. Often there is such an investment in data creation and curation needed that a computational approach becomes infeasible. Therefore, a comprehensive approach to accelerate corpus creation in early music is urgently needed, aiming at a better coverage of the repertoire in a form that ideally is ready to be used, or, more realistically, needs a limited amount of work to make it ready for use. We consider corpora and corpus creation to cover a broad range of products and activities, including the production of analytical encodings that contain only a few specific musical features, the encoding of the original notation of sources, and the creation of digital editions that meets scholarly and/or practical requirements. These activities have numerous commonalities in terms of how they are carried out on a day-to-day basis. They involve dealing with sources, notation forms, and editorial interventions, as well as working with similar technologies and encoding systems. Furthermore, they face the same problem of limited resources. There is so much work to be done with only limited financial means and human effort that sharing whatever can be shared between initiatives becomes imperative. Corpus creation must thus be an inclusive endeavour. This applies not just to exchanging encodings, but to participation and technology as well. Music lovers and citizen scientists put a tremendous amount of effort into encoding music for their editions, often with impressive results. Professionals and amateurs use a variety of technologies for creating their encodings that fit their working methods and aims well. Any effort to coordinate their activities should start with a recognition of their motivations, musical skills, and technical choices. The challenge is to devise coordination mechanisms on top of these that respect their autonomy, yet help to spend resources wisely and stimulate participation. Challenges do not end when the encoding is done. It must be possible to find encodings; they must possess attributes that document what they contain, how they were made and how they relate to other items in the digital ecosystem of early music. And they must remain accessible in the foreseeable future, on sustainable infrastructures. In the rest of this section, we will elaborate this vision in a number of 57 COST ACTION 21161 concrete aims. In the next section we will describe steps to realise some of these aims. Coordination • Create a group that acts as a glue between encoding and digital editing projects, collects experiences, advises ongoing projects and devises guidelines for the creation and sustainability of corpora. The current CORSICA team could form the core of this group. • Create an inventory of encodings that are suitable for musicological work or as input for corpus creation (provisionally called RIEM – Répertoire International des Encodages Musicaux). • Create a sustainable, decentralised platform with a small corpus of encoded scores, with the aim to grow little by little. This could also function as a repository for encodings created in short-term or small-scale projects lacking a publication platform. • Create a metadata service that is reliable, in one place and being maintained, that provides unique IDs for various entities and can be incorporated into IMSLP, CPDL, and other already existing platforms. • Create a unified user interface across the more important resources. Participation • Instead of building new communities, having a high risk of failure, identify the good (or promising) existing communities and build from there. • Encoding of early music sources is democratised and is public domain. • Mechanisms exist for acknowledging everyone’s contribution in shared work, such that they closely align with motivations of the participants. They should feel comfortable in contributing to the corpus. • Digital encoding/editing is part of the musicological curriculum: students learn to encode, to produce digital editions and to analyse encodings. • Enable creation of encoding campaigns: identify an encoding challenge, create a workflow, divide the tasks over participants and integrate the work in a collective result. Tools and support • Develop an easy-to use mobile-first editor for encoding to allow crowdsourcing efforts. • Design methods for a seamless and easy as possible conversion between encoding formats and notations. • Tech workflows, such as for the encoding campaigns described above, are low-threshold, easy to use and freely available. • Create a well-designed front end for distributing data and editions. • Successful prototypes for music encoding software can be turned into stable and usable products. 64 COST ACTION 21161 Notes in Computer Science, 14809. Springer, 2024. https://doi.org/10.1007/978-3-03170552-6_2. Robison, T.D. ‘IML-MIR: A Data-Processing System for the Analysis of Music.’ In Elektronische Datenverarbeitung in der Musikwissenschaft. Edited by H. Heckmann. Gustav Bosse Verlag, 1967. Rodríguez-García, E., and C. McKay. ‘Composer Attribution of Renaissance Motets: A Case Study Using Statistical Features and Machine Learning.’ In The Anatomy of Iberian Polyphony around 1500. Edited by E. Rodríguez-García and J. P. d’Alvarenga. Reichenberger, 2021. Rodríguez-García, E., and C. McKay. ‘Composer Attribution of Renaissance Motets (Iberian Polyphony around 1500): MIDIs and Extracted Features.’ Data set. https://doi.org/10.5281/zenodo.4027957. Roselló, A., E. Fuentes-Martínez, M. Alfaro-Contreras, D. Rizo, and J. Calvo-Zaragoza. ‘SourceFree Domain Adaptation for Optical Music Recognition.’ In Document Analysis and Recognition – ICDAR 2024. Edited by E.H. Barney Smith, M. Liwicki, and L. Peng. Lecture Notes in Computer Science, 14809. Springer, 2024. https://doi.org/10.1007/978-3-03170552-6_2. Selfridge-Field, E., ed. Beyond MIDI: The Handbook of Musical Codes. The MIT Press, 1998. Smith, D.J. ‘The Instrumental Music of Peter Philips.’ PhD diss., University of Oxford, 1994. https://ora.ox.ac.uk/objects/uuid:d1af3140-2553-4f58-8437a1eee66d7f13/files/m06f60f301b6cca42daeeca7b3508cd0a. Smith, D.J., and A. Woolley. ‘Editing Purcell’s Keyboard Music: Some Reflections on Collected Editions, Past and Future.’ Journal of New Music Research 53, no. 3–4 (2024): 277–296. https://doi.org/10.1080/09298215.2025.2472612. Stachowiak, H. Allgemeine Modelltheorie. Springer,1973. Tuggener, L., R. Emberger, A. Ghosh, et al. ‘Real World Music Object Recognition.’ Transactions of the International Society for Music Information Retrieval 7, no. 1 (2024): 1– 14. https://doi.org/10.5334/tismir.157. Tuppen, S., S. Rose, and L. Drosopoulou. ‘Library Catalogue Records as a Research Resource: Introducing “a Big Data History of Music”.‘ Fontes Artis Musicae 63 (2016): 67–88. https://doi.org/10.1353/fam.2016.0011. Umbreit, J., and S. Schumann. ‘OMR on Early Music Sources at the Bavarian State Library with MuRET – Prototyping, Automating, Scaling.’ In Proceedings of the 6th Workshop on Reading Music Systems. Edited by J. Calvo-Zaragoza, A. Pacha, and E. Shatri. https://doi.org/10.48550/arXiv.2411.15741. Wiering, F. ‘Digital Critical Editions of Music: A Multidimensional Model.’ In Modern Methods for Musicology: Prospects, Proposals, and Realities. Edited by T. Crawford and L. Gibson. Ashgate, 2009. Wiering F., and C. Inskip. ‘The Impact of the Pandemic on Musicologists’ Use of Technology.’ Digital Humanities Quarterly 19, no. 2 (2025). https://dhq.digitalhumanities.org/vol/19/2/000786/000786.html. Wikimedia Foundation. ‘Corpus Linguistics.’ https://en.wikipedia.org/wiki/Corpus_linguistics. 65 COST ACTION 21161 Appendices Appendix 1: CORSICA proposal Name of host Frans Wiering Host Institution Utrecht University, Department of Information and Computing Sciences, Music Information Computing group Role of host and host institution Host will guide the STSM project, plan activities, organise online preparation meetings, the workshop, and follow-up meeting(s) if these are needed. The host institution will provide working space for the workshop. In addition to the host, 1-2 researchers from the host institution will participate in the workshop. Time period Workshop will take place either in the week of 13-17 May 2024. Online activities will take place in the weeks before and immediately after the workshop. Number of guests c. 7 guests, from multiple countries Working group WG2, Sources; WG3, Publications Objective of STSM project Analysis of large textual corpora has revealed new insights in many historical fields. Projects such as ECOLM and the Stanford Josquin Project show that a corpus-based approach also works in the music domain. There are important differences though. Whereas textual corpora have been created at a large scale and are relatively homogeneous, musical corpora are comparatively small and show a considerable variety in encoding practices. For pre-1700 polyphonic music, the situation is especially complicated: few corpora contain more than several hundreds of items; different types of music notation such as mensural notation and tablature pose specific encoding problems; and transcription to modern notation can be supported in a number of ways. Some of the consequences of this state of affairs are: • Available corpora are not a representative selection of the known repertoire; • Off-the-beaten-track research requires an excessive 66 COST ACTION 21161 amount of encoding effort; • Quality and interoperability issues increase when using multiple corpora; • PhD students and early-career researchers do not have sufficient resources to embark on a career as digital musicologists. On the other hand, there are some important positive developments: • The emergence of MEI (Music Encoding Initiative) as a versatile ‘musicological’ encoding system with a growing community of users; • Improvements in Optical Music Recognition for early music; • The general practice of musicologists to use music notation software to create their transcriptions; • A growing practice of online sharing of editions created by both professional and amateur musicians. In the past, the solution for the lack of encodings has often been sought in standardisation and training. Such endeavours have often met with limited success for at least four interconnected reasons: limited tool support; required time investment for learning; loss of already created work; and most importantly loss of autonomy. The logical alternative is therefore to consider autonomy and variety of practices as starting points for a new paradigm of corpus creation. Some elements of this are: • Contribute encodings with minimal effort; • Metadata for effective documentation of encoding practices; • Interoperability rather than standardisation; • A set of intuitive tools for common tasks; • Practical guidelines for sharing and exchanging encodings; • Involving music professionals and lay experts in the encoding process; • Making participation rewarding and fun. Realising these (and other necessary) ambitions will take a multi-year effort, which could take its first steps towards 67 COST ACTION 21161 maturity in the environment of the EarlyMuse COST action. We propose this STSM project named ‘Creation Of eaRly muSIc CorporA’ (CORSICA) as the first of these steps. A small group of experts will come together to: • Survey the current state of the art in corpus creation of early music; • Identify bottlenecks and opportunities; • Create a vision document for accelerating corpus creation; • Create an implementation plan with practical recommendations. Relation to WG and Action goals The project is related to the following goals and objectives: MoU, Secondary objective 3: preservation of musical heritage and accessibility of resources. As a logical next step after digitisation, corpus creation increases deep access to sources as well as a way of preserving the content, not just the appearance, of a source. MoU, Secondary objective 3: creating a multidisciplinary research network. By creating larger corpora in a more systematic fashion, it becomes easier use these data in an interdisciplinary context, where researchers may not have the specialist knowledge that is currently necessary to employ corpora fruitfully. Larger corpora also allow a wider range of research problems to be addressed, so that a network of researchers with multiple interest can emerge around them. For early-career researchers this is particularly beneficial as they are no longer obliged to create their own digital research environment from scratch. GAPG 3 (WG2): describe ways of identifying and describing endangered musical sources (including in Ukraine). There may be tenuous relation to this goal, in that encoding can be seen as protection of the content. For example, community initiatives that aim to digitise their own musical heritage could be supported by guidelines that make this as straightforward as possible and at the same time compatible with scholarly approaches. GAPG 4 (WG3): Imagining innovative publication models. Several connections seem relevant: 68 COST ACTION 21161 • The increasing importance of reproduceability of research. Publishing a corpus, or referring to an existing corpus as underlying data makes research easier to reproduce; • Corpora could be seen as publications: one’s contributions to a corpus become part of one’s academic portfolio; • Encodings relate to digital editions in two ways: they can be used as input to an editorial process in which they are further refined, for example by collating multiple sources of the same work; or the edited work can be stored as an analysable encoding. GAPG 6 (WG2): establish an information exchange protocol with RISM. CORSICA has a deep connection to RISM. RISM may provide essential metadata that improve the quality, usability and accessibility of the compositions in a corpus. It may also play an essential role in the selection of items to be encoded and in addressing questions of representativeness. Conversely, the act of encoding of encoding will produce new metadata that need to be compliant with RISM and that may in some cases be used to correct or enhance the RISM data itself. In fact, corpus creation is an ideal case study for GAPG 6 itself. Expected outcomes / deliverables The following outcomes are expected: • A report describing the current state of the art in corpus creation of early music • A vision on how to accelerate corpus corpus creation • A plan for implementing this vision Short summary CORSICA (Creation Of eaRly muSIc CorporA) will address the shortage of large-scale corpora of encoded early music. Existing corpora are generally small (usually less than 1000 items) and display a variety of encoding practices. This makes large-scale computational analysis of the music currently unfeasible. Experts in the field will come together in a 1-week workshop in Utrecht (Netherlands) to survey the current state of early music encoding; to identify bottlenecks and opportunities (in terms of technologies as well as communities); to create a vision for the future and to define practical steps towards accelerating corpus creation in early music. 69 COST ACTION 21161 Appendix 2: Early music corpora CORPUS DESCRIPTIONS corpus name musical content data creation method url Accessible Lute Music Lute music from Renaissance and Baroque accessible corpus citizen science https://wp.lutemusic.org/ Archive of Iberian Polyphony Iberian polyphony from the Renaissance digital editions, nonpublic encodings academic research https://iberianpolyphony.fcsh.unl.pt/ Choral Public Domain Library (CPDL) Choral music of any age encodings available for many works citizen science https://www.cpdl.org/wiki/ Citations: The Renaissance Imitation Mass (CRIM) 16th c. imitation masses and their models accessible corpus academic research https://crimproject.org/ Classical Archives Western composers, 14-20th century accessible corpus citizen science https://www.classicalarchives.com/midi.html Computerized Mensural Music Editing (CMME) 15-16th c. polyphony, with works by Josquin, Berchem and others accessible corpus academic research http://www.cmme.org/; https://github.com/tdumitrescu/cmme-music Electronic Corpus of Lute Music (ECOLM) Lute music from Renaissance and Baroque accessible corpus, password protected academic research https://ecolm.org Electronic Linked Annotated Unified Tablature Edition (E-LAUTE) Lute music from the German-speaking areas, mainly in German lute tablature accessible corpus academic research https://e-laute.info/ Electronic Medieval Music Score Archive Project (EMMSAP) 14th century music accessible corpus academic research https://github.com/cuthbertLab/emmsap ELVIS database Contains multiple subcollections, some also existing as separate corpora. Includes compositions printed in Glarean’s Dodekachordon, works by Palestrina, Victoria and others accessible corpus academic research https://github.com/ELVIS-Project Furnace and Fugue Digital edition of Atalanta fugiens (1618), music composed by Michael Meier digital editions, nonpublic encodings academic research https://furnaceandfugue.org/ Gaffurius Codices Online (MCE) Polyphony from c. 1500, with works by Compere, Weerbeke, Gaffurius and others digital editions, nonpublic encodings academic research https://www.gaffurius-codices.ch/s/portal/page/editions 70 COST ACTION 21161 corpus name musical content data creation method url Gaspar Online Edition Collected works of Gaspar van Weerbeke accessible corpus academic research http://www.gaspar-van-weerbeke.sbg.ac.at/gaspar-onlineedition Gesualdo online Complete works of Gesualdo accessible corpus academic research https://ricercar.gesualdo-online.cesr.univ-tours.fr/ Goldberg Corpus 15th and early 16th century music non-public encodings citizen science https://www.goldbergstiftung.org/en/ Johannes Tinctoris Complete Practical Works Practical works of Johannes Tinctoris, 3 pieces by Dufay, and 2 by Busnois accessible corpus academic research https://earlymusictheory.org/Tinctoris/Music/ John Robinson/Lute Society editions Lute music from Renaissance and Baroque accessible corpus citizen science https://github.com/TimCrawford/jhr_repo Josquin La Rue Secure Duo Dataset Duos from masses by Josquin and La Rue accessible corpus academic research https://github.com/ELVIS-Project/mass-duos-corpusjosquin-larue/tree/Methodologies-for-Creating-SymbolicMusic-Corpora Josquin Research project (JRP) Works by Josquin and contemporaries accessible corpus academic research https://josquin.stanford.edu/ KernScores Western composers, 14-20th century, with works by Frescobaldi, Monteverdi, Corelli and others accessible corpus mix http://kern.ccarh.org/ Kunst der Fuge Western composers, 14-20th century accessible corpus citizen science https://kunstderfuge.com/midi.htm Lassus Tricinium Project Tricinia by Orlande and Rudolph de Lassus accessible corpus academic research https://lassus.mh-freiburg.de/; https://github.com/WolfgangDrescher/lassus-geistlichepsalmen Lost Voices 16th c. French chansons accessible corpus academic research http://digitalduchemin.org/ Marenzio Online Digital Edition (MODE) Works by Marenzio non-public encodings academic research http://www.marenzio.org/index.xhtml Measuring polyphony 14th century motets accessible corpus academic research https://measuringpolyphony.org/; https://github.com/MeasuringPolyphony/mp-music-files Miami Publication Server (University of Münster) Digital editions of early 17th c. sacred music by Crüger and Textorius accessible corpus academic research https://miami.uni-muenster.de/ Music21 corpus Western composers, 14-20th century accessible corpus mix https://www.music21.org/music21docs/about/referenceCor pus.html Neuma Contains multiple subcollections, with works by Josquin, Ockeghem, Dufay, Frescobaldi and others accessible corpus academic research http://neuma.huma-num.fr/ Polish Digital Scores Polish music, 16th-20th century accessible corpus academic research https://polishscores.org/ 71 COST ACTION 21161 corpus name musical content data creation method url Single Interface for Music Score Searching and Analysis (SIMMSA) Heterogeneous corpus, aggregated from various other corpora accessible corpus academic research https://simssa.ca/activities/corpora-and-datasets/ Symbolically Encoded Il Lauro Secco (SEILS) Madrigals by various composers published in Il Lauro Secco (1582) accessible corpus academic research https://github.com/SEILSdataset Tasso in Music Musical settings of Tasso’s poems, with works by Monteverdi, Marenzio, Wert, Cifra, others accessible corpus academic research https://www.tassomusic.org/ The 1520s Project 1520s polyphony, with works by Willaert, Senfl, Verdelot and others accessible corpus academic research https://1520s-project.org/about/ Tomás Luis de Victoria Spanish music from Morales onward, including works by Guerrero, Victoria, Ceballos, Vasquez accessible corpus citizen science https://www.uma.es/victoria/index.html Verovio Humdrum Viewer Various encodings, mostly from other projects such as 1, 8, 9, 25 accessible corpus academic research https://verovio.humdrum.org 72 COST ACTION 21161 CORPUS SIZES AND ENCODINGS Supported formats are indicated by ‘x’; storage formats used to derive other formats from are indicated by ‘o’. corpus name items before 1700 ASCII tablature CMME Humdrum Lilypond MEI MIDI MP3 MuseData MusicXML/ Finale PDF Sibelius Other Accessible Lute Music c. 17000 o x x Archive of Iberian Polyphony 130 x o x o Choral Public Domain Library (CPDL) c. 30000 x x x x x Citations: The Renaissance Imitation Mass (CRIM) c. 300 o x Classical Archives several 100s o Computerized Mensural Music Editing (CMME) c. 260 o Electronic Corpus of Lute Music (ECOLM) 1619 o Electronic Linked Annotated Unified Tablature Edition (E-LAUTE) 60, growing o Electronic Medieval Music Score Archive Project (EMMSAP) 3600 x ELVIS database c. 1000 x x x x Furnace and Fugue 50 o x o Gaffurius Codices Online (MCE) c. 55 o Gaspar Online Edition c. 100 x x o Gesualdo online 222 o x Goldberg Corpus 3600 o x Johannes Tinctoris Complete Practical Works c. 50 x o John Robinson/Lute Society editions c. 8000 o x x Josquin La Rue Secure Duo Dataset 77 x x x o Josquin Research project (JRP) 902 o x x x x x x x KernScores c. 350 o x Kunst der Fuge several 100s o Lassus Tricinium Project 50 o x Lost Voices c. 350 x x o Marenzio Online Digital Edition (MODE) 32 o x Measuring polyphony 61 o x x x 73 COST ACTION 21161 corpus name items before 1700 ASCII tablature CMME Humdrum Lilypond MEI MIDI MP3 MuseData MusicXML/ Finale PDF Sibelius Other Miami Publication Server (University of Münster) c. 230 x o Music21 corpus several 100s x x x Neuma c. 600 x x x Polish Digital Scores 1740 o x x x x x Single Interface for Music Score Searching and Analysis (SIMMSA) ??? x x x x x x Symbolically Encoded Il Lauro Secco (SEILS) 30 x o x x x Tasso in Music 778 o x x x x x x x The 1520s Project 491 x x x x x o Tomás Luis de Victoria c. 1000 o x x Verovio Humdrum Viewer c. 1500 o x x