scieee AI-readable full text Open interactive document viewer

Establishment of an open digital collection of historical music theory texts from German-speaking countries based on 19th-century examples (DigiMusTh) - Project Documentation and Guidelines

Moss, Fabian; Roeder, Torsten; Klinger, Jana; Keupp, Corinna; Roth, Janina Marie

Full text

Establishment of an open digital collection of historical music theory texts from German-speaking countries based on 19th-century examples (DigiMusTh) Project Documentation and Guidelines Fabian C. Moss1,(0000-0001-9377-2066), Torsten Roeder(0000-0001-7043-7820), Jana Klinger(0009-0008-5566-1101), Corinna Keupp(0009-0007-9300-7189), & Janina Marie Roth(0009-0004-7348-0903) Institut für Musikforschung & Zentrum für Philologie und Digitalität Julius-Maximilians-Universität Würzburg, Würzburg, Germany Abstract The DigiMusTh project focuses on the digital collection and analysis of historical music theory texts, with a particular emphasis on the 19th-century debate on major-minor dualism. While large digital collections of music theory texts already exist in other languages—such as the Thesaurus Musicarum Latinarum (TML) and the Thesaurus Musicarum Italicarum (TMI)—there is currently no comparable collection for German-language texts, despite their central role in European music theory since the 18th century. DigiMusTh aims to fill this gap by creating a sustainable and expandable digital edition that offers new perspectives for both musicology and the digital humanities. A key aspect of the project is the integration of different types of texts. In addition to monographs and treatises, theoretical journal articles will also be included to provide a more comprehensive picture of the historical debate. Relevant periodicals, such as the Neue Zeitschrift für Musik, the Allgemeine Musikzeitung, and the Neue Berliner Musikzeitung, are already available in digital form. Since a complete collection is not feasible, a targeted selection will be made: only articles that are thematically linked to the existing corpus—either through shared authorship or direct references—will be included. This approach will offer a broader perspective on the discourse of the time and its dissemination in the public sphere. Another goal is the technical and methodological advancement of digital music theory editing. The complex structure of the sources—including text, images, and musical elements—presents challenges that require innovative solutions. The continuous expansion of the collection will also contribute to the refinement of encoding guidelines and the development of music-theoretical ontologies. Additionally, the integration of individual research projects, such as master’s theses or doctoral dissertations on relevant texts, is possible after review, allowing for a collaborative expansion of the collection. In the long term, DigiMusTh aims to become an open, cooperatively expandable platform for historical music theory texts. Due to its interdisciplinary approach, the project is not only relevant to music theory and musicology but also to the digital humanities and the history of science, particularly concerning the physical and physiological foundations of harmonic theory. With its sustainable structure, DigiMusTh contributes to making historical music theory accessible and analyzable in modern digital formats. Funded by Text+, a consortium of the National Research Data Infrastructure (NFDI). Proposed citation: Moss, F. C., Roeder, T., Klinger, J., Keupp, C., & Roth, J. M. (2025). Establishment of an open digital collection of historical music theory texts from German-speaking countries based on 19th-century examples (DigiMusTh) – Project Documentation and Guidelines. Zenodo. https://doi.org/10.5281/zenodo.17974934 1 Contact: [email protected] Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project Table of Contents Table of Contents___________________________________________________________________________ 2 1 Introduction – digital editing of music-text documents__________________________________________3 1.1 Project background___________________________________________________________________ 3 1.2 Project goal_________________________________________________________________________ 3 1.3 Accessing the project materials_________________________________________________________ 3 1.4 Project team_________________________________________________________________________3 1.5 The texts____________________________________________________________________________4 2 Workflow – preparing and transforming the data______________________________________________ 4 2.1 Finding and scanning texts_____________________________________________________________4 2.2 Transcribing texts with OCR____________________________________________________________5 2.3 Transforming export to TEI-XML________________________________________________________ 5 2.4 Data Model_________________________________________________________________________ 6 2.6 Enhancing Data_____________________________________________________________________10 2.7 Integrating the collection to existing databases__________________________________________ 13 3 Website – publishing a digital edition online with minimal technical requirements________________14 3.1 Tech Stack__________________________________________________________________________14 3.2 Displaying the texts online____________________________________________________________14 3.3 Indices____________________________________________________________________________ 19 3.4 Search____________________________________________________________________________ 20 3.5 Whitespace issues___________________________________________________________________ 21 3.6 Adding texts to the collection_________________________________________________________ 22 4 Dissemination – communicating project outcomes to the scholarly community___________________ 23 4.1 Academic conferences_______________________________________________________________ 23 4.2 Invited talks and presentations________________________________________________________24 References________________________________________________________________________________ 25 Page 2 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project 1 Introduction – digital editing of music-text documents This report documents the rationale, workflow, and outputs of the scholarly research project Establishment of an open digital collection of historical music theory texts from German-speaking countries based on 19th-century examples (DigiMusTh) that involved the digital editing of combined music-text documents. Apart from being a documentation of the work accomplished within the project, this document is also meant to serve as guidelines for scholars with similar research interests. We thus adopt at times a writing style that directly addresses the reader in the second person, with the goal in mind to facilitate the adaptation of our workflow as much as possible. The procedures described below should thus aid in helping scholars add further texts concerning the dualism debate to this collection. Moreover, it can also provide a scaffolding for a vast range of other research interests based on multimodal texts, and can thus be useful for a broad community within the digital humanities. 1.1 Project background The DigiMusTh project ran from January to December 2025 at the Center for Philology and Digitality (Zentrum für Philologie und Digitalität, ZPD) of Julius-Maximilians-Universität Würzburg (JMU), Germany. It was funded by Text+ (Kett et al., 2022),2 a consortium within the German National Research Data Infrastructure (Nationale Forschungsdateninfrastruktur, NFDI; Kraft et al., 2021) as a so-called cooperation project within the Text+ Task Area “Collections”. 1.2 Project goal The main goal of the project was to provide an open collection of music theory-related texts digitally encoded according to modern encoding standards, collated around the topic of the so-called “dualism debate”. The DigiMusTh project is thus a direct successor of the previous project Digitizing the Dualism Debate (DDD; Moss et al., 2021)3 that ran in 2021 as a collaboration between the École Polytechnique Fédérale de Lausanne (EPFL) and Université de Lausanne (UNIL) in 2021. 1.3 Accessing the project materials During the year-long project duration, the team enhanced the texts taken from the DDD project and built a website displaying texts containing music notation, mathematical expressions, and diagrams. The texts are accessible with any web browser at the URL of the project website or via the “data” branch of the project GitHub repository. ● Project website: https://fabianmoss.github.io/digimusth/ ● Project GitHub repository: https://github.com/fabianmoss/digimusth/ 1.4 Project team Prior to the project, Felicitas Stickler4 developed a first prototype of the website as part of her bachelor’s thesis for the Digital Humanities programme at JMU.5 The project team consisted of three research associates sharing a 65% part-time position: Corinna Keupp, Jana Klinger, and Torsten Roeder. They were supported by one student assistant with 7 hours per week, Janina-Marie Roth. Project leader Fabian C. Moss coordinated their efforts and steered the project. 5 https://www.uni-wuerzburg.de/zpd/studium/bachelor/ 4 https://felicitasstickler.github.io/ba-thesis/ 3 https://dcmlab.github.io/ddd/ 2 https://text-plus.org/ Page 3 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project 1.5 The texts The project builds on the text collection provided by the DDD project, which had selected eleven texts, each of which constitutes an important contribution to the dualism debate. The following table gives an overview with the titles, authors, and years of publication of the texts. In the remainder of this documentation, we will frequently refer to the texts by their IDs stated in the last column. No. Title Author Year ID 1 Der accordliche Gegensatz und die Begründung der Scala Otto Kraushaar 1852 KRA1852 2 Die Natur der Harmonik und der Metrik: Zur Theorie der Musik Moritz Hauptmann 1843 HAU1853 3 Über die verschiedenen Bestimmungen der Tonverhältnisse und die Bedeutung des pythagoreischen oder reinen Quint-Systems für unsere heutige Musik Carl Ernst Naumann 1858 NAU1858 4 Harmoniesystem Carl Friedrich Weitzmann 1860 WEI1860 5 Die neue Harmonielehre im Streit mit der alten Carl Friedrich Weitzmann 1861 WEI1861 6 Kritische Beleuchtung des C. F. Weitzmann’schen Harmoniesystems Franz Joseph Kunkel 1863 KUN1863 7 Harmoniesystem in dualer Entwicklung Arthur von Öttingen 1866 OET1866 8 Die beiden Tongeschlechter und die neuere musikalische Theorie Adolf Thürlings 1877 THU1877 9 Die Lehre von den musikalischen Klängen: ein Beitrag zur aesthetischen Begründung der Harmonielehre Otakar Hostinský 1879 HOS1879 10 Die Zukunft der Musiktheorie Georg Capellen 1905 CAP1905 11 Das Problem des harmonischen Dualismus Hugo Riemann 1905 RIE1905 2 Workflow – preparing and transforming the data This section specifies the different steps within the workflow of finding a relevant text for a collection, scanning it, transcribing it, developing a data model that fits the needs of the text and its intended presentation, transforming it to this data model, and enhancing the data with additional information or through the explicit modelling of formerly untranscribed parts of the text (such as mathematical expressions, music notation, or diagrams). 2.1 Finding and scanning texts In the beginning of building a textual collection one has to select and find relevant texts. These texts can then be digitized by scanning the pages. If you are scanning the pages yourself you need to make sure that the quality of the scans adheres to typical standards. The German Research Community (Deutsche Forschungsgemeinschaft, DFG) published guidelines regulating the digitization of sources for humanities Page 4 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project research (Altenhöner, 2023). For information on technical standards for scanning texts refer to Section 3.2 of those guidelines. If you have found a library or another institution that does the digitization for you, make sure they know about your/your field’s technical requirements for the scans. If the text has already been scanned you need to find out whether these scans are free to use, or if a permission for usage in your project can be obtained from the rightholder. If you are planning to add a text to our DigiMusTh text collection, a list of possibly interesting texts can be found here: https://dcmlab.github.io/ddd/sources.html. 2.2 Transcribing texts with OCR Once you have scans of the selected texts, you need to transcribe them in order to make them computer-readable. There exist many tools that are able to do the bulk of the transcription process for you using automated procedures called Optical Character Recognition (OCR). OCR describes the process of automatically extracting textual information from visual input (e.g., scans) to a symbolic format (e.g., text). Depending on whether you need to transcribe hand-written or machine-written text some tools will work better than others. Mind that the used script (e.g. black letter vs roman) needs to be factored in. The DDD project used Transkribus (Kahle et al., 2017)6 for OCR, but there are other possible choices of tools, such as eScriptorium (Kiessling et al., 2019),7 OCR4all (Reul et al., 2019)8 or LAREX (Reul et al., 2017),9 which all have different advantages and disadvantages. It is highly recommended to try out different softwares and models to find out which one works best for you and your source material. To avoid too much manual work in later steps of the workflow, we recommend to mind a few caveats: Before transcribing the actual text, OCR software usually divides each individual page into a number of zones or regions (e.g. areas on the page that have text vs. areas that do not), a task that is commonly referred to as layout recognition. It has proven useful to supervise the automated process as it might be error-prone depending on the quality of the source scans. Moreover, it is beneficial to define and label a small set of different region types such as headings, paragraphs, footnotes, page numbers, music notation, graphics etc. For example, define a text-region in your OCR software called “heading” and use it to mark all headings. Marking regions that do not contain text allows one to mark places on a page that contain a graphical/pictorial/non-textual element for later reference, while also making it clear for the OCR software that it does not need to attempt to transcribe this part of the page, thus reducing the overall error rate. For each form of textual and non-textual phenomenon found in the texts there should be one and only one kind of zone. This is important, so that after exporting the transcription as XML the different types can be identified and handled differently without manually inspecting each instance and comparing it to the scans. Once the manual marking of the page areas is done, the transcription process can be conducted. Afterwards, the automated transcriptions should be corrected manually to fix (most) mistakes. When the quality of the text is satisfactory, they can be exported into the PAGE XML format that encodes the marked up page layout in a standardized manner. 2.3 Transforming export to TEI-XML Such a PAGE XML export is not easily human-readable and should thus be further transformed using, for example, eXtensible Stylesheet Language Transformations (XSLT) or general-purpose programming languages such as Python. When applying data transformations, it is important to keep in mind that one should keep as much information as possible because one might not anticipate all possible use cases of 9 https://www.uni-wuerzburg.de/zpd/larex/ 8 https://www.ocr4all.org/ 7 https://gitlab.com/scripta/escriptorium 6 https://www.transkribus.org/de Page 5 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project one’s own data. Others might be interested in aspects that one self would discard. This means concretely to keep all the information on the zones within the page that were defined earlier (coordinates, names, text zone or non-text zone and so on). Line beginnings and coordinates within larger zones might come in handy at a later point. The de facto scholarly standard for digital text encodings is the XML-based format developed by the Text Encoding Initiative (TEI).10 We thus transformed the PAGE XML files (in our case provided by the DDD project) into valid TEI-XML files.11 After transforming the PAGE XML to TEI, a central problem has to be tackled: When exporting a XML from an OCR software the typical structure is page-based, as the OCR transcription process is happening page by page. For the purposes of our text collection this page-based structure had to be changed to a structure that reflects the logical organization of the texts into sections, subsections, paragraphs, etc. This entails that paragraphs or footnotes spanning multiple pages have to be combined. Moreover, footnotes have to be attached to their position within the text so that they are coordinated in the absence of a page-based structure. Transforming the page-based structure into a content-based structure can be tricky, and we used a mix of XSLT-programming, regular expressions and manual adjustments for our project. If you want to transform an XML-file with the goal of adding it to our text collection, you need to be able to validate the XML against the TEI all-schema. Beyond that, the better your data model fits ours, the better displaying your text on our website will work (e.g. display of music examples, graphics, mathematical expressions). For a list of all the used elements and their meaning see the table in the following chapter. 2.4 Data Model All texts in our collection are encoded in TEI-XML. The data model follows the logical structure of the text instead of the structure of the individual pages, and the general structure of the documents is typical for TEI-XML: 1. Within the <TEI> root element the <teiHeader> with all the metadata on the text and its encoding can be found. 2. Following the <teiHeader> comes a section with a number of <facsimile> elements which contain information on the individual pages the text was transcribed from. For each scanned page there is a <facsimile> element with an unique @xml:id. It further contains a <surface>-element that is in turn comprised of a <graphic> element that links to the image of the scanned page as well as a number of <zone> elements, each one defining the area on the linked image in which different types of text or non-text are located. The @points within <zone> defines the area on the linked image and the @rendition defines whether the zone contains a text region, music notation or graphic. In @subtype all text regions are further defined as paragraphs, headings, lists, or footnotes. 3. After the <facsimile> elements comes the <text> element in which the actual transcribed and encoded text is placed. Here, the title pages, table of contents and preface can be found in the <front>, followed by the main content of each text in the <body> and possibly by a postface in the <back>. The rest of the elements and attributes used in the collection are summarized in the table below, along with examples and comments on specific use cases. 11 The transformed TEI-XML files from the DDD project might be a useful example to orient yourself towards. They can be found in the projects GitHub repository. There navigate to /data. In there you find a folder for each text (e.g. CAP1905). In each of these folders you can find a .xml-file with ‘_tei’ in the file name. If you want to use these files as orientation, also refer to the Data Model-chapter of this text. 10 https://tei-c.org/ Page 6 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project Use case Tag Example Comment title page <titlePage> Title pages were encoded following the rules of the DTABf. Their guidelines can be found here: https://www.deutschestextarchiv.de/doku/basisformat/tbAllg.html chapters <div> <div type=”chapter” n=”1”> The @n attribute refers to the level of a chapter, e.g. n=”1” for a first-level chapter, whereas subchapters would have n=”2” etc. headings <head> <head facs=”#zone-id”> Heading </head> The <head> element always immediately follows the div@type=”chapter”. The @facs contains a link to the <zone> of the corresponding text on the scanned page. line beginning <lb/> <lb facs=”#zone-id” n=”001” /> This empty element marks the beginning of a new line. @n contains the line number on the page. page beginning <pb> <pb facs=”#facs_13” n=”9”/> This empty element marks the beginning of a new page. @n contains the page number. If there is no page number printed on the page, then @n remains empty. hyphens <pc> vor<pc>¬</pc><lb/>geführt Hyphens at the end of a line, linking two parts of the same word, are marked with the <pc> (punctuation character) element and enable the display of the text in forms independent from the original line breaks. paragraphs <p> <p facs=”#zone-id”> text text text.</p> This element is used to mark each paragraph. It may contain multiple such links separated by a space, if the paragraph spans across multiple pages. bold, spaced, italic, subscript, superscript <hi> <hi rend=”bold:true;”>bold text</hi> @rend has five different values: “bold:true;”, “italic:true;”, “letterSpaced:true;”, “superscript:true;”, “subscript:true;” music notation <notatedMusic> <notatedMusic rend=”MEI” The element contains a <desc>-Element with fixed content “MUSIK” so that it will be displayed in case of issues with rendering the MEI file, making it easier to look Page 7 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project facs=”#zone-id” source=”zone-id.mei”> <desc>MUSIK</desc> </notatedMusic> for missing music notation in the frontend. With @rend=”MEI” the frontend knows to treat this element and its contents differently than the other elements. More on that in the chapter on displaying music notations below. The @source contains a link to the file containing MEI-encoded music notation, named after the music notation’s zone-id. graphics <figure> <figure facs=”#zone-id”> <graphic url=”zone-id.png”> </figure> Graphics from the text are captured as Portable Network Graphics (PNGs) and linked within the <figure> element as <graphic>. @url contains a link to the PNG file, named after the graphic’s zone-id. @rend=”inline” was used to differentiate between inline and block graphics. In rare cases of music notation not being transcribed but added as a graphic, those were marked with @type=”musicplaceholder”. mathematical expressions <formula> <formula notation="tex">$266 \frac{2}{3}$</formula> Mathematical expressions are transcribed using LaTeX notation. This change in encoding style is marked by @notation. footnotes <note> <note facs=”#zone-id” type=”foot” n=”*)”>footnotetext</note > Footnotes are encoded as <note type=”foot”>. The @n contains the characters used to reference the footnote in the text. lists <list> <list facs=”#zone-id”><item>ite m1 text</item><item>item2 text</item></list> Lists are encoded as <list> with an <item>-element containing the text of each point on the list. The @facs contains a link to the <zone> of the corresponding text on the scanned page. It may contain multiple such links separated by a space, if the list spans across multiple pages or was divided into different zones during the transcription process. music-theoret ical terms <rs> <rs ana=”#concept_mus>Tongest altungen</rs> The <rs> element encodes parts of the text of special interest for the dualism debate. The @ana has one of the following values: “#concept_mus”, “#concept_phy”, “#m_interval”, “#m_note”, “#m_chord”, “m_tonality”. Page 8 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project mentions of texts/works <title> <title ana="#text" ref="#W0042">Tonsetzkunst </title> When texts and musical works are mentioned, their titles are encoded as <title ana=”#text”>. @ref is used to link to another XML file containing a list with all mentioned works together with further metadata: This XML file can be found here: https://github.com/fabianmoss/digimusth/blob/data/tei_logische-struktur/works.x ml mentions of authors <persName> <persName role="author" ref="#P0041">Weitzmann</p ersName> When authors are mentioned, their names are encoded as <persName role=”author”>. @ref is used to link to another XML-file that contains a list with all mentioned authors together with further metadata. This XML-file can be found here: https://github.com/fabianmoss/digimusth/blob/data/tei_logische-struktur/people. xml mentions of places <placeName> <placeName ref="#L0015">Frankfurt a. M.</placeName> When places are mentioned, their names are encoded as <placeName>. @ref is used to link to another XML-file that contains a list with all mentioned locations together with further metadata. This XML-file can be found here: https://github.com/fabianmoss/digimusth/blob/data/tei_logische-struktur/locatio ns.xml Page 9 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project parent and format them in a smaller font size, which might also increase readability for users, as example see an extract form KUN1863 below: Rendering music notation To display music on our webpage we decided to use Verovio (Pugin et al., 2014). The basic implementation to display MEI files as Scalable Vector Graphics (SVG) is straightforward. However, in our case many examples were cut off as the size and scope of examples varies greatly: some spanning multiple pages while others only showcase a few notes. Therefore we needed a dynamic execution dependent on the specific files. In the first step all occurrences of notated music are collected. const musicElements = document.querySelectorAll('tei-notatedmusic[rend="MEI"]'); Out of each node the @source attribute is then extracted that stores the filename. After building the path to the MEI file, the file is loaded, analyzed, and—depending on the number of pages—height and scale are adjusted. vrvToolkit.loadData(data); const pageCount = vrvToolkit.getPageCount(); let scale = 80; let pageHeight = 12000; // Adjust scale/pageHeight for readability Page 16 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project if (pageCount === 1) { scale = 80; pageHeight = 6000; } else if (pageCount <= 3) { scale = 100; pageHeight = 12000; } else { scale = 120; pageHeight = 20000; } vrvToolkit.setOptions({ adjustPageHeight: true, breaks: "auto", pageWidth: 2100, pageHeight, scale }); The SVG is then rendered and Verovio produces an inner and an outer SVG that influence the sizing and need to be taken into consideration for the display. const svg = vrvToolkit.renderToSVG(pageNum, {}); When no MEI file can be found the code automatically falls back to display the PNG version. Therefore, naming needs to be exact and the folder structure needs to follow the pre-determined pattern. For users to be able to display the musical examples on a larger scale we added a function that opens the data in a new tab. Since the music is generated as SVG rather than simply loaded in, this data is stored in a blob.31 Displaying Mathematical Expressions Mathematical expressions are rendered using MathJax with some minor adjustments in the configuration (see code below). We changed characters marking the beginning and end of a formula, allowing the $ symbol as tag. Further, we disabled the built-in MathJax right-click menu and chose not to load this library over a Content Delivery Network, but save the data locally. MathJax = { tex: { inlineMath: { '[+]': [['$', '$']] } // adding the $ as a marker for formulas }, options: { enableMenu: false // disables the context menu } }; Additionally, the script checks for mathematical expressions as they are loaded onto the page. if (window.MathJax) { // Tell MathJax to process newly inserted content MathJax.typesetPromise(); } As there are some expressions that resemble mathematical fractions but that are actually figured-bass notation we chose to use CSS for visualization, e.g. KUN1863: 31 https://wiki.selfhtml.org/wiki/JavaScript/Blob Page 17 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project <lb facs="#facs_57_line_1614708594662_928" n="N003"/>dann in den durch zwei <rs ana="#concept_mus">Vorhalte</rs> <rs ana="#concept_mus">verzögerten</rs> <rs ana="#m_chord"><hi rend="italic:true; fontSize:0.0; kerning:0;">G</hi>-dur-Dreiklang</rs> (in der <rs ana="#concept_mus"><rs ana="terzsextlage"><span type="figured-bass" rend="stacked"><span>6</span><span>3</span></span>-Lage</rs></rs>) auf. Das <hi rend="italic:true; fontSize:0.0; kerning:0;"><rs ana="#m_note">h</rs></hi> des <rs ana="#concept_mus">Basses</rs> bleibt Displaying Graphics Other images that are saved as Portable Network Graphics (PNGs) are either loaded as image blocks (centered horizontally in the text) or inline. In the latter case, they are rendered with a maximum height of 3em. These are usually single symbols or text in a foreign language, mostly Greek (see CAP1905, p. 9 below). As images might be hard to read due to their size or position we added an onclick function that opens an image in a new tab. Since we have fallback images, images in footnotes, and regular images we need this functionally to work on multiple elements and therefore be included at different positions of the code. See below the code to add the function onto the SVGs and the images. const svgString = new XMLSerializer().serializeToString(svgClone); // Create a blob and open in new tab -> blob needed because svg gets generated and isnt saved as such const blob = new Blob([svgString], { type: 'image/svg+xml' }); const url = URL.createObjectURL(blob); window.open(url, '_blank'); Code for images: // Create a link wrapper so the image is clickable const link = document.createElement("a"); link.href = pngPath; link.target = "_blank"; // open in a new tab link.appendChild(img); // Replace the node contents with the clickable fallback image Page 18 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project node.innerHTML = ""; node.appendChild(link); Displaying page numbers and linking to facsimiles In the footer we added a button on the left-hand side displaying the current physical page number. Clicking on it opens the scan in a new window. For this we extract the <pb> elements. However, sometimes there is no value assigned to the @n attribute of a <pb> element, which can distort the ordering of displayed numbers. In OET1866 for example, the <facs> element with the number 42 can be found on the physical page with the number 34, which as a digital scan in JPG format has the number 51. We recommend citing using the facs element. Table of contents To navigate within the texts we use fist-level headings and implement them through a combination of the already existing structure in Ed and CETEIcean. Ed automatically adds a table of contents (toc) in the navigation bar. If the text is loaded from this markdown file the template automatically adds IDs to the text where headings are defined. It is therefore important that the spellings of headings in the text and the toc element are identical. To add the ID the anchor jumps to in our through XML loaded text we use a CETEIcean behaviour (see code below) that adds a <div> element in front of head elements with an ID. This ID is generated from the text in the head, therefore following the predetermined Ed logic. Further this structure technically allows the user to jump to all possible <head> elements. Due to a high number of subtitles we chose to only add fist-level heads. "head": function (elt) { // Clone the element and move it into a div let clone = document.createElement("div"); // Copy over the inner HTML so the TEI content displays clone.innerHTML = elt.innerHTML; // Find the first <hi> in the clone let hi = clone.querySelector("tei-hi"); if (hi) { // Extract and clean up its text, following the ed logic let text = hi.textContent.trim(); let cleaned = text .replace(/[.,\/#!$%\^&\*;:{}=\_`~()?"'””“„’]/g, "") .replace(/\s+/g, "-") .toLowerCase(); // Add the id attribute clone.setAttribute("id", cleaned); } // Return the new HTML structure so CETEIcean renders it return clone; }, 3.3 Indices In the texts terms have been annotated into various categories. While authors and works were already equipped with IDs, all the other classes were not. In the course of this project we developed a script to add Page 19 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project IDs to categories and chose to apply it on the locations since the number of occurrences is still manageable, for the concepts the occurrences are in the thousands and normalizing them was outside of a reasonable scope. Therefore the webpage offers three index pages, which presents users the occurrences in context and allows users to jump to the text extract to read more. From a technical standpoint, the indices are saved as XML files, following TEI conventions. These files are then loaded onto the page using JavaScript, an extension of these files therefore automatically updates the data on the online site. In the base view the user sees an alphabetically sorted list of the entities, a click on an entity opens a new page. This page displays the external norm database identifier, as well as notes connected to the entity. Further the occurrences of the entity is seen in context for all texts in the collection and through a link, the user can jump to the highlighted occurrence in the fulltext. The works can be found in their own index, but if they could be connected to an author, they can also be found in the people index. 3.4 Search The Ed template comes with an already configured search based on the elasticlunr implementation32 by Kathie Decora, which we have adapted for our project, as well as the lunr languages script33 by Mihai Valentin and the lunr stemmer by Oleg Mazko. However this search tool is not usable out of the box. Firstly the code expects the texts to be in the markdown files. For this project the text page is generated and various sources are loaded in, therefore the text content is in the XML files. To be able to use the search we have chosen to build the index manually and place it in the markdown files. When this initial search has found a string in the index, the full text then is searched to find the string in context. While this approach may seem overkill for our small collection it might come in handy when a lot of texts have been added. 33 https://github.com/MihaiValentin/lunr-languages 32 http://katydecorah.com/code/lunr-and-jekyll/ Page 20 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project Secondly the template initially could not handle German texts therefore we updated the code with a version that can support German. For example, when searching for “Piano,” several results are shown. Below you see the search interface and an excerpt of the result in KRA1852 that is also reachable via a direct link.34 When a term is found in the static index, the entire XML file is searched for its occurrences. These are then shown in context and a user can click on it to be guided to the actual text. 3.5 Whitespace issues The initial digital data was encoded true to the physical book including line breaks. However a presentation on a webpage allows multiple formats and users to view the page on diverse devices. The reading experience should be pleasant for all. During this project we had various issues with whitespaces: either missing whitespaces, leading to words being glued together, or too many whitespaces, ripping words apart. During close analysis different cases could be identified and fixed accordingly. A common cause could be found in missing white spaces between certain tags, this was fixed using various regular expressions that systematically added whitespace between the identified tags. A main issue were lines ending in a punctuation character element (<pc>) that led to inserted whitespaces. Punctuation marks are not displayed online and controlled via CSS rules. However if a whitespace character appeared it was fixed using CETEIcean behaviour. While it was important to remove the linebreak display in the fulltext, for some parts, such as lists, line breaks are very important for the semantics of the text. See for example NAU1858, p. 14: 34 https://fabianmoss.github.io/digimusth/texts/KRA1852/?q=Piano#facs=%23facs_18_TextRegion_1632819230547_899 Page 21 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project 3.6 Adding texts to the collection All following steps have to be done on the “page” branch35 for a new text to appear on the webpage. In order to add a text to the collection, a markdown file must be created for it, containing specific information, which is reused at certain points in the code. It must contain the following structure: - layout: music - title - author - name - year - search - (optional toc) An example would be: --- layout: music title: Die neue Harmonielehre im Streit mit der alten author: Carl Friedrich Weitzmann name: WEI1861 year: 1861 search: abgedruckt, abgegrenztes, abhandlung, abschlusse, ... übrigens --- The layout: music key-value pair must be kept, since it contains all functions needed to display the music data and correctly load the XMLs. The table of content can be extracted using a custom Jupyter notenbook,36 and a script automatically adds IDs to the elements marked as head in the text. As the strings must match for the logic to work we advise not to do this manually. Here is an example for RIE1905: toc: - I. Einleitung. - II. Sind die Obertöne der „Grund” der Konsonanz? - III. Intervallverschmelzung oder Klangvertretung? - IV. Die wahre Wurzel des harmonischen Dualismus. - V. Das Grundton-Problem. For the static search we need to have the content words in the markdown as the pages are generated from the XML files. You can use another custom notebook37 to extract unique words from XML files and then add them to the specific markdown file. Both scripts are found in the data branch. The markdown file then is placed in the _texts folder, while a folder named after the text’s ID is placed in data, following the structure in the following image: 37 https://github.com/fabianmoss/digimusth/blob/data/python/extract_searchterms.ipynb 36 https://github.com/fabianmoss/digimusth/blob/data/python/extract_heads.ipynb 35 https://github.com/fabianmoss/digimusth/tree/page Page 22 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project The MEI files are placed in data, and their fallback PNG images in music, while all other images are deposited in graphic. Facsimiles are found in scans. Some texts have an additional folder called table which contains PNG images of tables. When a text is added this way, it will automatically show up in the overview of the texts and will be searched through. Once you have followed all the steps open a pull request which can be merged after a review by the repository’s owner. 4 Dissemination – communicating project outcomes to the scholarly community The progress and the outcomes of the DigiMusTh project have been presented at various occasions in order to obtain feedback from the scholarly community and to disseminate our results. While the project focused on a particular collection of texts—19th century German music theory treatises—, we believe that many of its aspects, such as the multimodal nature of the documents with text, music, tabular, and mathematical elements, as well as our decision to work within a minimal editing framework based on static rendering of the website can be adapted for many other scenarios in the digital humanities. 4.1 Academic conferences TEI/MEC conference 2023 At the joint conference of the Text Encoding Initiative and the Music Encoding Initiative in 2023 in Paderborn, Germany, we presented preliminary work on text-music interlinking, which played an essential role in this project. The feedback obtained at the conference was instrumental in preparing the proposal. 1. Roeder, T., Köster, M., & Moss, F. C. (2023). Music-Text Interlinking as a Challenge for Digital Encodings of Music-Theoretical Writings. Encoding Cultures – Joint MEC and TEI Conference 2023, 4–8 September 2023, Zentrum Musik – Edition – Medien (ZenMEM), Paderborn, Germany. https://teimec2023.uni-paderborn.de/contributions/124.html MEC conference 2025 During the Music Encoding Conference 2025 in London (from June 3rd to June 6th) the preliminary studies and the current status of the project have been presented (“A Minimal Publishing Model for Text and Music Notation”). Minimal examples are available on the TEI Music SIG website: github.com/TEI-Music-SIG/examples Page 23 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project 2. Stickler, F., Roeder, T., & Moss, F. C. (2025) A Minimal Publishing Model for Text and Music Notation. Music Encoding Conference 2025. 3–6 June 2025, London, UK. TEI conference 2025 During the Text Encoding Initiative Conference 2025 in London (from September 16th to September 20th) the current status of the project was presented (“Multimodality and Minimal Publishing: TEI, MEI and more in 19th-Century Music Treatises”). Expanded minimal examples are available on the TEI Music SIG website: github.com/TEI-Music-SIG/examples 3. Roeder, T., Klinger, J., Stickler, F., Keupp, C., & Moss, F. C. (2025) Multimodality and Minimal Publishing: TEI,MEI and more in 19th-Century Music Treatises. 25th Annual Meeting of the Text Encoding Initiative. 16–17 September 2025, Kraków, Poland. https://10.5281/zenodo.17205769 4.2 Invited talks and presentations Coffee Talks at ZPD On Wednesday 21st of May 2025 Felicitas Stickler presented her work to colleagues at the Centre for Philology and Digitality in Würzburg. The presentation focused on her Bachelor thesis and a project realized during her studies. Following an insightful talk, she responded to questions from an engaged audience. 4. Stickler, F. (2025). Coffee Talks at ZPD, Würzburg, Germany. Text+ Kooperationsprojekte Fabian Moss introduced this project on the 28th of May 2025 during a meeting of all current Text+ projects in the field of organizing collections. 5. Moss, F. C. (2025, May 23) Vorstellung Kooperationsprojekt “Aufbau einer offenen digitalen Sammlung historischer musiktheoretischer Texte aus dem deutschsprachigen Raum anhand von Beispielen aus dem 19. Jahrhundert (DigiMusTh)”. 14. Gesamttreffen TA Collections Text+, online. DH Kolloquium der BBAW 6. Moss, F. C. (2025, June 30). Text+ Musik: Multimodale Kodierungsherausforderungen im DigiMusTh-Kooperationsprojekt. DH-Kolloquium an der Berlin-Brandenburgischen Akademie der Wissenschaften, online. Abschlussevent Text+-Kooperationsprojekte von 2025 7. Moss, F. C. & Roeder, T. (2026, February 17). DigiMusTh: Aufbau einer offenen digitalen Sammlung historischer musiktheoretischer Texte aus dem deutschsprachigen Raum anhand von Beispielen aus dem 19. Jahrhundert. Abschlussevent Text+-Kooperationsprojekte von 2025, online. Page 24 of 26 Moss, Roeder, Klinger, Keupp, & Roth (2025) – Project Documentation and Guidelines of the DigiMusTh project References Altenhöner, R., Berger, A., Bracht, C., Klimpel, P., Meyer, S., Neuburger, A., Stäcker, T., & Stein, R. (2023). DFG-Praxisregeln “Digitalisierung”. Aktualisierte Fassung 2022. Zenodo. https://doi.org/10.5281/zenodo.7435724 Caprotti, O., & Carlisle, D. (1999). OpenMath and MathML: Semantic markup for mathematics. XRDS, 6(2), 11–14. https://doi.org/10.1145/333104.333110 Cervone, D. (2012). MathJax: a platform for mathematics on the Web. Notices of the AMS, 59(2), 312-316. Huron, D. (1997). Humdrum and Kern: Selective Feature Encoding. In E. Selfridge-Field (Ed.), Beyond MIDI: The Handbook of Musical Codes (pp. 375–401). The MIT Press. Kahle, P., Colutto, S., Hackl, G., & Mühlberger, G. (2017). Transkribus—A Service Platform for Transcription, Recognition and Retrieval of Historical Documents. Proceedings of the 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), 19–24. https://doi.org/10.1109/ICDAR.2017.307 Kett, J., Kudella, C., Rapp, A., Stein, R., & Trippel, T. (2023). Text+ und die GND – Community-Hub und Wissensgraph. Zeitschrift für Bibliothekswesen und Bibliographie (ZfBB), 69(1–2), 37–47. https://doi.org/10.3196/1864295020691262 Kiessling, B., Tissot, R., Stokes, P., & Stökl Ben Ezra, D. (2019). eScriptorium: An Open Source Platform for Historical Document Analysis. International Conference on Document Analysis and Recognition Workshops (ICDARW), 19–19. https://doi.org/10.1109/ICDARW.2019.10032 Kraft, S., Schmalen, A., Seitz-Moskaliuk, H., Sure-Vetter, Y., Knebes, J., Lübke, E., & Wössner, E. (2021). Nationale Forschungsdateninfrastruktur (NFDI) e. V.: Aufbau und Ziele. Bausteine Forschungsdatenmanagement, 2, 1–9. https://doi.org/10.17192/bfdm.2021.2.8332 Miller, M., & Vielfaure, N. (2022). OpenRefine: An Approachable Open Tool to Clean Research Data. Bulletin - Association of Canadian Map Libraries and Archives (ACMLA), 170. https://doi.org/10.15353/acmla.n170.4873 Moss, F. C., Köster, M., Femminis, M., Métrailler, C., & Bavaud, F. (2021). Digitizing a 19th-Century Music Theory Debate for Computational Analysis. In M. Ehrmann, F. Karsdorp, M. Wevers, T. L. Andrews, M. Burghardt, M. Kestemont, E. Manjavacas, M. Piotrowski, & J. van Zundert (Eds.), CHR 2021: Computational Humanities Research Conference, November 17–19, 2021, Amsterdam, The Netherlands (pp. 159–170). CEUR. http://ceur-ws.org/Vol-2989/short_paper31.pdf Moss, F. C., Nápoles López, N., Köster, M., & Rizo, D. (2022a). Challenging sources: A new dataset for OMR of diverse 19th-century music theory examples. In J. Calvo-Zaragoza, A. Pacha, & E. Shatri (Eds.), Proceedings of the 4th International Workshop on Reading Music Systems (WoRMS 2022) (pp. 4–8). https://sites.google.com/view/worms2022/proceedings Moss, F. C., Nápoles López, N., Köster, M., & Rizo, D. (2022b). 19MT-OMR: A dataset for multimodal Optical Music Recognition [Data Report]. OSF. https://doi.org/10.17605/osf.io/QM9Z5 Pugin, Laurent, Zitellini, Rodolfo, & Roland, Perry. (2014). Verovio: A library for Engraving MEI Music Notation into SVG. In H.-M. Wang, Y.-H. Yang, & J. H. Lee (Eds.), Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR 2014) (pp. 107–112). Academia Sinica. https://doi.org/10.5281/zenodo.1417589 Reul, C., Springmann, U., & Puppe, F. (2017). LAREX: A semi-automatic open-source Tool for Layout Analysis and Region Extraction on Early Printed Books. Proceedings of the 2nd International Conference on Digital Access to Textual Cultural Heritage, 137–142. https://doi.org/10.1145/3078081.3078097 Reul, C., Christ, D., Hartelt, A., Balbach, N., Wehner, M., Springmann, U., Wick, C., Grundig, C., Büttner, A., & Puppe, F. (2019). OCR4all—An Open-Source Tool Providing a (Semi-)Automatic OCR Workflow for Historical Printings. Applied Sciences, 9(22), 4853. https://doi.org/10.3390/app9224853 Roeder, T., Klinger, J., Stickler, F., Keupp, C., & Moss, F. C. (2025) Multimodality and Minimal Publishing: TEI,MEI and more in 19th-Century Music Treatises. 25th Annual Meeting of the Text Encoding Initiative. 16–17 September 2025, Kraków, Poland. https://10.5281/zenodo.17205769 Roeder, T., Köster, M., & Moss, F. C. (2023). Music-Text Interlinking as a Challenge for Digital Encodings of Music-Theoretical Writings. Encoding Cultures – Joint MEC and TEI Conference 2023, 4–8 September Page 25 of 26