scieee AI-readable full text Open interactive document viewer

AnalysisGNN: Unified Music Analysis with Graph Neural Networks

Karystinaios, Emmanouil; Hentschel, Johannes; Neuwirth, Markus; Widmer, Gerhard

Abstract

Recent years have seen a boom in computational approaches to music analysis, yet each one is typically tailored to a specific analytical domain. In this work, we introduce AnalysisGNN, a novel graph neural network framework that leverages a data‐shuffling strategy with a custom weighted multi‐task loss and logit fusion between task‐specific classifiers to integrate heterogeneously annotated symbolic datasets for comprehensive score analysis. We further integrate a Non‐Chord‐Tone prediction module, which identifies and excludes passing and non‐functional notes from all tasks, improving the consistency of label signals. Experimental evaluations demonstrate that AnalysisGNN achieves performance comparable to traditional static‐dataset approaches, while showing increased resilience to domain shifts and annotation inconsistencies across multiple heterogeneous corpora. The full source code for this work is available at: [github.com/manoskary/AnalysisGNN]

Full text

AnalysisGNN: Unified Music Analysis with Graph Neural Networks Emmanouil Karystinaios1?,JohannesHentschel 2⇤, Markus Neuwirth2,GerhardWidmer 1 1Institute of Computational Perception, Johannes Kepler University Linz, Austria [email protected] 2Anton Brukner University Linz, Austria [email protected] Abstract. Recent years have seen a boom in computational approaches to music analysis, yet each one is typically tailored to a specific analytical domain. In this work, we introduce AnalysisGNN,anovelgraphneural network framework that leverages a data-shuffling strategy with a custom weighted multi-task loss and logit fusion between task-specific classifiers to integrate heterogeneously annotated symbolic datasets for comprehensive score analysis. We further integrate a Non-Chord-Tone prediction module, which identifies and excludes passing and non-functional notes from all tasks, improving the consistency of label signals. Experimental evaluations demonstrate that AnalysisGNN achieves performance comparable to traditional static-dataset approaches, while showing increased resilience to domain shifts and annotation inconsistencies across multiple heterogeneous corpora. The full source code for this work is available at: github.com/manoskary/analysisgnn Keywords: Music Analysis ·Graph Neural Networks ·Harmonic analysis. 1Introduction Symbolic music analysis is a cornerstone of music information retrieval and musicology. Traditional methods for harmonic analysis, cadence detection, and phrase segmentation often rely on rule-based systems or statistical models. While recent deep learning approaches have improved performance by leveraging relatively large annotated datasets or self-supervised pretraining, they typically treat each analysis problem separately, thereby missing the inherent interdependencies in musical structure. Multi-task learning offers a promising direction to exploit cross-task knowledge transfer—for example, features learned during harmonic analysis may enhance cadence detection—but in practice, available datasets focus on single musical elements (harmony, cadence, or hierarchical structure) and suffer from inconsistent ?*Equalcontribution. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 271 2 E. Karystinaios et al. annotations. To address these challenges, we replace task-incremental pipelines with a unified data-shuffling strategy: during training, mini-batches are sampled across all tasks, with a custom weighted multi-task loss that balances each objective, and logits from task-specific classifiers are fused to reinforce shared representations. Additionally, we incorporate a Non-Chord-Tone prediction branch that filters out passing and non-functional notes before downstream tasks, yielding cleaner input signals and reducing conflicting labels. Graph Neural Networks (GNNs) have emerged as a powerful solution for modeling the complex, non-sequential relationships inherent in music scores. For example, ChordGNN [9] represents notes as nodes in a graph to capture Roman numeral harmony, achieving competitive results compared to previous approaches. Similarly, graph-based methods have improved cadence detection by directly leveraging musical context. Our proposed framework, AnalysisGNN,buildsontheseideasbycombining a data-shuffling training regimen with weighted multi-task optimization and inter-classifier logit fusion, alongside a non-functional note exclusion mechanism. This design addresses the limitations of fragmented, specialized datasets and isolated models through a single, cohesive Graph Neural Network. The resulting system achieves good performance on each analysis task, including cadence detection, Roman numeral analysis, section and phrase identification, and metrical position estimation, while also maintaining robustness against domain shifts and annotation inconsistencies. Our contributions are four-fold: 1. We propose a training strategy and architecture for heterogeneous multi-task symbolic music analysis.. 2. We introduce a number of new music analysis tasks such as Non-Chord-Tone prediction module that identifies and filters passing and non-functional notes. 3. We assemble and preprocess the largest compilation of heterogeneously annotated symbolic music datasets, unified into a graph representation suitable for GNN processing. 4. We demonstrate that AnalysisGNN achieves competitive performance while exhibiting strong resilience to domain shifts and annotation variability. 2RelatedWork 2.1 Harmonic Analysis (Roman Numeral Labeling) Automatic analysis of functional harmony has long been studied in the fields of music information retrieval and music theory. Early systems employed rule-based grammars, probabilistic models, dynamic programming, and grammar induction for harmonic and metrical analysis of music [19,22,23]. The advent of deep learning brought significant improvements. Micchi et al. [15] trained one of the first neural networks (a convolutional recurrent model) for Roman numeral analysis, and Nápoles López et al. [18] later Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 272 AnalysisGNN: Unified Music Analysis with Graph Neural Networks 3 introduced AugmentedNet, a CNN-LSTM architecture that enhanced performance via data augmentation and multi-task learning. Subsequently, two more systems using multitask approaches for Roman numeral analysis were introduced using techniques to mitigate the interdependency between the Roman numeral prediction subtasks [14,16]. Despite these advances, purely sequential models struggle with the inherent complexity of polyphonic scores—they require the score to be serialized or binned into time slices, which risks losing voice-leading details and long-range dependencies. This limitation has motivated the use of graph neural networks. ChordGNN [9] constructs a graph of all notes in a piece and applies graph convolution to aggregate note features into chord predictions. It treats Roman numeral analysis as a multi-task classification problem (predicting multiple components of the Roman numeral label) similar to previous approaches and introduces an edge contraction algorithm to pool information from note-level to chord-level representations. As a result, ChordGNN achieved high performance on Roman numeral analysis, outperforming previous CNN/CRNN models on standard datasets. The most recent approach by Sailor [20] has used a transformer model and BERT-like pretraining to build a more robust encoder, which has in turn resulted in the current state-of-the-art performance. 2.2 Cadence Detection and Phrase Segmentation Cadences serve as musical punctuation, marking phrase or sectional boundaries. Traditional cadence detection methods relied on hand-crafted features and classifiers, such as SVMs trained on features representing intervallic sequences that form full (or authentic) or half cadences [2]. More recent approaches, such as those by Karystinaios and Widmer [8], reformulated cadence detection as a node classification task on a note-level graph representation of the score. In their work, aGNNwastrainedtopredictcadencelikelihoodforeachnote(orgroupof notes), capturing non-local context through message-passing. Derivative work has further refined the process by introducing more sophisticated training techniques and models that have resulted in better performance [10], although the lack of a unified benchmark prevents a direct comparison. 2.3 Graph Neural Networks in Music Information Retrieval Beyond analysis, GNNs have been applied in other MIR tasks, such as music recommendation, emotion modelling, music generation, and performance modeling. For example, Jeong et al. [7] combined a GNN with a hierarchical RNN to model expressive piano performance, capturing both local note interactions and broader musical context. Similarly, graph-based representations have been exploited in music generation, where hierarchical graphs of chord progressions and rhythmic patterns help maintain long-term coherence in generated compositions [3,12]. These developments underscore the flexibility of graph-based approaches and their potential to unify disparate music processing tasks under a single framework. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 273 4 E. Karystinaios et al. 2.4 Unified Multi-Task Learning Frameworks Multi-task learning (MTL) trains a single model on multiple related objectives simultaneously, leveraging shared representations to improve performance and generalization across tasks [24]. By interleaving training examples from heterogeneous datasets, rather than processing tasks sequentially, unified MTL avoids the catastrophic forgetting endemic to task-incremental pipelines and mitigates domain shifts between specialized corpora. A central challenge in unified MTL is balancing conflicting gradients and annotation schemes. Common solutions include: i) tailor optimization by assigning task-specific weights to losses or gradients, either heuristically or via uncertainty estimation, to modulate each objective’s contribution to the total loss [11,13]; ii) dynamic task sampling by shuffling mini-batches across tasks according to dataset size, task difficulty, or performance criteria, ensuring stable gradient updates and preventing domination by any single task [5], iii) logit-level fusion by combining the raw outputs of task-specific classifiers to reinforce shared features and encourage co-adaptation between tasks [17]. In symbolic music analysis, auxiliary prediction branches can further improve consistency by filtering noisy labels. For example, a Non-Chord-Tone (NCT) detection head can identify and exclude passing or non-functional notes before they propagate to downstream tasks, reducing label conflicts and sharpening the signal for harmony, cadence, and phrase analyses. Graph Neural Network architectures such as ChordGNN [9] have demonstrated the efficacy of combining shared graph encoders with modular task heads; our framework extends this paradigm with data-shuffling, custom weighting, logit fusion, and NCT filtering to achieve robust, unified score analysis. 3Methodology 3.1 Model Architecture For our model, we leverage previous graph-based architectures from the literature used for music analysis while integrating our contributions of logit fusion to architecture design. Input to our model is music scores represented as graphs, where vertices are notes and edges represent temporal relations between those notes, similar to [7,8]. The backbone of our model is a Hybrid Graph Neural Network similar to the one introduced in [10]. In more detail, the score graph enters a series of Heterogeneous Graph Convolutional blocks. In parallel, the note features are segmented per piece and padded where necessary and then they are passed through a sequential model such as a GRU with the same number of layers as the number of graph convolutions. The common representation is then used by a series of 2-layer MLP classifiers. Finally, a logit-fusion layer is added where the logit prediction of each task communicate with each other. A sketch of the architecture is shown in Figure 1. The HybridGNN encoder combines a sequential model (GRU) with a Graph Convolutional Network (GCN) to effectively capture the hierarchical structure Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 274 AnalysisGNN: Unified Music Analysis with Graph Neural Networks 5 inherent in musical scores. In this design, the GCN integrates information at multiple levels—notes, beats, and measures—thereby functioning as a heterogeneous graph neural network. For note-type nodes, the input features are distilled to essential attributes such as pitch spelling and duration. In contrast, the features for beats and measures are computed as the mean values of the features from their associated notes. Despite this layered encoding, all predictions are ultimately made at the note level. 3.2 Training Training AnalysisGNN proceeds by interleaving examples from all tasks via a data-shuffling strategy, combined with a custom weighted multi-task loss and logit-level fusion. At each update step, mini-batches are sampled across the current set of tasks T,ensuringthatnosingletaskdominatestheoptimization. We adopt a dynamically weighted cross-entropy loss similar to [11], but normalized by the number of tasks: Lclf =1 |T | X t2T✓Lt 22 t + log1+2 t◆,(1) where Lt is the cross-entropy for task t and each t is a learnable scale controlling the task’s weight. This formulation both balances gradients from heterogeneous annotation schemas and regularizes the learned weights to prevent any t from collapsing to zero. To encourage shared representations across tasks, we apply a logit-level fusion mechanism that refines each task’s raw outputs by integrating information from all other task heads. Concretely, for each task tin the current set T: zt=Clf t(h),(raw logits) pt=Proj t(zt),(projection into a common d-dimensional space) We then stack all projections into a matrix P=2 6 6 6 4 p1 p2 . . . p|T | 3 7 7 7 52R|T |⇥d and refine them via a transformer-style self-attention layer: e P= LayerNorm⇣P+ MultiHeadAttn(P, P, P)⌘, where MultiHeadAttn denotes standard multi-head attention and the LayerNorm adds a residual connection. Finally, each task’s refined logits are obtained by selecting its row from e P and passing it through the corresponding fusion head: ˆ zt= Fusionte Pt, Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 275 6 E. Karystinaios et al. where e Pt is the t -th row of e P .Wethencomputethecross-entropylosson ˆ zt for each task t . This attention-based fusion allows each task head to attend to and integrate information from all other tasks by sharing confidence and contextual cues, leading to more coherent multi-task representations. In addition to the main analysis tasks, we introduce a Non-Chord-Tone (NCT) prediction branch. This auxiliary head labels each note as functional or non-function in regards to the underlying harmony and structure, with its own cross-entropy loss LNCT .During training, we predict NCT labels but do not mask out non-chord-tones in the multi-task losses, since this preserves passing-note examples in the gradient signal and prevents the model from collapsing (by masking everything) or misclassifying functional notes. At inference time, however, we can leverage the NCT predictions as a gating mechanism by classifying each note as chord-tone or non-chord-tone, and pass only those identified as chord-tones to the task-specific heads. This selective inference reduces computational overhead and mitigates error propagation by focusing predictions on musically functional notes. Fig. 1. Architecture overview: the input score graph is processed by stacked Heterogeneous GCN blocks in parallel with note-sequence features through an equally deep GRU; their shared embedding feeds 2-layer MLP task heads, whose outputs are then refined via a cross-task logit-fusion layer. 4 Corpora 4.1 Overview and Preprocessing Our work combines multiple datasets with diverse analytical annotations to train a single model: the AugmentedNet dataset [18], the Distant Listening Corpus (DLC) [6], the Bach WTC cadence dataset [4], the Mozart string quartet cadence dataset [1], and the Haydn string quartet cadence dataset [21]. The DLC provides 1266 symbolic music scores in MuseScore format, featuring internally consistent annotations (i.e. chords, phrases, and cadences) integrated directly within the Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 276 AnalysisGNN: Unified Music Analysis with Graph Neural Networks 7 dataset Distant Listening Corpus (DLC) AugmentedNet Combined Cadence Datasets task pieces notes labels labels* pieces notes labels pieces notes labels note level 1266 2 060 662 353 758 555 100 131 418 chord 1266 2 034 151 318 781 234 667 353 757 989 88 559 00 0 phrase 1265 2 060 348 21 395 15 575 00000 0 cadence 916 1 334 356 13 908 9540 000100 131 418 4147 pedal 1266 2 034 151 3436 2632 00000 0 Table 1. Dimensions of the three datasets. The note-level row corresponds to the total number of graphs (pieces) and graph nodes (notes). This row does not report labels because the target labels on the note level either derive from the score (total number of notes) or from the harmony annotations (number of notes in the chord row). The note numbers in the rows below reflect for how many graph nodes a valid label was present for the corresponding task. The labels subcolumns correspond to the number of labels that have been included in each dataset’s graphs. labels* is present only for the DLC for which repeats have been expanded; this column indicates the number of labels present in the original files before the expansion. scores. AugmentedNet contributes 353 pieces, offering Roman numeral harmonic analyses in RomanText rntxt format linked to scores parsed from various formats, representing an aggregation of multiple analysis sources. The remaining datasets feature only cadence annotations which have been inserted into musicXML scores for ease of processing. For simplicity, hereforth we refer to the union of the three datasets that contain solely cadence annotations as the cadence dataset. In order to facilitate the compilation of a unified graph dataset, the note and annotation content of all datasets was processed into a shared tabular representation. Since the cadence dataset has no overlap with the DLC and AugmentedNet dataset, we focus on the processing of the two latter. Overlapping pieces between the two datasets were identified, and pieces present in the predefined AugmentedNet test set were excluded from our training data derived from the DLC to allow for later comparison. Key analytical information, such as Roman numerals (simplified to root and quality), local keys, and chord inversions, was unified across datasets into common formats. Additional features relevant for analysis, like note-level scale degrees and downbeat information, were derived for both datasets. The final combined dataset comprises the processed versions of all included pieces from all sources, retaining the distinct annotations for pieces present in all datasets (excluding the test set conflicts). This results in a total of 1719 annotated pieces available for our study (cf. Table 1). 4.2 Handling Data For all our datasets—including Cadence (which comprises several subdatasets), DCL, and AugNET—the tabular data (described in the previous section) is first converted into graph representations. In these graphs, each node corresponds to a musical note, and relevant analytical labels are propagated to the nodes using the GraphMuse Python package [10]. Furthermore, we apply transposition Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 277 8 E. Karystinaios et al. augmentation to our training data while preserving pitch spelling sensitivity, resulting in an approximate tenfold increase in available scores. During training, the model is exposed to the entire collection of datasets for a fixed number of iterations, with every batch containing a combination of all three datasets and all tasks. When it comes to validation and testing, we assign one test and validation split per dataset. Accordingly, each dataset is divided into training, validation, and test sets, with the test set covering roughly 20% of the total data. The AugmentedNet dataset provided predefined data splits for training, validation, and testing, while for other two datasets we perform a random split. As detailed in Table 1, each dataset presents a unique set of annotations and tasks. Certain datasets (particularly DLC) include pieces with partial, missing, or invalid annotations. To manage these issues, we frame each score graph as a semi-supervised node classification problem, masking out any invalid or missing labels at the note level. For example, in the DLC dataset, some pieces lack cadence annotations and certain notes may have missing or unencodable harmony labels. By masking these problematic samples during training, we prevent invalid data points from influencing the model while still leveraging all available notes for predicting the valid labels. 4.3 Tasks In our framework, we address a wide range of tasks that include both tasks already available in the literature and new tasks introduced in this work. Although our model makes predictions at the note level—assigning a label for each task to every note—the scope of these annotations varies: some apply uniformly to all notes sharing the same onset, others apply to individual notes, and some extend over longer segments of the score. Specifically, our GNN-based model predicts 20 distinct properties for each note (or graph node). However, the prediction can also be computed at the onset or beat level on demand. The predictions include detecting the presence and type of cadences, identifying phrase and section boundaries, flagging pedal points, assessing metrical strength, and marking harmony onsets/changes. In addition, we predict harmonic analysis features as defined in [18], such as local key, tonicization key, root, bass, harmonic rhythm, inversion, quality, pitch-class set, common Roman numeral, and chord degree. Beyond these established tasks, we introduce novel note-level tasks aimed at capturing the functional role of each note within the underlying harmony. For instance, our framework determines boolean properties indicating whether a note functions as the bass, the root, or is part of the expected chordal structure. For example, in a "C:I" annotation, the note C is identified as the root, whereas apassingnoteDisnotaconstituentofthechord—eventhoughbothmaybe consistently labeled as "C:I" by the model. By leveraging the granularity afforded by our GNN, we enrich the annotations with musicologically informed properties that offer deeper insight into the model’s understanding of harmonic structure. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 278 AnalysisGNN: Unified Music Analysis with Graph Neural Networks 9 Dataset Model Cadence Roman Phrase Pedal Metr. Section Cadence GraphMuse .516 – – – – – AnalysisGNN (multi-corpus) .497 – – – – – AugNet AugmentedNet –.464 – – – – ChordGNN+Post –.518 – – – – RNBert –.574 – – – – AnalysisGNN –.530 – – – – DLC RNBert –.301 – – – – AnalysisGNN .558 .516 .742 .771 .761 .768 Table 2. Overview of Models by Target Task. For each dataset, we compare baseline models—trained and evaluated on their respective datasets—with AnalysisGNN, which is trained on all three datasets. Cadence, phrase, pedal point, and section are evaluated using note-level macro F1 score; Roman numeral predictions are assessed with the CSR score [18]; and metrical strength is measured by accuracy. A Roman numeral is considered correct only when its local key, degree, quality, and inversion are all predicted accurately. Trained on Eval on Cadence Eval on AugNet Eval on DLC Cad. F1 RN Cad. F1 Phrase RN Cadence only .516 –––– AugNet only -.515 ––.441 DLC only .479 .503 .556 .751 .563 All corpora (combined) .497 .530 .558 .742 .516 Table 3. Cross-corpus evaluation of AnalysisGNN. Rows indicate the training set (single-corpus or combined) and columns report performance on each target corpus, broken down by task. 5 Experiments 5.1 Comparison to Single-Corpus Approaches To assess the benefits of our unified analysis framework, we compare AnalysisGNN against models trained on individual corpora. In our experiments, we train AnalysisGNN on the combination of all available datasets. Table 2 presents a comparative evaluation of AnalysisGNN alongside established baselines. For the AugmentedNet dataset, AnalysisGNN is benchmarked against state-ofthe-art models from the literature, including AugmentedNet [18], ChordGNN [9], and RNBert [20]. In the context of cadence detection, we compare against the HybridGNN variant from GraphMuse [10]. Notably, for the DLC dataset, characterized by heterogeneous annotations compared to the AugmentedNet, AnalysisGNN is uniquely capable of effectively handling and comparing the diverse labels. We observe that although AnalysisGNN does not perform as well with single-corpus models, it demonstrates robust, mean performance across tasks, underscoring the advantage of a unified approach in mitigating domain shifts and annotation discrepancies. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 279