scieee AI-readable full text Open interactive document viewer

Leveraging Melodic Context for Improved Svara Representation

Nuttall, Thomas; Vijayan, Vivek; Serra, Xavier; Pearson, Lara

Abstract

For the South Indian musical tradition known as Carnatic music, embeddings of svara (note) pitch time series have proven useful for tasks such as svara classification and performance analysis. In this paper, we extend an existing embedding method by incorporating findings from musicological research on the relationship between the performance of a svara and its immediate melodic context, in order to improve the learning of these embedding models. We present a context-aware GRU-based model, adapting the existing DeepGRU architecture to encode both svara and its surrounding melodic context, before combining them via a co-attention mechanism prior to classification. For a ground truth dataset of 2,077 expert svara annotations across two performances in raga Bhairavi, we observe that the inclusion of melodic context leads to a 6.6% absolute increase in F1 score for svara label classification (from 78.3% to 84.9%), and an 7.8% absolute increase (from 59.9% to 67.7%) for classification of svara-form: sub-svara clusters that capture gamaka (ornamentation) variations in the performed svara.

Full text

Leveraging Melodic Context for Improved Svara Representation Thomas Nuttall1[0000000163161424],VivekVijayan 1[0009000291537868], Xavier Serra1[0000000313952345],andLaraPearson 2[0000000250738738] 1Music Technology Group, Universitat Barcelona Fabra, Barcelona, Spain 2Institute of Musicology, University of Cologne, Cologne, Germany Corresponding author: [email protected] Abstract. For the South Indian musical tradition known as Carnatic music, embeddings of svara (note) pitch time series have proven useful for tasks such as svara classification and performance analysis. In this paper, we extend an existing embedding method by incorporating findings from musicological research on the relationship between the performance of a svara and its immediate melodic context, in order to improve the learning of these embedding models. We present a context-aware GRUbased model, adapting the existing DeepGRU architecture to encode both svara and its surrounding melodic context, before combining them via a co-attention mechanism prior to classification. For a ground truth dataset of 2,077 expert svara annotations across two performances in r¯aga Bhairavi, we observe that the inclusion of melodic context leads to a 6.6% absolute increase in F1 score for svara label classification (from 78.3% to 84.9%), and an 7.8% absolute increase (from 59.9% to 67.7%) for classification of svara-form:sub-svaraclustersthatcapturegamaka (ornamentation) variations in the performed svara. Keywords: Representation learning ·Time series analysis ·Carnatic music ·Coarticulation 1Introduction Carnatic music, an art music tradition from South India, has been the subject of considerable recent research interest in Music Information Retrieval (MIR). Active areas of research on this musical style include melodic pattern recognition [14–16,29,30], note transcription [13,40,49,52], music synthesis [39,48] and performance analysis [17,18,31,32]. These tasks often rely on latent representations, typically learnt by neural network models trained on expert-annotated data. Such representations offer a low-dimensional, length-invariant space that enables large-scale similarity comparisons across corpora. However, noteor pattern-level annotated data are scarce, costly to obtain, and require input from multiple experts, whose annotations may diverge due to valid differences in level of detail. Transcription from audio to symbolic notation Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 185 2T.Nuttalletal. is widely acknowledged in musicology as having a significant subjective component [11] with a typical issue in Carnatic music transcription being different possible approaches to the degree of detail represented in the notation [34]. Achallengeintheautomatedestimationofmelodicrepresentationsinthis style is that svaras (note-level units) are often performed with gamakas (ornamentation), some of which oscillate around the theoretical pitch position without resting on it [22]. Furthermore, in any given raga, a svara may be performed with different gamakas, depending on its melodic context. As a result, there is a great deal of variability in the ways that a svara can be performed, making automated transcription particularly challenging in this style. However, the number of ways that a svara may be performed are not limitless, but rather are constrained by the tradition. Indeed, ragas are often best identified by their characteristic gamakas and motifs, rather than by their scalar content [53]. Therefore, there is scope for applying this knowledge that lies within the tradition to assist with svara identification. In Carnatic music pedagogy, svaras and their associated gamakas are learnt in the context of musical phrases [53], and musicians themselves know that the way a svara is performed depends on its melodic context [48]. Following from these insights, the relationship between svara performance and immediate melodic context has been theorised using the concept of coarticulation: the tendency for the performance of a unit (in this case a svara) to be influenced by that which precedes or follows it [33]. Coarticulation is an important topic in phonetics, because phonemes differ in their realisation depending on their immediate context [24]. This indeed was one of the initial obstacles to automated speech transcription and synthesis [5,25], a situation that can be aptly compared to that facing the computational analysis of Carnatic music. The coarticulation hypothesis in Carnatic music has recently been investigated empirically using computational methods [31,32]. This existing work creates a small dataset of svara annotations, as well as groupings of the different svara realisations into svara-gamaka units that are referred to as svara-forms. A first paper demonstrated a significant relationship between neighbouring svara and svara performance, and subsequent work found that the relationship increases as context increases, and plateaus after around 2 svaras either side. In this paper, we present a novel methodology for improving svara representation learning by incorporating melodic context as an input to a svara embedding model. We demonstrate the value of this inclusion by evaluating the learned representations effectiveness for svara classification on an unseen ground-truth test set from the original annotations. The code to reproduce this analysis and access the Context-Aware DeepGRU model can be found in the accompanying Github repository3. 3https://github.com/thomasgnuttall/ContextGRU/ Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 186 Levaraging Melodic Context for Improved Svara Representation 3 2RelatedWork Latent representations of sequential data have proven to be valuable intermediates for various music analysis tasks, including transcription [13], motif finding [30, 56], classification [23, 26, 37, 44, 46], performance analysis [31, 32], and synthesis [38], both in Carnatic music and beyond. In many cases, the learned representation serves as a distilled encoding useful across multiple tasks [2,51,55]. Acommonlimitationintrainingthesemodelsisthelackofannotatedground truth data, and whilst some small datasets have been published [17, 30, 32], in Carnatic music, this remains a problem. A comprehensive review of deep learning techniques for time series with limited labels can be found in [10], and can be broadly grouped into three categories: developing model components with fewer parameters to improve performance on small datasets, such as the case with Gated Recurrent Units (GRUs) (shown to perform well in classifying svara with limited data) [27,31]; by pretraining in a self-supervised fashion on abroader,relateddataset[6,21,46];orbyleveragingmoreinformationfrom existing data, such as including additional contextual features [8] or modalities [7, 54] to contextualize the primary input. It is the latter that this paper is concerned with. Due to the ornamented nature of the style, Carnatic svaras are well characterized by segments of continuous pitch time series [1,17–20,22,29,35,36,41,42,53]. Deep learning models for time series representation do not typically incorporate explicitly defined contextual features [9,27,28]. However, for the case of pitch time series of svara, recent musicological and computational studies have shown that neighbouring melodic context provides important information about how individual svaras are performed [31–33]. In this paper we leverage the findings in [31–33] to improve svara representation. We do this by adapting an existing embeddings model [27,31] to incorporate contextual information provided by the time series surrounding ground truth svara annotations, demonstrating its effectiveness for the task of svara label classification. 3Dataset We work with two performances in r¯aga Bhairavi from the Carnatic corpus of the Saraga dataset: Kamakshi (composed by Syama Sastri), performed by Sanjay Subrahmanyan, and Raksha Bettare (composed by Tyagaraja), performed by Shruthi S. Bhat [43,47]. The total duration of these composition performances is approximately 25 minutes and they are accompanied by a total of 2,077 expert svara annotations, created collaboratively by two Carnatic musicians. The annotations are made in sargam notation, a form of notation traditionally used in Carnatic music that is similar to solfège (the seven svaras in this r¯aga are sa, ri, ga, ma, pa, dha, and ni). Of these 2,077 svara annotations, 804 svaras include an additional svara-form label, denoting sub-svara clusters that capture gamaka variations, such as the four ga variants in Fig. 1. All annotations were Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 187 4T.Nuttalletal. manually verified by a Carnatic music expert and have proven useful for music analysis [31,32]. Fig. 1. Predominant pitch curves for four variants of the svara ga in r¯aga Bhairavi. These plots illustrate how the same svara can be performed with a number of different gamakas - melodic movements - that traverse a range of pitches. 4Methodology In this section we detail the extraction of the f0 time series corresponding to the predominant vocal melody for each of the ground truth annotations and their surrounding context, and present a GRU-based recurrent neural network for encoding and classifying svara alongside its context. 4.1 Svara Time Series Each annotation in our dataset corresponds to a performed svara in one of two compositions in r¯aga Bhairavi. For both performances, we extract the f0 time series corresponding to the predominant vocal melody using the FTA-Net Carnatic pitch extraction model available in the compIAM package [12]. Short silences (up to 200 ms) are interpolated, and pitch is expressed in cents relative to the performer’s tonic, as provided by the Saraga dataset. We smooth the pitch curves using a cubic spline, preserving melodic peaks. Our classification model expects individual svara observations, which it encodes and classifies (Section 4.2). Each observation comprises three distinct time series: (1) the f0 pitch curve of the current svara annotation, (2) the f0 pitch curve of the preceding melodic context, and (3) the f0 pitch curve of the succeeding melodic context (see the upper section of Fig. 2). Current svaras containing intermediate silence (i.e., not at the boundaries) are excluded due to the likelihood that they have pitch extraction errors. The two contextual pitch curves are extracted from segments preceding and succeeding the current svara, with a duration twice that of the svara itself. This Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 188 Levaraging Melodic Context for Improved Svara Representation 5 follows [32], which shows diminishing correlation between svara performance and context beyond two adjacent svaras. If the contextual segments exceed one second, they are capped at one second. If silence is present, contextual segments are trimmed to start or end with silence for preceding and succeeding contexts, respectively. Whilst silence in the pitch curve of the central, current,annotationislikely due to pitch extraction errors and thus excluded, silence in the preceding and succeeding context is expected (e.g., ends of phrases) and treated as being meaningful (potentially having its own influence). We retain this silence by substituting it with a pitch value far outside the expected pitch range of a Carnatic vocalist (-4000 cents) and by introducing a binary feature indicating silence. Consequently, each of the three pitch curve time series per observation is twodimensional: one feature representing pitch (in cents), and one binary feature indicating silence. Each observation is associated with two labels: the svara label, indicating the performed svara (2{sa,ri,ga,ma,pa,dha,ni}), and the svara-form label, representing the cluster of unique gamaka (Section 3) [32]. Fig. 1 illustrates four observations for the svara ga. 4.2 Context-Aware DeepGRU Network Our model (Fig. 2) is an adaptation of the existing DeepGRU model [27], which has been demonstrated as effective in classifying svara in small datasets [31]. To allow the model to focus on temporally salient information within and beyond the primary sequence of interest, we alter the attention mechanism to apply it independently to the outputs of three GRU-based encoders: preceding,current, and succeeding. Consisting of stacked GRUs, each encoder processes the variablelength time series pitch curves corresponding to the preceding context, current svara, and succeeding context respectively (Section 4.1), to produce a sequence of hidden representations. Attention is computed on these representations with reference to the final hidden state of the current encoder. If H2RB⇥T⇥ddenotes the sequence of hidden states produced by an encoder - where Bis the batch size, Tthe sequence length, and dthe hidden dimensionality - and hc2RB⇥drepresents the final hidden state of the current encoder, the attention mechanism computes a compatibility score between each time step in Hand the reference state hc,projectedviaalearnablelinear transformation W2Rd⇥d: et=HtWh> c,for t=1,...,T These scores are normalized via a softmax function over the temporal dimension to obtain attention weights ↵2RB⇥T⇥1: ↵t=exp(et) PT k=1 exp(ek) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 189 6T.Nuttalletal. Fig. 2. Context-aware DeepGRU Network. Contextual and current pitch time series are encoded separately and combined through a co-attention mechanism before classification. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 190 Levaraging Melodic Context for Improved Svara Representation 7 And the context vector c2RB⇥dis computed as the weighted sum of the input sequence: c= T X t=1 ↵tHt As in the original DeepGRU model, to capture higher-order temporal interactions, this context vector is passed to an auxiliary single-layer GRU, initialized with hc,yieldingagatedcontextrepresentationc0. The final attention output is then formed by concatenating the raw context vector and the gated context: ˜c=[ckc0] This process is applied independently to the preceding, current, and succeeding encoders’ outputs, with all attention computations referencing the current encoder’s final hidden state, hc. This ensures that context integration remains dynamically aligned with the current sequence, and enables the model to selectively attend to temporally relevant cues from the two contextual sequences. The classification head of the original DeepGRU model remains unchanged - the three attention outputs are concatenated and passed to two fully connected layers with batch normalization, dropout, and ReLU activation functions, mapping the learned feature representations to the final output classes. 5ExperimentsandResults 5.1 Experiment 1: Ablation study To understand the impact of melodic context on svara classification, we train two models on the ground truth annotations: one with contextual data, using the model presented in Section 4.2, and one using only the central current time series, with the contextual encoders removed. In the latter case, the model reduces to the original DeepGRU architecture. Each model is trained twice—once to predict svara-form and once to predict svara—resulting in four experiments in total. Each encoder is configured with an input layer size of 32, halving in size at each subsequent layer. We use the Adam optimizer, a learning rate of 103, and a weight decay of 104.Allexperimentsuseamini-batchsizeof256anda dropout rate of 0.3. These training parameters follow the original DeepGRU PyTorch implementation. Models are evaluated using 3-fold cross-validation, and we report the average F1 score on the test set across all folds. To prevent overfitting, training data is augmented by balancing class distributions using time dilation with a factor of between 0.9 and 1.1. Table 1 displays the results. We note that for predicting both svara-form and svara,theinclusionof melodic context significantly improves model performance — a 6.6% absolute increase in F1 score for svara classification (from 78.3% (±2.6) to 84.9% (±2.5)), and a 7.8% absolute increase (from 59.9% (±0.6) to 67.7% (±3.3)) for svaraform. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 191 8T.Nuttalletal. Table 1. Effect of considering melodic context on Svara and Svara-form Classification Melodic Context Target F1 mean F1 std No Svara 0.783 0.026 Yes Svara 0.849 0.025 No Svara-form 0.599 0.006 Yes Svara-form 0.677 0.033 5.2 Experiment 2: Encoder Contributions To understand the relative impact of the preceding, succeeding, and current encoder on the final prediction, we conduct a gradient-based attribution analysis to measure how sensitive the model’s prediction is to small perturbations in specific encoder outputs, quantified via the L2 norm of the gradient. Specifically, we compute the gradient of the predicted class score with respect to the attention output of each encoder: preceding context, current input, and succeeding context. The L2 norm of each gradient is used as a proxy for the encoder’s influence on the model’s decision, with higher norms indicating greater sensitivity. This approach has been demonstrated to be useful for comparing internal components’ influence without retraining or ablation [3,4,45,50]. Averaged over the entirety of the ground truth data, the normalized gradient norms are 0.295, 0.402, 0.303 for the preceding, current, and succeeding encoder respectively, suggesting that whilst each encoder contributes meaningfully, the central input encoder has the highest influence on the classification decision, with the two contextual encoders demonstrating a smaller but equally balanced contribution. It is perhaps not surprising that the central input should dominate, with contextual information providing supportive but secondary contributions. 6Conclusion In this paper we present a context-aware, GRU-based time series classifier, adapting the existing DeepGRU to encode multiple time series separately and combine them via a co-attention mechanism before classification. We show the effectiveness of this model in contextualizing svara embeddings with their surrounding melodic information by improving the F1 score of svara classification by 6.6%, (from 78.3% (±2.6) to 84.9% (±2.5)), and 7.8% (from 59.9% (±0.6) to 67.7% (±3.3)) for svara-form, on a ground truth dataset of 2,077 expert annotations. This finding also corroborates previous musicological and computational studies on the relationship between svara performance and immediate melodic context, and explicates a characteristic of the style implicitly understood by practitioners. Given the widespread use of representation learning for MIR tasks, we anticipate that improvements to these models, such as that presented in this paper, will positively impact downstream applications such as transcription, pattern recognition, and synthesis, by more effectively capturing this implicit musical knowledge embedded within the Carnatic tradition. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 192 Levaraging Melodic Context for Improved Svara Representation 9 Acknowledgments. This research was carried out as part of the "IA y Música: Cátedra en Inteligencia Artificial y Música" (TSI-100929-2023-1), funded by the Secretaría de Estado de Digitalización e Inteligencia Artificial, and the European Union-Next Generation EU, under the program Cátedras ENIA 2022 para la creación de cátedras universidad-empresa en IA. References 1. Akant, K.: Measuring frequencies of shrutis in indian classical music. In: 2019 9th International Conference on Emerging Trends in Engineering and TechnologySignal and Information Processing (ICETET-SIP-19). pp. 1–4. IEEE (2019) 2. Alonso-Jiménez, P., Bogdanov, D., Pons, J., Serra, X.: Tensorflow audio models in Essentia. In: International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2020) 3. Ancona, M., Ceolini, E., Öztireli, C., Gross, M.: Towards better understanding of gradient-based attribution methods for deep neural networks. In: International Conference on Learning Representations (ICLR) (2018) 4. Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K.R., Müller, K.R.: How to explain individual classification decisions. Journal of Machine Learning Research 11,1803–1831(2010) 5. Birkholz, P.: Modeling Consonant-Vowel Coarticulation for Articulatory Speech Synthesis. PLOS ONE 8(4), e60603 (2013). https://doi.org/10.1371/journal.pone.0060603 6. Bošnjak, M., Richemond, P.H., Tomasev, N., Strub, F., Walker, J.C., Hill, F., Buesing, L.H., Pascanu, R., Blundell, C., Mitrovic, J.: Semppl: Predicting pseudolabels for better contrastive representations. arXiv preprint arXiv:2301.05158 (2023) 7. Clayton, M., Rao, P., Shikarpur, N.N., Roychowdhury, S., Li, J.: Raga classification from vocal performances using multimodal analysis. In: ISMIR. pp. 283–290 (2022) 8. Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., Zeghidour, N.: Moshi: a speech-text foundation model for real-time dialogue. arXiv preprint arXiv:2410.00037 (2024) 9. Ekambaram, V., Kumar, S., Jati, A., Mukherjee, S., Sakai, T., Dayama, P., Gifford, W.M., Kalagnanam, J.: Tspulse: Dual space tiny pre-trained models for rapid timeseries analysis. arXiv preprint arXiv:2505.13033 (2025) 10. Eldele, E., Ragab, M., Chen, Z., Wu, M., Kwoh, C.K., Li, X.: Label-efficient time series representation learning: A review. IEEE Transactions on Artificial Intelligence 5(12), 6027–6042 (2024). https://doi.org/10.1109/TAI.2024.3430236 11. Ellingson, T.: Transcription. In: Myers, H. (ed.) Ethnomusicology: an introduction, pp. 110–52. Ethnomusicology: an introduction, Norton, New York (1992) 12. Genís Plaja-Roglans and Thomas Nuttall and Xavier Serra: compiam (2023), https://mtg.github.io/compIAM/ 13. Gowrishankar, B., Bhajantri, N.U.: Deep learning long short-term memory based automatic music transcription system for carnatic music. In: 2022 IEEE International Conference on Distributed Computing and Electrical Circuits and Electronics (ICDCECE). pp. 1–6. IEEE (2022) 14. Gulati, S.: Computational approaches for melodic description in indian art music corpora. Ph.D. thesis, Universitat Pompeu Fabra (2016) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 193