Full text
Cluster Expressive Timing in Performed Music with VQ-VAE Zihan Chai1and Shengchen Li1[0000-0002-2488-298X] Xi’an Jiaotong-Liverpool University, No. 111, Ren’ai Road, Suzhou, China {[email protected], [email protected]} Abstract. There are many attempts on clustering expressive timing in performed classical piano music which suffer from variable lengths of phrases. This work uses VQ-VAE, a deep learning based method, to cluster expressive timing. The proposed method uses a codebook with acodecstructure,whereeachcodevectorcorrespondstoacluster.The code vectors that is very similar could be further merged, which gives a more flexible way to determine the number of clusters for expressive timing. To evaluate the proposed method, a model selection test with Gaussian Mixture Model (GMM) for expressive timing is repeated to compare the optimal number of clusters in expressive timing. The JS divergence between clusters resulted by both VQ-VAE and GMM is also tested to show the difference of cluster distribution. The result shows that the result of clustering by VQ-VAE is comparable with the results by GMMs. Keywords: Expressive Timing ·Performed Piano Music ·Phrasing · VQ-VAE ·Clustering 1Introduction In computational musicology, clustering expressive timing is one of the common ways to analyse expressiveness as demonstrated by Demos et al. [1]. Though the clustering of intra-phrase expressive timing proposed by Li et al. [2]enablesthe clustering over variable lengths of phrases. Many empirical parameters such music structure, the selection of music unit for analysis and the number of clusters precent automatic clustering of expressive timing. This paper aims to produce adeeplearningbasedsolutionthathasthepotentialforautomaticexpressive timing clustering that is capable of variable phrase lengths and flexible cluster number with the potential of automatic phrasing. Considering the expressive timing information is usually simple but with averycomplexdistributionandthecomputationalmusicologyrequiresinterpretability, VQ-VAE [3]ischosenamongstotherdeeplearning-basedclustering All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 973
Z. Chai and S. Li method due to better interpretability, stability and simplicity. The VQ-VAE has acodecstructurethatenablesproperstructurestocapturetemporaldependencies in expressive timing within a phrase. Unlike traditional VAE [4], the latent space in VQ-VAE is quantised and is represented by a codebook. The expressive timing is firstly encoded as a codebook and then mapped to a code in the codebook for clustering. The codes in codebook may be further merged with a given threshold of distance, which provides flexibility of cluster numbers. To validate the proposed model, a dataset combing the public shared Mazurka dataset and a private Islamey dataset are used for the proposed experiments. With the dataset, we firstly test the proposed model with different configurations of codebook in the latent space. The experiment reveals how the lengths of code and the size of codebook impact the number of clusters resulted. To further validate the resulted clusters of expressive timing, the distribution of the resulted cluster is compared. Given the cluster distribution resulted by GMM models and the proposed VQ-VAE models, the difference between resulting distributions is measured by Jensen-Shannon Divergence (JS Divergence) [5]. 2Background 2.1 The Evolution of Expressive Timing Analysis Expressive timing analysis aims to understand how expressiveness formed by continuous variations in the rhythm of a musical event [6]. Though there are a few attempts that synthesis expressiveness in the domain of waveform directly [7,8], the knowledge of expressiveness in performance is still able to provide expressiveness control that the end-to-end system [9]. The field subsequently shifted towards data-driven methods [10], where clustering is commonly used as a way to analyse expressive timing [1]. Legend analysis methods such Principle Component Analysis [11], Gaussian Mixture Model (GMM) [2]andSelf-OrganisingMap(SOM)[12]havebeenusedforexpressive timing analysis. Existing clustering methods for expressive timing suffer from a few more issues: the number of clusters, the variable lengths of phrases and the unit of analysis. First, the number clusters is a empirical decision in most studies. Though the model selection test [13]maygiveobjectiveanswertotheoptimalnumber of clusters, the model selection test requires a massive training process, which could be time-consuming and computational expensive. As a result, the proposed system for clustering expressive timing should produce a flexible number of clusters for different datasets. Second, given the phrase lengths is changed throughout a piece of music, the proposed system should be capable of processing phrase with different lengths. Though Li et al. [2]hasprovidedasolutionthatregressingintra-phraseexpressive timing with polynomial functions, the optimised order of polynomial functions is phrase-dependent. As a result, the proposed system should be able to adapt the inconsistent phrase lengths throughout the pieces of music. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 974
Cluster Expressive Timing with VQ-VAE Finally, as expressive timing is associated with music structure [14], it is hard to determine the most proper unit for analysis. As a result, the music structure is considered as an empirical decision for analysis in expressive timing. A possible solution for the problem is automatic phrasing which may be achieved by multitask learning. As a result, the proposed system should be capable of multi-task learning in the future, which requires a deep learning-based system. 2.2 Deep Generative Models for Unsupervised Clustering Deep learning techniques has been used for clustering sequential data [15]. Given the case of clustering expressive timing, which is represented as a vector of tempi, the proposed clustering algorithm needs to address the issue of low data availability, low dimensionality and complex distribution. As a result, a generative deep learning model needs to be selected. Among generative models, the generative adversarial network based solutions such as ClusterGAN [16] may not work well with low data availability. Traditional deep embedding for clustering [17] has low interpretability even with the sequence-to-sequence framework [18]used.Asaresult,thispaperselectsVariational Auto-Encoder (VAE) [4]astheframework. A standard VAE, which uses a simple unit Gaussian prior over the latent space, is not explicitly optimised for discovering cluster structure. VectorQuantised VAE (VQ-VAE) [3]quantisethefeaturevectorinthelatentspacesas a codebook. The quantised feature vectors,can be further merged if necessary. VQ-VAE has demonstrated its effectiveness in capturing complex musical features. For example, the Jukebox model by Dhariwal et al. employed [19] VQ-VAE to generate high-fidelity music in various genres, showcasing its capability to handle raw audio and symbolic music data. Moreover, VQ-VAE also demonstrates the ability to model discrete latent spaces for sequential data by time-series generation [20]. As a result, this paper proposes VQ-VAE for the study. 3Methods 3.1 Data Preparation Fig.1: Data preprocessing that converts original tempo curves to the input of the system. The expressive timing is represented in the form of tempo curves as the previous work [2]. However, to make the proposed model being capable of potential Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 975
Z. Chai and S. Li developments for automatic phrasing. The pipeline for the data preparation stage is shown in Fig. 1. Assuming a series of beat timing !T=(t1,t 2,...,t n,t n+1)where tn+1 represents the end of the last note. The tempo of the beat ⌧could be calculated as ⌧=60 ti+1ti. The raw vector of tempo T=(⌧1,⌧ 2,...,⌧ n)is smoothed by the moving average with a window of three samples. The tempo curve is then truncated according to phrase information and standardised with a regulated mean of 1. Aphrasingmaskinintroducedasapreparationofautomaticphrasingin the future. For the input smoothed and standardised tempo curves ˆ T,fourprior beats and four posterior beats of a phrase are also included. A phrasing mask M is applied with the beats in a phrase labelled as "1" and the other beats as "0". The expressive timing within a phrase can be taken as the element-wise product between the standardised tempo curve and the mask, i.e. ˘ T=ˆ TM. Classical data argumentation methods are then applied for better data availability and generalisation. The data augmentation methods used are: –Gaussian Noise:Addsnoisen⇠N(0,),⇠Uniform(0.05,0.15),clipped to ensure non-negativity: ˘ T0=clip(˘ T+n,0,max( ˘ T)). – Scaling:Multipliesbys⇠Uniform(0.8,1.2):˘ T0=s·˘ T. – Local Perturbation:Perturbsarandom10-stepsegmentstartingatk⇠ Uniform(0,T 10) by p⇠Uniform(0.2,0.2),clipped:˘ T0[k:k+ 10] = clip(˘ T[k:k+ 10] + p, 0,max( ˘ T)). –Sinusoidal Modulation:Adds0.05 ·sin(2⇡ft/T ),f⇠Uniform(0.5,1.5), t=[0,...,T 1], with clipping. Fig.2: The complete structure of the VQ-VAE model in the experiment. 3.2 VQ-VAE The proposed VQ-VAE system has an asymmetric code structure as shown in Fig. 2. In the encoder, the argumented masked standardised tempo curve is padded with zeros for the identical lengths of inputs within the same batch. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 976
Cluster Expressive Timing with VQ-VAE Encoder The padded tempo curves undergo an initial linear transformation to increase their dimensionality before further processing. The proposed system employs a two-stage strategy: a Bi-directional LSTM (BLSTM) layer to extract local temporal features from the transformed input, followed by a multi-head attention layer to capture global temporal features. The time step for the BLSTM is set equal to the padded tempo curve length. The subsequent multi-head attention layer processes the local temporal features extracted by the BLSTM, incorporating information from phrasing masks. The attention layer’s time step is aligned with the phrase lengths. Essentially, the attention layer inverts the mask to assign higher attention weights to interior beats, prioritising expressive timing gestures while ignoring boundary beats. The interaction between queries and keys in the attention layer allows indirect incorporation of contextual information from all beats, as the input features retain the full sequence context. The value emphasises interior patterns while maintaining global context, serving as input to both boundary prediction and latent encoding. With another layer of linear transform, the output of the attention layer, denoted as A,ismappedtoacontinuouslatentspace: ze=WfA+bf,(1) where Wf2RE⇥H,bf2RE,andEis the codebook embedding dimension, producing ze2RB⇥T⇥E.Eachze,b,t 2RErepresents a latent vector for time step t in batch b,encodingrefinedfeaturesthatcombineinterior-focusedattentionand full-sequence context. zeis used for computing the phrase-level representation and quantisation. Latent Space and Sampling As the classical VAE, the encoder integrates statistical features and temporal information to compute a phrase-level latent vector ze,phrase 2RE.First,statisticalfeaturesarecomputedfromthemasked sequence xm, which isolates interior beats to capture phrase-specific rhythmic characteristics: –Standard deviation: b=q1 Pmb,t Ptmb,t(xb,t µb)2,µb=Ptmb,txb,t Pmb,t , measuring the variability of tempo curve. –Peak count: pb=Pt1(xb,t>1+2b)mb,t Pmb,t ,countingsignificantpeaksontempo curves. –Skewness: sb=Ptmb,t(xb,t1)3 3 bPmb,t , assessing skewness of tempo curves. –Spectral energy: eb=1 5P5 k=1 |FFT(xbmb)k|,capturingfrequency-domain energy for tempo curves. These form sb=[b,p b,s b,e b]2R4are embedded via: se=Wssb+bs,(2) where Ws2RE⇥4,se2RE. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 977
Z. Chai and S. Li Next, the masked-weighted mean and maximum of zeare calculated to summarise the sequence while emphasising interior regions: ze,mean =Ptze,b,tmb,t Pmb,t ,ze,max = max t(ze,b,t ·mb,t),(3) where ze,mean,ze,max 2REcapture the average and peak latent features of interior beats, respectively, focusing on core musical content via the mask m. These are concatenated with seto form [ze,mean;ze,max;se]2R3E, which is projected via: ze,phrase =Wp[ze,mean;ze,max;se]+bp,(4) where Wp2RE⇥3E,bp2RE. This integrates temporal (from LSTM/attention) and statistical information into a single vector per phrase, connecting the sequencelevel zeto a compact representation for quantisation. Quantisation in Latent Space The phrase-level representation ze,phrase connects the encoder to the quantisation layer, where it is mapped to a discrete codebook vector: zq=C⇥arg min kkze,phrase ckk2 2⇤,(5) using a codebook C2RK⇥E, where Kis the number of embeddings. Indices are sampled via softmax with temperature ⌧=0.5: p(k)= exp(kze,phrase ckk2 2/⌧) Pk0exp(kze,phrase ck0k2 2/⌧).(6) The quantised vector is expanded to zq2RB⇥T⇥E,linkingthecompactrepresentation from the encoder to the decoder for reconstruction. Decoder The decoder mirrors the encoder in a symmetric structure, reversing the process to reconstruct tempo curves from quantised latent representations, using a multi-head attention layer followed by a Bi-directional LSTM layer, and concluding with a linear transform to output the final sequence. 3.3 Clustering The clustering method for tempo curve analysis employs a Vector Quantised Variational Auto-Encoder (VQ-VAE) with a discrete codebook. The codebook is initialised using mini-batch K-Means on tempo curve ze,sampledupto200,000 points with a batch size of 3,072 and 50 initialisations, minimising: J= n X i=1 min µj2C d X k=1 (zi,k µj,k)2.(7) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 978
Cluster Expressive Timing with VQ-VAE Codebook vectors wjare initialised uniformly in [1 k,1 k], where kis the number of embeddings, and standardised post-training to unit length, ensuring: d X i=1 w2 j,i =1,(8) achieved via wj wj/qPd i=1 w2 j,i,projectingvectorsontotheunithypersphere. Quantisation assigns zeto codebook entries using squared Euclidean distance: d(ze,w j)= d X i=1 (ze,i wj,i)2,(9) with a softmax probability (⌧:1.0to0.5overepochs): p(j|ze)= exp(d(ze,w j)/⌧) Pk l=1 exp(d(ze,w l)/⌧).(10) Assignments are processed in blocks of 1,000. Clusters are merged if Pd k=1(wi,k wj,k)2<threshold (e.g., 0.0016), where wi,wj2Rdrepresents codebook vectors, and Dis the embedding dimension. 4ExperimentsandResults 4.1 Datasets The datasets used in this paper is a combination of two datasets. The public Mazurka datasets and the private Islamey dataset. For Mazurka dataset, there are five pieces of Chopin Mazurka with multiple performances by different pianists: Op. 17/4 (13 phrases, 63 performances), Op. 24/2 (30 phrases, 64 performances), Op. 30/2 (8 phrases, 34 performances), Op. 63/3 (13 phrases, 95 performances) and Op. 68/3 (10 phrases, 51 performances). The private Islamey dataset contains 25 performances where each performance has 40 phrases. In total, there are 5756 tempo curves for intra-phrase expressive timing. 4.2 Training of VQ-VAE The training process is performed with a GPU of RTX 4060ti(8G) with a implementation with PyTorch with the hyper-parameter shown in Table 1. The VQ-VAE is trained with a composite loss with all items using mean squared error: L=Lrecon +Lvq +Lcommit,(11) where: –Lrecon =kˆ xxk2 2,isinmeansquarederror,ensuringˆ xmatches x. –Lvq =ksg[ze]zqk2 2,aligningzewith zq. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 979
Z. Chai and S. Li –Lcommit =kzesg[zq]k2 2,committingzeto the codebook. With the use of the AdamW optimiser with a cosine annealing scheduler, the training process stops when both training and validation losses do not decrease by a minimum threshold for five consecutive epochs. Table 1: Hyper-parameters of the VQ-VAE model that is used in the training process. The codebook used is specified in the context separately. Parameter Name Value Parameter Name Value Input Dimension 64 Weight Decay 0.01 Hidden Dimension 256 Commitment Loss Weight 2 Attention Heads 4 Codebook Diversity Weight 0.15 Batch Size 8 Initial Quantisation Temperature 1.0 Max Epochs 40 Final Quantisation Temperature 0.5 Learning Rate 0.0005 Quantisation Block Size 1000 4.3 Number of Clusters for Expressive Timing The resulting number of clusters for expressive timing in the proposed analysis will be presented first. There are three factors impact the resulting number of clusters: the size of codebook (i.e. the number of codes in a codebook), the length of code (i.e. the dimensional of embedded coding) and the threshold that merges similar vectors. According Gratton et al. [21], the Just Noticeable Difference (JND) in the circumstances of expressive timing could be around 12%. For a consistent sequence of stimuli, the JND may be significantly improved to 3%. Considering the expressive timing within a phrase could be steady (but no consistent) under special circumstances, the JND of expressive timing in the proposed experiment is taken as 4%, which corresponds to a threshold of 0.0016 due to the normalised codes. With the given threshold, the range of code length is decided. Given the fact that the phrase lengths in the dataset are usually no more than 24 beats, the candidate number of dimensions in the codebook is proposed to 32, 64, 128, 256 and 512, where the lowest dimension should be able to keep all information of tempo curves without loss. Moreover, the number of clusters for intra-phrase expressive timing is usually less than 10 in a piece of music [2]. As there are six pieces of music, the possible number of clusters for the whole dataset ranges from around 10 (assuming the same clusters of expressive timing across pieces) to 60 (if no clusters of expressive timing is shared between pieces). As a result, the size of the codebook in VQVAE is set to 8, 16, 32, 64 and 128, which set the maximum number of clusters in expressive timing from marginally fit to more than adequacy. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 980
Cluster Expressive Timing with VQ-VAE Table 2: The number of clusters resulted by VQ-VAE with different settings of tables. Ddim refers to the dimension of code (or code length) and Dsize refers the number of codes in codebook (i.e. codebook size). Ddim =32Ddim =64Ddim =128Ddim =256Ddim =512 Dsize =8 11 1 1 1 Dsize =16 11 2 3 7 Dsize =32 11 3 1220 Dsize =64 12 2 6126 Dsize =128 44125128127 As shown in Table 2,alargersizeofcodebookandalargerdimensionof code lead to more clusters in general. However, there are few exceptions: with Dsize = 64,Ddim = 256 leads to 61 clusters whereas Ddim = 512 leads to only 26 clusters. There are only a few minor cases (resulting 1 more or 1 less cluster). To select the number of clusters, the model selection test [2]isperformed as the baseline. In short, Assuming f(t)=P4 i=0 aixi,theregressedpolynomial coefficients A=(a0,a 1,a 2,a 3,a 4)for ˘ Tcan be fitted to a Gaussian Mixture Model (GMM) for clustering. The whole dataset used in the experiment is regressed by a polynomial function of a fourth order. A Gaussian Mixture Model (GMM) with the value of k set between 2 to 128 is trained. AIC (Akaike Information Criterion) measures the performance of resulting GMMs as demonstrated by Figure 3.Accordingto the model selection test, GMM with 12 Gaussian components outperforms other models, which leads to 12 clusters of expressive timing. Given the experiment presented in Table 2, we choose a codebook with Ddim = 256 and Dsize = 32. 4.4 Evaluation with Cluster Distribution The validation of expressive timing clustering involves comparing the cluster distributions from the VQ-VAE model against those from a classical Gaussian Mixture Model (GMM) [2]. The difference between the two distributions, resulted from GMMs (Q)andVQ-VAEsystems(P)isquantifiedusingJensenShannon (JS) Divergence. JS Divergence is a symmetrised measure based on Kullback-Leibler (KL) Divergence. The KL Divergence is defined as: KL(P||Q)= Px2P(x) log ⇣P(x) Q(x)⌘where xrepresents a specific cluster. The resulting JS Divergence is calculated as: JS(P||Q)=1 2(KL(P||Q)+KL(Q||P)). The clustering results are considered valid if the divergence between the VQVAE and GMM distributions is comparable to the natural variability found within GMMs run with different initial conditions. Ten VQ-VAE systems, using slightly varied thresholds for the same number of clusters (set between 0.0016 to 0.0020, 4% to 4.5% differences on average), are compared against ten GMMs with different initialisations. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 981