scieee AI-readable full text Open interactive document viewer

Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation

Ebrahimzadeh, Maral; Bernardes, Gilberto; Stober, Sebastian

Abstract

State-of-the-art symbolic music generation models have recently achieved remarkable output quality, yet explicit control over compositional features, such as tonal tension, remains challenging. We propose a novel approach that integrates a computational tonal tension model, based on tonal interval vector analysis, into a Transformer framework. Our method employs a two-level beam search strategy during inference. At the token level, generated candidates are re-ranked using model probability and diversity metrics to maintain overall quality. At the bar level, a tension-based re-ranking is applied to ensure that the generated music aligns with a desired tension curve. Objective evaluations indicate that our approach effectively modulates tonal tension, and subjective listening tests confirm that the system produces outputs that align with the target tension. These results demonstrate that explicit tension conditioning through a dual-level beam search provides a powerful and intuitive tool to guide AI-generated music. Furthermore, our experiments demonstrate that our method can generate multiple distinct musical interpretations under the same tension condition.

Full text

Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation Maral Ebrahimzadeh1,GilbertoBernardes 2,andSebastianStober 1 1Artificial Intelligence Lab, Otto-von-Guericke-University, Magdeburg, Germany 2University of Porto, Faculty of Engineering and INESC TEC, Porto, Portugal [email protected] Abstract. State-of-the-art symbolic music generation models have recently achieved remarkable output quality, yet explicit control over compositional features, such as tonal tension, remains challenging. We propose a novel approach that integrates a computational tonal tension model, based on tonal interval vector analysis, into a Transformer framework. Our method employs a two-level beam search strategy during inference. At the token level, generated candidates are re-ranked using model probability and diversity metrics to maintain overall quality. At the bar level, a tension-based re-ranking is applied to ensure that the generated music aligns with a desired tension curve. Objective evaluations indicate that our approach e!ectively modulates tonal tension, and subjective listening tests confirm that the system produces outputs that align with the target tension. These results demonstrate that explicit tension conditioning through a dual-level beam search provides a powerful and intuitive tool to guide AI-generated music. Furthermore, our experiments demonstrate that our method can generate multiple distinct musical interpretations under the same tension condition. Keywords: Symbolic Music Generation ·Beam Search ·Transformers. 1Introduction Automatic music generation has fascinated researchers for centuries, from early stochastic experiments such as musical dice games [21] to modern methods based on deep learning. Recently, Transformer models [27] have shown impressive musical coherence, mirroring successes in natural language processing [6] and computer vision [29]. Despite these advances, the issue of interactive control remains paramount. Musicians and composers require intuitive tools to e!ectively guide generative All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 795 M. Ebrahimzadeh et al. models, ensuring that the output reflects their unique artistic intentions. Previous works have explored controllability over style, instrumentation, and emotion [8,17,24,26], yet tonal tension, a key factor that influences perceived movement and expression in melody and harmony [16], is significantly underexplored. Several frameworks exist for modeling tonal tension, making it a viable direction for controllable music generation. Lerdahl’s influential model [16] is often impractical due to manual hierarchies and parameter tuning. MorpheuS [12,13] uses Spiral Array-based metrics (dissonance, momentum, tensile strain) but lacks sensitivity to hierarchical harmony. The Tonal Interval Vector (TIV) method [2,19] computes the tonal tension in the Tonal Interval Space (TIS) using a discrete Fourier transform of the chroma vector. It yields a computationally e"cient and perceptually relevant metric that captures multilevel pitch configurations and chord similarity, addressing the limitations of earlier models. To achieve controllability in music generation through tonal tension, previous studies primarily employed MorpheuS metrics (cloud diameter and tensile strain) [4,10,11], o!ering limited flexibility. The TIV-based method, despite its considerable potential, has been explored only once in fixed-voicing chord progression generation tasks (i.e., where the same chord cardinality is maintained throughout the progression), matching target tension curves through bio-inspired optimization techniques [20]. However, reliance on evolutionary search methods severely restricts real-time controllability and seamless integration with contemporary sequence generation models. Therefore, substantial unexplored potential remains for utilizing TIV-based tonal tension metrics. Controllability in music generation can be achieved through training-based or inference-time methods. Training-based approaches incorporate explicit control tokens for attributes like tempo or style [11,24,25], but require retraining and offer limited flexibility. In contrast, inference-time techniques allow dynamic plugand-play control without modifying the underlying model [7]. Beam search [22] can also support controllability by selecting candidate outputs based on external constraints, even when those constraints are non-di!erentiable. Motivated by these limitations and opportunities, we propose to integrate a TIV-based method into Transformer-based symbolic music generation. Implementation of this approach so far was limited to voice-leading calculations for chords with identical cardinality. We improve the implementation to handle chords of varying cardinalities, reflecting realistic musical scenarios. Our main contribution is a novel dual-level beam search strategy that explicitly guides music generation in two complementary stages. First, candidates generated at the token level are re-ranked based on model probability and diversity metrics, ensuring musical quality. Second, at the bar level, we apply tensionbased re-ranking rules to align the generated output with a user-specified tonal tension curve. Our comprehensive evaluations confirm the e!ectiveness of the approach in both objective measures and subjective listening tests. The implementation of the dual-level beam search and tonal tension computation described in this paper is available at https://github.com/MaraalE/tension-beamsearch. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 796 Tonal Tension Conditioning in Symbolic Music Generation 2RelatedWork 2.1 Controlling Tonal Tension in Symbolic Music Generation Prior works controlling tonal tension largely relied on Spiral Array metrics, as seen in MorpheuS [12, 13], which morphs music fragments via variable neighborhood search. Other works integrated Spiral Array metrics into VAEs [11], Transformer-based infilling models [10], or LSTM networks [28], generally emphasizing stylistic consistency over explicit tension shaping. MoodLoopGP [4] similarly employed Spiral Array metrics, primarily targeting emotion-driven music generation. Earlier work by Melo and Wiggins [18] demonstrated a neuralnetwork-based approach to chord progression generation driven by tension curves. While these methods often overlook critical dimensions like voice-leading, the TIV-based approach models voice-leading parsimony (i.e., minimizing motion between chords) and harmonic relationships [19], though it has so far seen limited use, only in o#ine fixed-voicing chord generation via bio-inspired optimization techniques [20]. Our work addresses this gap, integrating TIV metrics with modern Transformer-based models through dual-level beam search. 2.2 Inference-Time Controllability in Symbolic Music Generation While training-based methods embed control directly into model parameters, inference-time methods dynamically guide generation without retraining. Approaches like Bardo Composer [7] employ bi-objective beam search for real-time emotional control, whereas others use Monte Carlo Tree Search (PUCT) [8], hierarchical Transformers with DTW-based re-ranking [5], or text-prompt conditioned generation [1]. Bjare et al. [3] recently guided music generation toward surprisal profiles using beam search based on instantaneous information content. Most existing methods employ single-stage beam search or sampling-based adjustments. In contrast, our proposed dual-level beam search explicitly incorporates control at the token level (probability and diversity) and bar level (tonal tension alignment); This structured design facilitates more fine-grained control of musical properties. 3Method Our proposed method integrates a TIV-based tension metric into a Transformerbased music generation framework. Rather than only applying explicit tension conditioning during the training phase, our approach focuses mainly on inference-time control, enabling flexible and real-time adjustment of tonal tension without retraining the model. 3.1 Tonal Tension Metric To model tonal tension, we employ a computational method based on the TIV representation [2,19]. A fundamental building block of this model is the chroma Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 797 M. Ebrahimzadeh et al. vector, a twelve-dimensional vector representing pitch classes, commonly used to encode musical information across multiple pitch configurations such as individual pitches, chords, and keys. However, standard distance metrics applied directly to chroma vectors do not capture musically meaningful harmonic relationships. To overcome this, the TIV model applies the Discrete Fourier Transform (DFT) to chroma vectors, projecting them into a six-dimensional complexvalued space, termed the Tonal Interval Space (TIS), where intervallic structures become more distinguishable due to the DFT highlighting periodicities in pitchclass content [2]. Given two tonal interval vectors, T1and T2,representingchords or keys projected into the TIS, we compute their distance based on musical context using two core metrics. The Euclidean distance between two tonal interval vectors captures perceptual similarity and consonance; smaller distances indicate smoother harmonic transitions at the same musical level (e.g., two chords that are close together tend to preserve their harmonic role in functional harmony): d(T1,T2)=! " " # N→1 $ i=0 (T1,i →T2,i)2(1) In contrast, the angle between vectors captures harmonic alignment across different levels (e.g., chord-to-key), by assessing common tone retention: d(T1,T2) = arccos %↑T1,T2↓ ↔T1↔↔T2↔&(2) We use three main components from a tonal tension model [19], which is based on these distance metrics. 1. Tonal Distance: This includes three distinct subcomponents: (a) Distance from the current chord to the previous chord (same musical level, thus using the Euclidean distance). (b) Distance from the current chord to the overall key (across musical levels, thus using angle). (c) Distance from the current chord to its tonal function (tonic I, subdominant IV, or dominant V; also angle-based). 2. Tonal Dissonance: It is computed as the normalized Euclidean norm subtracted from unity, aligning higher values with greater internal dissonance. ddiss =1→↔Ti↔ ↔Tmax↔(3) Here, Tidenotes the tonal interval vector of the current chord, and ↔Ti↔ its Euclidean norm. ↔Tmax↔is the maximum observed TIV norm across all chords, used to normalize the dissonance value between 0 and 1. 3. Voice Leading: This is captured by evaluating melodic stability between notes of consecutive chords. The voice-leading tension for each chord Tin progression Pis computed as: m(Ti,P)= V $ l=1 e 1 0.05sµ(nli,nli→1)(4) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 798 Tonal Tension Conditioning in Symbolic Music Generation Here, Vrepresents the number of voices (notes) in the chord. For each voice l,thetermsdenotes the melodic interval between consecutive notes nl→iand nli→1, measured in semitones. The term µ(nli,n li→1)measures the perceptual distance between the notes nli(note lin chord Ti)anditscorresponding voice nli→1in the previous chord Ti→1, calculated in the TIS. Thus, smaller melodic intervals and lower perceptual distances result in reduced voiceleading tension, aligning with established perceptual principles of musical stability. Notably, the original implementation assumes chords with equal numbers of pitch classes, limiting its applicability to realistic musical contexts. To address this limitation, we use a publicly available implementation by Dmitri Tymoczko3to compute the melodic interval s(number of semitones), which is capable of handling chords with di!ering cardinalities. This modification enables more flexible and realistic voice-leading computation without altering the theoretical foundation of tonal tension. We compute a weighted combination of these components using scalar weights adopted from the original implementation [19], where dissonance and voice leading are scaled by 30.3 and 2.71 respectively. 3.2 Training Phase We adopt the REMI+ [24] representation, an extension of the Revamped MIDI (REMI) [15]. REMI represents each note through four consecutive tokens encoding position, pitch, velocity, and duration, and additionally includes chord and tempo tokens. REMI+ further extends this representation by adding barlevel time-signature tokens and per-note instrument tokens, supporting multiinstrument compositions with variable time signatures. We use a standard Transformer model in a translation-style setup. Specifically, during training, our Transformer encoder receives bar-level control tokens, while the decoder outputs the corresponding REMI+ token sequences. Following the setup in Figaro [24], we adopt control tokens for time signature, instrument list, and note density, to which we add tonal tension as an additional conditioning feature. These attributes are encoded as discrete tokens selected from a predefined dictionary, clearly defining each bar’s musical context for the Transformer to learn from. For instance, the pitch of a note might be represented by a token PITCH_32, which is then mapped to its corresponding index in the vocabulary. The model predicts each REMI+ token sequentially conditioned on previous tokens and the provided control tokens. 3.3 Generation Phase: Dual-Level Beam Search Our core contribution is a dual-level beam search strategy applied at inference (Algorithm 1), simultaneously enforcing local musical quality and diversity at the token level, and global tonal tension alignment at the bar level. 3https://dmitri.mycpanel.princeton.edu/voiceleading_utilities.py Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 799 M. Ebrahimzadeh et al. Algorithm 1 Dual-Level Beam Search (Inference) 1: procedure DualLevelBeamSearch(M,R,T,K) 2: Input: Transformer model M, reference piece R,targettensioncurveT,beam width K 3: beams →{(BOS, 0)} ωInitialize beam with start token 4: while generation incomplete do 5: candidates →{} 6: for all (seq, score) in beams do 7: tokens →NucleusSample(M, seq, p=0.9,K) 8: for all token in tokens do 9: new_seq →append(seq, token) 10: bar_cand →current bar segment of new_seq 11: bar_ref →corresponding bar from R 12: div →DiversityMetric(bar_cand, bar_ref) 13: token_score →LMnorm(new_seq)+ε·div 14: candidates.add((new_seq, token_score)) 15: end for 16: end for 17: beams →TopK(candidates, K) 18: if any beam completes a bar then 19: for all (seq, score) in beams do 20: tcand →tensions of completed bars in seq 21: ttarget →corresponding values from T 22: tscore →TensionSimilarity(tcand,ttarget) 23: score →LMnorm(seq)+tension_weight ·tscore 24: end for 25: beams →TopK(beams, K) 26: end if 27: end while 28: return Top Ksequences in beams 29: end procedure Token-Level Beam Expansion (Algorithm 1, lines 6–16) At each token generation step, we use nucleus (top-p)sampling[14]tosampleKcandidate tokens from the current Transformer context. For each active beam (partial sequence), these candidate tokens are appended to form new sequences. Each candidate sequence is scored by combining the length-normalized log-probability of the Transformer, denoted LMnorm, with a diversity penalty derived from its current bar segment: Score(x)=LMnorm(x)+ω·(PV(x)+DV(x)+PE(x)) (5) where ωis a weighting factor (0.7 in our experiments), and PV, DV, and PE represent pitch variety, duration variety, and pitch 3-gram entropy. We compare the diversity metrics of each candidate bar segment with a corresponding bar from a reference piece R,encouragingstylisticallycoherentdiversity.After scoring, we prune the beam, retaining the top Kcandidates. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 800 Tonal Tension Conditioning in Symbolic Music Generation Bar-Level Re-ranking (Algorithm 1, lines 17–26) Whenever a candidate sequence completes a bar, a bar-level re-ranking step is applied. For each candidate, we calculate the average tonal tension for all completed bars, forming acandidatetensioncurve(tcand), which is compared with the corresponding portion of a user-specified target tension curve T(ttarget). Similarity between the two tension curves is computed using Pearson correlation when the variance of ttarget exceeds a threshold (0.001); otherwise, absolute di!erence is used to ensure stability, as correlation becomes unreliable on nearflat curves. Candidates are then re-ranked using a combined score that balances the same length-normalized language-model score, LMnorm, with the tension similarity tsimilarity: Final Score =LMnorm +tension_weight ·tsimilarity (6) We again prune beams to retain the top Ksequences based on the updated final scores. Upon completion, the top-ranked sequence is considered the primary output, while remaining candidates provide alternative musical realizations under the same tension profile, highlighting our method’s flexibility. 4ExperimentalSetup 4.1 Dataset We utilize the Lakh MIDI-Matched dataset [23], a large-scale collection of symbolic MIDI files. Since our approach requires chord information to compute tonal tension, we preprocess all MIDI files using Midi Miner [9] to automatically identify and extract tracks containing chords. After filtering out MIDI files without chords, we obtain a final dataset of 25,555 MIDI files suitable for computing bar-level tension tokens, which we then split into training, validation, and test sets at proportions of 0.85, 0.10, and 0.05, respectively. 4.2 Model and Training We use a standard Transformer architecture with a model dimension of 512, employing 12 attention heads, 4 encoder layers, and 6 decoder layers, with a maximum sequence length of 256 tokens [24]. Models are trained using crossentropy loss for 12 epochs over approximately 28 hours using an NVIDIA A40 GPU. 4.3 Inference During inference, we employ nucleus sampling with a threshold of 0.9, combined with a dual-level beam search. We set the beam width to 8, generating 8candidatesequencesateachtokenstep.Atthetokenlevel,were-rankthese eight candidates using a diversity weight of 0.7 and model probability scores. Subsequently, at the bar-level re-ranking stage, we select the top three beam candidates based on a tension weight of 4.0, with a sampling temperature of 0.9, thus guiding generation explicitly toward the desired tension curve. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 801 M. Ebrahimzadeh et al. 5Evaluation 5.1 Objective Evaluation Our evaluation assesses both the musical quality and the tonal tension accuracy of generated samples under varied training and inference conditions. Experiments isolate the impact of explicitly conditioning tension during training and inference. We measure how closely the generated output aligns with our conditioning targets using Instrument F1,Note Density Accuracy,andTension Correlation, which correspond directly to our control tokens. Additionally, we measure Groove Similarity as an indicator of rhythmic coherence, thus reflecting overall musical quality [24]. Evaluations are conducted primarily on 8-bar samples, with additional tests on 16-bar sequences to assess scalability. Table 1. Objective evaluation metrics comparing baseline Transformer models (with and without tension conditioning) against our dual-level beam search strategy. Results for 8-bar samples represent the full test set; the 16-bar result (Dual Beam 1) uses 200 randomly selected samples. “Dual Beam 1–3” denote the first-, second-, and third-best candidates from our inference method (not separate models). Model Inference Bars Instr. F1 Note Density Groove Sim Tension Corr Baseline Normal 80.82 0.88 0.52 0.16 Baseline + Tension Normal 80.83 0.62 0.54 0.18 Baseline + Tension Dual Beam 1 80.86 0.85 0.56 0.50 Baseline + Tension Dual Beam 2 80.86 0.72 0.55 0.48 Baseline + Tension Dual Beam 3 80.86 0.70 0.56 0.45 Baseline + Tension Dual Beam 1 16 0.85 0.62 0.56 0.42 We consider three experimental conditions: – Baseline Model: Transformer trained using Figaro expert settings [24] with control tokens (time signature, instrument list, note density), without tension control. – Baseline + Tension: Baseline model augmented with an additional tension control token during training. – Dual Beam Inference (Our Main Method):Baseline+Tensionmodel combined with our dual-level beam search at inference, explicitly aligning generated tension with target profiles. Table 1 summarizes the key evaluation results. The baseline model (without tension control) achieves a relatively low tension correlation (0.16), indicating limited alignment with the target tension curves. Adding a tension control token during training (Baseline + Tension, normal inference) only slightly improves the tension correlation (0.18), suggesting that simply adding a control token without specialized inference is insu"cient. Our proposed dual-level beam search substantially enhances the tension correlation. The top beam candidate (Dual Beam 1) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 802 Tonal Tension Conditioning in Symbolic Music Generation achieves a significantly higher tension correlation (0.50), indicating a good alignment with the target tension profiles. Alternative beam candidates (Dual Beam 2andDualBeam3)similarlydemonstratemeaningfulimprovements(0.48and 0.45 respectively), confirming that our inference method consistently provides multiple viable interpretations aligned closely to target tension. Furthermore, extending the generation length to 16 bars shows a slight reduction in tension correlation (0.42), highlighting the room for future improvements to maintain control over longer musical sequences. Table 2. Detailed tension correlation analysis of dual-level beam candidates, showing averages with and without filtering negative and low-variance cases. Model Avg. Tension Corr Avg. Corr (Filtered) Median Corr Dual Beam 1 0.50 0.62 0.57 Dual Beam 2 0.48 0.61 0.56 Dual Beam 3 0.45 0.58 0.52 Also, Table 2 examines tension correlation by filtering out negative and lowvariance cases (approximately 12% of samples), since correlation becomes less informative for nearly constant curves. Under this analysis, Dual Beam 1 achieves the highest average correlation (0.62), demonstrating strong alignment in perceptually meaningful cases. Unfiltered results still show Dual Beam 1 leading (0.50), confirming consistent tension control across diverse outputs. Overall, these results demonstrate that our dual-level beam search inference e!ectively and consistently improves tonal tension control without sacrificing other aspects of musical quality such as instrument accuracy, groove similarity, and note density. In addition, generation times average 7 minutes for an 8-bar sample, confirming feasibility for practical, o#ine use cases despite being slower than simple sampling methods. 5.2 Subjective Evaluation We conducted a listening study to evaluate the perceptual e!ectiveness of our method specifically in terms of tonal tension alignment, which is the primary focus of this work. Five representative tension curves were selected from computed tension values of samples in the dataset (Figure 1). We generated two music samples per target tension curve, resulting in ten stimuli (Samples A–J). 18 participants, mostly musically experienced, were asked to identify the tension curve that best matched each sample. As shown in Figure 2, 4 out of 10 samples (A, E, G, and J) were clearly identified by most listeners, indicating strong perceptual alignment. Many mismatches among the remaining samples were likely due to similar rising shapes, as Curves 1, 2, 3, and partially Curve 4 all exhibit upward tension profiles with variations in timing and slope. For instance, samples from Curve 1 (B, H) and Curve 3 (F, I) were often misclassified as Curve 2, suggesting listeners perceived the general rise but Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 803