Full text
GuitarFlow: Realistic Electric Guitar Synthesis From Tablatures via Flow Matching and Style Transfer Jackson Loth1[0009000287091218],PedroSarmento 1,2[0000000245180194], Mark Sandler1[0000000256918107],andMathieu Barthet1,3[0000000298691668] 1Centre for Digital Music, Queen Mary University of London {j.j.loth,p.p.sarmento,mark.sandler,m.barthet}@qmul.ac.uk 2Music.AI 3Aix-Marseille Univ CNRS PRISM Abstract. Music generation in the audio domain using artificial intelligence (AI) has witnessed steady progress in recent years. However for some instruments, particularly the guitar, controllable instrument synthesis remains limited in expressivity. We introduce GuitarFlow, a model designed specifically for electric guitar synthesis. The generative process is guided using tablatures, an ubiquitous and intuitive guitar-specific symbolic format. The tablature format easily represents guitar-specific playing techniques (e.g. bends, muted strings and legatos), which are more difficult to represent in other common music notation formats such as MIDI. Our model relies on an intermediary step of first rendering the tablature to audio using a simple sample-based virtual instrument, then performing style transfer using Flow Matching in order to transform the virtual instrument audio into more realistic sounding examples. This results in a model that is quick to train and to perform inference, requiring less than 6 hours of training data. We present the results of objective evaluation metrics, together with a listening test, in which we show significant improvement in the realism of the generated guitar audio from tablatures. Keywords: Flow matching ·Style transfer ·Guitar ·Synthesis ·Audio effects. 1Introduction Recent advances in generative audio systems have given rise to impressive textto-music systems that allow users to generate full songs from simple text prompts [14] [15]. While this is very interesting from a science and technology perspective, it is arguably lacking as a creative tool due to a lack of fine-grained control over the music missing. Music generation systems with full control over the notes do exist, but typically in the symbolic domain [39] [30], requiring a way to transform a symbolic music score into audio. While this can be achieved Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 46
2 J. Loth et al. through virtual instruments to varying degrees of success, guitars are particularly expressive instruments which are difficult to represent through the standard Musical Instrument Digital Interface (MIDI) [33]. Guitar timbre itself can even be perceived differently when performed in different playing styles [32]. Previous attempts at synthesising expressive music instruments have largely focused on MIDI or representations derived from MIDI [45] [24]. This paper explores the task of synthesising electric guitar by processing synthetic audio rendered from an expressive symbolic musical representation. Instead of MIDI, we use guitar tablatures (see Figure 1) to represent the musical content of a desired audio synthesis. This allows us to incorporate expressive playing techniques such as slides, bends, hammer-ons, etc, something that other music instrument synthesis models struggle to do without additional conditioning mechanisms. We also introduce an intermediary step of first rendering a guitar tablature to audio using a quick and simple sample-based virtual instrument. Our model performs style transfer on this simple audio rendering and transforms it into more realistic sounding audio. We use ‘style‘ here to refer to various performance-related qualities which distinguish natural and more synthetic performances. The style transfer is achieved using Flow Matching [28], a recent generative modeling paradigm. This method allows us to greatly simplify the complexity of the training and inference pipeline, requiring less data and less time to train. The contributions of this paper are summarized as: (1) GuitarFlow, a novel model and methodology for realistic electric guitar synthesis from guitar tablatures using Flow Matching and style transfer; (2) an evaluation of the model using both objective metrics and a subjective listening test which shows the success of GuitarFlow in transforming audio rendered with a virtual instrument to sound more realistic; (3) a public repository4of code to allow other researchers to replicate and extend the research. This paper demonstrates the potential of the technique in greatly lowering data and computational requirements when training generative audio models. 2Background 2.1 Guitar Tablatures Guitar tablatures (refer to Figure 1), also known as tabs, are symbolic representations of guitar music and have seen increased attention in recent years due to their ability to easily represent guitar-specific expressions [38]. In contrast to MIDI, which simply represents a note’s pitch and velocity over time, tabs represent both the fret and string number of a guitar. They can also support expressive playing techniques such as bends, hammer-ons, pull-offs, strum directions, and more. Tabs have seen an increase in attention from the MIR research community in the past few years in areas such as guitar tablature generation [39] [40], automatic guitar transcription [46] [8] and tablature prediction from MIDI [11]. 4https://github.com/JackJamesLoth/GuitarFlow Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 47
GuitarFlow 3 Fig. 1. Example of a guitar tablature, from the Guitar Pro editing software. 2.2 Style Transfer Broadly speaking, style transfer involves transforming the “style” of a signal while preserving the “content” of signal. The task is popular in the image domain [16] [49] . A common assumption is that the audio of an instrument can be broken up into “content”, referring to the pitch, length and arguably loudness or intensity of the notes, and “style”, referring to the actual sound of the instrument and expressiveness of the performer playing those notes. Some musical style transfer works use a GAN-based approach, training a generator to convert an audio input to the desired timbre [22] [48]. Others use an autoencoder to produce disentangled representations of content and style [35] [1]. As with generative tasks recently, Diffusion has also become a popular choice [6] [21], while differential digital signal processing (DDSP) [13] provides an alternative to all of these models by allowing trainable DSP functions. 2.3 Music Instrument Synthesis Instruments such as guitar have traditionally been simulated by modeling the physics of the instrument and strings [27]. More recently, models such as WaveNet [43] opened the door to neural generation conditioned on symbolic musical representations [19] [25]. DDSP has seen a lot of use creating audio synthesis models [2] [41] due to its flexibility. Despite its large data requirements, diffusion has also become a popular method for neural audio generation. While much work is focused on generating full song mixes [15] [14], synthesising symbolic musical representations such as MIDI allows much finer musical control over the synthesised audio. Hawthorne et al. [18] trained a Transformer-based Diffusion model for multi instrument synthesis, which was in turn expanded on by Kim et al. [24] for specifically acoustic guitar. While the models showed promising results, they were held back by lack of annotated data. Maman et al. [34] used automatic MIDI transcription systems to help address this problem while adding conditioning on instrument timbre. All of these methods have the drawback of requiring large amounts of data and necessitating significant computational resources to train. For example, the training in [34] took 350 hours over three powerful Nvidia A100 GPUs. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 48
4 J. Loth et al. 3Methodology 3.1 Flow Matching Flow Matching (FM) [28] is a generative modeling paradigm which provides a way to efficiently train continuous normalizing flows (CNF). It has recently picked up interest among audio researchers [17] [44]. Intuitively, the idea is to learn the direction to move a point sampled from some data distribution over time in order to transform it into a sample from a different data distribution. More formally, we would like to transform a data distribution p0(x0)to another data distribution p1(x1)by learning a velocity field vtwhich points in the direction in which we would like to transform a point at a given time t2[0,1]. Given a vector field utwhich generates a target probability flow pt,welearn a time-dependent flow which matches pt.Todothis,weconsidertheordinary differential equation (ODE): dxt=vt(xt)dt (1) Flow Matching attempts to learn this time-dependent vector field vtby minimising a conditional Flow Matching (CFM) loss [42]. This is generalized to arbitrary source and target distributions, making CFM an intuitive and powerful method for transforming data distributions. We adopt rectified flow [29] by simply sampling x0and x1from our data and calculating the interpolated point x=(1t)x0+tx1and target velocity ut(x|z)=x1x0.Wethentrainaneural network ✓by minimising the following: L=||v✓(t, x)(x1x0)||2 2(2) Once we have learned our v✓(x, t),wecanuseanODEsolver[3]toapproximate the solution to Equation 1 over some set of discrete time steps to transform source data to the target distribution. 3.2 Method Our approach focuses on modeling the direct input (DI) signal, the raw output of an electric guitar, rather than the amplified and distorted signal which is typically heard in recordings. This allows us to simplify the task and offload distortion processing to digital amplifier models [47] [4]. Our training pipeline involves first rendering synthetic audio from a guitar tablature using a virtual instrument. This and corresponding real guitar DI audio are then encoded into a latent space using Music2Latent [36], a pretrained autoencoder. The audio is broken into four second chunks in order to keep the audio length consistent and keep the data size reasonably small. Our model then learns a mapping from a latent distribution p0(x0)of synthetic guitar DI rendered by our virtual instrument to a latent distribution p1(x1)of real guitar DI. We can then use x0⇠p0and x1⇠p1,correspondingtothesynthetic and real guitar DI respectively, for training. This makes the model effectively Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 49
GuitarFlow 5 aone-to-onestyletransfermodel,astherealandsyntheticguitarrecordings are required to contain exactly the same musical content (i.e. notes, timings, expressive techniques, etc.). 3.3 Model Our model, titled GuitarFlow, is primarily based on the UNet [37] architecture, consisting of four downsampling and upsampling layers5.Toconditionthemodel on t,x0and tare concatenated together in the feature level. The model then outputs predicted flow velocities vt, which are used alongside real flow velocities utto calculate mean squared error (MSE) loss. Figure 2 shows the training and inference pipeline. Fig. 2. Training and inference using GuitarFlow. For inference, an ODE solver [3] is used to integrate across the learned velocity field across 100 discrete time steps. We found the Dormand-Prince method [9] to yield significantly higher quality results compared to a more standard Euler ODE solver. 4 Experiments 4.1 Data and Training The GOAT dataset [31] was used to train and evaluate our model. This dataset contained roughly 5.75 hours of real guitar DI audio recorded at 44.1kHz and 5https://github.com/clemkoa/u-net Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 50
6 J. Loth et al. the corresponding tablature annotations. The performances were performed by three guitarists on four different guitars, largely cover rock and indie rock genres and include a wide variety of expressive techniques. The guitar tablatures were rendered into audio using the Realistic Sound Engine (RSE) virtual instrument in Guitar Pro 76. The data was first preprocessed using Music2Latent into individual foursecond audio chunks prior to training. The model was then trained on a NVIDIA RTX A5000 GPU for 50 epochs (3,900 total steps) with a batch size of 64 and learning rate of 0.0001,whichtookroughly12minutesintotal.Whilethisis asurprisinglysmallnumberoftrainingsteps,wefoundittobesufficientfor convergence. 4.2 Evaluation Both the Guitar Pro and GuitarFlow audio are evaluated against the real DI audio in the test split of GOAT. Because electric guitar is almost always heard through the distortion of a guitar amplifier (whether subtle or heavy distortion) rather than the pure DI, all of the evaluation audio was separately rendered using a digital guitar amplifier7plugin. We then evaluate both the DI and amplifier conditions of the evaluation audio. Through this, we test two hypotheses: H1:audiofromGuitarFlowsoundsmorerealisticthantheoriginalGuitarPro rendered audio as DI; and H2:audiofromGuitarFlowsoundsmorerealistic than the original Guitar Pro rendered audio when rendered using a distorted guitar amplifier. Fréchet Audio Distance (FAD) [23] is a commonly used metric to measure audio similarity, while Kernel Audio Distance (KAD) [5] was recently proposed as a more flexible alternative which does not rely on a normality assumption of the audio embeddings. We calculate both FAD and KAD to evaluate the closeness of the synthetic and transformed audio to the real guitar DI recordings, using the real audio as the parent distribution. Several different embeddings [20] [12] [26] [10] [7] are used in these calculations. Following Hawthorne et al. [18], we also calculate the reconstruction embedding distance between the real DI audio and both the Guitar Pro audio and GuitarFlow output audio. This gives a better metric of how closely the actual audio is to the intended target, unlike FAD and KAD which work over a distribution of audio embeddings. This distance is obtained by calculating the Frobenius norm for both embeddings, and averaged over all time frames. A simple listening test was also conducted with 16 participants (13 male, 3 female, with an average age of 27.8 years) to evaluate the model outputs. We selected 15 four-second outputs which covered single notes, chords, and playing techniques such as bends and muted strings. Participants were asked to rate each 6https://www.guitar-pro.com/blog/p/14545-signature-sounds-explained-guitar-pro7 7The “crunch” amplifier and default cabinet IR were used from https://neuraldsp.com/plugins/archetype-nolly Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 51
GuitarFlow 7 Table 1. Comparison of the original synthetic Guitar Pro virtual instrument (GP) and GuitarFlow. We calculate Fréchet Audio Distance (FAD), Kernel Audio Distance (KAD) and reconstruction distance (Recon. Dist.) on both the DI and the DI rendered through a guitar amplfier. Best values for each metric in each row marked in bold. Condition Embedding Model FAD #KAD #Recon. Dist. # GP GuitarFlow GP GuitarFlow GP GuitarFlow DI VGGish 2.71 2.35 8.16 5.72 0.79 0.90 CLAP 207.26 120.74 17.65 4.65 0.09 0.08 PANNs 16.83 8.48 17.30 4.84 0.96 0.66 EnCodec 39.12 14.74 22.89 9.70 1.82 1.40 OpenL3 64.71 37.98 8.20 3.33 1.56 1.72 Amplifier VGGish 1.96 0.77 4.17 5.28 1.09 1.10 CLAP 80.25 33.41 8.73 8.61 0.05 0.04 PANNs 13.90 5.18 6.51 9.99 1.00 0.74 EnCodec 11.76 4.39 8.20 13.77 1.57 0.93 OpenL3 48.93 20.62 5.29 5.06 1.77 1.42 in terms of realism, which we define as how close the audio resembles a human playing guitar. All audio examples were normalised to -9dB RMS prior to the amplifier in order to ensure that the gain staging was consistent. Participants were paid with a £10 Amazon voucher. Finding a baseline model to compare against is difficult as most synthesis models use MIDI as a input sequence which cannot replicate any of the numerous expressive techniques, creating an unfair comparison with GuitarFlow. Many common style transfer approaches are also difficult to compare due to monophonic constraints [1] [13] or large computational and data requirements [21] [6]. Since our primary goal in this work is to establish the feasibility of flow-matchingbased transfer in the context of symbolic-to-real audio synthesis, we focus on evaluating the improvement that the model makes compared to the intermediary synthetic audio within that pipeline. We leave a comprehensive comparison with other style transfer methods as an important direction for future work. 5ResultsandDiscussion 5.1 Objective Metrics The FAD, KAD and reconstruction distances are presented in Table 1. GuitarFlow generally performs much better than Guitar Pro, particularly in the DI condition. The amplifier condition is a bit more mixed, with GuitarFlow struggling to improve on Guitar Pro in the KAD metric. However, the FAD and reconstruction distance results are still strong in the amplifier condition. It is possible that the clipping and distortion from the amplifier removed some information from the DI which affected the calculated embeddings in a way that the KAD metric is sensitive to. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 52
8 J. Loth et al. 5.2 Listening Test The listening test mean opinion score (MOS) results are shown in Figure 3. Friedman tests revealed significant differences between groups in both the DI (2(3) = 267.114,p<.001)andamplifier(2(3) = 107.153,p<.001)conditions. A pairwise Wilcoxon signed-rank test was then performed for both conditions using a Bonferroni-corrected ↵=0.0167.FortheDIstimuli,significant differences were found between the real and GP (p<.001) groups and the real and GuitarFlow (p<.001) groups. For the amplifier stimuli, significant differences were found between the real and GP (p<.001) groups, the real and GuitarFlow (p<.001) groups and the the GP and GuitarFlow (p<.001) groups. Fig. 3. Boxpot with mean indicators for the MOS results of both the DI and amplifier conditions of the listening test. Mean scores in each group marked by white dots. While the GuitarFlow model is only barely perceived as more realistic than the Guitar Pro audio in the DI scenario (H1), it is perceived as significantly more realistic when run through the distortion of a guitar amplifier (H2). It is interesting that the realism MOS would increase after the distortion is applied to the signal, as this distortion clips the signal and loses information. This also seems to contrast the KAD results of the amplifier condition, though this discrepancy could be due to a limitation of the KAD metric, the embeddings used, or simply a difference between what each result is measuring. We also see the MOS scores of the Guitar Pro and GuitarFlow stimuli increase and the real stimuli decrease in the amplifier condition. In their post-survey remarks, one participant noted that some of the examples had “characteristic sounds of neural synthesis”. We theorise that this amplifier distortion process helps the participants to focus more on the realism of the example, in the context of its content and timbre. This is because the amplifier distortion process hides the aforementioned neural artifacts that would otherwise cause the MOS to be lowered, as in the DI case, and distract listeners from the task at hand. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 53
GuitarFlow 9 5.3 Subjective Analysis AcarefulsubjectivelisteninganalysisoftheaudioexamplesrevealedthatGuitarFlow seems to particularly excel at recreating strumming chords. This is particularly relevant given that most commercially available guitar virtual instrument software notoriously struggle with this technique. However, the model appears to struggle much more when generating single notes, creating obvious neural artifacts which are not present during strummed chords. Due to this discrepancy, we hypothesize that this could be due to the notes of the real guitar DI not being perfectly aligned to the notes in the Guitar Pro rendering. As chords are strummed, they inherently have a looser timing window compared to single notes. This should be addressed in a future listening study which focuses on strumming vs. single notes. 5.4 Limitations & Future Work As in many deep learning related works, data is the main limiting factor in our approach. While the experiment showed great promise despite its data constraints, additional data points from a more varied pool of guitars and guitarists could potentially allow for better sound quality and generalisability, as well as explicit style controllability. Unfortunately, obtaining paired examples of guitar audio and tablatures is an expensive and time consuming process. Pretraining on synthetic data or using unpaired training methods may help address this. The annotation alignment issue may have also adversely affected the final audio quality. However, the experiment undertaken alone is not sufficient to clarify this hypothesis. Additionally, the listening test only measures a broad “realism" of the synthesis quality, and thus we are unable to gain any insight to any more finegrained aspects of the results such as the timbre, note accuracy or human-like variability. 6Conclusion In this paper we presented GuitarFlow, a novel methodology and model for synthesising realistic electric guitar from guitar tablatures. This model makes use of latent Flow Matching to perform style transfer on a basic audio render of the tablature, allowing the model to be trained quickly on an extremely small amount of data while still generalising to unseen data. This approach was justified through several objective metrics and a listening test. We hope that our work will inspire more researchers to investigate Flow Matching for generative audio, as well as work on generative systems which allow for increased musical control. Acknowledgments. This work is supported by the EPSRC UKRI Centre for Doctoral Training in Artificial Intelligence and Music (Grant no. EP/S022694/1) and UKRI - Innovate UK (Project no. 10102804). Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 54