Full text
id188440 GENERATIVE MUSIC TECHNIQUES: SAMPLE SYNTHESIS USING AUTOENCODERS CLARA RIVADULLA DURÓ Thesis supervisor SERGIOÁLVAREZNAPAGAO(DepartmentofComputerScience) Degree Master'sDegreeinArtificialIntelligence Master's thesis School of Engineering Universitat Rovira i Virgili (URV) Faculty of Mathematics Universitat de Barcelona (UB) Barcelona School of Informatics (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech 23/10/2024
Acknowledgements First and foremost, I would like to express my gratitude to my original project supervisor, Miquel Sànchez-Marrè, for his guidance in helping me select the topic and for his support during the early stages of the project. I wish him a happy and well-deserved retirement. My sincere thanks also go to Sergio Alvarez-Napagao, who took over as my supervisor in the final months and provided the crucial feedback and encouragement needed to complete this work. I am deeply grateful to my partner and my family for their endless love, support, and belief in me throughout this journey. Finally, I would like to acknowledge all those whose work is referenced either directly or indirectly in this thesis, for their generosity and their contributions to the field.
Abstract This thesis investigates the intersection of generative artificial intelligence and music creation, with a focus on drum sample synthesis. It begins with a comprehensive literature review that examines various music representation techniques and traces the evolution of music generation methods, from early algorithmic composition to contemporary deep learning approaches. The core of the research involves the design, implementation, and evaluation of two generative models: a Convolutional Autoencoder (CAE) and a Convolutional Variational Autoencoder (CVAE). These models are developed with the primary aim of synthesizing high-quality drum samples from existing audio data. The study compares the performance of these architectures across different datasets, analyzing their capabilities in sample reconstruction, latent space representation, and novel sound generation. By exploring the strengths and limitations of each approach, this research contributes to the growing field of AI-driven music production tools and provides insights into the potential of deep learning for creative audio synthesis.
Table of contents 1. INTRODUCTION ................................................................................................................. 1 2. LITERATURE REVIEW ....................................................................................................... 3 2.1. Representation ........................................................................................................ 3 2.1.1. Sheet Music ....................................................................................................... 4 2.1.2. Symbolic ........................................................................................................... 6 2.1.2.1. Piano roll .................................................................................................... 6 2.1.2.2. MIDI ........................................................................................................... 7 2.1.2.3. ABC Notation ............................................................................................. 7 2.1.2.4. MusicXML ................................................................................................. 8 2.1.3. Audio ................................................................................................................ 8 2.1.3.1. Waveform .................................................................................................. 8 2.1.3.2. Spectrogram .............................................................................................. 9 2.2. Generation ............................................................................................................. 10 2.2.1. Brief History .................................................................................................... 11 2.2.2. Algorithmic Composition ............................................................................... 12 2.2.2.1. Markov Chains ......................................................................................... 12 2.2.2.2. Generative Grammars ............................................................................. 13 2.2.2.3. Genetic Algorithms .................................................................................. 14 2.2.3. Deep Learning ................................................................................................ 15 2.2.3.1. First Works ............................................................................................... 15 2.2.3.2. Recurrent Neural Networks (RNNs) ....................................................... 15 2.2.3.3. Long Short-Term Memory Networks (LSTMs) ...................................... 16 2.2.3.4. Convolutional Neural Networks (CNNs) ................................................ 16 2.2.3.5. Autoencoders ........................................................................................... 17 2.2.3.6. Generative Adversarial Networks (GANs) .............................................. 19 2.2.3.7. Transformers .......................................................................................... 20 2.2.3.8. Diffusion models .................................................................................... 20 2.2.3.9. Combined architectures .......................................................................... 21 Summary of the chapter .............................................................................................. 21 3. METHODOLOGY ............................................................................................................. 22 3.1. Research Design .................................................................................................... 22 3.1.1. Objectives ........................................................................................................ 22
3.1.2. Strategy ........................................................................................................... 23 3.2. Data Preparation ................................................................................................... 24 3.2.1 Dataset Selection ............................................................................................. 24 3.2.2. Data Processing ............................................................................................. 26 3.3. Model Architecture ................................................................................................ 31 3.3.1. Autoencoder Design ....................................................................................... 32 3.3.2. Requirements and Tools ................................................................................ 34 3.4. Experiments .......................................................................................................... 35 3.4.1. Convolutional Autoencoder ........................................................................... 36 3.4.1. Convolutional Variational Autoencoder ........................................................ 40 Summary of the chapter .............................................................................................. 41 4. RESULTS ........................................................................................................................ 43 4.1. Evaluation ............................................................................................................. 43 4.1.1. Latent space visualization .............................................................................. 44 4.1.2. Sample reconstruction ................................................................................... 47 4.1.3. New Sample Synthesis ................................................................................... 52 Summary of the chapter ............................................................................................. 58 5. CONCLUSIONS ................................................................................................................ 60 5.1. Findings ................................................................................................................. 60 5.2. Future work ........................................................................................................... 62
1 1. Introduction This thesis is motivated by a strong personal interest, and seeks to address an opportunity within the Master’s program to further explore audio and music applied techniques within the context of generative artificial intelligence. As an amateur music producer, I became intrigued by the potential of generative AI to aid the musical creative process, as it has been achieved in fields involving text or image generation, which have represented a breakthrough for both companies and individuals in terms of creativity, efficiency, and cost-effectiveness. To address this, I aim to explore the following research questions: Can music be generated through computational techniques? If so, how? This investigation starts with a comprehensive literature review (2. Literature Review), including an exploration of music representation techniques—sheet music, symbolic and audio—, a brief history of generative music, and an examination of methods ranging from algorithmic composition techniques in the 19th century—such as Markov chains, generative grammars, and genetic algorithms—to contemporary deep learning approaches like Transformers, LSTM networks, GANs, diffusion models, and autoencoders. The chosen method for the empirical part of this research focuses on autoencoders, and leads to the second question: Can we perform sample synthesis by training a DL model on an audio dataset? Sample synthesis involves generating new sounds by manipulating or combining pre-recorded audio samples, and the use of AI for this task
2 could either inspire or provide music producers with new, interesting sounds to enhance their creative process. Specifically, we train a Convolutional Autoencoder (CAE) and its variation, the Convolutional Variational Autoencoder (CVAE), on a dataset of drum samples, represented as Log Mel Spectrograms (LMS). The models are evaluated using several subjective methods, including latent space visualization (to assess whether the model can distinguish between different drum classes), sample reconstruction (comparing original and generated LMS), and new sample generation. The generated spectrograms are then converted to audio and analyzed aurally. The design, implementation, and evaluation processes of the model are discussed in chapters 3. Methodology and 4. Results. Finally, in chapter 5. Conclusions, we draw conclusions regarding the outcomes of the implemented generative model, along with the insights gleaned from its development and assessment, and discuss potential avenues for future research. Disclaimer: It’s important to note that this study does not engage in the philosophical debate over the creativity of AI-generated content, as we see AI as a tool to support human creativity rather than replace it. Such discussion is beyond the scope of this research and is better addressed within disciplines such as philosophy or psychology. For further reading on the topic, we recommend [1], [2], [3].
9 2.1.3.2. Spectrogram A spectrogram is a visual representation of the spectrum of frequencies of a signal, where the x-axis represents time in seconds, the y-axis represents frequency in kHz and the third axis (color or shading) represents intensity in dBFS. As spectrogram representations resemble images, they can be used along with CNN architectures for sound applications. A spectrogram is obtained by dividing a signal into a sequence of short duration sub-signals and performing a Fast Fourier Transformation (FFT) on each sub-signal. Fourier Transform (FT) can decompose a signal into its constituent frequencies and gives the magnitude of each frequency present in the signal. The FTT algorithm calculates Discrete Fourier Transform (DFT) of a given sequence, defined as !(#)= &'(())*!"#$ %&' %!( &)* where '(()) corresponds to equally spaced samples of an analog time function '(+) [7]. There are several variations to the spectrogram: Mel Spectrogram A visualization of the spectrum of frequencies in an audio signal over time, where frequencies are converted into the Mel scale . The Mel scale is a perceptual scale of pitches that approximates the human auditory system's response to different frequencies. MFCCs Originally developed for automated speech recognition, they capture timbral properties of the signal [4] . MFCCs are a compact representation of the spectral envelope of an audio signal and are derived from the Mel spectrogram.
10 Chromagram Chroma features aggregate all spectral information that relates to a given pitch class into a single coefficient, and a chromagram can be obtained summing up all pitch coefficients that belong to the same chroma [4] . The chromagram discretizes sound onto the tempered scale, disregarding octave distinctions and limited to pitch classes. Tempogram Similarly to a spectrogram, for each time instance, the tempogram encodes the local relevance of a specific tempo for a given music recording. 2.2. Generation Having explored the various methods of representing music –from traditional sheet music to symbolic notations and audio representations–, we now turn our attention to the generation of music itself, and we'll see how these representations suit specific generation techniques: symbolic representations (e.g., MIDI) work well with rule-based and early machine learning models, while audio formats fit deep learning models best. Music generation techniques can focus on one or more elements such as melody, harmony, rhythm, and timbral properties, and can also be categorized into note-based (learning from music scores or similar representations) and signal-based (learning from audio signals) methods [8]. Moreover, these generative techniques can be applied to various domains: style transfer (adapting the style of one piece or sample to another), voice cloning (replicating a specific voice’s characteristics), sample synthesis (creating or modifying audio samples), melody creation (generating new melodic lines), or accompaniment (producing harmonies or background music to support a lead melody), among others. Lastly, music generation may serve two distinct purposes: to design autonomous systems or to assist musicians in
11 production, arrangement, composition, sound synthesis, or orchestration tasks using digital environments that allow human-level control and interaction [5]. In this section, we dig into the origin and evolution of music generation techniques by means of a brief historical overview, and we introduce different techniques that have been used over the years: from algorithmic composition to deep learning-based systems. Each method is explained in broad strokes, and illustrated through examples of several domains. 2.2.1. Brief History Around AD 1000, Guido d'Arezzo, who also contributed considerably to music notation, developed a system to generate melodies from text materials, mapping letters, syllables and verses to tone pitches and melodic phrases [9]. In the eighteenth century, classical composers like J. P. Kirnberger, W. A. Mozart and C. P. E. Bach conceived and played musical dice games to generate music by throwing a dice or choosing a random number [10]. The idea of musical automatas, closely connected with the idea of generative music, is also centuries old. Jacquet- Droz and his son designed The Musician, a female organ player that played a programmed piece by pressing keys with its fingers [11]. In 1821, D. N. Winkel created the Componium, an automatic organ consisting of two barrels that revolve simultaneously capable of producing variations on a theme programmed into it [12] [13]. One of the first computer-generated pieces was Pinkerton's Banal Tune Maker [14], from 1956, calculating transition probabilities from 39 nursery songs. A year later, Guttman came up with The Silver Scale, a 17-second-long melody generated by Music I (Bell Laboratories), a software for sound synthesis. Its contemporary The Illiac Suite, by Hiller and Isaacson, was the first score composed by a computer using algorithmic composition techniques such as stochastic models (Markov chains) and a set of rules. In 1962, Xenakis also experimented with stochastic composition and created Atrées. In 1971, S. Smoliar developed a programming language called EUTERPE to model musical structures in real time, and used it to generate Gregorian chant, medieval polyphony, Bach counterpoint, and sonata-form examples [12]. In the 1980s, Ebcioğlu designed CHORAL, a composition software for four-part chorale composition in the style of Bach, with more than 350 handcrafted rules. From 1983 to 1989, David Cope developed Experiments in Musical Intelligence (EMI), a system able to learn from a corpus of scores of a composer to create its own grammar and database of rules [5]. On the other hand, the first deep learning-based approaches emerged with the experiments of Todd (1989) [15] or Mozer (1994) [16], which used artificial neural networks to generate
12 melodies. These early endeavors laid the groundwork for subsequent advancements. Over time, these initial approaches evolved into more sophisticated architectures such as Recurrent Neural Networks (RNNs), Long Short-Term Memory networks (LSTMs), Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), among others. These advanced models enable a deeper comprehension of musical relationships, consequently leading to the production of higher-quality compositions. In opposition to grammar-based or rule-based systems, which are difficult to specify and errorprone [17], machine learning and deep learning approaches allow for generality, as they automatically learn from a corpus and can be used for several music genres. Major contributions to the field came later with big companies' projects such as Google's MusicLM [18], OpenAI's Jukebox [19] or Meta's MusicGen [20] and AudioGen [21] (2023). 2.2.2. Algorithmic Composition Algorithmic composition methods allow for the creation of music by identifying specific rules, patterns, or statistical properties. Over the 20th century, various approaches to algorithmic composition emerged, including probabilistic models such as Markov chains, rule-based systems like generative grammars, and evolutionary methods inspired by genetic algorithms. These methods have often been used to model and replicate musical styles or to generate monophonic melodies. 2.2.2.1. Markov Chains Markov chains were introduced by Russian mathematicians Andrey Andreyevich Markov and Sergey Natanovich Bernstein, and are rooted in the theory of stochastic processes, which describe a sequence of random events dependent on the time parameter t. In a Markov chain, the probability of the future state Xt+1, where X is the random variable and t+1 is the given time, depends on the current state Xt. For tm and tm+1, the transition probability is: ,(-+,-( = ./|/-+, =1)=/2."(+,,+,-() A Markov chain can also be represented by a state transition graph, or by a transition matrix. In the context of algorithmic composition [9], transition probabilities are generated either by adjusting individual structural settings or by analyzing existing music to mimic its style. In 1950, Harry F. Olson analyzed eleven melodies by Stephen Foster and developed first and
13 second-order Markov models to analyze their pitch and rhythm patterns. In Hiller and Isaacson's Illiac Suite (1956), the fourth "experiment" (movement) was composed using Markov models of variable order for the generation of musical structure. In Iannis Xenakis' Analogique A, Markov models are used to organize segments with varying densities. Later, in 2002, Pachet developed the Continuator [22], an improvisation system in which musical styles are learned automatically, not requiring any symbolic information, and adapting to the musician's playing mode. Apart from using Markov models, the Continuator was augmented with a hierarchical model and a facility for biasing the Markovian generation to handle changing harmony. In 2019, D. Williams et al. [23] presented an architecture for the creation of emotionally congruent music using machine learning aided sound synthesis, generating a small corpus of music using Hidden Markov Models. 2.2.2.2. Generative Grammars Generative grammars have their origin in the linguistic model developed by Noam Chomsky in 1957 and have been applied to the field of generative music by several researchers and musicians, such as Roads, Steedman, Sundberg, Lerdahl, or Jackendoff, since the 1970s, mostly based on the Chomsky hierarchy model [9]. They consist of a set of formal rules that systematically define how symbols or elements in a language, or in this case, a musical composition, can be combined to generate valid sequences. These rules operate within different levels of the Chomsky hierarchy—ranging from unrestricted (type-0) and contextsensitive (type-1) grammars to context-free (type-2) and regular (type-3) grammars [9]— allowing for varying degrees of complexity in the generated output. Figure 4. Jazz pianist Albert van Veenendaal playing with a grand piano and Pachet's Continuator [74].
14 Generative grammars are well-suited for music analysis, style imitation, and composition, as they provide a hierarchical and context-sensitive organization of musical material. However, their sequential nature limits their effectiveness with more complex compositions, such as polyphonic music. Algorithmic composition using Generative grammars has two main approaches: knowledgebased, if explicitly formulated rules are assumed; and non-knowledge-based or grammatical inference, if a system automatically generates rewriting rules out of a corpus [9]. As of 1979, Roads and Wieneke [24] highlighted contributions from Ruwet, Nattiez, Laske, Smoliar, Winograd, Moorer, Jackendoff and LerdahlIn. In 1981, Holtzman [25] introduced generative grammars for music composition with use of the GGDL (Generative Grammar Definition Language) compiler, which allowed the description of musical languages, and gave the compositional rules of a Schoenberg Trio as an example of a grammar description. 2.2.2.3. Genetic Algorithms Genetic algorithms, primarily developed as an optimized search tool, mimic natural evolutionary processes, where a population of individuals, represented by chromosomes, evolves through selection, mutation, and crossover to find optimal solutions. Musical composition using GAs can be seen as a three-step procedure: 1. Creating the starting idea 2. Mutating the piece 3. Assessing the results which somehow emulates the human creative process of generating musical pieces [26]. There are two common approaches to assess fitness in a GA: the use of a human critic (Interactive Genetic Algorithm), which can take too much time due to the temporal nature of the medium [26], or automatic assessment. Gartland [26] provided a generative algorithm which consisted of a five-step procedure to evolve from starting to target musical fragments: 1. Create an initial population providing MIDI files; 2. Select a population member and perform mutation and crossover (add a note, transform a note, etc.); 3. Assess the fitness of the mutated population member; 4. Accept into the population or discard; 5. Repeat until the target is reached.
15 In [8], a three-phase melody generation method using genetic algorithms along with LSTM networks is proposed. In the first phase, an initial GGA is used to generate enough sample data for training; in the second phase, this data is used to train an LSTM neural network; finally, the trained LSTM network serves as the objective function for the main GGA process. The resulting model, trained on ABC notation melodies, can generate melodies 70% like human ones. 2.2.3. Deep Learning In the past few years, some of the most used NN architectures for music generation have been Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), Long Short-Term Memory (LSTM) and Transformers [27]. In this section, we provide an overview of each of these techniques. However, before delving into more contemporary approaches, it's essential to revisit the initial works that established the groundwork in the field. 2.2.3.1. First Works Todd's [15] first generative method for a monophonic melody was a Time-Windowed architecture, where generation was conducted iteratively, segment by segment, and recursively, as the predicted output of one segment was used in the prediction of the next. Unfortunately, this architecture couldn't handle with long-term correlations, but only with pairwise relations between successive notes. A later design of his, the Sequential architecture, one of the first examples of a recurrent architecture and an iterative strategy, improved this aspect by dividing the input layer into two parts: the context, which contains the melody generated so far, and the plan, a particular melody learned by the network. Training is done by choosing a plan repeatedly, with as many melodies as needed. Generation is performed with a new input plan (not seen in the training set), and time step after time step as in the Time-Windowed approach [17]. Todd's Sequential architecture can be seen as a first approach to the use of RNNs for music generation, which, as we'll see, would later evolve into more sophisticated methods. 2.2.3.2. Recurrent Neural Networks (RNNs) RNNs [28] are valid for training sequences of data where the order of it matters. The key feature of RNNs is their ability to maintain a state or memory of previous inputs (in the hidden
16 state) while processing new inputs. This recurrent nature allows them to handle sequences of varying lengths and capture dependencies over time. However, they fail to store long-term dependencies due to the vanishing gradient problem, where gradients diminish exponentially as they propagate back through time during training. To address this issue, variants of RNNs such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) architectures have been developed. 2.2.3.3. Long Short-Term Memory Networks (LSTMs) LSTMs [29] improve RNNs by handling long-term dependencies and addressing the vanishing gradient problem by introducing a memory cell with a gating mechanism. In a traditional RNN, the hidden state at each time step is updated based on the current input and the previous hidden state using a simple transformation, typically a tanh or sigmoid activation function. LSTMs have a more complex architecture, consisting of a memory cell and three gating mechanisms: input gate, forget gate, and output gate. These gates control the flow of information into and out of the memory cell, enabling LSTMs to selectively retain and update information over time. The first LSTM application for music generation was by Eck et al. in 2002 [30], who used such architecture to learn blues music and compose novel melodies in that style. Melody RNN is Google Magenta's LSTM network for music generation, which already has several variations: the simple dual layer LSTM network; Lookback RNN, with custom inputs and labels; and Attention RNN, with the use of attention to access past information without having to store it in the RNN cell’s state [31]. The inputs and outputs of the baseline model are in MIDI format, but reduced to the range of pitches [48, 84]. 2.2.3.4. Convolutional Neural Networks (CNNs) Google Deepmind's WaveNet [32] is a fully probabilistic and autoregressive deep neural network for generating raw audio waveforms based on PixelCNN [33], where the probability distribution is modelled by a stack of dilated causal convolutional layers, with no pooling layers and the output has the same dimensionality of the input. WaveNet has been proven to be faster and cheaper than most RNN or LSTM architectures, and that serves both for text-to- speech and music generation if trained accordingly. WaveNet uses audio representations, taking the original waveforms as inputs.
17 2.2.3.5. Autoencoders An autoencoder (first introduced in [34]) is an unsupervised learning algorithm which is great at representation learning (learning patterns in data). Its architecture (as seen in Figure 5) consists of an encoder and a decoder. The encoder leads to a bottleneck, compressing data into a lower-dimensional representation in the latent space, which focuses on representing the most important features of data. This is suitable for data that has dependencies across dimensions. The decoder decompresses the representation back to its original domain. Formally, the autoencoder must learn the functions 4:/ℝ&→/ℝ/ (encoder) and 8:/ℝ/→/ℝ& (decoder), satisfying 9:;<1(0,2/=[?(@,8∘4(@)]. A and B are most commonly neural networks, but they can also be linear operations. In this case, the autoencoder becomes a linear autoencoder, where the latent representation corresponds to what would be obtained using Principal Component Analysis (PCA) [35]. The reconstruction loss of an autoencoder quantifies the difference between the input data and the reconstructed data, and it’s typically measured using MSE. Variational Autoencoders (VAEs) [36] extend the capabilities of autoencoders by describing data generation through a probabilistic distribution (usually Gaussian). The latent vector C is a random variable distributed according to some prior 2(C), and data generation is [37]. In an autoencoder, each input is directly mapped to a single point in the latent space, whereas in a VAE, each input is mapped to a multivariate normal distribution, centered around a point in the latent space (as seen in Figure 6) [38]. To sample a point C in the normal distribution, we compute C =/D+FG Figure 5. Illustrative architecture of a vanilla autoencoder.
18 where D is the mean, F is the standard deviation and G is a sampled point from the standard normal distribution. To sample a point C in a multivariate normal distribution, as in a VAE, we compute C =DH+ΣG/ where DH is the mean point of the distribution, Σ =*!"# $ % % & ' ' and G is a sampled point from the standard normal distribution. The loss of a VAE requires not only the reconstruction loss of an Autoencoder, but also the Kullback-Leibler (KL) Divergence: J34[K(D,F)∥K(0,1)]=/1 2&(1+log(F#)−D#−F#) / and so, the loss is calculated as the sum of the MSE and the KL Divergence. In 2019, Bitton et al. [39] presented a real-time sound synthesis non-autoregressive model using Wasserstein Auto-Encoders (WAEs) with Adaptive Instance Normalization, Feature Wise Linear Modulation, and adversarial training with a Fader latent discriminator to generate individual musical notes with high-level control over attributes like timbre, note class, and playing style. Figure 6. Comparison of the latent space mappings in an Autoencoder versus a Variational Autoencoder. Reproduced from [38].
25 The audio samples needed to meet specific criteria: they had to be short in duration, cleanly produced, and of a quality high enough to capture the intricate details necessary for effective model training. As mentioned, the focus of the research was specifically narrowed to drum sample synthesis, so only datasets comprising various classes of drum sounds (kicks, snares, hi-hats, etc.) were considered. After careful evaluation, two datasets were chosen as the primary sources: tiny-audio- diffusion-drums [58], available in Hugging Face Datasets, and 29kSamplesDrums [59], available in Zenodo. The tiny-audio-diffusion-drums dataset was selected for its smaller size and limited number of classes, offering a more compact and manageable set of samples, while the 29kSamplesDrums dataset provided larger size and greater diversity of classes. By training the model on datasets of varying dimensions we’ll be able to see how the model performs with different levels of data diversity and size. Their details and characteristics are discussed in the following lines. 1. tiny-audio-diffusion-drums The original dataset consists of four distinct classes: hi-hat, snares, kicks, and percussion. However, the percussion class was excluded from this study due to its inclusion of mixed drum elements, which could introduce inconsistencies in the data. The remaining three classes are more uniformly distributed, as illustrated in Figure 9. Their average durations are shown in Figure 8. The samples range in length from 0.02 seconds to 5.76 seconds, with a mean duration of 0.73 seconds, a median of 0.5 seconds, and a standard deviation of 0.67 seconds. The distribution of sample durations is depicted in Figure 10. All samples were recorded at a sampling rate of 16,000 Hz. Figure 9. Distribution of audio samples across different drum classes of the tiny-audio- diffusion-drums dataset. Figure 8. Average duration (in seconds) of audio samples for each drum class of the tinyaudio-diffusion-drums dataset.
26 2. 29kSamplesDrums The dataset contains 29,000 audio samples across 22 different classes. However, for this study, we focused exclusively on the eight classes corresponding to single drum instruments, omitting the remaining classes that consist of instrument combinations. The selected classes include ride cymbal (cy), crash cymbal (cr), hi-hat (hh), snare drum (sd), kick drum (kd), floor tom (ft), mid tom (mt), and high tom (ht). The distribution of these classes is relatively uniform, as Figure 12 illustrates. The samples have an average duration of 0.42 seconds, with a median of 0.25 seconds and a standard deviation of 0.25 seconds. The shortest sample is 0.25 seconds, while the longest is 1.00 second. The distribution of sample durations is shown in Figure 11. All samples were recorded at a sampling rate of 44,100 Hz. 3.2.2. Data Processing Figure 10. Histogram of sample durations of the tiny-audio-diffusion-drums dataset, with bins representing time intervals in seconds (e.g., bin 0 corresponds to durations between 0 and 1 seconds). The y-axis indicates the number of samples within each duration range. Figure 12. Distribution of audio samples across different drum classes of the 29kSamplesDrums dataset. Figure 11. Average duration (in seconds) of audio samples for each drum class of the 29kSamplesDrums dataset.
27 In this section, we describe the data processing steps involved both before and after training the model. The process begins by loading and transforming the audio samples into the selected representation form. In our case, we choose Mel Spectrograms above other audio representations for their ability to capture the perceptual characteristics of sound while significantly reducing the dimensionality of the data. After training the model, inferred Mel Spectrograms must be reversible to audio to allow for playback. a. Loading samples Audio samples in .wav format are loaded in mono mode using the librosa library, and with a specified sample rate (16,000 Hz for the tiny-audio-diffusion-drums dataset and 44,100Hz for the 29kSamplesDrums dataset). These samples are then adjusted to match a predefined target length, measured in terms of time steps (or frames). For instance, with a sample rate of 16,000 Hz, an audio sample of 16,000 time steps corresponds to 1 second of audio. If a sample’s length is shorter than the target length, it’s padded with zeros. For both datasets, the target length in time steps is 22,400, which corresponds to 1.4 seconds for tiny-audio-diffusion-drums and 0.5 seconds for 29kSamplesDrums. Although this decision may affect some samples, it has been made to preserve the same spectrogram size between datasets so we can reuse the model’s convolutional layers’ configuration, and mainly because of our lack of computational resources (with bigger values, we’d get an error every time). Figure 13. Padded waveform (sample rate = 16,000 Hz; target length = 20,400) of the snare_103.wav sample from the tinyaudio-diffusion-drums dataset.
28 b. Computing Log Mel Spectrograms In our implementation, Mel Spectrograms are obtained using the MelSpectrogram function from the nnAudio [60] library with the parameters specified: - Sample rate (sr): The sample rate determines the number of audio samples captured each second and is expressed in Hertz (Hz). - Window length (win_length): The size of window frame and STFT filter. - Window size (n_fft): The window size for the STFT. - Hop length (hop_length): The hop (or stride) size. - Number of mels (n_mels): The number of Mel filter banks. The filter banks map the n_fft to Mel bins. The resulting Mel Spectrograms are then transformed to Log Mel Spectrograms (LMS) TUV by adding a small constant to avoid taking the logarithm of zero and applying the natural logarithm to the adjusted values. @5=ln(@+X) Finally, LMS are normalized using standard normalization. The normalization performed can be expressed by the following formula for each element x’ in the set of LMS: @′&67, =@′−D F where x’ is the original value in the set of LMS; D is the mean of values in the set of LMS; F is the standard deviation of the values in the set of LMS; and @′&67, is the normalized value.
29 Both the mean and the standard deviation are kept to be able to reconstruct the original audio at a later stage. Table 1. Configuration used for obtaining Log Mel Spectrograms. The mean D and standard deviation F of both datasets used for normalizing the LMSs are shown in the following Table: tiny-audio-diffusion-drums 29kSamplesDrums Z [ Z [ -13.11 5.05 -7.66 6.50 Table 2. Normalization stats (mean ! and standard deviation ") for the tiny-audio-diffusion- drums and the 29kSamplesDrums datasets. c. Reconstructing audio After inference, the Log Mel Spectrograms (LMS) need to be reverted into an audio waveform for playback. This process involves two key steps: first, denormalizing the LMS, and second, converting the Mel Spectrogram back to a time-domain waveform. hop_length n_ftt n_mels target_length win_length shape 80 2,048 128 20,400 2,048 [128, 256] Figure 14. Normalized Log Mel Spectrogram (hop_length = 160; n_fft = 400; n_mels = 80; sample_rate = 16000; win_length = 400) of the snare_103.wav sample from the tiny-audio-diffusion-drums dataset.
30 Step 1: Denormalizing the Log Mel Spectrograms During preprocessing, the LMS are normalized using the mean D and standard deviation F of the dataset. To revert this, the LMS values are first denormalized by applying the inverse of the normalization operation: @′=@′&67,8×F+/D where @&67,8 represent the normalized LMS, and the corresponding D and F are retrieved from the dataset statistics used during normalization. After denormalizing, the exponential function is applied to reverse the log transformation performed earlier, which yields the original Mel Spectrogram: @ =*95 Step 2: Converting the Mel Spectrogram back to audio Once the Mel Spectrogram is restored, the next step is to convert it back to a waveform. This is done using the mel_to_audio() [61] function from librosa, which implements an inverse Mel transform to recover the original time-domain audio using Griffin-Lim [62] [63]. The function takes the denormalized Mel Spectrogram x along with the exact parameters used during the initial conversion. It’s important to note that after converting audio to a Log Mel Spectrogram (LMS) and subsequently reconstructing it back to audio using the described method, the reconstructed audio does not perfectly match the original. Some variations and distortions may occur, as illustrated by the waveform visualizations in Figure 15. This discrepancy should be considered when evaluating the results of training the model and generating samples, as it may affect the quality of the final audio, even though the generated LMS may not necessarily exhibit these variations.
31 Figure 15. Waveforms of the original hihat_097.wav audio sample from the tinyaudio-diffusion-drums dataset and the audio reconstructed from its Log Mel Spectrogram. The comparison highlights the discrepancies between the original and the reverted audio. 3.3. Model Architecture The proposed architecture for learning Log Mel Spectrograms representations of drum samples and generating new ones after is the Convolutional Autoencoder (an Autoencoder with convolutional layers, which is also a CNN). Given the 2D nature of the chosen representation, our inputs can be treated similarly to images, and convolutions can help capture spatial patterns effectively. We discarded the use of RNNs and LSTMs because one-hit drum samples do not particularly require the ability to learn temporal relationships, and while GANs, Transformers, and diffusion models could have been viable alternatives, we chose the Convolutional Autoencoder (CAE) for its simplicity and reduced demand for computational resources. In this section, we provide an outline of its architecture, as well as the resources needed to carry out its development and training.
32 3.3.1. Autoencoder Design As seen in 2.2.3.5. Autoencoders, an autoencoder consists of two main components: an encoder that reduces the dimensionality of the input, and a decoder that reconstructs the input from the compressed latent space. Figure 16. Architecture of the CAE for learning LMS representations. In our implementation (illustrated in Figure 16), the encoder compresses the input Log Mel Spectrograms through a series of 2D convolutional layers. Each convolutional layer applies a kernel to the input, progressively reducing its spatial dimensions while capturing important features of the audio, and it’s followed by a ReLU activation function. After the convolutional operations, the final feature map is flattened into a 1D vector, and a fully connected layer maps it into a latent space of a specified dimension. This latent space is a compressed representation of the input, from which the decoder will attempt to reconstruct the original audio. The decoder starts with a fully connected linear layer and an unflatten layer to expand the latent space vector back into the 3D shape of the final convolutional layer from the encoder. It then applies a series of transposed convolutional layers, each reversing the operation of the corresponding convolutional layer in the encoder. This process involves upsampling the spatial dimensions while reducing the number of channels. Each transposed convolution is also followed by a ReLU activation function, except for the final layer, which outputs the reconstructed Log Mel Spectrogram.
33 Finally, the reconstruction loss (MSE) measures the difference between the original input and the reconstructed output. In addition, we also work with a derived architecture, the Convolutional Variational Autoencoder (CVAE). The CVAE extends the CAE by incorporating probabilistic modeling into the latent space representation. In this architecture, the encoder not only compresses the input Log Mel Spectrograms but also learns to parameterize a multivariate normal distribution defined by two outputs: the mean and the log variance, as seen in Figure 18. During training, we use the reparameterization trick, allowing us to sample latent variables in a differentiable manner, which is essential for optimizing the model using gradient descent. Figure 18. Architecture of the CVAE for learning LMS representations. The CVAE also modifies the loss function (see Figure 19). In addition to the mean squared error (MSE) reconstruction loss, we include a Kullback-Leibler (KL) divergence term. This term measures the difference between the learned latent distribution and a Figure 17. Loss function of the CAE.
34 standard normal distribution, encouraging the model to learn a more structured latent space [64]. Figure 19. Loss function for the CVAE. The strategy pipeline for our research on training CAE and CVAE architectures for drum sample synthesis has been further refined (see Figure 20), incorporating data gathering, preprocessing, model configuration, and training. This pipeline provides the foundation for evaluating the models’ performance, conducting inference, and generating outputs, which be addressed in detail in later chapters. Figure 20. Detailed architecture of the strategy pipeline of the project. 3.3.2. Requirements and Tools The implementation of a deep learning model for audio sample synthesis requires specific computational resources and software tools to ensure efficient training and
41 Figure 23. Training and validation losses during 150 epochs of the CVAE for the tiny-audio- diffusion-drums dataset. On the other hand, we've trained the model on the 29kSamplesDrums dataset and the configuration #3 of our previous experiment (Table 4), and have achieved a vl=213.49 and a tl=197.28 at the best epoch (n=143), and a vl=213.57 and a tl= 196.52 at the last epoch (n=150), with minimal overfitting, as illustrated in Figure 24. Figure 24. Training and validation losses during 150 epochs of the CVAE for the 29kSamplesDrums dataset. Summary of the chapter In this chapter, we outlined our comprehensive methodology for drum sample synthesis using deep learning techniques. We began by detailing our research design, including specific objectives and hypotheses. Next, we described the data preparation process,
42 where we selected and processed audio datasets, converting them into Log Mel Spectrograms. We then introduced our model architectures, focusing on the Convolutional Autoencoder (CAE) and its variant, the Convolutional Variational Autoencoder (CVAE). A series of experiments were conducted to explore various configurations of these models, optimizing their performance on the tiny-audio- diffusion-drums and 29kSamplesDrums datasets. The preliminary results, however, offer limited insight into the models’ overall performance, due to the complexity of the data and the subtle variations across configurations. Therefore, further evaluation is necessary to conduct a more in-depth analysis, both visually—through latent space and LMS visualizations—and auditorily, by examining the audio outputs. The next chapter will delve deeper into these analyses, providing a more comprehensive exploration of the models’ capabilities.
43 4. Results This chapter presents the results of our music generation model, focusing on drum sound synthesis. We evaluate the performance of Convolutional Autoencoders (CAE) and Variational Autoencoders (CVAE) using the two selected datasets: tiny-audio-diffusion- drums and 29kSamplesDrums. 4.1. Evaluation Evaluating the results obtained with a music generation model is not straightforward, nor is it for any other generative model. On the assumption that such models imply a creative output (in the sense that they create something, and that something is often within a creative context) and that the ultimate judge of creative output is the human, subjective evaluation is generally preferred [65]. Our approach includes both visual and auditory analyses. The visual analysis consists of visualizations of the latent space, as well as visualizations of spectrograms—both reconstructed from the training dataset and generated from scratch—along with the waveforms of the reverted audio. However, due to the nature of this paper, the auditory observations can only be left in writing.
44 4.1.1. Latent space visualization To assess how well the model organizes and represents data internally, we analyze the latent space. This is achieved by applying a dimensionality reduction technique to the high-dimensional latent space generated by the model by passing all the training data through the model’s encoder, which transforms each input into a 64-dimension latent vector. Specifically, we use t-SNE (t-Distributed Stochastic Neighbor Embedding) to reduce the dimensionality of the latent vectors from 64 to 2 and 3 dimensions. Clear clusters in the latent space, where samples from the same class are grouped together, will suggest that the model effectively captures the underlying structure of the data. Why t-SNE? t-SNE (t-Distributed Stochastic Neighbor Embedding) is particularly effective for visualizing high-dimensional data because it preserves local structures and relationships among data points. In our analysis of the latent space, t-SNE and UMAP (Uniform Manifold Approximation and Projection) both revealed clear clusters, indicating effective class separation, whereas PCA (Principal Component Analysis) showed overlapping clusters that obscured distinct classes due to its focus on maximizing global variance. This comparison highlights t-SNE’s ability to effectively capture nuanced relationships within data points, making it the preferred choice for our analysis, as illustrated in the accompanying figures. Figure 25. Latent space visualization of the latent space of the CAE trained on the 29kSamplesDrums dataset, reducing dimensionality with t-SNE (left), UMAP (center) and PCA (right). Figure 26. Latent space visualization of the latent space of the CAE trained on the tiny-audio- diffusion-drums dataset, reducing dimensionality with t-SNE (left), UMAP (center) and PCA (right).
45 Let’s start by assessing how well the CAE model organizes and separates different classes of drum sounds in the latent space. Figure 27 shows the 2D and 3D visualizations of the latent space of the CAE trained with tiny-audio-diffusion-drums dataset samples, and reveals distinct clustering patterns for the three drum sound types: kicks, hi-hats, and snares. Kick drums exhibit the most cohesive grouping, predominantly occupying the upper region of the plot, indicating strong similarity among samples. Hi-hats display more variability, clustering in the lower left and middle right areas, while snares show the highest degree of dispersion across the middle and right side of the plot, occasionally intermixing with hi-hat samples. This overlap suggests potential shared characteristics between certain snare and hi-hat sounds. The 2D and 3D visualizations of the 29kSamplesDrums dataset's latent space in Figure 28 reveal distinct clustering patterns for different percussion instruments, demonstrating the autoencoder's ability to differentiate between various sounds. Ride and crash cymbals are closely positioned, indicating similar tonal characteristics, while hi-hats form a separate cluster. The tom-toms (high, mid, and floor) show a gradient-like arrangement, reflecting their pitch progression. Kick and snare drums form welldefined, separate clusters in both visualizations. Figure 27. 2D and 3D visualizations of the tiny-audio-diffusion-drums latent space learnt by the CAE.
46 In general, the visualizations of the latent spaces for both datasets denote that the CAE is indeed capable of effectively capturing and distinguishing the underlying structures of various percussion sounds. Lastly, we evaluate the CVAE latent space with the tiny-audio-diffusion-drums and the 29kSamplesDrums dataset and find no visible clustering, which is expected due to its stochastic nature. Unlike the CAE, which deterministically maps inputs to fixed points in the latent space, the CVAE encodes inputs as probability distributions, leading to more dispersed representations. The KL Divergence regularization encourages a smooth, continuous latent space, but this comes at the cost of clear separation between sound classes, making it harder to observe distinct clusters. Figure 28. 2D and 3D visualizations of the 29kSamplesDrums latent space learnt by the CAE. Figure 29. 2D and 3D visualizations of the tiny-audio-diffusion-drums latent space learnt by the CVAE.
47 Figure 30. 2D visualization of the 29kSamplesDrums latent space learnt by the CVAE. 4.1.2. Sample reconstruction Next, we evaluate the model’s performance in reconstructing Log Mel Spectrograms (LMSs) from the latent vectors, which involves decoding the compressed latent representations back into their original form. Ideally, when a sample from the training set is passed through the Autoencoder (encoded into a latent vector and then decoded), the resulting LMS should closely resemble the original. In Table 5 and Table 6, we present the original and reconstructed LMSs for two samples of each class for the tiny-audio-diffusion-drums dataset and for one sample of each class for the 29kSamplesDrums dataset with the CAE. Overall, the CAE manages to preserve the primary structure and features of the original spectrograms across all classes, and distinct patterns corresponding to each sound class are recognizable in the reconstructions. However, there is some degree of degradation in the finer details. The reconstructions appear slightly blurred compared to the original spectrograms, suggesting that while the model captures the essential information necessary to reproduce the sounds, it does lose some of the more intricate nuances during the encoding-decoding process. In the case of the CVAE applied to the tiny-audio-diffusion-drums dataset, the reconstruction of LMSs is similar to the CAE, but even more definition is lost, as seen in Table 7. Same thing happens with the 29kSamplesDrums dataset. The reconstructed LMSs for each drum class maintain the general structure of the original spectrograms, but with a more pronounced blurring effect.
48 For both architectures, when converting the reconstructed LMSs back to audio, we find that the results are generally close to the original samples but with a noticeable reduction in sound quality. This degradation can be attributed to two main factors: first, the reconstructed spectrograms, while capturing the essential features of the original, lack some finer details, appearing slightly blurred and less defined; second, the process of converting spectrograms back to audio, as outlined in 3.2.2. Data Processing, introduces further distortions.
49 Table 5. Original and reconstructed LMSs for the tiny-audio-diffusion-drums dataset. hihat_184 hihat_128 snare_005 snare_054 kick_140 kick_051 Original Original Reconstructed Reconstructed
50 sd_528 mt_905 kd_63 cr_1529 ht_1273 hh_97 cy_831 ft_398 Original Reconstructed Original Reconstructed
57 The same occurs with the 29kSamplesDrums dataset when generating unconditioned samples (Table 15); and conditioned-generated samples, both globally (Table 16) and class-specific (Table 17), improve considerably with respect to the CAE. Unconditioned Table 15. LMS generated from unconditioned latent noise using the CVAE trained on the 29kSamplesDrums dataset. Conditioned on global statistical information Table 16. LMS generated by the CVAE with latent noise conditioned on global statistics from the 29kSamplesDrums training set.
58 Conditioned on class statistics cr cy ft hh ht kd mt sd Table 17. LMS generated by the CVAE with latent noise conditioned on class-specific statistics from the 29kSamplesDrums dataset. Summary of the chapter In conclusion, our evaluation of the CAE and CVAE models for drum sound synthesis has revealed both promising results and areas for improvement. The models demonstrated the ability to effectively capture and represent the underlying structure of
59 drum sounds in their latent spaces, particularly evident in the CAE's distinct clustering patterns. While both architectures showed potential in reconstructing and generating drum samples, the quality and coherence of the outputs varied depending on the dataset and generation method used. The CVAE exhibited superior performance in unconditioned sample generation, whereas the CAE excelled in reconstruction tasks. However, both models faced challenges in consistently producing high-quality, recognizable drum sounds, especially with the larger 29kSamplesDrums dataset. These findings set the stage for a deeper analysis and discussion in the following chapter. Chapter 5 will delve into the implications of these results, examining the strengths and limitations of our approach. It will also explore potential avenues for improvement and discuss the broader implications of this research for the field of AI-driven music synthesis.
60 5. Conclusions 5.1. Findings Our study has successfully demonstrated the feasibility of using deep learning approaches, specifically Convolutional Autoencoders (CAE) and Convolutional Variational Autoencoders (CVAE), for drum sample synthesis. The research provides strong support for several of our initial hypotheses, while also revealing nuanced insights into the strengths and limitations of these architectures. Log Mel Spectrograms (LMS) proved to be an adequate and efficient format for representing audio samples and training the models, supporting our second hypothesis. This representation allowed the models to capture essential characteristics of drum sounds effectively. The CAE architecture demonstrated the ability to differentiate between distinct drum sound classes through its latent space representations, as evidenced by the latent space visualizations. This finding strongly supports our third hypothesis and highlights the models' capacity to learn meaningful representations of different drum types. The study also explored the potential for controlled generation by conditioning the latent space with statistical information during inference. While this approach showed promise, particularly with the CAE model and the smaller dataset, the level of control and quality varied between datasets and model architectures. This partially supports our
61 fourth hypothesis, suggesting that further refinement of conditioning techniques could yield more consistent results across different scenarios. Comparing the CAE and CVAE architectures revealed interesting trade-offs. The CVAE showed improvements in sample generation diversity compared to the CAE, especially with unconditioned latent vectors, partially supporting our fifth hypothesis. However, this enhanced generative capability came at the cost of reduced reconstruction fidelity. The CAE, on the other hand, excelled in reconstruction tasks and showed clearer clustering in the latent space, but struggled more with unconditioned generation. The research also highlighted significant differences in model performance between the two datasets used. With the smaller tiny-audio-diffusion-drums dataset, both CAE and CVAE models achieved better results in terms of sample reconstruction and generation. However, the larger 29kSamplesDrums dataset posed more challenges, especially for the CAE model in generating recognizable drum sounds. This observation suggests that the relationship between dataset size, complexity, and model performance is not straightforward and warrants further investigation. An important finding of the study was the impact of the audio reconstruction process on the final output quality. The conversion of reconstructed or generated LMS back to audio introduced additional distortions, affecting the fidelity of the produced sounds. This highlights the critical role of the entire pipeline, from data representation to final audio output, in achieving high-quality results. While the models demonstrated the ability to generate drum-like sounds, especially when conditioned, the quality and recognizability of the generated samples varied. This suggests that our first hypothesis is partially supported, but there is still room for improvement in generating consistently high-quality, realistic drum samples. Nevertheless, the study revealed an interesting potential application in experimental music production, where even imperfect or unconventional drum sounds could find creative use.
62 5.2. Future work Building on these findings, several avenues for future research emerge. A primary focus should be on improving the audio reconstruction process from LMS to minimize distortions and enhance the quality of the final output. This could involve exploring advanced signal processing techniques or neural network architectures specifically designed for high-fidelity audio reconstruction. Another crucial area for investigation is the preservation of fine spectral details during the encoding-decoding process. This might be achieved through more sophisticated network architectures, novel loss functions, or perhaps a combination of both. Additionally, for datasets with numerous classes like 29kSamplesDrums, exploring class-specific model training could potentially yield improved results. Future research could also explore hybrid architectures that combine autoencoders with other deep learning approaches, such as adversarial training or attention mechanisms, to achieve more robust and versatile drum synthesis systems. These hybrid models could potentially overcome the limitations of individual architectures, leading to improved latent space representations, more realistic audio generation, and greater control over the synthesis process. To provide finer control over the generated output, more advanced conditioning techniques could be investigated. This could potentially allow for more nuanced manipulation of drum sound characteristics, giving users greater creative control. Developing interactive tools for real-time latent space exploration could also provide an intuitive interface for sound designers to create novel drum samples. Lastly, conducting extensive user studies with music producers and sound designers would be valuable to evaluate the practical applicability and perceived quality of the generated samples in real-world music production scenarios. This user-centric approach could provide crucial insights for further refining and tailoring the models to meet the needs of creative professionals.
63 Bibliography [1] E. Zhou and D. Lee, "Generative artificial intelligence, human creativity, and art," PNAS Nexus, vol. 3, no. 3, p. 52, 2024. [2] M. A. Boden, "Creativity in a Nutshell," in The Creative Mind: Myths and Mechanisms, London, Routledge, 2004, pp. 1-10. [3] Z. Wu, D. Ji, K. Yu, X. Zeng, D. Wu and M. Shidujaman, "AI Creativity and the Human-AI Co-creation Model," in Human-Computer Interaction. Theory, Methods and Tools, Springer International Publishing, 2021, pp. 171-190. [4] M. Müller, Fundamentals of Music Processing, 2015. [5] J. P. Briot, G. Hadjeres and F. D. Pachet, "Deep Learning Techniques for Music Generation -- A Survey," 2019. [6] S. Ji, J. Luo and X. Yang, "A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions," 2020. [7] A. V. Oppenheim, "Speech spectrograms using the fast Fourier transform," IEEE Spectrum, vol. 7, no. 8, pp. 57-62, 1970. [8] M. Farzaneh and R. M. Toroghi, "GGA-MG: Generative Genetic Algorithm for Music Generation," Journal of Evolutionary Intelligence, 2020. [9] G. Nierhaus, Algorithmic Composition: Paradigms of Automated Music Generation, Springer Vienna, 2009. [10] S. A. Hedges, "Dice Music in the Eighteenth Century," Music & Letters, vol. 59, no. 2, pp. 180-187, 1978. [11] Wikipedia, "Jaquet-Droz automata," [Online]. Available: https://en.wikipedia.org/wiki/Jaquet-Droz_automata. [Accessed 26 March 2024]. [12] C. Roads, "Artificial Intelligence and Music," Computer Music Journal, vol. 4, no. 2, pp. 13-25, 1980.
64 [13] Wikipedia, "Componium," [Online]. Available: https://en.wikipedia.org/wiki/Componium. [Accessed 26 March 2024]. [14] R. C. Pinkerton, "Information Theory and Melody," Scientific American, vol. 194, no. 2, pp. 77-87, 1956. [15] P. M. Todd, "A Connectionist Approach to Algorithmic Composition," Computer Music Journal, vol. 13, no. 4, pp. 27-43, 1989. [16] M. C. Mozer, "Neural Network Music Composition by Prediction: Exploring the Benefits of Psychoacoustic Constraints and Multi-scale Processing," Connection Science, vol. 6, no. 2-3, pp. 247-280, 1994. [17] J. P. Briot, "From Artificial Neural Networks to Deep Learning for Music Generation - - History, Concepts and Trends," 2020. [18] A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour and C. Frank, "MusicLM: Generating Music From Text," 2023. [19] P. Dhariwal, H. Jun, C. Pay, J. W. Kim, A. Radford and I. Sutskever, "Jukebox: A Generative Model for Music," 2020. [20] J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi and A. Défossez, "Simple and Controllable Music Generation," 2024. [21] F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman and Y. Adi, "AudioGen: Textually Guided Audio Generation," 2023. [22] F. Pachet, "The Continuator: Musical Interaction With Style," 2002. [23] D. Williams, V. J. Hodge, L. Gega, P. I. Cowling, D. Murphy and A. Drachen, "AI and Automatic Music Generation for Mindfulness," 2019. [24] C. Roads and P. Wieneke, "Grammars as Representations for Music," Computer Music Journal, vol. 3, no. 1, pp. 48-55, 1979.
65 [25] S. R. Holtzman, "Using Generative Grammars for Music Composition," Computer Music Journal, vol. 5, no. 1, pp. 51-64, 1981. [26] A. Gartland-Jones, "Can a Genetic Algorithm Think Like a Composer?," 2003. [27] C. Hernandez-Olivan and J. R. Beltran, "Music Composition with Deep Learning: A Review," 2021. [28] R. M. Schmidt, "Recurrent Neural Networks (RNNs): A gentle Introduction and Overview," 2019. [29] S. Hochreiter and J. Schmidhuber, "Long Short-term Memory," Neural computation, vol. 9, pp. 1735-1780, 1997. [30] D. Eck and J. Schmidhuber, "Finding temporal structure in music: blues improvisation with LSTM recurrent networks," Proceedings of the 12th IEEE Workshop on Neural Networks for Signal Processing, pp. 747-756, 2002. [31] Q. Lou, "Music Generation Using Neural Networks," 2016. [32] A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior and K. Kavukcuoglu, "WaveNet: A Generative Model for Raw Audio," 2016. [33] A. van den Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves and K. Kavukcuoglu, "Conditional Image Generation with PixelCNN Decoders," 2016. [34] D. E. Rumelhart and J. L. McClelland, "Learning Internal Representations by Error Propagation," in Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Foundations, MIT Press, 1987, pp. 318-362. [35] D. Bank, N. Koenigstein and R. Giryes, "Autoencoders," 2021. [36] D. P. Kingma and M. Welling, "Auto-encoding variational bayes," 2013. [37] A. Roberts, J. Engel, C. Raffel, C. Hawthorne and D. Eck, "A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music," 2019.
66 [38] D. Foster, Generative Deep Learning: Teaching Machines to Paint, Write, Compose, and Play, O'Reilly Media, Inc, 2019. [39] A. Bitton, P. Esling, A. Caillon and M. Fouilleul, "Assisted Sound Sample Generation with Musical Conditioning in Adversarial Auto-Encoders," 2019. [40] K. N. Haque, . R. Rana and B. W. Schuller, "High-Fidelity Audio Generation and Representation Learning With Guided Adversarial Autoencoder," vol. 8, pp. 223509- 223528, 2022. [41] A. Caillon and P. Esling, "RAVE: A variational autoencoder for fast and high-quality neural audio synthesis," 2021. [42] A. Roberts, J. Engel, C. Raffel, C. Hawthorne and D. Eck, "A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music," 2019. [43] I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville and Y. Bengio, "Generative Adversarial Networks," 2014. [44] W. Y. Hsiao, H. W. Dong, L. C. Yang and Y. H. Yang, "MuseGAN: Multi-track Sequential Generative Adversarial Networks for Symbolic Music Generation and Accompaniment," 2017. [45] S. Ji, J. Luo and X. Yang, "A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions," 2020. [46] C. Donahue, J. McAuley and M. Puckette, "Adversarial Audio Synthesis," 2019. [47] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser and I. Polosukhin, "Attention Is All You Need," 2023. [48] C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A. M. Dai, M. D. Hoffman, M. Dinculescu and D. Eck, "Music Transformer," 2018. [49] B. Haki, M. Nieto, T. Pelinski and S. Jordà, "Real-Time Drum Accompaniment Using Transformer Architecture," in Proceedings of the 3rd Conference on AI Music Creativity, 2022.