scieee AI-readable full text Open interactive document viewer

Deep learning applied to speech synthesis

Pascual de la Puente, Santiago

Abstract

Deep Learning has been applied successfully to speech processing problems. In this work we explore its capabilities, focusing concretely in recurrent neural architectures to build a state of the art Text-To-Speech system from scratch. The different steps to make the full TTS system are shown. Also, a post-filtering method to improve the generated speech naturalness is applied and evaluated. The objective results show which architecture fits better our problem, achieving low error rates in term of cepstral distortion, pitch estimation error and voiced/unvoiced classification error. Also, subjective results suggest that the model achieves a state of the art quality in the synthesis, where the post-filtering factor seems to be a key component to get a good level of naturalness. A novel architecture called Multi-Output TTS is also proposed to hold multiple speakers inside the same structure. Some hidden layers are shared by all the speakers, while there is a specific output layer for each speaker. Objective and perceptual experiments prove that this scheme produces much better results in comparison with single speaker models. Moreover, we also tackle the problem of speaker adaptation by adding a new output branch to the model and successfully training it without the need of modifying the base optimized model. This fine tuning method achieves better results than training the new speaker from scratch with its own model. Finally, we also tackle the problem of speaker interpolation by adding a new output layer (alpha-layer) on top of the Multi-Output branches. An identifying code is injected into the layer together with acoustic features of many speakers. Experiments show that the alpha-layer can effectively learn to interpolate the acoustic features between speakers.

Full text

UNIVERSITAT POLITÈCNICA DE CATALUNYA MASTER THESIS Deep learning applied to Speech Synthesis Author: Santiago PASCUAL DE LA PUENTE Supervisor: Dr. Antonio BONAFONTE CAVEZ A thesis submitted in fulfillment of the requirements for the degree of Master in Telecommunications Engineering in the TALP Research Center Signal Theory and Communications Department Escola Tècnica Superior d’Enginyeria de Telecomunicació de Barcelona June 30, 2016 iii “We are drowning in information but starved for knowledge.” John Naisbitt v Abstract Deep Learning has been applied successfully to speech processing problems. In this work we explore its capabilities, focusing concretely in recurrent neural architectures to build a state of the art Text-To-Speech system from scratch. The different steps to make the full TTS system are shown. Also, a post-filtering method to improve the generated speech naturalness is applied and evaluated. The objective results show which architecture fits better our problem, achieving low error rates in term of cepstral distortion, pitch estimation error and voiced/unvoiced classification error. Also, subjective results suggest that the model achieves a state of the art quality in the synthesis, where the post-filtering factor seems to be a key component to get a good level of naturalness. A novel architecture called Multi-Output TTS is also proposed to hold multiple speakers inside the same structure. Some hidden layers are shared by all the speakers, while there is a specific output layer for each speaker. Objective and perceptual experiments prove that this scheme produces much better results in comparison with single speaker models. Moreover, we also tackle the problem of speaker adaptation by adding a new output branch to the model and successfully training it without the need of modifying the base optimized model. This fine tuning method achieves better results than training the new speaker from scratch with its own model. Finally, we also tackle the problem of speaker interpolation by adding a new output layer (α-layer) on top of the Multi-Output branches. An identifying code is injected into the layer together with acoustic features of many speakers. Experiments show that the α-layer can effectively learn to interpolate the acoustic features between speakers. vii Resum El Deep Learning s’ha aplicat amb èxit a problemes de processament de la parla. En aquest treball explorem les capacitats d’aquesta disciplina, fent especial èmfasi en les arquitectures recurrents per a construir un sistema de síntesi de veu des de zero. Es mostren les diferents etapes per fer el sistema de síntesi complet. A més, s’aplica i s’avalua un mètode de post-processament per tal de millorar la naturalitat de la veu generada. Els resultats objectius mostren quina arquitectura encaixa més amb el nostre problema, aconseguint errors baixos en termes de distorsió cepstral, error d’estimació de pitch i error de classificació sonor/sord. També els resultats subjectius indiquen que el model arriba a tenir una qualitat de síntesi comparable amb la de les últimes tecnologíes, on el fet de fer post-processament sembla ser una peça clau per obtenir un bon nivell de naturalitat. També es proposa una arquitectura novedosa anomenada Multi-Output TTS, la qual conté diferents parlants dins la mateixa estructura. Algunes capes ocultes es comparteixen entre tots els parlants, mentres que hi ha una capa de sortida específica per a cada un d’ells. Els experiments perceptuals i objectius mostren que aquest esquema produeix força millors resultats en comparació amb els models de parlants sols. També abordem el problema d’adaptació de parlants afegint una nova capa de sortida al model i entrenant-la sense necessitat de modificar el sistema base ja optimitzat. Aquest mètode d’afinament del model a l’última capa permet obtenir millors resultats que entrenant el model del nou parlant des de zero amb el seu propi model sol. Finalment també abordem el problema d’interpolació de parlants afegint una nova capa sobre les sortides del Multi-Output, la qual es diu capa-α. A la nova capa se li insereix un codi d’identificació juntament amb les característiques acústiques dels diferents parlants. Els experiments mostren que la capa-αpot aprendre, en efecte, a interpolar valors intermitjos respecte els parlants modelats. ix Resumen El Deep Learning se ha aplicado con éxito a problemas de procesado del habla. En éste trabajo exploramos las capacidades de ésta disciplina, haciendo especial énfasis en las arquitecturas recurrentes para construir un sistema de síntesis de voz desde cero. Se muestran las distintas etapas para hacer el sistema de síntesis completo. Además se aplica y se evalúa un método de post-procesado con tal de mejorar la naturalidad de la voz generada. Los resultados objetivos muestran qué arquitectura encaja más con nuestro problema, consiguiendo errores bajos en términos de distorsión cepstral, error de estimación de pitch y error de clasificación sonoro/sordo. También los resultados subjetivos indican que el modelo llega a tener una calidad de voz comparable con la de las últimas tecnologías, donde el hecho de aplicar el post-procesado parece ser una pieza clave para obtener un buen nivel de naturalidad. También se propone una arquitectura innovadora llamada Multi-Output TTS, la cual contiene diferentes hablantes dentro de la misma estructura. Algunas capas ocultas se comparten entre todos los hablantes, mientras que hay una capa de salida específica para cada uno de ellos. Los experimentos perceptuales y objetivos muestran que éste esquema produce resultados bastante mejores en comparación con los modelos de hablantes solos. También abordamos el problema de adaptación de hablantes añadiendo una nueva capa de salida al modelo y entrenándola sin necesidad de modificar el sistema base ya optimizado. Éste método de afinado del modelo en la última capa permite obtener mejores resultados que entrenando el modelo del nuevo hablante desde cero con su propio modelo. Finalmente también abordamos el problema de interpolación de hablantes añadiendo una nueva capa sobre las salidas del Multi-Output, la cual se llama capa-α. A la nueva capa se le introduce un código de identificación del hablante junto con las características acústicas de los distintos hablantes. Los experimentos muestran que la capa-αpuede aprender, en efecto, a interpolar valores en un rango intermedio entre los dos hablantes modelados. xvi 3.12 Architecture of an LSTM cell. i(t): input gate at time-step t. o(t): output gate at time-step t. f(t): forget gate at time-step t. c(t): cell state at time-step t.................... 30 3.13 LSTM cell unfolded in time. The red arrows depict the inference flow of data between time-steps. ............. 32 4.1 Schematic of the developed two stage Text-To-Speech system. 34 4.2 Label file example resulting from processing the Spanish text "buenos días"............................. 37 4.3 Schematic representation of the Vocoder for encoding the input voice with windowed frames into acoustic parameters. . 39 4.4 Histogram of WAV files’ durations for speaker M1. It is shown in a scale of seconds. ....................... 43 4.5 Data generation architecture with parallelized pipeline. Every pipeline processes a tuple (lab,wav) and accumulates the result to be given to the data generator. ............ 44 4.6 WAV files’ durations histograms for M1 and F1 subsets of 100 files each. ............................. 45 4.7 Results of data generation acceleration for speakers M1 and F1. M1-1x: 63 min. M1-32x: 3min. F1-1x: 22 min. F1-32x: 1 min. ................................. 45 4.8 Vocoder sliding window example for an arbitrary phoneme. The stride is the increment in time taken by the next window ∆t=ti+1 −ti. The red zone shows a hypothetical empty zone out of the phoneme, so the frame duration is including information from the next phoneme or a silence. The grey arrow is the direction of the windowing analysis. ....... 47 4.9 Duration histograms for speaker M1 phonemes. Top plot shows the real durations in milliseconds, bottom plot shows the log-compressed durations. ................. 48 4.10 Plot of |ˆy−y|2against |ˆy−y|and the derivative with respect to the error:2|ˆy−y|......................... 49 4.11 Example of a regression estimation where we see how the outlier produces a high deviation over the correct estimation. (a) Gaussian distributed data with µ= 0 and σ= 4 over the line y= 1.2x. (b) Outlier artificially inserted to the data in (a). 49 4.12 Top: F0 contour example for male speaker M1. Bottom: F0 contour (blue line) example for male speaker M1 with interpolation contour (green line). .................. 51 4.13 Histogram of voiced frequency values in the training data. . 51 4.14 Final setup of the architecture, where duration prediction and acoustic prediction work together in a pipeline fashion. 53 4.15 Ratio σi g σi pper phoneme and coefficient, where σi gstands for ground-truth standard deviation at i-th coefficient, and σi p stands for prediction standard deviation at i-th coefficient. There are 32 phonemes analyzed. ................ 54 xvii 4.16 Geometric mean of σg σpcompared to the post-filtering curves for values pf = 1.04,pf = 1.05,pf = 1.06............ 55 4.17 Geometric mean of σg σpbefore and after post-filtering with p= 1.04. Green: post-filtered. Blue: raw prediction. ........ 55 4.18 Evolution in time of the 5th cepstral coefficient for two test files. Blue: natural speech. Green: post-filtered prediction. Red: raw prediction. ....................... 56 4.19 Evolution in time of the 10th cepstral coefficient for two test files. Blue: natural speech. Green: post-filtered prediction. Red: raw prediction. ....................... 56 4.20 Evolution in time of the 15th cepstral coefficient for two test files. Blue: natural speech. Green: post-filtered prediction. Red: raw prediction. ....................... 57 4.21 Box-plot of the subjective naturalness test results relative to real human voice. SPSS: Statistical Parametric Speech Synthesis. US: Unit Selection. LSTM-raw: Two-stage TTS without post-filtering. LSTM-pf: Two-stage TTS post-filtered with pf = 1.04. Ahocoded: Natural speech parameterized with the Ahocoder and reconstructed. Natural: real human voice. Red line: median. Red dot: mean. ................ 62 4.22 Activations map for the forget gates of the first 20 hidden cells in the first LSTM hidden layer. Red regions are high activations (thus preserve the past) and yellow means forget the past. .............................. 63 4.23 Activations map for the input gates of the first 20 hidden cells in the first LSTM hidden layer. Red regions mean updating the cell state a lot with the new candidate. On the other hand, yellow means not letting the information in. ......... 64 4.24 Activations map for the output gates of the first 20 hidden cells in the first LSTM hidden layer. Red regions mean letting the cell state flow out of the cell. ................. 64 4.25 Average activations for the first LSTM hidden layer with 512 cells. Input gate, forget gate and output gate are shown. Green dashed lines are the phoneme boundaries. ....... 65 4.26 Averaged activations of the input gates for the first hidden LSTM layer in blue line and 1−ftaverage activation in red. 65 4.27 Activations map for the forget gates of the 43 output LSTM cells. Red regions mean letting the cell state flow out of the cell. ................................. 66 4.28 Activations map for the input gates of the 43 output LSTM cells. Red regions mean letting the cell state flow out of the cell. ................................. 66 4.29 Activations map for the output gates of the 43 output LSTM cells. Red regions mean letting the cell state flow out of the cell. ................................. 67 xviii 4.30 Average activations for the output LSTM layer with 43 cells. Input gate, forget gate and output gate are shown. Green dashed lines are the phoneme boundaries. ........... 67 5.1 Proposed architecture using regular feed forward (dense) layers and recurrent LSTM layers. There are N outputs belonging to N different speakers. ................... 71 5.2 Exemplified training round for the N mini-batches. Dashed lines represent the corresponding output error back-propagation. The numbers in brackets express the order of that mini-batch inside the round. ......................... 72 5.3 Speaker adaptation system by fine-tuning a pre-trained Multi- Output model. The new layer can be trained in two ways: solid-line: 1) fine-tune only new branch with frozen model in the lower layers. 2) Fine-tune the whole model, thus propagating the error until the first hidden layer. .......... 73 5.4 Speaker interpolation system by training a new mixing layer, the α-layer. The new layer uses the input αcodes to learn to interpolate between the extreme examples given. ....... 73 5.5 α-interpolation training method for an example with M= 3 for 3batches of examples. The one-hot code expresses the identity of the currently shown speaker. Each smis the output prediction of the corresponding multi-output branch for the m-th speaker. ......................... 75 5.6 Training loss evolution comparison. Speakers F1 and M1 decrease the learning cost when trained with other speakers altogether. .............................. 76 5.7 Box plot of preference test scores. Scores range from −2(multiple output model is preferred) to 2(single output trained model is preferred). Both is the summary of all the answers, joining both speaker results. Red lines: medians. Blue dots: means. ............................... 77 5.8 Validation loss evolution comparison of different batch sizes, with frozen shared layers and fine-tuned shared layers. . . . 78 5.9 MCD when varying αvalues. The variation is made for speaker F1 and it is (1−α)for M1. All others speakers remain 0.M= 6.79 5.10 F0 RMSE when varying αvalues. The variation is made for speaker F1 and it is (1−α)for M1. All others speakers remain 0.M= 6............................... 80 5.11 F0 Histograms: original M1 and F1 speakers in blue. αF1= (0.25,0.5,0.75) and αM1= (0.75,0.5,0.25) interpolations in green. M= 2............................ 81 5.12 F0 Histograms: original M1 and F1 speakers in blue. αF1= (0.25,0.5,0.75) and αM1= (0.75,0.5,0.25) interpolations in green. M= 6............................ 81 xix List of Tables 4.1 Context-dependent label format. ................ 35 4.2 Number of questions per Entity. LL:Left-Left, L:Left, C:Central, R:Right, RR:Right-Right ..................... 38 4.3 Label symbol types and amount of classes per symbol. See Table 4.1 for a description of each symbol. ........... 41 4.4 Comparison of different architectures for the duration model with their objective result. FC: Fully-Connected layer. The best performing model is in bold text. ............. 58 4.5 Comparison of different architectures for the acoustic model with their objective results. LSTM: Hidden LSTM layer and units. Params: Number of parameters of the network. Embeddings: Number of input Fully Connected layers for first projections of the data. F0 RMSE: Root Mean Square Error of the F0. MCD: Mel Cepstral Distortion. UV Acc: Accuracy of Voiced/Unvoiced flag prediction. The best performing model is in bold text. ..................... 60 4.6 Statistics of the subjective results for the 6systems. ...... 62 5.1 Objective evaluation for M1 and F1 trained alone with a single output model and together with other speakers (mixed) in the multiple output architecture. ............... 77 5.2 Objective evaluation for F3 as an adaptation subject. Full: all layers are fine-tuned. Frozen: only new output branch is fine-tuned. ............................. 78 xxi List of Abbreviations ANN Artificial Neural Network BP Back Propagation BPTT Back Propagation Through Time DBN Deep Belief Network DNN Deep Neural Network FC Fully Connected FIFO First InFirst Out GMM Gaussian Mixture Model GPOS Guess Part OfSpeech GPU Graphics Processing Unit GRU Gated Recurrent Unit HMM Hidden Markov Model LSTM Long Short Term Memory MCD Mel Cepstral Distortion MDN Mixture Density Network MLP Multi Layered Perceptron MO Multi-Output MTL Multi Task Learning MSE Mean Squared Error NMT Neural Machine Translation NN Neural Network RBM Restricted Boltzman Machine RMSE Root Mean Squared Error RNN Recurrent Neural Network SGD Stochastic Gradient Descent SPSS Statistical Parametric Speech Synthesis US Unit Selection UV Unvoiced Voiced 1 Chapter 1 Introduction Speech synthesis or Text-To-Speech is the process of converting text into a voice signal. Previous to Deep Learning, existing text to speech technologies included the unit selection speech synthesis (Hunt and Black, 1996) and the statistical parametric speech synthesis (SPSS) (Zen, Tokuda, and Black, 2009). Unit selection analyzed the set of phonemes contained in a sentence and their context, and those features were mapped into pieces of recorded natural speech, all being concatenated to produce a continuous stream of voice signal. SPSS introduced the concept of learning a speaker model from data with parametric representations and then throw away the data once speaker characteristics were learned. Some remarkable differences between both was that, although SPSS could not reproduce the same level of naturalness (Zen, Tokuda, and Black, 2009) as unit selection did, it had much less footprint in memory, and it also let the user transform any speaker model to adapt the voice to different requirements in speed, intonation, etc. Two important techniques developed within the framework of SPSS were the speaker adaptation and the speaker interpolation. In the former case we could add the voice of someone that was not previously in the system for whom we have a low amount of data, and extracting characteristics from the other speaker models it could reach a good level of similarity to that of the original speaker. In the case of interpolation, a new voice model could be build from scratch without any data for a new speaker by combining the existing speaker models available in the system. Deep Learning has been applied successfully to different kinds of tasks such as computer vision, natural language processing or speech processing (Deng and Yu, 2014), outperforming the existing systems in many cases. In the case of speech synthesis, many works included Deep Neural Networks and Deep Belief Networks to perform acoustic mappings in the SPSS framework and prosody prediction. Also, Recurrent Neural Networks and their variants, like the Long Short Term Memory architecture, have leveraged completely the sequences processing and prediction problem, which makes them lead to interesting results in the speech synthesis field, where an acoustic signal of variable length has to be generated out of a set of textual entities. Despite Deep Learning TTS has improved the speech quality generated by HMM-based SPSS, it lost some of the flexibility offered by these models. First, the research in speaker adaptation for the DNN/RNN approach is very recent and not many techniques are proposed in comparison to the SPSS. Moreover and to the best of our knowledge, there are no works exploring the speaker interpolation capabilities of these systems, whilst this 2Chapter 1. Introduction was a remarkable feature in the SPSS. The first goal of this work is to build a state of the art TTS in UPC with Deep Learning techniques and analyze the insights of the developed architecture. After this, we want to achieve new flexible models capable of doing speaker adaptation and interpolation without sacrificing the generated speech quality, in addition to represent many speaker models inside the same structure. The structure of this work is the following: Chapter 2begins exploring the state of the art techniques for speech synthesis, from Unit Selection systems to the latest Deep Learning TTS architectures and methods. Then in Chapter 3there is an introduction to the Deep Learning topic, its elements, techniques and terminology. It is a guide to follow what comes in Chapter 4, which is the research and development of the TTS system made from scratch with RNN-LSTM. Afterwards, in Chapter 5a proposed multiple speaker TTS is shown and analyzed in detail, as well as the proposed adaptation and interpolation methods built on top of it. Finally, the conclusions, future work and research contributions can be found in Chapter 6. 3 Chapter 2 State of the Art Speech synthesis, also known as Text To Speech, is the technique with which computers can speak. These systems have gone through a great evolution during the past two decades, and in this chapter some currently existing systems are introduced with their most used techniques. First, we review the most used in comercial environments which are the so called Unit Selection systems. Secondly, the Statistical Parametric Speech Synthesis is seen, which has leveraged the speech synthesis research during the last decade. The chapter then concludes presenting the state of the art techniques of Deep Learning applied to speech synthesis. 2.1 Unit Selection Speech Synthesis This type of synthesis has been operative during many years because it offers the best naturalness level, as it is based on real recorded speech (Hunt and Black, 1996). The way in which this system works is by concatenating segments of speech, which are usually the so called diphone. A diphone is a voice unit of the same size as a phoneme, defined in between of two phonemes (i.e. from the middle of a phoneme to the middle of the next one). Figure 2.1 exemplifies some hypothetic diphone boundaries compared to those of phonemes. The reason to do the division at the middle point of the phoneme is because it is the more stable point and the one least influenced by the co-articulation, which is the influence of neighboring phonemes to the current one. FIGURE 2.1: Voice stream where blue dashed lines show hypothetic phoneme divisions and green lines show hypothetic diphoneme divisions. During concatenation to make the speech reconstruction the following issues need to be considered: discontinuities between the speech segments in 10 Chapter 2. State of the Art the one for the predictions of the DNN. Given this approach, their objective results indicated that the perceptual domain system achieved the best quality. Kang, Qian, and Meng (2013) followed a similar approach with a Deep Belief Network (DBN) generative model to represent the dependencies between linguistic and acoustic features. They also showed how with these RBM-DBN systems outperformed the classical HMM approach used in SPSS with objective and subjective results. In Zen and Senior (2014) the authors claim that the previous approaches where a DNN was used to model the acoustic mapping had some limitations, and they addressed the following ones with a Deep Mixture Density Network (MDN): First, the regular DNNs do not have the power to model distributions of outputs that are more complex that a unimodal Gaussian distribution. Secondly, The outputs of an ANN only provide the mean values, whilst variance has been proven to be an important property to achieve good naturalness in speech synthesis. Their MDN then provides a set of outputs that model Gaussian Mixture Models of every output acoustic parameter, with means and variances. Their results, both objective and subjective, reflect that their model can relax the limitations in the DNN- based acoustic modeling. Similarly, the work in Uria et al. (2015) use a Realvalued Neural Autoregressive Density Estimator (RNADE) (Uria, Murray, and Larochelle, 2013), which is a similar approach to that of MDN. RNADE also uses a neural network to predict a distribution of acoustic features conditioned on a set of phonetic labels by outputting the parameters of a realvalued distribution. The main difference with MDN however is the fact that RNADE predicts each dimension within an acoustic frame sequentially, thus the values of features already predicted are also input to the network, which allows RNADE to capture dependencies between the different acoustic features in a frame. In Wu et al. (2015b) the authors address two problems for the way in which DNNs are applied to speech synthesis: Perceptual sub-optimality and frame-by-frame independence. They claim that the first problem comes from the training criterion, which typically aims to maximize the likelihood of acoustic features which are a rather poor representation of human speech perception. Besides, the error in the speech feature space is not an accurate reflection of the expected perceptual error, and they propose a Multi-Task Learning (MTL) procedure to get around it, where the DNN learns to predict a perceptual representation of the target speech as a secondary task besides predicting the typical invertible vocoder parameters as the main task. These secondary task predictions are discarded during synthesis, such that they only serve as "hints" during the main task training. Regarding the second problem they refer to Recurrent Neural Networks as effective models to treat with sequential data, but they are difficult to optimize as well as computationally expensive. Their solution is simpler than any recurrent topology by means of a technique called bottleneck feature stacking, where they train a first DNN with a bottleneck hidden layer (a hidden layer with a smaller set of units than the ones preceding it). Afterwards they take the activations of the bottleneck (which yield a compact representation of both acoustic and linguistic information for each frame) of many contiguous frames (e.g. ht−1,ht,ht+1). These activations are then stacked together and joint with the linguistic features to 2.3. Deep Learning in Speech Synthesis 11 get into another DNN stage that makes the acoustic mapping, thus considering the dependencies between frames in this case. In Wu and King (2015) they take advantage of the stacked bottleneck activations to have a wide linguistic context with a training criterion that minimizes the speech parameter trajectory errors taking dynamic constraints from a wide acoustic context. This way they minimize the utterance-level trajectory error instead of the frame-by-frame error, and they achieve better naturalness than previous approaches with the proposed training criterion. In Fan et al. (2015) they propose a model to hold many speakers out of the same shared DNN structure with a specific training mechanism where they back-propagate all the speakers information in the same mini-batch. They achieve better results with the multitask approach than learning a single speaker parameters isolated. They further transfer the learning of the base shared structure for a new speaker to achieve speaker adaptation with limited training data, achieving good results in naturalness and similarity to the original speaker. In Wu et al. (2015a) they perform DNN speaker adaptation with three types of techniques: adding identity information to the input features, Learning Hidden Unit Contribution (LHUC) (Swietojanski and Renals, 2014), and making output feature space transformations. On the other hand, RNNs and their variants, like the LSTM architecture introduced in Hochreiter and Schmidhuber (1997) (see Chapter 3, sections 3.4 and 3.6), have leveraged completely the sequences processing and prediction problem, which makes them lead to interesting results in the speech synthesis field, where an acoustic signal of variable length has to be generated out of a set of textual entities. Regarding RNNs in Chen, Hwang, and Wang (1998) they explored the usage of standard RNN architectures with many hidden layers synchronized by different timings (at syllable level and at word level) for prosodic parameter prediction, such as syllable pitch contours, syllable energy levels, syllable initial and final durations, as well as intersyllable pause durations. In Achanta, Godambe, and Gangashetty (2015) they investigate two variants of RNNs applied to acoustic parameter generation: Elman-RNN and Clockwork-RNN. They show that Clockwork is equivalent to an Elman-RNN with a particular form of Leaky Integration (LI) (Bengio, Boulanger-Lewandowski, and Pascanu, 2013). Even though, the most widely used RNN in speech processing applications is the LSTM aforementioned. In Fernandez et al. (2014) they use a bidirectional LSTM architecture to predict F0 contours. In Zen and Sak (2015) they employed a unidirectional LSTM architecture to make a low-latency speech generation model. An interesting result is that using a recurrent output layer they obtained better results than making dynamic parameters prediction and deriving the trajectory using Tokuda’s algorithm (Tokuda et al., 2000). In Wu and King (2016) the authors explore the effectiveness of LSTM architectures for speech generation purposes. Concretely, they explore the necessity of the different gating mechanisms (see section 3.6) and come up with a more efficient solution that only requires the forget gate and the input gate, which in turn is the inverse of the forget gate value and does not need parameters to be learned. Finally, in Coto-Jiménez and Goddard- Close (2016) they proposed a post-filtering methodology to be carried out in SPSS where an LSTM learns its parameters to perform enhancement of 12 Chapter 2. State of the Art the predicted speech in order to be closer to a natural voice than what is obtained with HMM synthesis. 2.4 Summary In this chapter the different currently-in-use speech synthesis systems have been reviewed, where Statistical Parametric Speech Synthesis and Deep Learning Speech Synthesis are the ones under more research nowadays, whilst Unit Selection systems are the ones raising the highest-quality voice, achieving more naturalness owing to the fact that they treat with natural voice directly rather than generating acoustic trajectories with statistical models. Regarding SPSS there are explanations about the training procedure as well as the synthesis one, briefly explaining the concepts of states clustering with decision trees and linguistic contexts, which will be useful to understand the contents in Chapter 4. Finally, the set of state of the art Deep Learning techniques applied to the speech generation problem have been shown, which contain many acoustic mapping mechanisms by means of DNN and DBN architectures. Also, speaker adaptation and multitask methods are shown which raise better results than those with direct acoustic mappings. RNN architectures are also reviewed, and more concretely advanced variants like LSTM extensions, which achieve state of the art results thanks to their implicit temporal-dependency modeling of sequences. 13 Chapter 3 Introduction to Deep Learning Deep Learning is an area of Machine Learning that has been recently very explored, composed of a set of tools and techniques that let powerful models learn complex patterns automatically from data. This learning process includes the multiple transformations of the data within the model itself to get from low level processing functions to more abstract ones. This means that, when we inject information for a specific task, the model makes the data flow through many stages, extracting different sub-types of information that interact and produce the final prediction. We will see some examples of these levels decomposition, picturing some internal structures of the learned models. The models we use in this framework are the Artificial Neural Networks (ANN), and the number of transformation stages, called layers, are what conform the depth of these networks. To be more specific, lets review what an Artificial Neural Network is and what types of layer do compose it, and later some other types of ANN that get us closer to the specific tasks of this thesis will be seen. From now on we may refer to an ANN as a Neural Network (NN). 3.1 Artificial Neural Network An Artificial Neural Network is a composition of elementary units called neurons. First, we can see what a neuron is and what computation does it perform and later go to the stacking process to obtain the full network. A neuron, then, is a basic computation unit and also an analogy of what a biological neuron is. A schematic of it is depicted in Figure 3.1. There we can see some operations performed to get an output scalar yout of an input vector x={x1, x2, x3}. Equation 3.1 is the operation made in this unit. There we can see that, when the input vector xis injected into the neuron inputs, they are multiplied by a set of weights (arranged as a vector w), such that each input link has a weight, and then we sum up these products together with a bias term. Finally, the result of this addition of products is passed through a function fshown in Equation 3.2, which can be of many types. To emulate the biological neurons, which fire an electrical impulse or not depending on a 14 Chapter 3. Introduction to Deep Learning x2w2Σf Activation function y Output x1w1 x3w3 Weights Bias b Inputs FIGURE 3.1: Artificial Neuron threshold on the input sum, it is usually exemplified with the Sigmoid function σ(x), shown in Figure 3.2 in the scalar case. This unit is also called the perceptron. a= (wT·x+b)(3.1) y=f(a)(3.2) −1−0.8−0.6−0.4−0.2 0.2 0.4 0.6 0.8 1 0.2 0.4 0.6 0.8 1 x y σ1(x) = 1 1+e−5x σ2(x) = 1 1+e−10x σ3(x) = 1 1+e−15x+9 FIGURE 3.2: Sigmoid function with different scalar weightings and bias: red w= 5, b= 0; blue w= 10,b= 0; green w= 15,b= 9. Equation 3.3 shows the σ(x)function’s form. This is a thresholding function, such that when the sum in 3.1 goes beyond the saturation point, the output is constantly 1, and the same happens in the negative region where the output goes to 0. We could obtain the same behavior with a step function, but an important feature of σ(x)is that it is differentiable, something required for the learning process as we will see later in this chapter. σ(x) = 1 1 + e−wT·x+b(3.3) 3.1. Artificial Neural Network 15 Now that the artificial neuron has been introduced along with the operation it performs, an explanation of what this means geometrically will further help to interpret what happens in the stacking process of these elements to form a Neural Network. The equation 3.1 is the expression of a hyperplane, where the set of weights w={w1, w2,··· , wN}control the rotation and skew of it, and the bias term bcontrols its translation from the origin. This is then a linear operation dividing the hyper-space Sinto two-splits, and the application of the Sigmoid function lets us interpret the regions in those splits as probabilities of pertaining to a class conditioned on the input. The neuron’s operation after the σ(x)is called the Logistic Regression, a technique to perform binary classification tasks modeling the posterior probabilities with the hyper-plane equation we have seen earlier (Hastie et al., 2005). So if we interpret the σ(x)as being the probability, as mentioned earlier: P(Y= 1|X) = 1 1 + e−wT·x+b=ewT·x+b 1 + ewT·x+b(3.4) This brings the idea of the logit, which is the inverse of the logistic function and expresses that this probability is derived from a linear regression of the boundary between the classes, which is in turn the neuron operation. logit(P(Y= 1|X)) = log( P(Y= 1|X) 1−P(Y= 1|X)) = wT·x+b(3.5) Equation 3.5 is a Linear Regression (Hastie et al., 2005), and it lets us approximate functions depending on the w={w1, w2,··· , wN}predictors (e.g. Price of a house based on its location, year of construction and area). So this shows how the neuron first estimates an approximating linear function to build a hyper-plane, and then the element-wise Sigmoid turns the estimation into a posterior probability. The following explanation will be focused on classification tasks, but it is a natural extension to an N-dimensional regression as well. The neuron is the beginning of what a neural network will be, and now we will see the limitations of the single perceptron in order to get to a more complex model. First, the perceptron only lets us classify in a binary fashion, because there is only one neuron to fire an output. Second, it can only work with a single hyper-plane to discriminate highly complex patterns. The natural extension to overcome the binary limitation will be to put many neurons in parallel, each processing its binary output ynfrom a set of inputs: x={x1, x2,··· , xM}. When dealing with classification all the binary contributions are normalized to sum up to one, such that we obtain a probability distribution out of the Noutput neurons. The output activation becomes the one in equation 3.6, and its name is the Softmax function. P(y=k|x) = exp xTwk PN n=1 exp xTwn (3.6) Where kstands for the k-class and Nis the total output neurons as previously stated. This topology is depicted in Figure 3.5. For convention, the Input Layer is drawn like a set of Mneurons, but they are not real neurons in the sense of computation units, just each input xm. 16 Chapter 3. Introduction to Deep Learning . . . . . . x1 x2 x3 xM y1 y2 yN Input layer Ouput layer FIGURE 3.3: Logistic Regression of Noutputs. This is letting us make a many-to-many mapping operation: RN→RM Now the other limitation was the computational power of this system. This can be seen with the XOR example, which is frequently used to illustrate the shortcomings of the Logistic Regression with respect to the Neural Network. This is again a binary classification problem (i.e. single perceptron). The neuron has to learn the XOR operation for two inputs, such that we have: X={(0,0),(0,1),(1,0),(1,1)} Note the Xnotation meaning we can arrange the values as a matrix of values, where we have 4rows and 2columns. X=     0 0 0 1 1 0 1 1     Each row is called an input sample to our model, and we want the following predictions: Y=     0 1 1 0     Where we have 1prediction per input sample. So in this example M= 2 and N= 1, following our previously established notation. If we try to plot this in a 3Dspace, we would get what is depicted in Figure 3.4. 3.1. Artificial Neural Network 17 00.20.40.60.810 0.5 1 0 0.5 1 x1 x2 y FIGURE 3.4: XOR resulting approximate surface. From Figure 3.4 we can intuitively get to the conclusion that a single neuron could separate one of the two four existing regions (two regions with 0 and two regions with 1). Thus we need to add at least the capacity to define two discriminating units, which brings the idea of the Hidden Layer, an intermediate set of neurons that will first map the input space to a linearly separable representation where the final decision will be taken (either classification or regression). In this case, the following Neural Network has the weights trained to make the XOR function: 1 1 1/0 1/0 x1 x2 h1 h2 y1 −5 −5 10 −10 −10 10 −5 10 10 Input layer Hidden layer Ouput layer FIGURE 3.5: XOR Neural Network with Sigmoid perceptrons. This is the so called Artificial Neural Network, a set of artificial neurons linked together in order to perform an arbitrary function, as in this example where the XOR is implemented with the depicted trained weights. An important and hard task for these systems is the learning process, that will be reviewed later in this chapter. The learning process is what makes the 18 Chapter 3. Introduction to Deep Learning network set its weights to perform the desired function. Another important thing to point out is that connections between layers are depicted with arrows, which indicates that this model is directed, and this is why this is also called a feed-forward Neural Network, thus making the inputs flow up to the output layer in what we call the inference operation. To inspect a little bit the final components and equations of this system, we see that in this XOR toy example the inputs are x1and x2, which can have binary values 0or 1, and there is one output y1, binary as well. The unlabeled neurons are the biases, drawn here as inputs in every layer such that the bias values can be appreciated together with the weights. This is incorporated in the equation 3.1 by appending the scalar value 1as an input in x, and letting the value wi0=biin the set of weights on every neuron. So now, as we have many neurons in every layer, the set of all links from layer ito layer jcan be written as a matrix Wji, where each wji is the link weight between neuron/input in layer iand neuron in layer j, which turns to be the layer i+ 1. Also, Iis the number of neurons/inputs in layer i, and Jis the number of neurons in layer j, leading to a matrix WJ×I. Now, the computation of a layer are called activations (e.g. for the hidden layer in Figure 3.4) and the operation can be written as: z=σ(Wji ·x)where x={1, x1,··· , xN}(3.7) Later, when going in depth with many architectures, we may refer to a feed-forward layer as a hidden or output layer composed of this directed connections and activation functions (such as the Sigmoid). Yet another plausible name will be Fully Connected (FC) layer. Depending on the type of input/output values the sigmoidal units may implement the tanh as a non-linearity instead of the Sigmoid, due to its extended range between {−1,1}. Nonetheless, in regression tasks the output might be directly the operation in equation 3.1 for all neurons in the output layer, such that the layer operation is purely linear and returns a real value. Regression also works whilst predictions are made within the linear region of the function with normalized outputs. Also, in terms of notation there is a variant name for the whole feed-forward architecture which might be used in this work as well, the Multiple Layer Perceptron (MLP). 3.2 Training the network: Back-Propagation algorithm In this section it is discussed the way in which the network is trained, thus the way in which it learns the weight value of every link and the bias value of every neuron. The book of Bishop (1995) is a very good reference for understanding this methodology, and many parts of this section will be based on it. Also the work from LeCun et al. (2012) is a thorough analysis of good practices to make this algorithm work well, but it is out of the scope of this work to discuss every aspect with the tricks and tweaks of the learning procedure. The learning process is based on a cost function (also called error function) that we define depending on the problem we are facing, whether it is classification or regression. Given this cost the network’s task is to minimize it 3.2. Training the network: Back-Propagation algorithm 19 with respect to its weights and biases, so the cost function must be differentiable. The algorithm to evaluate the derivatives for the cost function of our choice is called the Back-Propagation (BP) algorithm, and it is based on the propagation of errors backwards through the network structure, from the output layer to the input layer, in order to correct the weights that provoke these mistakes with small modifications proportional to the committed errors. It is important to recall that there can be many non-linearities happening layer by layer at this point (see equation 3.7), and this provokes the lack of a closed solution for estimating the parameters. This algorithm then behaves in an iterative manner, such that it learns step-by-step for certain inputs shown many times to the model. As posed in Bishop (1995), each step can be divided in two stages: 1. First compute the derivatives from the error and back-propagate them. 2. Secondly update the weights of each layer. It is worth mentioning that the second stage can be implemented in a variety of ways, and some reference to modern techniques will be given later in this section. Now it is good to have a more in-depth analysis of the BP algorithm for a general network, with an arbitrary amount of feed forward layers, with an arbitrary non-linearity and with an arbitrary error function. In a general feed-forward neural network, each unit computes its activation like shown in equation 3.7 previous to applying σ, which can be rewritten as: aj=X i wjizi(3.8) where ziis the activation of a unit or an input (equation 3.7 is written with xi, which has been changed by zito better conceive it as an arbitrary input, either from the input neurons or from an intermediate layer). ziis then injected through the connection wji, meaning it goes from all neurons ito the receiving neuron j. After applying the f(.)non-linearity to the activation we have the final activation of the jneuron as: zj=f(aj)(3.9) In equation 3.2 the output unit was denoted with a y, but again zjgeneralizes to show it could be the activation of an intermediate hidden layer. However, when referring to the network outputs the notation will still be yn. It has been mentioned that what we want are a set of weights optimized by means of an error function. We can take into account an error defined for each (input,output) training tuple of examples with the following expression: E=X t Et(3.10) Where tenumerates the input/output pair. The error Etis a differentiable function of the outputs, so that: Et=Et(y1···yN)(3.11) 26 Chapter 3. Introduction to Deep Learning Dropout. 3.3 Deep Neural Network The concept of Deep Neural Network (DNN) is based on stacking many hidden layers to make the model deeper. Neural Networks appeared decades ago, but building deep models was not a simple task by then, because the BP algorithm did not work well when dealing with deep architectures as the gradients vanished easily, and it was also usual falling at a local minima, because increasing the model complexity makes it highly non-linear so that local minima might be more frequent. Training deep architectures was also quite slow because of the lack of computational resources. And another very important drawback was the lack of data. Nowadays it is different though, because the Internet provoked a huge increase in the amount of data available (images, videos, speech, audio, text, etc.), and the Graphical Processing Units (GPUs) make the matrix operations really fast and efficiently (for instance, the equation 3.7). Also, new learning algorithms like Adam (Kingma and Ba, 2014), make the learning faster with a better propagation of the gradients through the model. Furthermore, many initialization schemes for the weights have been proposed so that it gets harder to fall into a local minima, and this has been a trendy area of research over the last years in Deep Learning. These methods are mostly based on pretraining the models (i.e. give some meaningful values to the weights in order to make them not random and close to what we want to achieve). It is out of the scope to explain how these mechanisms work, however the Restricted Boltzman Machines (Tutorial on Restricted Boltzman Machines (RBM) 2010), the Denoising Autoencoders (Tutorial on Denoising Autoencoders (DA) 2010) and the Deep Belief Networks (Tutorial on Deep Belief Networks 2010) are referenced if the readers want to get further details. Special attention is given to recurrent architectures as they are the main type of model considered in this work. 3.4 Recurrent Neural Network Recurrent Neural Networks (RNNs) are a special type of Neural Network topology well suited for processing sequences. The recurrent keyword stands for the fact that these networks perform the same task for every element of a sequence (Recurrent Neural Networks Tutorial 2015), having an output that depends on previous computations. This characteristic can be seen as a memory feature that these models have, as they "remember" the past flows of information. When we talk about a unidirectional recurrent layer we can say that it has memory about the past, so at time t they have an input vector xt∈Rnand the memory state at time t−1ht−1∈Rm, producing the new memory state ht, also called hidden state, with the following set of operations: ht=g(W·xt+U·ht−1+bh)(3.22) 3.4. Recurrent Neural Network 27 where Wis the input-to-hidden weights matrix (feed forward behavior), U is the hidden-to-hidden weights matrix where the feedback is made, bhis the bias vector and gis a specified element-wise non-linear transformation, such as the hyperbolic tangent or the sigmoid seen previously. As we can see, for every input sequence X={x1,x2,...,xT}we obtain an output sequence H={h1,h2,...,hT}, where each output from the layer keeps track of dynamic changes in time, and this is what makes the recurrent model a really powerful option for sequences. These models then keep track of the context for the input features, and a very interesting property they have is that they share the same parameters for every time-step of the process, thus reducing the amount of weights needed to take into account a large context. It is usual to visualize an RNN unfolded in time, like in Figure 3.10. There the output is also shown as the result of another matrix computation, though it is not part of the recurrent layer: yt=f(V·ht+bv)(3.23) FIGURE 3.10: Example RNN neuron unfolded in time. The training for RNNs is similar to that of feed forward ones by means of using a BP algorithm, but here the time dimension also needs to be taken into account. This is important because as mentioned previously, the parameters of the recurrent matrix are shared between time-steps, thus the gradient depends not only on the current time step, but also on the previous ones. For instance as it is exemplified in (Recurrent Neural Networks Tutorial 2015) if we wanted to compute the gradient at t= 4 we would need to back-propagate 3steps and sum up the gradients. This technique is called Back-Propagation Through Time (BPTT), and a quick review to take a glance at the different with the standard one is given in the next section. 28 Chapter 3. Introduction to Deep Learning 3.5 Back-Propagation Through Time First and as previously settled for standard BP (see section 3.2), a cost function is defined to train our RNN, and in this case the total error at the output of the network is the sum of the errors at each time-step: E(y,ˆ y) = T X t=1 Et(yt,ˆ yt)(3.24) Where ytis the ground-truth example we show to the network at time-step tand ˆ ytis its prediction. For the case of the network shown in Figure 3.10 the goal of BPTT is to compute the gradients with respect to our parameters W,Uand Vto learn the appropriate weight values, so it is the same idea as in the feed forward BP. Just as it is done in equation 3.24 to sum up the errors, the gradients also get summed through time as: ∂E ∂W= T−1 X t=0 ∂Et ∂W(3.25) The following procedure can be found in more detail in (Recurrent Neural Networks Tutorial 2015). The chain rule is applied again here, and the first gradients we refer to are the ones for Vwhich are quite straightforward: ∂E3 ∂V=∂E3 ∂ˆ y3 ∂ˆ y3 ∂V=∂E3 ∂ˆ y3 ∂ˆ y3 ∂z3 ∂z3 ∂V(3.26) Where z3=V·h3. Beware that the time-step 3has been considered to continue with the first example that was mentioned with the RNN gradient computations. An important thing to note here is how ∂E3 ∂Vdoes not depend on other values than the current time-step, as the dependency is built upon ( ˆ y3,y3,h3). It is not like so with the matrices within the recurrent layer however, because for Umatrix for example: ∂E3 ∂U=∂E3 ∂ˆ y3 ˆ y3 ∂h3 ∂h3 ∂U(3.27) Now from equation 3.22 we see that h3depends on h2, which in turn depends on h1and so forth. This means that taking the derivative is applying the chain rule for as many time-steps as we have: ∂E3 ∂U= 3 X t=0 ∂E3 ∂ˆ y3 ˆ y3 ∂h3 ∂h3 ∂ht ∂ht ∂U(3.28) This is expressing that the gradient is accumulated through time and not only in the feed-forward way. Figure 3.11 illustrates the gradient flowing backwards from the current time-step. As previously seen with the feedforward BP-algorithm, a delta can also be defined here, being it: δ(t) 2=∂Et ∂zt−1 =∂Et ∂ht ∂ht ∂ht−1 ∂ht−1 ∂zt−1 (3.29) 3.5. Back-Propagation Through Time 29 Where zt−1is defined to be the activation (previous to applying any nonlinearity) of the recurrent layer: zt−1=W·xt−1+U·h1(3.30) FIGURE 3.11: Gradient flow at the fourth time-step 3. Violet arrows: inference flow of the input data. Red arrows: back-propagated gradients/errors. In practice the distance-in-time can get too large and that makes the gradient propagation difficult, because the back-propagation in time can be seen as back-propagating through many layers (as many as time-steps), so the number of time-steps is truncated to have a maximum length of the training sequences, and then the algorithm is sometimes called Truncated Back-Propagation Through Time. This type of RNN are usually called the vanilla ones, which stands for the basic model with tanh/sigmoidal units. In theory they should work well to model any time dependencies so that they recall any context within a range of Ttime-steps, but in practice they suffer from two problems: The vanishing gradients and The exploding gradients. Both of these problems are based on the inherent behavior of the equation 3.22 applied time-step by time-step: ht=g(W·xt+U·g(···g(W·xt−T+U·ht−T+bh)···) + bh)(3.31) The process of multiplying the matrix Uat every time-step is what makes the gradients so unstable, rapidly diminishing to zero or exploding depending on the value of the highest eigenvalue of the matrix. This makes it really complicated to train these models in an effective manner and also provokes the lost of long-term dependencies in memory when doing inference, because the multiplicative behavior of these neurons do not preserve well the information through long sequences. To overcome these effects there are extensions of the vanilla RNN, leading to what are called memory cells. Currently there are two main types of these cells: Long Short Term Memory (LSTM) cells and Gated Recurrent Units (GRU) (Cho et al., 2014b). The 30 Chapter 3. Introduction to Deep Learning former ones have been extensively used in this work and this is why the next section concludes the background by presenting them. Nevertheless, there are some really interesting works exploring the capabilities of some of these architectures empirically, showing that they perform quite similar most of the time (Chung et al., 2014; Wu and King, 2016). 3.6 Long Short Term Memory In this work we used LSTM layers, as they cope better with the vanishing gradient problems (Hochreiter, 1998) that appeared when training regular RNNs as mentioned previously. They also model the long term dependencies in a better way than the simple RNNs do because of their gating mechanisms (Hochreiter and Schmidhuber, 1997). The structure of an LSTM cell is shown in Figure 3.12. There are three gates in this structure: •Input gate: control the flow of information coming in. •Forget gate: control which components of the cell state are forgotten (i.e. multiplying by zero to delete from memory). •Output gate: control the flow of information going out. FIGURE 3.12: Architecture of an LSTM cell. i(t): input gate at time-step t. o(t): output gate at time-step t. f(t): forget gate at time-step t. c(t): cell state at timestep t. Figure 3.13 shows the unfolded version of the LSTM cell, where we can see two main flows of information through time: htand ct. The equations that describe this architecture are the following ones: it=σ(Wixt+Uiht−1+bi)(3.32) ˆ Ct= tanh(Wcxt+Ucht−1+bc)(3.33) 3.6. Long Short Term Memory 31 ft=σ(Wfxt+Ufht−1+bf)(3.34) Ct=itˆ Ct+ftCt−1(3.35) ot=σ(Woxt+Uoht−1+bo)(3.36) ht=ottanh(Ct)(3.37) The symbol means element-wise product here. We can see how there are many matrices now, as well as bias vectors: Wi,Ui,Wf,Uf,Wo,Uo,Wc,Uc, bi,bf,boand bc. All these weights are now parameters to be learned for the LSTM layer, so the architecture has become way more complex that the vanilla one, but the interesting thing is that the gating mechanisms control what to keep and what to forget about the inputs and the past. There is a good analysis in Understanding LSTM Networks (2015) about it, summarized here. The gates are seen as soft-switches (because of the sigmoid that is bounded σ∈ {0,1}, but with a transition of real values between the boundaries): •The forget gate then decides whether to forget contents or not for each cell in the LSTM layer (i.e. we have a vector of forget activations per time-step) by multiplying the past cell states by its activation value: 0means forget completely about what was seen, 1means keep it all. The forget activation is obtained by looking at the past ht−1and at the current input xt. •The input is on behalf of deciding whether the new information available at the input gets into the memory state or not. This process is divided in two steps: –Get the activation of the soft-switch by looking at the past ht−1 and the input xt. –A candidate vector ˆ Ctis computed (equation 3.33) that could be added to the memory cell state. Then the results of these two steps are combined to update the state with the right amount and type of information. •Next, the new cell states Ctare computed. It is important to note that the computation (equation 3.35) includes applying the forget elementwise product to every past cell state. Also, the combination of "What information should get into the memory" and the candidate memory are multiplied, and this way of updating the information stored in the cell is the main difference with the vanilla RNN. Here the updates are computed through a summation instead of a multiplication, and also the regulation of input and forgot flows makes these systems able to store long-term dependencies. 32 Chapter 3. Introduction to Deep Learning •To finish, the output gate activation is computed (from ht−1and xt again) and applied to the new cell state to finally get the right components to generate (i.e. "what is actually required from the different cells in the layer to generate the right information?"). FIGURE 3.13: LSTM cell unfolded in time. The red arrows depict the inference flow of data between time-steps. 3.7 Summary In section 3.1 we have went through the basics of what an Artificial Neural Network (ANN, NN, MLP) is, from its elemental unit (the neuron), to the full architecture of the feed-forward construction. The idea of the different layers (Input, Hidden and Output) has been explained, detailing why any hidden units are required to make a more powerful model with a toy example, which was the XOR. Also, the layer activation equation was shown and explained, looking at the components that compose it. Following to the basics, the Back-Propagation algorithm has been discussed in section 3.2 in order to see the learning details of the neural networks. There some terminology has been introduced to be understood in later chapters when dealing with the models developed for this work. The feed forward architectures explanations conclude with a brief description of the Deep Neural Network architectures, where some past issues are exposed regarding the difficulties of training these models and some current solutions the problems are mentioned. The last part of the chapter is focused on Recurrent Neural Networks and the way in which they are trained, and more concretely Long Short Term Memory networks and their sequence modeling capabilities are explored. Also, a detailed explanation of the LSTM gating mechanisms is given, something that is useful for the gating analysis performed in Chapter 4. 33 Chapter 4 Two stage Text-to-Speech with RNN-LSTM 4.1 Introduction The main purpose of this work has been building a state of the art Text To Speech (TTS) system with Deep Learning techniques. As seen in Chapter 2, some previous works performed quite well with Deep Neural Network variants and also with recurrent architectures, either as standalone systems for mapping linguistic contents to acoustic frames, or to replace some parts of the SPSS pipeline. In this section, the process to build the full TTS system for this thesis is explained, which is fully built with RNN architectures as later will be seen. Figure 4.1 represents the general scheme of the developed work. The system is subdivided in two main steps: training and synthesis, similar to the SPSS system seen in Chapter 2. During training there is a Speech Database available that contains contextual features and speech recordings. The contextual features are generated from analyzing the textual transcriptions of those recordings, and they will be presented in section 4.2.1. It is basically a representation of the linguistic complexities that must be derived into speech signal. Out of the recordings of that database many acoustic features are also extracted by means of a Vocoder, as explained in section 4.2.2, which are used to train two models: duration RNN and acoustic RNN. This two stage architecture was influenced by the work in Zen and Sak (2015). Both training processes are depicted in Figure 4.1 in the Duration RNN training and Acoustic RNN training blocks. Once the models are trained the weights are saved for later usage. The second part of the system is the synthesis stage, where both trained models are used in a pipeline fashion. First, raw text is inserted by the user and the text analysis front-end converts it to contextual features. Then these features are injected into the duration model that will predict the duration of each phoneme one by one, sending its predictions to the acoustic model, which will generate the acoustic parameter trajectories for a Vocoder. In the end, the Vocoder will convert the parameterization of the speech into the waveform again. This chapter is structured in the following manner: Section 4.2 describes in detail what data is required to train this model and how it is prepared, with a parallelized pipeline architecture to speed things up. The Two-stage-TTS modules are described in Section 4.3, and a post-processing improvement to 34 Chapter 4. Two stage Text-to-Speech with RNN-LSTM FIGURE 4.1: Schematic of the developed two stage Text-To-Speech system. 4.2. Data Preparation 35 achieve better naturalness in the synthesis is applied in section 4.4. Finally, results are shown and discussed in section 4.5. 4.2 Data Preparation In this section it is explained the way in which data is generated. This is basically the process to go from raw data to features that will be processed either as predictors or as predictions. To clarify things and avoid confusion, input features are called predictors and outputs are called predictions. First, the textual features, which will be mainly predictors, are explained. Their types and the process they go through is depicted, such that it is understood why are these features chosen for the synthesis purpose. Then, acoustic features will be described, as well as the process they go through to produce the speech stream with the Vocoder. A parallelization mechanism to speed up the generation of features is also briefly shown. It is considered to be an important part of the system and the project, owing to the fact that it lets us make more experiments and quicker if we want to generate different sets of speakers’ features. We will see in the next chapter that we require data from many speakers to be generated, so this tool is useful for that. 4.2.1 Text to Label process As mentioned earlier, raw text is processed into a more convenient representation that we call label. This representation is composed of a set of contextualized prosodic and phonetic features. The features are a phonetic transcription of a few windowed phonemes, so that the synthesis of the current phoneme takes into account the surrounding phonemes for coarticulation purposes. Also, information about stressed syllables, position of the phoneme inside the current syllable, the position of the syllable in the word, etc. is embedded in these features. Table 4.1 shows the label format features/ symbols and their descriptions. TABLE 4.1: Context-dependent label format. label format Symbol Description p1 phoneme identity before the previous phoneme p2 previous phoneme identity p3 current phoneme identity p4 next phoneme identity p5 the phoneme after the next phoneme identity p6 position of the current phoneme identity in the current syllable (forward) p7 position of the current phoneme identity in the current syllable (backward) a1 whether the previous syllable is stressed or not (0; not, 1: yes) a2 whether the previous syllable is accented or not (0; not, 1: yes) 42 Chapter 4. Two stage Text-to-Speech with RNN-LSTM Table 4.3 (continued) c2 Boolean 2 c3 Real - d1 Categorical 47 d2 Real - e1 Categorical 47 e2 Real - e3 Real - e4 Real - e5 Real - e6 Real - e7 Real - e8 Real - f1 Categorical 47 f2 Real - g1 Real - g2 Real - h1 Real - h2 Real - h3 Real - h4 Real - h5 Categorical 6 i1 Real - i2 Real - j1 Real - j2 Real - j3 Real - As we will see soon (section 4.3), in the acoustic prediction there are two additional inputs derived from the duration model. Those inputs are the duration of the current phoneme to synthesize, and the relative position of the current frame within the phoneme. We will see in detail what do they mean exactly, but a thing to point out here is that duration is log-normalized and then the max-min (between {0,1}) is applied. On the other hand, the relative duration is normalized by the absolute duration. Equations 4.5 and 4.6 express these operations. Note that the relative duration is computed with the absolute duration in the truth time dimension, not the log-compressed range. ˆ d=ln d−ln dmin ln dmax −ln dmin (4.5) ˆrd=rd d(4.6) The reason to make the log-compression will be seen with the explanation about the duration model. 4.2. Data Preparation 43 4.2.4 Parallelizing the pipeline: Speeding up the data generation The label generation is made with Ogmios from Bonafonte et al. (2006a) as mentioned before. Once all text is converted to label files, which is a quick task, all recordings have to be processed to get the Ahocoder (Erro et al., 2011) features. This process is time-consuming, especially given the many files of long duration available, as shown in the durations histogram in Figure 4.4. FIGURE 4.4: Histogram of WAV files’ durations for speaker M1. It is shown in a scale of seconds. To accelerate the Ahocoding, a parallelization approach has been taken. The data generation system is then composed of the following components: •FIFO Queue: to keep the audio files to be processed ready to go into the computation flow. •Pipeline: A pipeline is composed of two modules: –Audio trimmer: trim the first and last silence regions to a maximum length of 100 ms each. This is done to avoid having an overhead of silence samples that can bias the network learning procedure, as it could get to learn how to produce the silence to much at the expense of doing worse with other phonemes. Moreover, the length of the silences at the beginning and end of recordings are more erratic than the ones inside the sentence, which were tied to a prosodic pattern. –Ahocoder: make the spectral estimation through Ahocoder. •Datagen: the data generator, which converts all the resulting lab files and acoustic parameters into the tables that the TTS will process. The data generation task is fast enough to not require parallelization, as it is only gathering all data into the table structure with some intermediate transformation (i.e. the aforementioned log-compression of some features). 44 Chapter 4. Two stage Text-to-Speech with RNN-LSTM The architecture shown in Figure 4.5 is thus the one for accelerating the data generation process. There we can see how, having the WAV and label files, a process queue is built to hold up to Nparallel ahocodings. Later, all those results are used by the data generator to build the tables and files: •Training split for duration prediction: 362 predictors and 1prediction. •Test split for duration prediction: 362 predictors and 1prediction. •Validation split for duration prediction: 362 predictors and 1predic- tion. •Training split for acoustic prediction: 364 predictors and 43 predictions. •Test split for acoustic prediction: 364 predictors and 43 predictions. •Validation split for acoustic prediction: 364 predictors and 43 predictions. •Acoustic statistics for normalization •Duration statistics for normalization FIGURE 4.5: Data generation architecture with parallelized pipeline. Every pipeline processes a tuple (lab,wav) and accumulates the result to be given to the data generator. Some tests have been made in order to evaluate the acceleration provoked by this system. For this, we picked the two speakers: M1 (male) and F1 (female). See section 4.5.1 for further details of the speakers. Then, we took 100 audio/label files for each speaker and ran the data generation pipelines with 1and 32 parallel processes. The histograms for the duration of the 100-files-subsets per speaker can be seen in Figure 4.6. There we see that the mode of the histograms (both of them) is around 15 seconds, but there are many files ranging from 15 to 45 seconds. The speed improvement can be clearly appreciated in Figure 4.7. In the most extreme case (i.e. the male voice), the elapsed time got reduced from 1hour to 3minutes. 4.2. Data Preparation 45 FIGURE 4.6: WAV files’ durations histograms for M1 and F1 subsets of 100 files each. FIGURE 4.7: Results of data generation acceleration for speakers M1 and F1. M1-1x: 63 min. M1-32x: 3min. F1-1x: 22 min. F1-32x: 1min. 46 Chapter 4. Two stage Text-to-Speech with RNN-LSTM 4.3 Two-stage RNN-LSTM model The are two types of information required to produce a voice with good quality: •The prosodic information: It considers intonation, phoneme duration, pauses between words, etc. characteristics that can make a huge effect on the voice naturalness. •The acoustic information: Spectral estimation processed by the Vocoder system to generate the waveform. A good estimation is required for naturalness and also intelligibility. The prosodic prediction is the first problem we tackle in this TTS design, and to be more specific we begin with special focus on the phoneme duration prediction. We basically need to know the amount of frames to be generated with the Vocoder, and then those frames will be generated out of the acoustic prediction system. This is where the "Two-stage" name comes from: 1. Predict the duration for the current phoneme out of the encoded input linguistic features. 2. Predict the acoustic frame coefficients, for as many frames as dictated by the duration prediction, also taking the linguistic features. Beware of two things at this point: there is no specific pause prediction system, because the pause annotations are given by the front-end that generates the labels. Secondly, the F0 contour which carries the intonation behavior is not predicted at this stage but with the acoustic model. Each of the two stages is performing a mapping, where the chosen base model for both tasks is a Recurrent Neural Network (RNN), and more concretely a Long Short Term Memory (LSTM), yet the two tasks end up having different architectures (we will see both later) as we have two different problems, despite working in a pipeline fashion. The Recurrent Neural Network is a well suited model for these tasks, because it grants the following properties: •It keeps track of the sequential context, as it reads the input phonemes by time-steps (phoneme by phoneme). Each time-step is thus an encoded label. •In the case of the acoustic prediction, having a recurrent output layer improves the continuous prediction between frames (Zen and Sak, 2015), getting even better results than predicting static and dynamic features, as it was used to be done with SPSS (see Chapter 2). In the acoustic prediction we want to work with frames (like the Vocoder does). Figure 4.8 exemplifies the hypothetical analysis for a phoneme that lasts less than 3windows. It helps in the reasoning for the way in which 4.3. Two-stage RNN-LSTM model 47 duration is predicted. It could be done by predicting the amount of frames that the current phoneme lasts, owing to the fact that the acoustic RNN will work frame-wise. Although predicting the amount of frames as an integer is a possibility, it imposes an error for the duration RNN, and it is that we show it examples biased to a quantization error. In Figure 4.8 there is a red region containing a not-current-phoneme region that fell inside the third window, and in an frame-wise duration prediction fashion that would be mistaken with a rounding error of 3windows. FIGURE 4.8: Vocoder sliding window example for an arbitrary phoneme. The stride is the increment in time taken by the next window ∆t=ti+1 −ti. The red zone shows a hypothetical empty zone out of the phoneme, so the frame duration is including information from the next phoneme or a silence. The grey arrow is the direction of the windowing analysis. This brings the idea of making the duration RNN predict the real amount of time that the current phoneme lasts (normalized), and with this we avoid making bigger errors in the duration prediction stage produced by unnecessary round operations. Thus, the output of the duration model is a linear feed forward layer, also called Fully Connected layer (FC), such that: y=w·x+b Some experiments were made to train the model with a recurrent output. This idea comes from the acoustic model, as will be seen soon, and we wanted to see if there was any improvement with this task, but there was not any besides increasing the number of parameters of the network. Moreover, it has been mentioned how the RNN models keep track of the past context, going phoneme by phoneme, and this makes us discard some unnecessary information embedded in the classic label format. Hence features referring to past-context are removed, as commented previously in section 4.2.3, going from 405 initial features to 362. These conform what is called the linguistic inputs, which is a 362-dimensional vector that is fed into the duration RNN. An important issue here is the way in which the duration is normalized. Previously we mentioned that there is a log-compression of the real duration per phoneme. This is to smooth the effect of the outliers that the model may encounter during training. 48 Chapter 4. Two stage Text-to-Speech with RNN-LSTM FIGURE 4.9: Duration histograms for speaker M1 phonemes. Top plot shows the real durations in milliseconds, bottom plot shows the log-compressed durations. Figure 4.9 depicts the log-compression result in the histograms for the male speaker (M1) phonemes. In the real durations there are many examples provoking a long-tailed distribution of the data, something that distorts the regression training with the Mean Squared Error (MSE) minimization. This is why the logarithm is applied. To be more specific, the natural logarithm has been applied in all the experiments for this work. This is then finally normalized in a min-max range, as mentioned in section 4.2.3, making the values fall in the range {0.01,0.99}. To get the duration in milliseconds during prediction we just take the exponential of the prediction, after denormalizing the min-max range. Let’s analyze why are the outliers compressed with this method: the sum of squares error that is used to train this model for the regression purpose comes from the supposition that the target data (i.e. the duration examples shown to the network for back-propagation of the errors) follows a Gaussian distribution, and the distribution of the duration values then looks closer to a Gaussian after compressing the long tail we have seen. This means that when the logarithm is applied the outliers turn into inliers. To justify this transformation we can see what would be the effect of outliers during training if we analyze the behavior of MSE in equation 4.7 (Bishop, 1995). E=X t N X n=1 |ˆyn(xt;w)−yt n|2(4.7) In our case there is only one output to compute the cost so we can express it like the sum over the training samples: E=X t|ˆy(xt;w)−yt|2(4.8) 4.3. Two-stage RNN-LSTM model 49 Figure 4.10 shows the cost for a single sample, related to equation 4.8. Observing |ˆy−y|2against |ˆy−y|we can see the effect of a big outlier by means of the derivative of the function, which is very significant, as it means a very high error and a correction towards it during minimization. FIGURE 4.10: Plot of |ˆy−y|2against |ˆy−y|and the derivative with respect to the error:2|ˆy−y|. Figure 4.11 illustrates an example of the effect that an outlier may produce over our model, where the highly distant sample deviates the line a lot. This is why our main interest is reducing this effect avoiding the long-tailed behavior. FIGURE 4.11: Example of a regression estimation where we see how the outlier produces a high deviation over the correct estimation. (a) Gaussian distributed data with µ= 0 and σ= 4 over the line y= 1.2x. (b) Outlier artificially inserted to the data in (a). 50 Chapter 4. Two stage Text-to-Speech with RNN-LSTM Now that the duration model has been discussed the acoustic model is explained, which is the second stage of the full TTS system. This model is also built with Recurrent Neural Network architectures, as mentioned previously, yet the output layer is also recurrent in this case as aforementioned. The normalization of the features for this stage is a min-max for all the outputs, which are: •Mel Cepstral Coefficients. •Voiced Frequency. •log-F0 contour. •Voiced/Unvoiced flag. We have seen that the F0 contour is one of the acoustic parameters to be predicted, but this parameter has a special behavior depending on whether the current frame contains voiced or unvoiced sounds. For voiced frames it behaves as a continuous signal, but for unvoiced frames its value is zero (thus indicating that there is no periodic behavior in the frame). There is an example of pitch contour in Figure 4.12 to observe this mixture of discrete and continuous values. However, the network predicts the log of the F0 at each frame, because it is the type of data with which the Ahocoder deals. It is then important to get rid of the discrete symbol somehow to be able to normalize the pitch with the min-max range, because we will have a very frequent outlier with a very large negative value (Ahocoder encodes the unvoiced symbol with a value of −1000000 to indicate it goes to −∞). Otherwise, if we had the unvoiced symbol in the training set, the min-max operation would compress all the continuous values to a tiny range, which wouldn’t let the network learn the real distribution of the pitch. The way to solve this is by means of an interpolation of the original pitch (also shown in Figure 4.12), where the following operation has been performed during the unvoiced frames: •If the unvoiced frame is at the beginning of the contour where no previous voiced value appeared, put the first next voiced F0 value as a constant value •If the unvoiced frame is at the end of the contour where there are no more voiced values put the last previous voiced F0 value as a constant value (inverse of the first operation) •In all intermediate unvoiced regions, interpolate linearly in the logdomain between right previous and right next voiced values. Equation 4.9 shows the operation. log Fi 0= log Fp 0+ (log Fn 0−log Fp 0)·i−p n−p(4.9) Where nis the next voiced value’s frame index, Fn 0is the next voiced value, pis the previous voiced value’s frame index, Fp 0is the previous voiced value and Fi 0is the i-th pitch value we want to get at frame index i. 4.3. Two-stage RNN-LSTM model 51 FIGURE 4.12: Top: F0 contour example for male speaker M1. Bottom: F0 contour (blue line) example for male speaker M1 with interpolation contour (green line). On the other hand, the acoustic model must determine when is it that there is an unvoiced frame in order to mask out the interpolated values, which are false information given for the sake of good learning. There is then the voiced/unvoiced flag for this task, and it is a learned output as well. Finally, all other parameters are normalized in a min-max fashion with the reduced range {0.01,0.99}, but the Voiced Frequency parameter (discussed in section 4.2.2) is also log-compressed to avoid a long tailed distribution. The original distribution of values can be appreciated in the histogram in Figure 4.13. FIGURE 4.13: Histogram of voiced frequency values in the training data. In order to predict all the acoustic parameters we use the same linguistic inputs as before to take into account not only the phoneme, but also the context that surrounds it. In addition, we must take care with the duration 58 Chapter 4. Two stage Text-to-Speech with RNN-LSTM 4.5.2 Objective Evaluation To see the effect of different architectures we made an objective evaluation by means of specific metrics for each kind of predicted feature. In the case of duration, we chose the Root Mean Squared Error (RMSE) in millisecond scale, as it is a well-accepted error measure in the literature for speech synthesis tasks, defined as: RMSE [ms] = v u u t N−1 X t=0 (durt−ˆ durt)2(4.13) where Nis the number of test phonemes used for the evaluation. It is important to mention that silence phonemes were predicted but not included in the evaluation computations because they have a large variance and they would distort the prediction of the regular phonemes duration. Table 4.4 shows the result for different architectures. All those models had Dropout between the Output Layer and the previous Hidden Layer with p= 0.5 (probability of being activated). Also, when the output layer is recurrent (LSTM) the activation is sigmoid (the outputs are normalized to work in the linear region) to gain stability whilst training. Moreover, the batch size was 64, the maximum sequence length to propagate the gradient through time was 10 time-steps, the optimizer was Adam (see Chapter 3), with learning rate lr = 0.001 and, during training, a validation set was used to do earlystopping, having a maximum of 100 epochs per model. The early-stopping mechanism had a patience of 10 epochs per model, which means that if the validation loss does not improve at a certain point during training along 10 epochs, the training is aborted and the best validated model is the one stored, and it is also the one with which results are computed. TABLE 4.4: Comparison of different architectures for the duration model with their objective result. FC: Fully-Connected layer. The best performing model is in bold text. Embeddings Hidden LSTM Output Type Num. params. RMSE [ms] 0 1 ×256 FC 634K18.51 0 1 ×64 FC 110K18.71 0 1 ×256 LSTM 635K18.73 1×256 1 ×256 FC 618K18.77 2×256 1 ×512 FC 1.7M18.95 2×128 1 ×256 FC 457K19.01 0 1 ×64 LSTM 110K19.06 0 1 ×1024 FC 5.7M19.34 0 2 ×256 FC 1.1M20.04 0 2 ×256 LSTM 1.2M22.77 We can see how increasing the number of parameters tends to give worse results when we go over 256 for our search, probably owing to over-fitting effects as the amount of data is not very high. Also, going under this amount of units does not provide better result either, thus losing some representation capabilities for the provided data, though the difference is not 4.5. Experimental Design and Results 59 high. Even though, we picked the best performing model as we considered it has a reduced set of parameters considering the state of the art models. Some embedding layers were also tested with no success for this model. The same happened when inserting a recurrent output. It is worth mentioning that this is not the most exhaustive search that could be done, but it is considered to be good enough to get a first approach for the purpose of this work. Regarding the acoustic model other tests were performed for the same purpose, and the results can be seen in Table 4.5. In this case we have the following metrics for each kind of predicted feature: •Mel Cepstral Distortion (MCD) for the MFCC predictions in Decibel [dB] scale. •RMSE of the F0 prediction in Hertz [Hz] scale. •Accuracy metric for the Voiced/Unvoiced flag prediction in %. Some authors claim that the MCD (Mashimo et al., 2001) is correlated to subjective evaluations (Kubichek, 1993) and it is defined as: MCD = (10√2)/(Tln 10) T−1 X t=0 v u u t 39 X n=0 (ct,n −ˆct,n)2(4.14) where T is the number of test frames, ct,n are the real cepstral coefficients and ˆct,n are the predicted cepstral coefficients. The MCD is computed without applying post-filtering. The RMSE of the F0 is computed as the duration one: RMSE [Hz] = v u u t T−1 X t=0 (f0t−ˆ f0t)2(4.15) Finally, we consider the Accuracy of the Voiced/Unvoiced flag prediction, which is defined as: Acc[%] = TP +TN TP +FP +TN +FN ×100 (4.16) Where TP stands for True Positives, TN are the True Negatives, FN are the False Negatives and FP are the False Positives. The silence frames were also dropped for the acoustic evaluation, and the training conditions for all the acoustic models were the same as in the duration models except that the batch size was 256 and the maximum sequence length to back propagate the gradients was 30. Here we can see how including some embedding layers decreases the error for some metrics, possibly due to the mix of data types that we have at the inputs (duration and contextual features). Also, we can see how increasing the number of cells and layers gives the lowest results, as we have way more examples than the ones for the duration model (see section 4.5.1). Nonetheless, it is clearly 60 Chapter 4. Two stage Text-to-Speech with RNN-LSTM seen that there is no much variation in the acoustic prediction results for the tested parameters. We ended up picking the one with lowest error in all measurements again, like in the duration model search, and this fits perfectly the purpose of this work as the subjective results will suggest. TABLE 4.5: Comparison of different architectures for the acoustic model with their objective results. LSTM: Hidden LSTM layer and units. Params: Number of parameters of the network. Embeddings: Number of input Fully Connected layers for first projections of the data. F0 RMSE: Root Mean Square Error of the F0. MCD: Mel Cepstral Distortion. UV Acc: Accuracy of Voiced/Unvoiced flag prediction. The best performing model is in bold text. Embeddings LSTM Params. F0 RMSE[Hz] MCD [dB] UV Acc[%] 2×256 2 ×512 3.9M14.90 6.09 95.00 2×256 1 ×512 1.8M14.90 6.17 94.86 2×256 2 ×256 1.2M15.00 6.09 94.86 0 2 ×256 1.2M15.03 6.22 94.60 2×256 1 ×256 736K15.22 6.18 94.96 3×512 1 ×256 1.55M15.99 6.29 94.70 4.5.3 Subjective Evaluation Objective tests performed well to compare different architectures and to get some clue of how well does the TTS work. To get a more in-depth evaluation of the model however, a subjective test is conducted to evaluate the naturalness of the TTS developed in this work and also to make a comparison with the other currently used TTS techniques. The platform and synthesized examples can be found in UPC TTS Benchmark (2016). In the test 17 listeners were given 5sentences to each one, randomly selected from a set of 15 short sentences. For every sentence, the listeners evaluate 6differ- ent versions generated with the following 6different systems, all built from data of the same speaker: •Natural voice: original speaker recording. •Ahocoded version: the original recording was parameterized with the Ahocoder and constructed back into the waveform to evaluate the amount of naturalness that the developed TTS loses already in the waveform generation part, which is extrinsic to the neural network. •LSTM raw: Raw prediction from the two stage LSTM TTS without applying any post-filtering. •LSTM pf: Prediction from the two stage LSTM TTS with a post-filtering factor of pf = 1.04. •US: Unit Selection system of UPC (Bonafonte et al., 2006a). •SPSS: Statistical Parametric Speech Synthesis generated with the HTS framework (Zen et al., 2007). 4.5. Experimental Design and Results 61 The users are then asked to evaluate the hidden natural voice with a 100 and to set the other ones inside the 0−100 scale of naturalness (thus they have to set values relative to that of the natural voice) (ITU-R, 2003). The participants can listen the different recordings as many times as required to make comparisons between the different systems. All the TTS systems predict the duration, pitch and the acoustic parameters. As opposed to the acoustic objective evaluations, here there is no forced alignment to evaluate the full TTS capability. In order to proceed with the comparisons a normalization technique was applied to all the recordings based on ITU-T P.56 Recommendation (ITU-T, 2011). In this way all files are equalized to have the same power level and we make sure that there is no difference because of a masking effect over noise artifacts (for instance caused by the multiplying factor of the post-filtering curve, that would increase the energy of the signal otherwise). Figure 4.21 and Table 4.6 show the results of the subjective test. Figure 4.21 is a box-plot, where each system’s data distribution is drawn in quartiles. The boxes extension is the distance between the first and the third quartiles, and the red line depicts the median of the data. The vertical dashed lines extending up and down from the boxes are called the whiskers. In the edge case that the first and third quartiles are equivalent, the whiskers extension will be set to the min-max range of values in the distribution (for instance in the Natural case). Otherwise, the whiskers show the range 1.5times the first or third quartile, as in the Ahocoded case. The points outside of the whiskers range are outliers found in the distribution of the data resulting from evaluation. In the box-plot we can see that the results are good for the TTS developed in this work, as it is the one right after the Ahocoded voice in terms of naturalness. We can note there that the post-filtering technique indeed improves the perception of naturalness, thus confirming the effectiveness of this method. It is important to see that the Ahocoder already introduces some artifacts that diminish the naturalness of speech, getting a mean evaluated value of 82.65. The Ahocoder score is the best that our model could achieve, so we set it as a reference for main comparisons. The two stage TTS system with post-filtering achieves a mean value of 60.80, so 21.85 points is the mean loss in naturalness introduced by the neural network with respect to the Ahocoder. In addition to this, we can also see that the mean value of the TTS without post-filtering is 47.13, which means that the mean improvement achieved by adding the post-filtering mechanism is 13.70. The case of Natural speech is the one with highest naturalness, as expected. Actually, the distribution is so condensed near 100 that the quartiles coincide. In the case of US system we can see how the variance is quite higher than in other systems, and this is mainly because the US performs quite well for very studied contexts in the dataset of speech samples, but when there is a rare context appearing in the test text to be synthesized the system performs a lot worse than normally (Hunt and Black, 1996), and thus the resulting variance in the evaluation is higher because of those noticeable failures. The SPSS used for this comparison performed clearly worse, but it is important to mention that no post-filtering methodologies were used in this 62 Chapter 4. Two stage Text-to-Speech with RNN-LSTM case to combat the over-smoothing effect that also appears in HMM-based synthesis. For both US and SPSS systems the already available systems were used, but perhaps they could be optimized to get a closer result to that of our optimized TTS. FIGURE 4.21: Box-plot of the subjective naturalness test results relative to real human voice. SPSS: Statistical Parametric Speech Synthesis. US: Unit Selection. LSTM-raw: Two-stage TTS without post-filtering. LSTM-pf: Two-stage TTS postfiltered with pf = 1.04. Ahocoded: Natural speech parameterized with the Ahocoder and reconstructed. Natural: real human voice. Red line: median. Red dot: mean. TABLE 4.6: Statistics of the subjective results for the 6systems. System µ σ Natural 97.69 6.61 Ahocoded 82.65 19.85 LSTM raw 47.13 24.63 LSTM pf 1.04 60.80 22.20 US 49.20 30.80 SPSS 32.54 21.21 4.5. Experimental Design and Results 63 4.5.4 Gate activations analysis Once the acoustic model is trained we can look at its inner structure to check what is activated through time depending on the input sequence of phonemes. Here an analysis of the input, output and forget gate activations is shown graphically for an arbitrary test file (see section 3.6 in Chapter 3 for a description of the gating mechanisms). The chosen sentence is "Llamó desde la recepción con su voz lúgubre.", where the first LSTM hidden layer and the output layer gates have been analyzed. Figures 4.22,4.23 and 4.24 depict the heatmap of activations of 20 LSTM cells in the hidden layer. These plots can look confusing as we see a lot of activities happening through the different cells (plus only 20 cells are shown from the total amount of 512, so we don’t have the full picture for better visualization purposes). In order to get a closer look at what’s going on, Figure 4.25 illustrates the averaged activities for the three types of gates over all the hidden units. FIGURE 4.22: Activations map for the forget gates of the first 20 hidden cells in the first LSTM hidden layer. Red regions are high activations (thus preserve the past) and yellow means forget the past. Something quite interesting can be observed in Figure 4.25: the forget gate removes more content as the input phoneme is still repeating, though it is done in a patient manner (close to a linear behavior). However, when a new phoneme gets into the network there is a peak in the gate, thus it preserves as much as possible the information about its past. On the other hand, the input gate follows what looks like the inverse behavior, meaning that it lets in as much new information as possible (maybe to try to get new clues when the changes in input features are not quite relevant). This "inverse" behavior with respect to the forget gate was already mentioned in the work by Wu and King (2016). There the authors proposed an interesting efficient recurrent architecture based on this fact, where they only keep a forget gate as a memory control mechanism and get rid of the output gate, 64 Chapter 4. Two stage Text-to-Speech with RNN-LSTM FIGURE 4.23: Activations map for the input gates of the first 20 hidden cells in the first LSTM hidden layer. Red regions mean updating the cell state a lot with the new candidate. On the other hand, yellow means not letting the information in. FIGURE 4.24: Activations map for the output gates of the first 20 hidden cells in the first LSTM hidden layer. Red regions mean letting the cell state flow out of the cell. 4.5. Experimental Design and Results 65 but preserving a special type of input gate which is: it= 1 −ft, whilst the forget gate ftis still parameterized with the set of learnable weights. FIGURE 4.25: Average activations for the first LSTM hidden layer with 512 cells. Input gate, forget gate and output gate are shown. Green dashed lines are the phoneme boundaries. In Figure 4.26 it is shown the comparison between the learned input gate activations itand the type of input gate proposed in Wu and King (2016) 1−ft. There we can indeed see that the curves are correlated. FIGURE 4.26: Averaged activations of the input gates for the first hidden LSTM layer in blue line and 1−ftaverage activation in red. The same analysis can be done in the output recurrent layer, where a different behavior is observed. The heatmaps for the 43 output LSTM cells gate activations are shown in Figures 4.27 4.28 and 4.29. There it is depicted the behavior of the cepstral predictions (the first 40 rows beginning at the bottom of the figure), the voiced frequency, the pitch and the top row is the voiced/unvoiced flag. In the three heatmaps it can be appreciated how the voiced/unvoiced prediction get extreme values (either close to one or to zero) as expected, and an important thing to note is that those values match their value to the type of phoneme pretty accurately, so for instance 66 Chapter 4. Two stage Text-to-Speech with RNN-LSTM phonemes like "pau", "T", "k" or "s" 1(Wells et al., 1997) that are unvoiced, get a low level in the output gate which is blocking the cell output to not make a prediction there. Also, a closely related behavior happens in the voiced frequency value as expected, which has a wider range of intermediate values but is correlated to the type of phoneme injected. FIGURE 4.27: Activations map for the forget gates of the 43 output LSTM cells. Red regions mean letting the cell state flow out of the cell. FIGURE 4.28: Activations map for the input gates of the 43 output LSTM cells. Red regions mean letting the cell state flow out of the cell. 1SAMPA phonetic notation is used. 4.5. Experimental Design and Results 67 FIGURE 4.29: Activations map for the output gates of the 43 output LSTM cells. Red regions mean letting the cell state flow out of the cell. Figure 4.30 depicts the averaged gate activations in a similar analysis to that of Figure 4.25 for the hidden layer. Here it is appreciated how the output activation is usually high due to the functionality that this layer is performing, but there appear some valleys in the unvoiced regions which are probably due to the loss of energy, so some output has to be blocked, as well as the U/V and FV values commented previously. The behavior of the input and forget gates seem to be more erratic than the linear decaying found in the hidden layer, but both gates seem to be again closely related as previously seen in Figure 4.26. FIGURE 4.30: Average activations for the output LSTM layer with 43 cells. Input gate, forget gate and output gate are shown. Green dashed lines are the phoneme boundaries. 74 Chapter 5. Multiple Output Acoustic Mapping The α-layer is trained by freezing the multi-output model weights. As mentioned previously, the layer has Mspeaker branches as input, thus raising M×O+Minput units, where Ois the acoustic vector dimension. The M added inputs come from another input vector which is inserted to control the weight that every speaker has in the interpolation, called the αvector (which gives its name to the layer), which is a one-hot code of dimension M. During training, linguistic inputs are injected into the multi-output model and then the inference is made, which turns to be the inputs for the interpolation layer, which also get the one-hot αconcatenated, expressing the identity of the current shown speaker at the interpolation layer output as mentioned before. The same training data used to train the multi-output model is used for the Mspeakers. Figure 5.5 depicts the training procedure for an example with 3speakers and for 3batches. This methodology then expects the layer to learn not only each extreme case (i.e. each one-hot case shown during training), but it is also expected to infer intermediate values for the acoustic outputs, and as it is seen in section 5.6.3, it actually learns to interpolate the features. During synthesis, the αcode is replaced by a probability distribution among the speakers, thus expressing the percentage to synthesize for every speaker, so for instance we could have the following α-vector in synthesis: α= (0.5,0.5,0,0) This is a code for an interpolation layer made out of M= 4 speakers, and it means that the mixture is made with 50% of speaker 1and 50% of speaker 2, while the other two remain 0, so no information from them is required to generate the acoustic predictions. This type of code is the first approach for making an interpolation architecture that will be improved in a future line of research with identity related to acoustic characteristics of speakers, such as the i-vectors. 5.6 Results The results for the different architectures proposed are explained here, so that we see the effect of: •How is the overall quality of the multi-output model compared to the individual models? Are the speakers distorting each other or are they helping each other? •How does the speaker adaptation perform when we back-propagate different amounts of the available data? How does affect the backpropagation through the whole model in comparison to freezing the base model and updating only the new output? How does it do comparing to the single speaker model of the new speaker? •Regarding the interpolation, is the α-layer learning a suitable and meaningful intermediate range of values for the different speakers? How is the M≥µaffecting the resulting prediction of µspeakers? 5.6. Results 75 FIGURE 5.5: α-interpolation training method for an example with M= 3 for 3batches of examples. The one-hot code expresses the identity of the currently shown speaker. Each smis the output prediction of the corresponding multioutput branch for the m-th speaker. 76 Chapter 5. Multiple Output Acoustic Mapping The following sections analyze these questions given the objective and subjective results obtained during the course of this work. 5.6.1 Results: Multi-output We make a first analysis by looking at the training loss evolution of the different speaker outputs, and concretely focusing on two speakers: M1 and F1. To establish a reference, we trained M1 and F1 with a single output architecture and multiple output one. The results can be seen in Figure 5.6, where the 7speaker learning curves are shown, depicting that all output converge with a noisy behavior, given by the training methodology where every speaker distorts each other’s learning process for mini-batches of data. FIGURE 5.6: Training loss evolution comparison. Speakers F1 and M1 decrease the learning cost when trained with other speakers altogether. There we can see how speakers F1 and M1 get to a lower training loss when they are trained with the multiple output mechanism. This is normally related to a better training procedure where they reach a better point in the optimization. To really see this effect, we first make an objective evaluation by means of specific metrics for each kind of predicted feature, as it was previously done in Chapter 4. Concretely we used the Mel Cepstral Distortion (Mashimo et al., 2001) (MCD), the RMSE of the predicted F0 (Hertz scale) and the Accuracy in UV flag prediction. Table 5.1 shows the results for F1 and M2, which show the alone models (i.e. the speaker trained with its own acoustic model isolated from others) and the mixed models (i.e. the speaker trained in a multi-output fashion with the other 6speakers). The alone results for M1 differ a little bit from the ones obtained in Chapter 4because we have different data splits, which at same time come from a much less amount of samples. The objective results suggest the improvement in the features estimation when the speaker models are trained in the multi-output fashion, as they 5.6. Results 77 share knowledge in the lowest layers of the network to transfer learning about the final mappings. TABLE 5.1: Objective evaluation for M1 and F1 trained alone with a single output model and together with other speakers (mixed) in the multiple output architecture. Model MCD[dB] F0[Hz] UV[%] M1 alone 7.6 14.4 92.3 M1 mixed 7.2 13.8 94.2 F1 alone 7.0 17.3 95.2 F1 mixed 6.5 17.3 96.2 A subjective evaluation has been carried out as well with a preference test made by 16 subjects. For both F1 and M1 speakers, 5sentences are selected and evaluated. The listeners can choose a declining score between two synthesized utterances; one generated by the single output model and another one by the multiple output one. Listeners then find five options available from −2(multiple output is much preferred) to 2(single output is much preferred). The type of system per utterance is hidden for the listeners and randomly ordered. The results are depicted in Figure 5.7. It can be seen that the testing subjects have all rather preferred the multiple output model in most of the cases. We also made a Wilcoxon test for the subjective evaluation to find out how statistically meaningful are these results, obtaining the following p-values: pF1= 5.3·10−7and pM1= 2.2·10−5. FIGURE 5.7: Box plot of preference test scores. Scores range from −2(multiple output model is preferred) to 2(single output trained model is preferred). Both is the summary of all the answers, joining both speaker results. Red lines: medians. Blue dots: means. 78 Chapter 5. Multiple Output Acoustic Mapping 5.6.2 Results: Adaptation In this section we analyze the effect of an adaptation layer constructed on top of the MO architecture. In Figure 5.8 we can see how the validation cost for the F3 speaker improves when we add more data, something we could expect, as it learns better with the more data it gets to fine-tune the new output branch. It is interesting to see how freezing the shared layers and training only the new output branch we get to a very similar result, and more smoothly. Table 5.2 summarizes the objective evaluation for this fine-tuning, getting a good result with respect to the single output model of F3 when only the last layer is trained on top of the shared parts of the model. The error values are quite high in comparison with the previous ones (F1, M1), because this speaker was taken from an expressive subset of data, being it quite different from that of F1 and M1. An informal listening test suggested that the adaptation sounded close to the original speaker, thus validating this approach. TABLE 5.2: Objective evaluation for F3 as an adaptation subject. Full: all layers are fine-tuned. Frozen: only new output branch is fine-tuned. Model MCD[dB] RMSE F0[Hz] UV[%] F3 alone 8.11 28.07 9.00 F3 fine-tuned full 100% data 7.96 26.96 7.74 F3 fine-tuned frozen 100% data 7.90 26.44 6.73 FIGURE 5.8: Validation loss evolution comparison of different batch sizes, with frozen shared layers and fine-tuned shared layers. 5.6. Results 79 5.6.3 Results: α-interpolation Finally, we designed the experiments to evaluate the α-interpolation proposal with 2speakers out of the 6mentioned previously from the TCSTAR database. For the interpolation, two configurations were trained as will be seen, M= (2,6). This means that the interpolation was carried out between speakers F1 and M1, but there is also a configuration where the α-layer is trained will all the MO speakers, thus M= 6. An informal subjective test clearly showed that increasing Mimproved the naturalness of the output speech although only 2speakers are interpolated in the evaluation. Objective tests have been performed to evaluate the performance of the α interpolation. These consist in analyzing the evolution of the MCD between the interpolation output and each of the Mbranches, and also the evolution of the F0 RMSE. Figures 5.9 and 5.10 show these results, where the αvariation is made for speaker F1, so it is αF1, speaker M1 has αM1= (1 −αF1) and all others are αm= 0. We may refer to α= 0.5for the point at αF1=αM1= 0.5. FIGURE 5.9: MCD when varying αvalues. The variation is made for speaker F1 and it is (1 −α)for M1. All others speakers remain 0.M= 6. From the curves we see how, although we only show to the network the extreme values with an orthogonal code, it learns the intermediate representations effectively. The MCD values vary smoothly between the interpolated speakers F1 and M1, whilst other speakers’ MCD remain with a short variation. It is interesting the fact that the crossing point is very close to α= 0.5. 80 Chapter 5. Multiple Output Acoustic Mapping Note that the values may differ from those in Table 5.1 because the distances are not computed to natural speech but to the multi-output predictions in this case. FIGURE 5.10: F0 RMSE when varying αvalues. The variation is made for speaker F1 and it is (1 −α)for M1. All others speakers remain 0.M= 6. Regarding the F0 RMSE evolution, there is a biasing of the crossing point, which shows us how the F0 prediction is biased towards the male speaker, as it is the one getting less error for α= 0.5. These interpolation results are coherent with perceptual impression. As previously mentioned, increasing Mhelped in the naturalness of the 2-speaker interpolation, so an analysis of the F0 distributions is also made for the cases M= 2 and M= 6. These analysis are shown in Figures 5.11 and 5.12 respectively. First, we can confirm the biasing towards the male speaker when α= 0.5 (M1 50% F1 50%) in both cases. Nevertheless an important difference is the fact that training the layer with a higher Mincreases the distributions variance, which turns out to be a less monotonous sound at the output, and thus more natural. 5.7 Discussion A novel architecture for modeling the acoustic mapping for many speakers at once has been proposed and studied in section 5.3. The model lets us represent Nspeakers with a very reduced set of parameters in comparison to making Nseparated models, once per speaker. Moreover, we have seen how the different speakers help each other in the learning process by transferring their knowledge in the linguistic mappings done in the earlier layers. Objective and subjective evaluations confirm the learning advantage 5.7. Discussion 81 FIGURE 5.11: F0 Histograms: original M1 and F1 speakers in blue. αF1= (0.25,0.5,0.75) and αM1= (0.75,0.5,0.25) interpolations in green. M= 2. FIGURE 5.12: F0 Histograms: original M1 and F1 speakers in blue. αF1= (0.25,0.5,0.75) and αM1= (0.75,0.5,0.25) interpolations in green. M= 6. 82 Chapter 5. Multiple Output Acoustic Mapping that the different speakers obtain by being trained jointly with the multioutput architecture. On top of this, a speaker adaptation technique is presented in section 5.4, where a new output branch is attached to the multi-output architecture and two back-propagataion approaches have been studied. The objective evaluation showed the improvement of this technique over training the speaker isolated as well. Finally, an interpolation model has been built also taking advantage of this novel architecture by means of the so called α-layer, which is trained with an orthogonal code expressing speaker identities during training and speaker portions during synthesis. The results suggest that the layer can effectively learn intermediate ranges of speaker representations by only being trained with extreme cases (each speaker’s voice examples). Furthermore, inserting as many speakers as possible to the α-layer increases the variability of the predictions in the output, which turns out to be a more natural result. 83 Chapter 6 Conclusions This chapter is devoted to make a review of this work, discussing the implemented architectures and achieved results. Lines of future work are also explained such that the reader can get to see the possibilities opened by this contribution. In the end of the chapter the research contribution of the work is shown. 6.1 Thesis Review A speech synthesis system made from scratch with RNN-LSTM architectures has been built in this work. The system is based on a two stage architecture, where first the duration of the phonemes to generate are predicted, and then the acoustic parameters to be passed to a Vocoder are generated frame by frame up to the corresponding duration. The types of features with which the network makes the predictions are explained in detail: the context labels that contain phonetic and prosodic information about the input text and the acoustic features predicted. To generate the data used by the neural networks a parallelized framework is made, such that the Vocoding and text-to-label processes are sped up to be able to deal with many speakers’ data quickly and with large amounts of data. It has also been shown how the acoustic model suffers an over-smoothing effect of the generated parameters, because it tends to predict the means and not the variances for the inherent behavior of the regression training function, the MSE. A post-filtering mechanism is then applied in the generation of acoustic parameters to overcome this issue and enhance the naturalness of the synthesis. After presenting all the methodologies behind the Text-To-Speech (TTS) system there are different types of evaluations made for the duration and acoustic models: some brief architecture search to tune the different parameters of the network in an objective way is performed, with which we obtained a deeper architecture in the acoustic model than in the duration one. The duration model architecture seeking process shows that it suffers from variability in the results when we vary the network depth (amount of hidden layers) and width (amount of hidden units/cells), so the variation of topology in this case is meaningful to achieve better results. On the other hand the acoustic model is not perturbed very much and the results are just slightly different when tuning the model for the chosen parameters. Nonetheless. the best performing model in objective terms was picked as the representative one in both cases. 90 BIBLIOGRAPHY Tutorial on Restricted Boltzman Machines (RBM) (2010). http://deeplearning. net/tutorial/rbm.html. [Online; accessed May-2015]. Understanding LSTM Networks (2015). http : / / colah . github . io / posts/2015-08-Understanding-LSTMs/. [Online; accessed February- 2016]. UPC TTS Benchmark (2016). http://veu.talp.cat/neural_eval/. [Online; accessed June-2016]. Uria, Benigno, Iain Murray, and Hugo Larochelle (2013). “RNADE: The real-valued neural autoregressive density-estimator”. In: Advances in Neural Information Processing Systems, pp. 2175–2183. Uria, Benigno et al. (2015). “Modelling acoustic feature dependencies with artificial neural networks: Trajectory-RNADE”. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 4465–4469. Valentini-Botinhao, Cassia, Zhizheng Wu, and Simon King (2015). “Towards minimum perceptual error training for DNN-based speech synthesis”. In: Proc. Interspeech. Wells, John C et al. (1997). “SAMPA computer readable phonetic alphabet”. In: Handbook of standards and resources for spoken language systems 4. Wu, Zhizheng and Simon King (2015). “Minimum trajectory error training for deep neural networks, combined with stacked bottleneck features”. In: Proc. Interspeech, pp. 309–313. — (2016). “Investigating gated recurrent networks for speech synthesis”. In: 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 5140–5144. Wu, Zhizheng et al. (2015a). “A study of speaker adaptation for DNN-based speech synthesis”. In: Proceedings interspeech. Wu, Zhizheng et al. (2015b). “Deep neural networks employing multi-task learning and stacked bottleneck features for speech synthesis”. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 4460–4464. Ze, Heiga, Andrew Senior, and Mike Schuster (2013). “Statistical parametric speech synthesis using deep neural networks”. In: 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, pp. 7962– 7966. Zen, Heiga and Hasim Sak (2015). “Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis”. In: Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. IEEE, pp. 4470–4474. Zen, Heiga and Andrew Senior (2014). “Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis”. In: 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 3844–3848. Zen, Heiga, Keiichi Tokuda, and Alan W Black (2009). “Statistical parametric speech synthesis”. In: Speech Communication 51.11, pp. 1039–1064. Zen, Heiga et al. (2007). “The HMM-based speech synthesis system (HTS) version 2.0.” In: 6th ISCA Workshop on Speech Synthesis. ISCA, pp. 294– 299.