scieee AI-readable full text Open interactive document viewer

A Spanish multispeaker database of esophageal speech

Serrano García, Luis,Raman, Sneha,Hernáez Rioja, Inmaculada,Navas Cordón, Eva,Sánchez de la Fuente, Jon,Saratxaga Couceiro, Ibon

Abstract

This work was partially funded by the Spanish Ministry of Economy and Competitiveness with FEDER support (RESTORE project, TEC2015-67163-C2-1-R), the Basque Government (PIBA-018-0035) and by the European Union’s H2020 research and innovation program under the Marie Curie European Training Network ENRICH (675324)

Full text

A Spanish Multispeaker Database of Esophageal Speech Luis Serrano Garc´ıa, Sneha Raman, Inma Hern´aez Rioja, Eva Navas Cord´on, Jon Sanchez, Ibon Saratxaga HiTZ Basque Center for Language Technology, University of the Basque Country (UPV/EHU), Bilbao, Spain Abstract A laryngectomee is a person whose larynx has been removed by surgery, usually due to laryngeal cancer. After surgery, most laryngectomees are able to speak again, using techniques that are learned with the help of a speech therapist. This is termed as alaryngeal speech, and esophageal speech (ES) is one of the several alaryngeal speech production modes. A considerable amount of research has been dedicated to the study of alaryngeal speech, with a wide range of aims such as helping speech therapists with evaluation and diagnosis, and improving its quality and intelligibility using digital signal processing techniques. We present to you a database of Spanish ES voices, named AhoSLABI, which is designed to allow the development of new support technologies for this speech impairment. The database primarily consists of recordings of 31 laryngectomees (27 males and 4 females) pronouncing phonetically balanced sentences. Additionally, it includes parallel recordings of the sentences by 9 healthy speakers (6 males and 3 females) to facilitate speech processing tasks that require small parallel corpora, such as voice conversion or synthetic speech adaptation. Apart from the sentences, the database includes sustained vowels and a small set of isolated words, which can be valuable for research on ES analysis, diagnosis and evaluation. The paper describes the main contents of the database, the recording protocols and procedure, as well as the labeling process. The main acoustic 1Current Address: Communications Engineering Dept., Faculty of Engineering of Bilbao, University of the Basque Country (UPV/EHU) Spain Preprint submitted to Computer Speech and Language September 23, 2020 This is the acepted version of an article published by Elsevier . The final version of Luis Serrano García, Sneha Raman, Inma Hernáez Rioja, Eva Navas Cordón, Jon Sanchez, Ibon Saratxaga. A Spanish multispeaker database of esophageal speech. Computer Speech & Language 66 : (2021) // Article ID 101168 is available at https:// doi.org/10.1016/j.csl.2020.101168 ©2020 Elsevier Ltd. All rights reserved.This manuscript version is made available under the CC-BY-NC-ND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/ characteristics of the voices, such as speaking rate, durations of the recordings, phones and silences, and other such characteristics are compared with those of a reduced set of healthy voices. In addition, we describe an experiment using the database to improve the performance of an ASR system for ES speakers. This new resource will be made available to the scientific community with the hope that it will be used to improve the quality of life of the laryngectomees. Keywords: esophageal speech, voice conversion, speech databases, speech intelligibility, speech analysis 2010 MSC: 00-01, 99-00 1. Introduction Esophageal speech (ES) is a type of speech produced by laryngectomees, which are people whose larynx has been surgically removed. The larynx is a fundamental organ in the speech production mechanism. It contains the vocal folds, which are responsible for generating the air vibrations that are perceived5 as a sound. In addition to the removal of the larynx, the laryngectomy separates the nasal cavity and the vocal tract. As a result, the laryngectomees breathe through a hole (called the stoma) which lets the outside air directly into the trachea. Despite the removal of the vocal folds, it is still possible for people who have undergone a total laryngectomy to produce intelligible speech via one of10 the three main types of alaryngeal speech: using an electro-larynx (EL Speech), tracheoesophageal speech (TES) and ES. EL speech uses an external vibration device which is placed in contact with the throat. This device generates an acoustic buzz that can be modulated by the movement of the articulators. Intelligible speech is obtained with this method,15 but the quality is poor, mainly due to the dominant buzzing. The main and perhaps the only advantage of this method is that no learning is required. In Healthy Laryngeal Speech (HS), the air that flows through the lungs and the trachea vibrates the vocal folds to create sound. This is not anatomically possible for a laryngectomee. Therefore, airflow is produced using other strate-20 2 gies, the effectiveness of which depends on the characteristics and anatomy of each person. A surgical solution is to create a fistula allowing air to pass between the trachea and the esophagus. A valve is placed in the fistula so that no food or liquid can pass to the trachea. The airflow aided by this valve produces vibrations of the esophageal sphincter, which generates TES. TES is more intel-25 ligible and requires less effort from the speaker than other methods of obtaining the air [1][2][3][4][5][6]. However, the valves must be changed periodically (requiring surgery) and there are other possible medical complications associated with the implant [7][8]. Unlike TES speakers, ES speakers do not have the valve that allows the30 controlled entry of air. ES speakers achieve this function by swallowing air and expelling it, very much like the production of a burp. Like TES, the pharyngoesophageal segment is used as a substitute vibratory element instead of the vocal folds. Learning to produce speech in this manner requires long periods of training (usually months) with the assistance of a speech therapist. Due to35 the difficulties in the production method [9], some individuals never manage to learn ES. However, despite the long periods of learning, ES has the advantage of not requiring a device or periodic surgeries. Therefore, we consider there is a clear advantage in promoting the learning of ES. As ES and HS production mechanisms are very different, their speech signal40 characteristics differ greatly too [4][10]. The main consequence for ES is a dramatic reduction in naturalness and intelligibility [11][5][12]. A considerable amount of research effort has been devoted to overcome these limitations of ES, some of which involve artificially modifying its characteristics. There have been several approaches to enhance the quality and intelligibility45 of alaryngeal voices. Some studies use source-filter analysis of the pathological signal and focuses on modifying the source, the filter, or both. An example of this approach can be found in [13], where an adaptive gain equalizer algorithm was used to modify the ES source; or in [14], where the reconstruction of normal sounding speech for laryngectomy patients was attempted through50 a modified CELP (Code Excited Linear Prediction) codec. In [15], different 3 manipulations of both source and filter were evaluated. Another approach to improve intelligibility and quality is to work with the prosodic elements. In [16], the pitch information extracted from an electroglottograph (EGG) was used to create a synthetic glottal signal which reduced jitter and shimmer. Addition-55 ally, spectral smoothing and tilt correction were applied. These modifications reduced the harshness and breathiness of the TES speech. The same authors describe a method for rectifying the duration of pathological phones in [17]. Along the same lines, [18] presents a system where concatenation of randomly chosen healthy reference patterns replaces the pathological excitation, adjusting60 the short, medium and long-term variability of the pitch. A different approach to the problem is to use Voice Conversion (VC) techniques. VC aims to modify the characteristics of the voice of an input speaker, making them sound like those of a target speaker. In the classical approach [19], a conversion function is trained using data from both source and target65 speaker. Although non-parallel training is also possible [20], in VC, a set of parallel source-target sentences is desirable. A set of 50 phonemically balanced sentences in Japanese were used to evaluate the performance and capability of the different VC strategies that were aimed at improving the quality and intelligibility of alaryngeal voices [21][22][23][24][25][26][27].70 In this paper, we present an acoustic speech database specifically designed for the research of speech conversion techniques, applied to ES. The purpose of developing the AhoSLABI database2was to compile acoustic data which would allow us to investigate the use of VC techniques in improving the intelligibility and quality of ES. Some of our previous work in the field of personalized syn-75 thetic voices [28][29] has revealed that laryngectomees are a highly interested user group of the technological developments in speech synthesis, speech recognition and VC techniques. Although no official statistics have been published, in 2018, an estimated 1200 laryngectomies were performed in Spain [8]. We 2The name is a combination of the laboratory name of the authors (Aholab), and the name of the Biscayan Association of Laryngectomees ASLABI 4 aim to provide some useful tools for these laryngectomees, and for social and80 geographical reasons, we have developed the database in Spanish. To promote research and the comparison of techniques and results, we have provided an open access phonetically labeled database. First, we present an overview of the existing databases in the following section. Thereafter, we describe the contents of the database and the processes85 performed. Section “Design of the AhoSLABI database” describes the corpus contents and characteristics of the speakers, as well as the recording setup. The “Results” section presents some metrics of the database and provides some linguistic and acoustic statistics about its contents. In this section, we also report the process of extraction of phonetic labels and the evaluation of the automatic90 labeling procedure. In addition, we give some preliminary results of ASR and VC experiments performed with the database. The final section presents the conclusions and discusses possible future uses of the database. 2. Existing Related Material In some types of pathological speech such as dysarthric speech, certain95 databases have been extensively used and have become a de facto standard [30][31][32]. The same cannot be said for alaryngeal speech databases. For ES, many different recordings of varied characteristics have been performed, each adapted to the purpose of the study. In this section, we review the research publications in the field and give an overview of the existing recordings and100 their characteristics. Research on alaryngeal speech has traditionally focused on the production of sustained vowels. Vowels allow easy measurement of fundamental frequency, harmonic properties, and intensity and duration of phonation, which are basic features in assessing the speaker’s voice quality and speaking proficiency. Vowels105 based analysis were performed in a number of studies [33][4][34][35][36][37][38][39]. Some studies used recordings of words and sentences to measure the speaking rate [40][2][41], to study pauses [42] or both [43]. Recordings of words and sen5 tences have also been used in perceptual evaluations [44][45][46][47][48][49][50], and to evaluate synthetic manipulations [15].110 Automatic speech recognition (ASR) is also problematic for alaryngeal voices. Some ASR experiments use only vowels [51][52]. Typically, hundreds of sentences are used to train such ASR systems. In [53], a parallel database of 500 sentences pronounced by seven EL and seven HS German speakers were used to evaluate an ASR designed for HS speakers. In [54] 480 sentences produced115 by one French ES speaker were recorded with the purpose of improving the performance of an existing ASR system. The statistical VC experiments described in [21][22][55][23][24][25][26][27] use 50 parallel HS-ES sentences, but in Japanese. In order to facilitate the alignment procedure, the HS speaker tried to imitate the rhythm of the ES120 speakers’ utterances. Such a parallel HS-ES database is desirable for VC. In conclusion, to the best of our knowledge, no standard database exists to perform comparable research of Spanish ES, let alone to carry out VC experiments. We hope to fill this void with the database described in this paper. 3. Design of the AhoSLABI database125 3.1. Text content We selected the Spanish text corpus called ZureTTS described in [28] for the recordings. This corpus contains 100 phonetically balanced sentences encompassing all the phonemes in Castilian Spanish. The phoneme frequency distribution is shown in Table 1. The phoneme codes follow the Spanish SAMPA130 convention3. The total number of phones is 5625. This distribution is consistent with other previous Spanish corpora (see for example [56]). The sentences in the corpus are semantically relatively complex. As we already have HS recordings of this corpus, it made sense to record the ES database with the same corpus. 3https://www.phon.ucl.ac.uk/home/sampa/spanish.htm 6 This allowed us to have a parallel ES-HS corpus which is useful for tasks such135 as parallel VC. For a healthy speaker, the recording process usually takes between 30 and 40 minutes. For an ES speaker, the same task takes longer (see subsection Recorded Material and Durations) and for novice ES speakers, it can be quite exhausting. This is why the 100 sentences recorded were further divided in140 three blocks of 33, 33 and 34 sentences respectively. Each one of these blocks was phonetically balanced within itself. Therefore, if a speaker was tired and decided to not continue with the recording process after the first or the second block, the collected material would still be useful. Table 1: Percentage of phonemes in the AhoSLABI corpus. Phoneme Occurrences (%) Phoneme Occurrences (%) a 12.71 b 2.83 e 13.17 d 4.98 i 8.69 g 1.44 o 9.76 p 1.92 u 4.43 t 4.48 m 2.52 k 3.47 n 7.13 f 1.08 J 0.30 s 5.99 l 4.96 T 1.99 L 0.69 x 0.82 jj 0.41 tS 0.44 r 4.75 rr 1.03 In addition to the 100 sentences, each ES speaker recorded 4 instances of145 the sustained articulation of all five Spanish vowels. Four words containing diphthongs were also recorded (murci´elago,acu´ıfero,ayuntamiento,aceituno). Ten isolated words, which are also present in the ZureTTS corpus, were included in the recordings, to enable future evaluations of spoken term detection tasks 7 and the like.150 3.2. Characteristics of the Speakers All the ES speakers who participated in the recording process are members of the Association of Laryngectomees of Biscay (AhoSLABI). The speakers underwent speech therapy sessions after the laryngectomy to learn ES production techniques.155 Most candidates performed the recordings months after having finished the speech therapy sessions. We call these speakers ’proficient’ ES speakers. On the other hand, 4 of them were still attending the therapy sessions and their speech had very low intelligibility. We call these speakers ’non-proficient’ speakers. Out of the 4 non-proficient speakers, 2 returned after finishing the therapy and160 performed the recordings again. We have kept all these sessions in the database. The database contains recordings from 31 speakers (27 male and 4 female). It is composed of 34 different sessions as follows:  26 proficient ES speakers with one recording session each  2 non-proficient ES speakers with one recording session each165  2 ES speakers with one recording session each when they were non-proficient and one when they were proficient (in total 4 sessions)  1 speaker’s recordings in both TES and ES (in total 2 sessions) In summary, out of the 34 sessions, 29 correspond to proficient ES speakers, one to a proficient TES speaker and the remaining four to non-proficient ES170 speakers. The mean age of the speakers was 65 years and 4 months, but with large variation. The youngest was 51 years and 4 months old at the time of recording, and the oldest was 82 years and 5 months old. In order to identify each session, a four character code is used:175 8  The first two numbers identify the speaker (01 to 32)4  One character specifies the speaker’s gender M or F.  One character specifies the kind of speaker: ”3” for the proficient speakers and ”2” for the non-proficient speakers. For the TES speaker a ”T” has been used.180 The majority of sessions (25) feature proficient male speakers. Table 2 lists all the session identifiers. Table 2: Session identifiers. Session identifier Non-proficient, male 13M2, 14M2, 16M2 Non-proficient, female 15F2 Proficient, female 11F3, 15F3, 25F3, 28F3 Tracheoesophageal speaker, male 09MT Proficient, male All the others In addition to the ES speakers, recordings of the 100 sentences from 9 healthy speakers (6 males, 3 females, average age: 36 years and 3 months) are provided. These speakers were selected because of their availability and willingness to be185 part of the public database, and no criteria of age balance was considered. 3.3. Recording protocol The database recording protocol and procedures were approved by the ethics committee of the University of the Basque Country (UPV/EHU) (signed on 26th February 2017). The recordings were made in the soundproofed record-190 ing room at the Faculty of Engineering (UPV/EHU). Four different microphones (studio microphone - Neumann TLM 103, instrumentation microphone 4Recordings from speaker number 27 are not included in the database 9 syllables per second obtained for each session, ordered by mean. The results for the set of 9 HS speakers are also shown. As expected, HS showed a higher speaking rate than ES. It can also be seen that the TES speaker (session 09MT)310 achieved a speaking rate which does not differ from that of a healthy speaker which corroborates previous analysis on TES and ES [44]. Moreover, when the same speaker did not use the valve (session 09M3), his speaking rate slowed considerably. Another interesting result is that 3 out of the 4 non-proficient speakers had the slowest speaking rates. Two of these non-proficient speak-315 ers repeated the recordings three months later, after gaining more control and speech proficiency. While speaker 15F increased her speed, speaker 16M was speaking even slower. However, based on only these two speakers, we cannot generalise these observations. Figure 4: Speaking rate. Speaking rate calculated for 34 sessions of esophageal speakers (blue) and 9 of healthy speakers (green). In each box, the center line is the median, the edges of the box represent the 25th and 75th percentiles, the whiskers extend to the most extreme values not considered outliers, and the outliers are shown individually with a red cross. 16 4.4. ASR experiments320 Standard ASR systems normally use healthy speech as training material and therefore perform poorly for ES. In this subsection, we describe an experiment where we compare the results of two ASR systems, one trained with HS and the other one with ES from the AhoSLABI database. The starting point for both ASR is a standard Spanish ASR built using325 the Kaldi toolkit [58]. The specific implementation for Spanish is described in [59] and it is implemented following the recipe s5 for the Wall Street Journal database. The training begins with a flat-start initialization of contextindependent phonetic Hidden Markov Models (HMM), and then a series of accumulative trainings are done. For the final step of the recognizer, a neural330 network is trained. The input features to the neural network consist of a series of 40-dimensional features. The network sees a window of these features, with 4 frames on each side of the central frame. The features are derived by processing the conventional 13-dimensional Mel-Frequency Cepstral Coefficients (MFCCS) to which a process of mean and variance normalization (CMVN) is applied to335 mitigate the effects of the channel. The necessary steps are described in [60] and basically consist in applying a series of transformations to the normalized cepstra: first linear discriminant analysis (LDA), then maximum likelihood linear transform (MLLT) and global feature-space maximum likelihood linear regression (fMLLR). At the recognition stage, the same transformations are applied340 to the test data, handling them as a block. The main corpus used for the training of the acoustic models is the Spanish section of a subset of the Basque Parliament database. This subset contains the recordings of 47 parliamentary sessions of the Basque Parliament in both Basque and Spanish, together with their correspondent transcriptions 5. Some345 preliminary work has been done to separate the Spanish interventions from the Basque ones. As a result, there are more than 124 hours of speech in Spanish 5This database is presently being developed by the GTTS research group of the UPV/EHU, contact [email protected] 17 uttered by 84 different speakers, 45 male and 39 female. Additionally to the Basque Parliament database, about 4 hours of speech extracted from 5 audio files in Spanish extracted from the Spanish MAVIR workshops held in 2006,350 2007 and 2008 was also used to train the acoustic models (see [61] for more details). To avoid the effects of Out Of Vocabulary (OOV) words, the lexicon for both ASR systems has been reduced to the vocabulary of the 100 sentences of the database and unigram models are used. For the ASR trained with HS, the355 healthy speakers of the database had a mean WER score of 15.8±3.9, while the ES speakers had a mean WER score of 68.7±16.9. These results show how problematic generic ASR trained with HS can be for ES. To train the system with ES, we used all the ES speakers for which the complete set of 100 sentences was available. The speakers were divided into 3360 blocks of 10 speakers each. The sentences were divided into 10 blocks. A two level cross validation was performed, one at the speaker level and the other at the sentence level. In total 10 (sentence blocks) times 3 (speaker blocks) i.e., 30 cross-validations were performed to ensure that the test data was not seen in the training phase. In each of these cross-validations, 90 sentences from all the365 speakers of 2 blocks were used as training material and the 10 test sentences of the 3rd block of speakers were evaluated. When done 30 times, all the sentences from all the speakers were covered. The ASR scores for the 29 proficient speakers from both systems (ASR trained with HS and ASR trained with ES) are presented in figure 5. The370 non-proficient speaker (14M2) has been removed from the global results due to their poor performance (WER higher than 100%). The WER scores from the ASR trained with HS were significantly higher than the ASR trained with ES (t(28)=16.14, p<0.001). As can be observed, some speakers benefit more than others from the ES training. The mean improvement in WER is 23.2±7.7.375 This result demonstrates that generic Spanish ASR sytems can be made more ES inclusive by using the AhoSLABI database. 18 Figure 5: ASR Results. Mean speaker-wise Word Error Rates (in %) for ASR trained with HS and ASR trained with ES 5. Conclusions In this article we have described a database of Spanish ES called AhoSLABI. The database comprises mainly male ES, although it also contains the recordings380 of four female ES speakers and one male TES speaker. The main content of the database is the recordings of a set of 100 phonetically balanced sentences. The database also contains parallel recordings of 9 healthy speakers. We performed segmentation and labeling on the data. We have described the main aspects of the experimental setup, speaker characteristics and the acoustic properties of385 the recordings. The primary motivation for creating this database was the authors’ desire to have the laryngectomees benefit from recent advances in speech technologies, specifically in the field of VC. In particular, as reported in section 1, VC techniques have been proposed in the literature to improve the intelligibility of these390 voices. This was the main reason to record the sentences, as most VC systems need parallel source-target utterances to train the conversion function. Some of our VC work ([62] and [63]) demonstrates how ES can be made more intelligible 19 or more preferable to listeners using VC techniques. Although VC was our main intended application, there are many other areas395 of study where these recordings could be of interest. The sustained vowels recordings are helpful in the evaluation of fundamental frequency, shimmer, jitter, and intensity and duration of phonation. The signals can be used to train and test the performance of ASR systems with ES as shown in section 4.4 of this paper. Additionally, a small set of isolated words is also available which400 can be useful to test ASR systems in a spoken term detection task. Another research area is related to the loss of identity in the laryngectomees voices. One’s voice is a very important personality trait which is lost with laryngectomy. The recordings available could be useful in the emulation of prelaryngectomy speech characteristics. Investigating ways to restore this identity405 could be more feasible if pre-surgery recordings were available. In the future, the authors intend to also record voices of pre-laryngectomy patients. Subjective evaluation of the quality and intelligibility of alaryngeal speech to improve diagnosis and therapy is also possible with these recordings, because the number and variety of individuals is considerably high. A preliminary study410 of the intelligibility and listening effort for AhoSLABI was conducted in [64]. We believe that it is not only speech engineers but also researchers in speech therapy who can benefit from this database 6. 6. Acknowledgments This work was partially funded by the Spanish Ministry of Economy and415 Competitiveness with FEDER support (RESTORE project, TEC2015-67163C2-1-R), the Basque Government (PIBA-018-0035) and by the European Union’s H2020 research and innovation program under the Marie Curie European Training Network ENRICH (675324). 6The database is available for researchers through the European Language Resources Agency repository. 20 The authors want to thank the Asociaci´on Bizkaina de Laringectomizados for420 their valuable collaboration and all the laryngectomees for their voice donations. We also would like to thank the reviewers for their fruitful comments that have contributed greatly to the value of the paper. References [1] S. E. Williams, J. B. Watson, Speaking proficiency variations according to425 method of alaryngeal voicing, Laryngoscope 97 (1987) 737–739. [2] R. H. Pindzola, B. H. Cain, Acceptability ratings of tracheoesophageal speech, Laryngoscope 98 (1988) 394–397. [3] W. Ainsworth, S. W., Perceptual comparison of neoglottal, oesophageal and normal speech., Folia Phoniatr (Basel) 44 (6) (1992) 297–307.430 [4] F. Debruyne, P. Delaere, J. Wouters, P. Uwents, Acoustic analysis of tracheo-oesophageal versus oesophageal speech, The Journal of Laryngology & Otology 108 (4) (1994) 325–328. [5] T. Most, Y. Tobin, R. C. Mimran, Acoustic and perceptual characteristics of esophageal and tracheoesophageal speech production, Journal of com-435 munication disorders 33 (2) (2000) 165–181. [6] L. ˇ Siri´c, D. ˇ Soˇs, M. Rosso, S. Stevanovi´c, Objective assessment of tracheoesophageal and esophageal speech using acoustic analysis of voice, Collegium antropologicum 36 (2) (2013) 111–114. [7] B. M. Op de Coul, F. J. Hilgers, a. J. Balm, I. B. Tan, F. J. van den Hoogen,440 H. van Tinteren, A decade of postlaryngectomy vocal rehabilitation in 318 patients: a single Institution’s experience with consistent application of provox indwelling voice prostheses., Archives of otolaryngology–head & neck surgery 126 (11) (2000) 1320–8. doi:10.1001/archotol.126.11. 1320.445 URL http://www.ncbi.nlm.nih.gov/pubmed/11074828 21 [8] P. D´ıaz de Cerio Canduela, I. Ar´an Gonz´alez, R. Barber´a Durban, A. Sistiaga Su´arez, M. Tobed Secall, P. L. Parente Arias, Rehabilitation of the laryngectomised patient. Recommendations of the Spanish Society of Otolaryngology and Head and Neck Surgery, Acta Otorrinolaringol´ogica450 Espa˜nola (2018) 1–6doi:10.1016/j.otorri.2018.01.003. URL https://doi.org/10.1016/j.otorri.2018.01.003 [9] E. Lundstr¨om, Voice Function and Quality of Life in Laryngectomees, in: Studies in Logopedics and Phoniatrics, 13, Karolinska Institutet, Stockholm, 2009.455 [10] W. Wszolek, M. Modrzejewski, M. Przysiezny, Acoustic analysis of esophageal speech in patients after total laryngectomy, Archives of Acoustics 32 (4 (Supplement)) (2007) 151–158. [11] B. Weinberg, Acoustical properties of esophageal and tracheoesophageal speech, Laryngectomee rehabilitation (1986) 113–127.460 [12] T. Drugman, M. Rijckaert, C. Janssens, M. Remacle, Tracheoesophageal speech: A dedicated objective acoustic assessment, Computer Speech & Language 30 (1) (2015) 16–31. [13] R. Ishaq, B. G. Zapirain, Esophageal speech enhancement using modified voicing source, in: Signal Processing and Information Technology (ISSPIT),465 2013 IEEE International Symposium on, IEEE, 2013, pp. 000210–000214. [14] H. R. Sharifzadeh, I. V. McLoughlin, F. Ahmadi, Reconstruction of normal sounding speech for laryngectomy patients through a modified celp codec, IEEE Transactions on Biomedical Engineering 57 (10) (2010) 2448–2458. [15] R. van Son, I. Jacobi, F. J. Hilgers, et al., Manipulating tracheoesophageal470 speech., in: Interspeech, 2010, pp. 274–277. [16] A. Del Pozo, S. Young, Continuous tracheoesophageal speech repair, in: Signal Processing Conference, 2006 14th European, Citeseer, 2006, pp. 1– 5. 22 [17] A. Del Pozo, S. Young, Repairing tracheoesophageal speech duration, in:475 Proc Speech Prosody, Citeseer, 2008, pp. 187–190. [18] O. Schleusing, R. Vetter, P. Renevey, J.-M. Vesin, V. Schweizer, Prosodic speech restoration device: Glottal excitation restoration using a multiresolution approach, in: International Joint Conference on Biomedical Engineering Systems and Technologies, Springer, 2010, pp. 177–188.480 [19] Y. Stylianou, O. Capp´e, E. Moulines, Continuous probabilistic transform for voice conversion, IEEE Transactions on Speech and Audio Processing 6 (2) (1998) 131–142. doi:10.1109/89.661472. [20] D. Erro, A. Moreno, A. Bonafonte, Inca algorithm for training voice conversion systems from nonparallel corpora, IEEE Transactions on Audio,485 Speech, and Language Processing 18 (5) (2009) 944–953. [21] M. Kishimoto, T. Toda, H. Doi, S. Sakti, S. Nakamura, Model training using parallel data with mismatched pause positions in statistical esophageal speech enhancement, in: Signal Processing (ICSP), 2012 IEEE 11th International Conference on, Vol. 1, IEEE, 2012, pp. 590–594.490 [22] H. Doi, K. Nakamura, T. Toda, H. Saruwatari, K. Shikano, Esophageal speech enhancement based on statistical voice conversion with gaussian mixture models, IEICE TRANSACTIONS on Information and Systems 93 (9) (2010) 2472–2482. [23] H. Doi, K. Nakamura, T. Toda, H. Saruwatari, K. Shikano, Statistical ap-495 proach to enhancing esophageal speech based on gaussian mixture models, in: Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on, IEEE, 2010, pp. 4250–4253. [24] H. Doi, K. Nakamura, T. Toda, H. Saruwatari, K. Shikano, Speaking-aid systems based on one-to-many eigenvoice conversion for total laryngec-500 tomees, APSIPA ASC 2010 - Asia-Pacific Signal and Information Processing Association Annual Summit and Conference. 23 [25] H. Doi, K. Nakamura, T. Toda, H. Saruwatari, K. Shikano, An evaluation of alaryngeal speech enhancement methods based on voice conversion techniques, in: Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE505 International Conference on, IEEE, 2011, pp. 5136–5139. [26] K. Yamamoto, T. Toda, H. Doi, H. Saruwatari, K. Shikano, Statistical approach to voice quality control in esophageal speech enhancement, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2012, pp. 4497–4500.510 [27] H. Doi, Augmented speech production beyond physical constraints using statistical voice conversion – alaryngeal speech enhancement and singing voice quality control, Ph.D. thesis, Nara Institute of Science and Technology (2013). [28] D. Erro, I. Hern´aez, E. Navas, A. Alonso, H. Arzelus, I. Jauk, N. Q. Hy,515 C. Magarinos, R. P´erez-Ram´on, M. Sulır, et al., Zuretts: online platform for obtaining personalized synthetic voices, Proceedings of eNTERFACE (2014) 1178–1193. [29] D. Erro, I. Hernaez, A. Alonso, D. Garc´ıa-Lorenzo, E. Navas, J. Ye, H. Arzelus, I. Jauk, N. Q. Hy, C. Magari˜nos, R. P´erez-Ram´on, M. Sul´ır,520 X. Tian, X. Wang, Personalized synthetic voices for speaking impaired: Website and app., in: Interspeech, 2015, pp. 1251–1254. [30] M. Eye, E. Infirmary, Voice disorders database, version. 1.03 (cd-rom) (1994). [31] X. Menendez-Pidal, J. B. Polikoff, S. M. Peters, J. E. Leonzio, H. T. Bun-525 nell, The nemours database of dysarthric speech, in: Spoken Language, 1996. ICSLP 96. Proceedings., Fourth International Conference on, Vol. 3, IEEE, 1996, pp. 1962–1965. [32] H. Kim, M. Hasegawa-Johnson, A. Perlman, J. Gunderson, T. S. Huang, 24 K. Watkin, S. Frame, Dysarthric speech database for universal access re-530 search, in: Proceedings of Interspeech, 2008, pp. 1741–1744. [33] M. Kinishi, M. Amatsu, Pitch perturbation measures of voice production of laryngectomees after the amatsu tracheoesophageal shunt operation, Auris Nasus Larynx 13 (1) (1986) 53–62. [34] M. R. Arias, J. L. Ram´on, M. Campos, J. J. Cervantes, Acoustic analysis535 of the voice in phonatory fistuloplasty after total laryngectomy, Otolaryngology—Head and Neck Surgery 122 (5) (2000) 743–747. [35] C. J. van As-Brooks, F. J. Koopmans-van Beinum, L. C. Pols, F. J. Hilgers, Acoustic signal typing for evaluation of voice quality in tracheoesophageal speech, Journal of Voice 20 (3) (2006) 355–368.540 [36] M. Carello, M. Magnano, A first comparative study of oesophageal and voice prosthesis speech production, EURASIP Journal on Advances in Signal Processing 2009 (1) (2009) 821304. [37] J. K. MacCallum, L. Cai, L. Zhou, Y. Zhang, J. J. Jiang, Acoustic analysis of aperiodic voice: perturbation and nonlinear dynamic properties in545 esophageal phonation, Journal of Voice 23 (3) (2009) 283–290. [38] N. Deore, S. Datta, R. Dwivedi, R. Palav, R. Shah, S. Sayed, M. Jagde, R. Kazi, Acoustic analysis of tracheo-oesophageal voice in male total laryngectomy patients, The Annals of The Royal College of Surgeons of England 93 (7) (2011) 523–527.550 [39] H.-J. Shim, H. R. Jang, H. B. Shin, D.-H. Ko, Cepstral, spectral and timebased analysis of voices of esophageal speakers, Folia Phoniatrica et Logopaedica 67 (2) (2015) 90–96. [40] J. Robbins, H. B. Fisher, E. C. Blom, M. I. Singer, A comparative acoustic study of normal, esophageal, and tracheoesophageal speech production,555 Journal of Speech and Hearing disorders 49 (2) (1984) 202–210. 25