Design and evaluation of mobile computer-assisted pronunciation training tools for second language learning
Abstract
Departamento de Informática (Arquitectura y Tecnología de Computadores, Ciencias de la Computación e Inteligencia Artificial, Lenguajes y Sistemas Informáticos)
Full text
DOCTORAL PROGRAM IN COMPUTER SCIENCE DOCTORAL THESIS: DESIGN AND EVALUATION OF MOBILE COMPUTER-ASSISTED PRONUNCIATION TRAINING TOOLS FOR SECOND LANGUAGE LEARNING Thesis submitted by Cristian TEJEDOR-GARCÍA as part of the requirements for the degree of DOCTOR OF COMPUTER SCIENCE with mention of INTERNATIONAL DOCTORATE by the University of Valladolid Supervised by: Dr. Valentín CARDEÑOSO-PAYO Dr. David ESCUDERO-MANCEBO Valladolid, 2020
PROGRAMA DE DOCTORADO EN INFORMÁTICA TESIS DOCTORAL: DISEÑO Y EVALUACIÓN DE HERRAMIENTAS MÓVILES PARA EL ENTRENAMIENTO ASISTIDO POR ORDENADOR DE LA PRONUNICACIÓN PARA EL APRENDIZAJE DE IDIOMAS Presentada por Cristian TEJEDOR-GARCÍA para optar al grado de DOCTOR EN INFORMÁTICA con MENCIÓN DE DOCTORADO INTERNACIONAL por la Universidad de Valladolid Dirigida por: Dr. Valentín CARDEÑOSO-PAYO Dr. David ESCUDERO-MANCEBO Valladolid, 2020
iii Abstract The quality of speech technology (automatic speech recognition, ASR, and textto-speech, TTS) has considerably improved and, consequently, an increasing number of computer-assisted pronunciation (CAPT) tools has included it. However, pronunciation is one area of teaching that has not been developed enough since there is scarce empirical evidence assessing the effectiveness of tools and games that include speech technology in the field of pronunciation training and teaching. This PhD thesis addresses the design and validation of an innovative CAPT system for smart devices for training second language (L2) pronunciation. Particularly, it aims to improve learner’s L2 pronunciation at the segmental level with a specific set of methodological choices, such as learner’s first and second language connection (L1– L2), minimal pairs, a training cycle of exposure–perception–production, individualistic and social approaches, and the inclusion of ASR and TTS technology. The experimental research conducted applying these methodological choices with real users validates the efficiency of the CAPT prototypes developed for the four main experiments of this dissertation. Data is automatically gathered by the CAPT systems to give an immediate specific feedback to users and to analyze all results. The protocols, metrics, algorithms, and methods necessary to statistically analyze and discuss the results are also detailed. The two main L2 tested during the experimental procedure are American English and Spanish. The different CAPT prototypes designed and validated in this thesis, and the methodological choices that they implement, allow to accurately measuring the relative pronunciation improvement of the individuals who trained with them. Both rater’s subjective scores and CAPT’s objective scores show a strong correlation, being useful in the future to be able to assess a large amount of data and reducing human costs. Results also show an intensive practice supported by a significant number of activities carried out. In the case of the controlled experiments, students who worked with the CAPT tool achieved better pronunciation improvement values than their peers in the traditional in-classroom instruction group. In the case of the challenge-based CAPT learning game proposed, the most active players in the competition kept on playing until the end and achieved significant pronunciation improvement results. Keywords:Computer-assisted pronunciation training (CAPT), second language (L2) pronunciation, automatic speech recognition (ASR), text-to-speech (TTS), autonomous learning, automatic assessment tools, learning environments, mobile learning game, minimal pairs.
v Resumen El aumento de la mejora de la calidad de las tecnologías del habla (reconocimiento automático y síntesis del habla, ASR y TTS, respectivamente) trae consigo el incremento del número de herramientas para el entrenamiento de la pronunciación asistida por ordenador (CAPT). Sin embargo el uso de la tecnología para el entrenamiento de la pronunciación no está aún extendido de forma masiva debido a, entre otras causas, la falta de evidencia empírica sobre la eficacia de las aplicaciones que incluyen tecnología del habla para la enseñanza y entrenamiento de la pronunciación. Esta tesis doctoral aborda el diseño y la validación de un innovador sistema CAPT para dispositivos inteligentes para el entrenamiento de la pronunciación de lengua extranjera (L2). En concreto, tiene como objetivo mejorar la pronunciación L2 a nivel segmental del alumno mediante un conjunto específico de opciones metodológicas de carácter individual y social, como la conexión entre la lengua materna (L1) y la L2, un ciclo de entrenamiento de exposición–percepción–producción con pares mínimos, y la inclusión de tecnología ASR y TTS. La investigación experimental realizada aplicando estas opciones metodológicas con usuarios reales valida la eficacia de los prototipos CAPT desarrollados para los cuatro experimentos principales de esta tesis. Dichos sistemas CAPT recogen y analizan la información de interacción del usuario para proporcionarle una retroalimentación específica e inmediata. Gracias a los protocolos, métricas, algoritmos y métodos descritos en este trabajo, los resultados se analizan estadísticamente y discuten. Los dos principales L2 probados durante el procedimiento experimental son el inglés americano y español. Los diferentes prototipos CAPT diseñados y validados en esta tesis, y las opciones metodológicas que implementan, permiten medir con precisión la mejora de la pronunciación relativa de los estudiantes que entrenaron con ellos. Tanto las puntuaciones subjetivas de los evaluadores, como las objetivas de los sistemas CAPT, muestran una alta correlación, siendo estas últimas útiles en el futuro para la evaluación de una gran cantidad de datos y la reducción de tareas para los evaluadores. El número significativo de actividades llevadas a cabo por los participantes respalda una práctica intensa en la experimentación. En el caso de los experimentos controlados, los estudiantes que trabajaron con la herramienta CAPT lograron mayores valores de mejora de la pronunciación que sus compañeros en el grupo de aprendizaje tradicional en el aula. En el caso del juego educativo CAPT basado en desafíos, los jugadores más activos en la competición lograron resultados de mejora de pronunciación significativos. Palabras clave:Entrenamiento de la pronunciación asistida por ordenador (CAPT), pronunciación de segunda lengua (L2), reconocimiento automático del habla (ASR), síntesis de habla (TTS), aprendizaje autónomo, entornos educativos, juego educativo móvil, pares mínimos.
vii Acknowledgements In the following paragraphs I would like to thank the people who have accompanied and supported me in this exciting and intensive chapter of my life to obtain the PhD degree. It has not only scientifically educated me, but also developed myself personally, making me realize my (various) limitations and taking on new responsibilities. This PhD thesis has been carried out in the ECA-SIMM research group with the Department of Computer Science of the University of Valladolid. My first contact with the research group was when I was finishing my master’s degree in Computer Science at the middle of the year 2015. They were looking for one candidate to start researching in foreign pronunciation teaching with speech technology and mobile applications within a research project. I was really interested in what I was doing from the very first day in the laboratory. First of all, I am deeply indebted to my two supervisors for guiding me throughout the process of becoming a PhD graduate. Dr. Valentín Cardeñoso-Payo, there are not enough words to describe your contribution to my life during these last five years. We have shared many trips, anecdotes, and hard days of work. Among your innumerable virtues I would like to point out your capacity to synthesize knowledge and for finding solutions to almost every problem. I would also like to extend my deepest gratitude to Dr. David Escudero-Mancebo. Your proactive personality and your unique ability for finding significant results have helped me to reach the objectives of this thesis. You are such a nice person who has not only earned a successful research career, but is well-aware of the environmental issues. I am also extremely grateful to my colleague Dr. César González-Ferreras. I cannot thank you enough for everything that you have done for making this work a possible one. I really appreciate the pragmatic and helpful way that you have proceeded whenever I needed help. I would also like to take this opportunity to thank Dr. Enrique Cámara-Arenas for sharing his ideas and helpful suggestions for this work. Your love for writing and endless creativity have been essential to this thesis to succeed. I want to thank Dr. Mario Corrales-Astorgano, my colleague from the ECA-SIMM group. Sharing my time and experiences with you has been very rewarding. I am also grateful for your selfless help. Although we started and almost finished the predoctoral program at the same time, our research careers have just started. I also want to thank Dr. María Jesús Rodríguez-Triana, Dr. Tobias Ley, and Dr. Luis Pablo Prieto Santos for giving me the opportunity of being part of your research group during three months at the Centre of Excellence in Educational Innovation at the University of Tallinn, Estonia. I am so grateful for letting me into your international family and in what I consider my second home. Traveling to such a wonderful (and different) city is exciting but also a little overwhelming. Thanks also to Mr.
xv List of Tables 4.1 Comparison of CAPT experiments in the literature. ........... 39 7.1 Main elements included in each prototype of the experimentation. . . 74 7.3 ASR-related results gathered with the Minimal Pairs CAPT system, adapted from [17]. ............................... 79 7.4 ASR-related metrics gathered with the Minimal Pairs CAPT system, adapted from [17]. ............................... 80 7.5 Mean distribution of the target word in each recognized utterance with the Minimal Pairs CAPT system, adapted from [17]. ........ 80 7.6 Most frequently unrecognized words by the ASR system in the Minimal Pairs prototype (in percentage), adapted from [17]. ......... 81 7.7 TTS-related results gathered with the Minimal Pairs CAPT System, adapted from [17]. ............................... 82 7.8 User’s behavior according to the number of times an activity type performed in the TipTopTalk! prototype. .................... 89 7.9 Average time (s) spent by users in each activity type of the TipTopTalk! prototype. ................................... 90 7.10 Average number of discrimination and production events per participant of the TipTopTalk! prototype. ..................... 91 7.11 Success rate in each activity type of the TipTopTalk! prototype. ..... 93 7.12 Number of tasks of each training mode of the Guided Learning experiment, adapted from [23]. ..........................101 7.13 ABX questions and answers of the English Vowels prototype. . . . . . 105 7.14 User’s performance with the CAPT system of the English Vowels prototype, adapted from [23]. ..........................108 7.15 Right, wrong, and listening events categorized by phoneme of the English Vowels prototype, adapted from [23]. ................109 7.16 Confusion matrices of the English Vowels prototype, adapted from [23].109 7.17 Comparison between following recommended feedback or not of the English Vowels prototype, adapted from [24]. ...............110 7.18 Sequences of wrong production, listen, and repeat of the English Vowels prototype. .................................111 7.19 Pre-test and post-test mean production scores of the English Vowels prototype, adapted from [23]. ........................112 7.20 ABX test results of the English Vowels prototype. .............112 7.21 Correlation between the software and human raters post-test scores of the English Vowels prototype, adapted from [23]. ...........113 7.22 User’s performance with the CAPT system of the Japañol prototype. . 115 7.23 Right, wrong, and listening events as a function of phonemes of the Japañol prototype ...............................115 7.24 Confusion matrix of discrimination tasks of the Japañol prototype. . . 116 7.25 Confusion matrix of production tasks of the Japañol prototype. . . . . 117
xvi 7.26 Scores at different stages of the Japañol prototype, adapted from [25]. . 118 7.27 WER values (%) of the six models tested for the Kaldi ASR system. . . 119 7.28 Google and Kaldi results of the tests utterances of the Japañol prototype.119 7.29 Extra points scoring system of COP (ExtraScore value), adapted from [31]. .......................................127 7.30 Average number of discrimination and production events per participant of the COP prototype, adapted from [31]. ..............130 7.31 Indicators of activity per declared level of English of the COP prototype.130 7.32 Kruskal–Wallis test results of indicators of activity per declared level of English of Table 7.31 of the COP prototype. ...............131 7.33 Mann–Whitney Utest results by declared level of English of Table 7.31 in the COP prototype. ............................131 7.34 Indicators of activity per type of user of the COP prototype, adapted from [31]. ....................................132 7.35 Kruskal–Wallis test results of indicators of activity of Table 7.34 in the COP prototype, adapted from [31]. .....................133 7.36 Mann–Whitney Utest results for the three group pairs of Table 7.34 in the COP prototype, adapted from [31]. ...................133 7.37 Success rates of discrimination and production events at the beginning and at the end of the COP prototype, adapted from [31]. . . . . . 134 7.38 Declared reasons for playing of the COP prototype, adapted from [31]. 135 7.39 Attitude toward competition of the COP prototype, adapted from [31]. 136 7.40 Early abandonment questionnaire results of the COP prototype, adapted from [31]. ....................................136 7.41 Notes gathered from the intrinsic and extrinsic focus group sessions of the COP prototype. ............................137 7.42 Notes gathered from the English proficiency level focus group session of the COP prototype. ............................138 7.43 Notes gathered from the degree of competitiveness focus group session of the COP prototype. ..........................139 A.1 Comparative of time and space complexities of the brute force-based and tree-based algorithm for elaborating minimal pairs lists. ......173 B.1 Number of development, recruitment, and testing days of each one of the prototypes of the experiments. .....................181 B.2 Comparative among experiments’ training methodology. ........182 B.3 Comparative among experiments’ pronunciation assessment approach.183 B.4 Comparative among experiments’ integrated technology. ........184 B.5 Comparative among experiments’ gamification instruments (I). . . . . 185 B.6 Comparative among experiments’ gamification instruments (II). . . . . 186 B.7 Comparative among experimentation participants’ demographics. . . 187 B.8 Comparative among experiments’ CF strategies. .............188 C.1 Pre-test and post-test words list of the English Vowels prototype. . . . 189 C.2 Pre-test and post-test words list of the Japañol prototype. ........190 D.1 Descriptive statistics of the speech data gathered from the Japañol and COP prototypes. ................................191
xvii List of Acronyms and Abbreviations 3-D (3)Three Dimensional App Application (software) ASR Automatic Speech Recognition AVR Automatic Voice Recognition CALL Computer-Aided Language Learning CAPT Computer-Assisted Pronunciation Training CEFR Common European Framework Reference CF Corrective Fedback CMVN Cepstral Mean and Variance Normalization CN Chinese cn_ZH Simplified Chinese (Mainland China) COP Clash ofPronunciations (prototype) CPU Central Processing Unit DE German de_DE German (Germany) DNN Deep Neural Networks ECA-SIMM Entornos de Computación Avanzada y Sistemas de Interacción Multimodal EFL English as Foreign Language EME-E Escala de Motivación Educativa (España) EN English en_US American English EÑ Estoñol (prototype) ES Spanish es_ES Castilian Spanish ET Estonian et_EE Estonian (Estonia) EVow English Vowels (prototype) fMLLR Feature Space Maximum Likelihood Linear Regression FST Finite-State Transducers GCSTT Google Cloud Speech-To-Text GMM Gaussian Mixture Model GOP Goodness ofPronunciation GUI Graphical User Interface HMM Hidden Markov Model ICT Information and Communications Technology IEEE Institute of Electrical and Electronics Engineers IPA International Phonetic Alphabet JCR Journal Citation Report JÑ Japañol (prototype) JP Japanese
xviii jp_JP Japanese (Japan) JSON JavaScript Object Notation L1 First(1)Language L2 Second/Foreign(2)Language LDA Linear Discriminat Analysis LDC Linguistic Data Consortium LL Language Learning MALL Mobile-Assisted Language Learning MFCC Mel Frequency Cepstral Coefficients MP Minimal Pairs (prototype) NCM Native Cardinality Method OOV Out-of-vocabulary (context) OS Operating System PC Personal Computer PhD Doctor ofPhilosophy PLP Perceptual Linear Prediction PPV Positive Predictive Value PT Portuguese pt_BR Brazilian Portuguese pt_PT European Portuguese RO Research Objective RQ Research Question SAMPA Speech Assessment Methods Phonetic Alphabet SAT Speaker Adaptive Training SCI Social Citation Index SGMM Subspace Gaussian Mixture Model SLA Second Language Acquisition SVN Apache Subversion TPR True-Positive-Recall TTS Text-To-Speech TTT TipTopTalk! (prototype) UK United Kingdom USA United States of America UX User Experience VTLN Vocal Tract Length Normalization vs. Versus WER Word Error Rate WFSTs Weighted Finite State Transducers
xix To my beloved Teresa, Bernardo and Raúl. You were, are, and will always be by my side.
1 Capítulo R1 Resumen de la Tesis Doctoral El éxito de la comunicación en un idioma extranjero depende en gran medida de la inteligibilidad, comprensión, acento y fluidez del habla. Los cursos de aprendizaje de idiomas se han centrado tradicionalmente, sin embargo, en otras áreas de habilidades lingüísticas, como la gramática o la comprensión escrita. Por un lado, los nuevos enfoques que ofrecen las recientes innovaciones en las tecnologías del habla mejoran significativamente el rendimiento del reconocimiento y síntesis del habla. Dicha tecnología se puede integrar en sistemas pedagógicos para dispositivos inteligentes actuales para el entrenamiento de la pronunciación mediante aplicaciones que complementen el aprendizaje. Esto permite a los estudiantes usarlas de forma continua y autónoma. Por otro lado, las aplicaciones de juegos educativos tienen un enorme potencial para la educación y, en particular, para el aprendizaje de idiomas. El proceso de aprendizaje se ve afectado por la participación social que implican dichos juegos, cuya utilidad y eficacia deben ser evaluadas. No obstante, existe escasa evidencia experimental de la eficacia del uso aplicaciones tecnológicas para el entrenamiento y mejora de la pronunciación extranjera. Este capítulo plantea el contexto y motivación de este trabajo de tesis doctoral en relación a los temas mencionados anteriormente. Además, se describen de forma general las contribuciones aportadas en la tesis doctoral, al resolver las cuestiones planteadas en las preguntas de investigación, y conseguir los objetivos descritos. R1.1 Motivación La demanda actual de aprendizaje de segunda lengua (SLA) es muy alta. A finales de 2016 existían 912 millones de estudiantes de segunda lengua (L2) en todo el mundo [1], una séptima parte de la población mundial. Esto es, en cierta medida, por la necesidad de comunicación entre personas de cualquier lugar del mundo por medio de la tecnología actual que permite este proceso. No obstante, la gran cantidad y diferencias entre idiomas y culturas pueden ser una barrera para conseguir una comunicación exitosa. Se estima que cada persona tiene acceso a alrededor de 6 dispositivos inteligentes en el año 2020 [1]. Dichos dispositivos forman parte de la tecnología educativa (e-learnig) y autoaprendizaje; alternativas interesantes a los cursos tradicionales en el aula. En particular, los sistemas de aprendizaje de idiomas asistido por ordenador (CALL) y por dispositivos móviles (MALL) integran tecnología avanzada muy atractiva para el aprendizaje de idiomas y que puede ayudar en el proceso de aprendizaje y enseñanza de manera eficiente. Sin llegar a reemplazar a los tutores humanos, pueden desempeñar un papel complementario en la educación al aumentar la eficiencia y la motivación del proceso de aprendizaje, se pueden utilizar en cualquier lugar, momento y tantas veces como se desee.
2Capítulo R1. Resumen de la Tesis Doctoral Actualmente, el número de juegos educativos para el aprendizaje de idiomas está aumentando, dado que la inclusión de elementos de juego en herramientas educativas favorece un mejor rendimiento individual [2]. Estudios recientes detallan que la motivación y el compromiso del alumno mejoran no solo dentro sino también fuera del aula [3]. Aunque la inclusión de elementos sociales y competitivos en cualquier sistema pedagógico debe hacerse con precaución, existen estudios que indican que la competitividad en el contexto del aprendizaje basado en juegos facilita el logro de objetivos educativos [4] y fomenta la cooperación como un elemento de apoyo al trabajo en clase [5]. Los sistemas para el entrenamiento de la pronunciación asistida por computador (CAPT) dan soporte a investigaciones y prácticas innovadoras que favorecen a la transformación del aprendizaje de idiomas, creando oportunidades para revisar las viejas ideas y desafiar las creencias establecidas [6]. CAPT es una subárea importante de CALL y MALL en constante cambio, que combina la retroalimentación correctiva y la evaluación automática de la calidad de la pronunciación, entre otras funcionalidades proporcionadas por las tecnologías del habla incorporadas. Las más comunes son el reconocimiento automático de voz (ASR) y la síntesis de habla (TTS), que transforman la voz en texto escrito, y viceversa, respectivamente. Hoy en día, estos sistemas están respaldados por una enorme cantidad de datos y algoritmos complejos que mejoran significativamente su calidad. Por ejemplo, Google reportó que sus recientes avances en el aprendizaje automático aplicado a TTS han ayudado a generar formas de onda de voz 1000 veces más rápido que antes (generar un segundo de audio solo tarda 50 milisegundos), y han logrado calificaciones más de un 20% mejores que las voces estándar [7]. Además, en el campo del reconocimiento de voz, Google también ha alcanzado una tasa de precisión de palabras del 95% para el idioma inglés, por lo tanto, alcanzando el umbral de precisión humana [8]. Es probable que se obtengan mejores tasas en el futuro cercano con el uso de técnicas de redes neuronales profundas (DNN) más sofisticadas, y mayores unidades de procesamiento central (CPU) [9], [10]. Aunque hasta 2014 se han reportado pocos estudios de investigación revisados por pares sobre CAPT (solo un 26.9% de los 75 estudios del estado de la cuestión resumidos en [11]), los sistemas CAPT están evolucionando y apareciendo cada vez más debido a las mejoras y las nuevas posibilidades que ofrecen [12]. Una metodología de entrenamiento correcta debe abordar adecuadamente los aspectos de la retroalimentación automática instantánea y el diseño de actividades y elementos de enseñanza de acuerdo con la lengua materna (L1) y L2 del alumno, a fin de optimizar la eficacia de las herramientas CAPT y el tiempo de uso [13]. El proceso de aprendizaje para la adquisición de L2 se ve muy afectado por una percepción habitual bien establecida de los movimientos y sonidos articulatorios L1. A menudo conduce a errores e imprecisiones en la pronunciación L2 de los alumnos (es decir, una transferencia negativa del idioma [14]). En esta tesis se ha utilizado la técnica de pares mínimos, pares de palabras que varían en un solo sonido. El uso de pares mínimos puede aportar grandes beneficios en el aprendizaje y la enseñanza de la pronunciación, ya que aparecen en casi todos los idiomas y pueden contrastar los sonidos L1 y L2 [15]. En resumen, el aprendizaje de L2, y más precisamente, el entrenamiento de pronunciación de L2, está abierto a nuevos paradigmas de enseñanza. La posibilidad de que los alumnos entrenen en cualquier momento y en cualquier lugar, a su propio
R1.2. El Problema 3 ritmo, permite a los maestros proporcionar una instrucción individualizada en grupos pequeños en lugar de los tradicionales grupos de mayor tamaño. El hecho de que los estudiantes de hoy en día estén acostumbrados a la tecnología digital motiva el desarrollo de sistemas de aprendizaje CAPT para dispositivos inteligentes. Los alumnos están acostumbrados a utilizarlos para interactuar en entornos digitales para la comunicación, información, contacto social, reunión y análisis. Aunque pueden ser nativos digitales y estar cómodos e inmersos en la tecnología, dependen de maestros y expertos para aprender a través de los medios digitales. Además, las tecnologías ASR y TTS han mejorado drásticamente su rendimiento en los últimos años, pudiendo integrarse en los recursos educativos. Por lo tanto, el desafío actual es diseñar y adaptar cuidadosamente un sistema CAPT efectivo con no solo dicha tecnología, sino también con una metodología de entrenamiento, una evaluación de mejora de la pronunciación y una estrategia de retroalimentación correctiva, de acuerdo con las L1 y L2 del alumno. R1.2 El Problema Proporcionar un conjunto adecuado de actividades de entrenamiento para la pronunciación no es una tarea fácil ya que hay varios factores a tener en cuenta: •L1 del alumno. Las similitudes y diferencias entre la primera y segunda lengua del estudiante varían la dificultad del proceso de aprendizaje. •El conjunto de actividades de entrenamiento personalizadas. Dependiendo del nivel de pronunciación L2 del alumno y su desempeño, las actividades recomendadas deben ser individualizadas y adaptadas para que sean efectivas. •Evaluación de los resultados del alumno. Se pueden proporcionar valoraciones subjetivas y objetivas a los usuarios, no solo al final de la experimentación, sino también durante el entrenamiento. •La retroalimentación proporcionada a los estudiantes. Se necesita más retroalimentación de la que un maestro puede ofrecer en clase. Sin embargo, en algunos casos esta retroalimentación es insuficiente o demasiado difícil de entender para los alumnos. •La tecnología incluida en la metodología de entrenamiento debe seleccionarse cuidadosamente ya que las puntuaciones de evaluación deben ser lo más precisas posible y orientarse al objetivo del entrenamiento. •Elementos motivacionales, como los elementos de juego o la interacción con otros alumnos pueden influir en los resultados del entrenamiento, desviando a los estudiantes del objetivo real de mejora de la pronunciación o incluso desanimándolos. Un trabajo de investigación que plantee los problemas descritos anteriormente debe abordarse desde una perspectiva multidisciplinar que incluya: metodología educativa, diseño de herramientas educativas, modelado de datos y técnicas de evaluación, entre otros. Finalmente, es necesario establecer protocolos adecuados para recopilar y analizar los resultados de los experimentos, ya que la eficacia de la herramienta CAPT está influenciada por su escalabilidad y rendimiento. En primer lugar, la posibilidad de ampliar el conjunto de idiomas conduce a generalizar los conceptos y estrategias de entrenamiento, y a seleccionar correctamente la tecnología de voz necesaria. En
10 Capítulo R1. Resumen de la Tesis Doctoral • Las decisiones metodológicas llevadas a cabo en las diferentes versiones de las herramientas CAPT diseñadas y validadas en este trabajo han permitido medir la mejora de la pronunciación relativa de las personas que entrenaron con ellas. –Se han utilizado listas de pares mínimos elaboradas mediante un novedoso protocolo semi-automático propuesto en esta tesis doctoral, que tiene en cuenta la L1 y L2 del participante y la tecnologías ASR y TTS. –Se han incluido dichas listas en ejercicios de exposición, percepción y producción en diferentes modos de entrenamiento, sonidos e idiomas, en los que se ha utilizado tecnología ASR y TTS. –Se han empleado diferentes técnicas de retroalimentación correctiva que han demostrado ser útiles y efectivas. Con ellas, los usuarios han podido superar los ejercicios de entrenamiento propuestos; y nosotros hemos sido capaces de averiguar sus mayores dificultades en cuanto a sonidos y actividades de entrenamiento. –Se han reportado no solo resultados positivos de mejora de habilidades de percepción y producción de manera objetiva y subjetiva, sino que los participantes de grupos que utilizaron las herramientas CAPT lograron una mejora mayor que la lograda en los grupos de instrucción con el profesor en el aula. • Por último, los elementos de juego han tenido una influencia positiva en la motivación, el rendimiento y el aprendizaje de los participantes en los diferentes sistemas CAPT desarrollados en esta tesis. En concreto, la competición de COP ha demostrado ser un factor motivacional positivo, especialmente para los usuarios más activos, cuya participación intensiva en el juego les permitió lograr una mejora significativa de la pronunciación L2 al final del experimento. Además de las colaboraciones que se mantienen relacionadas con esta tesis doctoral, la gran cantidad de datos recopilados y que los resultados de esta tesis son satisfactorios, hay algunos aspectos que pueden mejorarse y dar paso a nuevas líneas de trabajo futuro, como son: • Analizar y diseñar algoritmos específicos de reconocimiento de voz para la identificación de errores de pronunciación que permitan caracterizar el nivel de habilidad de pronunciación. Con ello, se determinará el conjunto de características clave obtenidas al correlacionar los errores de pronunciación con las valoraciones de expertos humanos para que un sistema de clasificación automática permita formular recomendaciones personalizadas sobre el modo y lugar de articulación de la pronunciación. • Encontrar nuevas técnicas para adaptar el sistema CAPT al usuario de una manera más personalizada e individualizada, ayudará aún más a mejorar no solo sus resultados de mejora de la pronunciación, sino también su grado de motivación durante el entrenamiento de pronunciación. • Analizar la relación entre el diseño del sistema CAPT, la estrategia seguida por los participantes durante el entrenamiento y sus resultados finales será útil para clasificar y predecir el comportamiento de los usuarios con dicho sistema.
11 Parte I Introduction
13 Chapter 1 Introduction Acquiring a proper communication level in any foreign language is mainly affected by intelligibility, nativelikeness, comprehensibility, and fluency of speech. However, traditional language learning courses and systems often focus on other language skill areas, such as grammar or writing. On the one hand, recent advances in speech technology have reported new approaches that improve significantly the performance of voice recognition and speech synthesis. Consequently, these technologies are integrated into state-of-the-art pedagogical systems for pronunciation training as complementary tools through applications for smart devices, allowing learners to use them continuously and autonomously. On the other hand, learning games have a remarkable potential for education, and in particular, for language learning. They provide an emergent form of social participation that deserves the assessment of their usefulness and efficiency in the learning process. Nevertheless, there is still scarce empirical evidence about the effectiveness of CAPT systems with speech technology. This chapter discusses the feasibility of the topics mentioned above for pronunciation training, introducing the reader in the field, and showing both the context and motivation of this thesis work. Furthermore, the specific problems and challenges that lead to the objectives and research questions defined for this dissertation are identified. 1.1 Motivation There is currently a growing demand on second language acquisition (SLA). A recent study at the end of 2016 informed that there were approximately 912 million second language (L2) learners worldwide [1], a seventh part of the global population in that year. Besides, communication between people from different places of the world is no longer a problem since the advancements in technology ease this process. However, the variety of languages and cultural environments of individuals might be a barrier to develop a successful communication. The predictions for the year 2020 advanced that each person would have access to 6.58 smart devices [1]. These devices are present in e-learning and one-to-one tutoring, interesting alternatives to traditional in-classroom courses. In particular, computer-assisted language learning (CALL) and mobile-assisted language learning (MALL) systems integrate advanced technology that become very attractive to language learning and can help in the process of learning and teaching in an efficient way. Even though such devices and technology cannot serve as human tutors, they can perform a complementary role in education by increasing efficiency and
14 Chapter 1. Introduction motivation of the learning process, being used anywhere at any time, and repeated as many times as desired. Including game elements in educational tools favors a better individual performance [2]. Currently, the number of learning games for language learning is increasing. Recent studies report that learner’s motivation and engagement are enhanced not only inside but also outside the classroom [3]. Although the inclusion of social and competitive elements in any pedagogical system must be done with caution, there are studies that indicate that competitiveness in the context of game-based learning facilitates the achievement of learning objectives [4] and encourages cooperation as an articulating element of class work [5]. Computer-assisted (aided) pronunciation training (CAPT) systems support innovative research and practices which lead to transform language learning, creating opportunities to revisit old ideas and challenge established beliefs [6]. CAPT is an important sub-area of CALL and MALL constantly undergoing change, which combines corrective feedback and automatic pronunciation quality assessment, among other functionalities provided by the speech technologies incorporated. They are often automatic speech recognition (ASR) and text-to-speech (TTS), which transform speech into written text, and vice versa, respectively. Nowadays, these systems are supported by enormous quantity of data and complex algorithms that improve their quality significantly. For instance, the Google company reported that their recent advances in machine learning applied to TTS have helped to generate speech waveforms 1000 times faster than before (generating one second of audio only takes 50 milliseconds), and have achieved over 20% better quality ratings than standard voices [7]. Besides, in the field of speech recognition, Google have also achieved a word accuracy rate of 95% for the English language, therefore reaching the threshold of human accuracy [8]. Better rates are likely to be obtained in the near future with the use of more sophisticated deep neural network (DNN) techniques and greater central processing units (CPUs) [9], [10]. Although until 2014 there were scarce peerreviewed research investigations about CAPT (only a 26.9% of the 75 state-of-the-art studies surveyed in [11]), CAPT systems are being incorporated in recent experiments more frequently due to the improvements and new possibilities they offer [12]. A correct training methodology must adequately address the aspects of instant automatic feedback and the design of activities and teaching elements according to learner’s L1 and L2, in order to optimize CAPT tools’ efficiency and use time [13]. The L2 acquisition learning process is intensely affected by a well-established habitual perception of articulatory motions and sounds in the learner’s mother language (L1). It often leads to mistakes and inaccuracy in speech production of the L2 learners (i.e., a negative language transfer [14]). For these reasons, in this thesis the minimal pairs technique has been used. Minimal pairs (pairs of words that vary by only a single sound) bear great benefits in pronunciation learning and teaching since they appear in almost all languages and can contrast L1 and L2 sounds [15]. They are often used for teaching L2 segmental pronunciation (i.e., the teaching of single speech sounds, such as vowels, disregarding intonation, and other suprasegmental aspects of connected speech [32]). In summary, L2 learning, and more precisely, L2 pronunciation training, is opened to new teaching paradigms. The possibility of learners training anytime anywhere,
1.2. The Problem 15 at their own pace, allows teachers to provide small group and individualized instruction rather than lecturing to an entire class. The fact that today’s students are digitally literate motivates the development of learning CAPT systems for smart devices. Learners are used to interact into digital environments for communication, information, social contact, gathering, and analysis. Although they might be digital natives, comfortable with, and immersed in technology, they depend on teachers and experts to learn through digital means. Besides, ASR and TTS technologies have drastically enhanced their performance in recent years, being possible to be integrated into educational resources. Therefore, the current challenge is to carefully design and adapt an effective CAPT system with not only such technology, but also with a training methodology, a pronunciation improvement assessment, and a corrective feedback strategy, according to learner’s L1 and L2. 1.2 The Problem Providing an effective set of pronunciation training activities is not an easy task since there are several factors to take into account: •Learner’s L1. The similarities and differences between mother and target languages vary the difficulty of the learning process. •The set of personalized training activities. Depending on the student’s L2 pronunciation level and her/his performance, the activities recommended must be individualized and adapted to each learner in order to be effective. •Assessment of learner’s results. Both subjective and objective scores can be provided to users, not only at the end of the experimentation, but also during the training. •The feedback provided to the students. More feedback than a teacher alone can give in class is needed. However, in some cases this feedback is insufficient or too difficult to understand by the learners. •The technology included in the training methodology must be carefully selected since these assessment scores must be as precise as possible. •Motivational elements, such as game elements or interaction with other learners might influence the results of the training, deviating students from the actual pronunciation improvement goal or even discouraging them. A research challenge that raises the problems described above must be addressed from a multidisciplinary perspective which includes: learning methodologies, design of learning tools, data-based modeling, and evaluation techniques, among others. Finally, establishing proper protocols for gathering and analyzing results from the experiments is necessary since the effectiveness of the CAPT tool is influenced by its scalability and performance. Firstly, the possibility to extend the range of languages leads to generalize training concepts and strategies, and to correctly select the necessary speech technology. Secondly, the user’s interaction with a CAPT system tends to be massive in terms of speech data and log activity. In some cases an important investment of money is necessary in devices and computer servers.
16 Chapter 1. Introduction 1.3 Objectives and Research Questions The main objective of this thesis is defined as: To design and evaluate a CAPT tool for smart devices which incorporates current TTS and ASR technology; helping students to work autonomously, at their own pace, and with the possibility of providing real-time feedback. This main objective is divided into four specific research objectives: •RO1. To analyze and define a set of activities, protocols, and motivational elements for the improvement of L2 pronunciation with a CAPT system which integrates TTS and ASR technology. •RO2. To select the most appropriate metrics for the assessment of the speaker’s pronunciation level. •RO3. To design a semi-automatic method supervised by experts for obtaining a specific set of minimal pairs adapted to L2 pronunciation problems, according to the speaker’s L1 and to the limitations of the TTS and ASR technology. •RO4. To select and design a CAPT system with current TTS and ASR technology that provides an individualized feedback to the speaker for improving L2 pronunciation. In order to carry out the experimental procedure of this thesis, three research questions are identified to validate the research objectives, categorized by topics. The first topic is related to the feasibility of current speech technology (TTS and ASR systems) integration in CAPT tools: •RQ1. Can current TTS and ASR systems be successfully used in a non-obstructive way in the CAPT tool developed? –Issue 1.1. Can current TTS and ASR systems help to assess different groups of speakers according to their L2 pronunciation level in the CAPT tool developed? The second topic refers to the implications of the training methodology with CAPT tools in learner’s pronunciation improvement: •RQ2. To what extent can methodologically sensitive design issues, such as the use of exercises based on minimal pairs within the training activities cycle proposed in the CAPT tool developed affect user’s pronunciation improvement? –Issue 2.1. Can a relative improvement in the student’s pronunciation be assessed after using the CAPT tool? –Issue 2.2. If any, is there a relevant pronunciation improvement from a quantitative point of view? –Issue 2.3. Does the tool reveal what the real difficulties of the users are (most difficult sounds and most difficult training activities)? Finally, the last research question aims at answering how game elements and social approaches affect learner’s implication in pronunciation training with CAPT tools: •RQ3. To what extent can gamified versions of the tool affect user’s motivation, performance, and learning?
1.4. Research Methodology 17 1.4 Research Methodology In order to accomplish the objectives and give answers to the research questions proposed in this thesis, an experimental research [16] is conducted for the whole experimentation process with a multidisciplinary group of researchers and experts. Five phases can be defined in each experimental iteration: 1. Identifying the research problem. The process starts by clearly identifying the problems that will be addressed during the research process, starting with the existing solutions in the state-of-the-art, and considering what possible methods will affect a solution. 2. Planning the experimental research study. An experiment is carefully devised to test the research objectives and questions. (a) Selection of participants. The target population, enrollment rules, sample size, and groups are defined. (b) Variables. Different metrics are defined to measure the research variables from the data results gathered from the instruments. (c) Assessment protocol. The research variables are measured before, during, and after performing the training activities. (d) CAPT tool development. For each experiment in this thesis a novel CAPT tool is developed. 3. Conducting the experiment. At the beginning, the participants’ groups must be established. Then, each user performs the activities defined for her/his group in the previous phase, and the experimental data related to the variables of the study is collected with specific instruments for each experiment. 4. Analyzing the data. The data gathered is analyzed. It must be decided which indicators will be, and will not be, important, in order to corroborate how the experiment is successful. 5. Publication of findings. The most relevant results are shared and published in scientific journals and conferences by means of articles, abstracts, show and tell demonstrations, and presentations. 1.5 Outline This document is structured in six parts. In the first chapter of the first part (Chapter 1), the main topic of this thesis has been presented and motivated, the research objectives and questions have been settled, and a global vision of the research methodology carried out for this dissertation has also been given. In the second part, a deep revision of related work in the state-of-the art is presented and discussed on the light of the main characteristics of systems and experiments related to the objectives of this thesis. In particular, in Chapter 2a review of traditional pronunciation training activities and CAPT with the possible integration of ASR and TTS systems in pronunciation instruction are reviewed. In Chapter 3the state-of-the-art pronunciation improving assessment strategies for CAPT are discussed. In Chapter 4the corrective feedback strategies adopted by the state-ofthe-art CAPT studies are examined. In Chapter 5the fundamentals of individualistic
18 Chapter 1. Introduction and social learning applied to pronunciation training are described. The implications of gamification elements in learning contexts are also stated. The specific details about the experimental framework’s concepts, strategies, and elements necessary for the experimentation are specified in the third part of this document. In particular, in Chapter 6an exhaustive review of the common dimensions of the experimental procedure is included, in relation to the state-of-the-art. In Chapter 7each experiment of this dissertation is detailed in depth, giving an evolutive vision of the work carried out along this thesis. In Chapter 8the results obtained in all experiments are discussed to give answer to the research questions of this thesis. In the fourth part (Chapter 9) the conclusions are summarized and some future directions of this thesis work are defined. Furthermore, the publications, research funding, achievements, and attributions obtained during the course of this dissertation are enumerated. The appendices are included in the fifth part of this document. Appendix Aexplains in depth the algorithm for elaborating minimal pairs lists designed in this thesis. In Appendix Ban overview about the similarities and differences of all the experiments of this thesis is presented via comparative tables. Appendix Cshows the pre-test and post-test list of words given to users in the experiments which included them. Appendix Drepresents the main characteristics of the speech corpus data gathered in two experiments of this thesis. Appendix Esheds light to a standard Kaldi project structure and a list of steps for developing an ASR system. Finally, in the sixth and last part of this thesis the references in the bibliography chapter are included. They are compliant with the IEEE Reference Guide1. 1http://ieeeauthorcenter.ieee.org/wp-content/uploads/IEEE-Reference-Guide.pdf
19 Part II Literature Review
26 Chapter 2. CAPT Methodologies for Second Language Learning of the first computer tutors and the employment of ASR for LL and speech therapy. Although Dragon Naturally Speaking was designed by native speakers and was not developed for error detection, [52] reported that it might have pedagogical value in the future as a means of giving corrective feedback and identify the problems in pronunciation that affect humans’ understanding of non-native speech. [53] complemented the idea of [52], by affirming that ASR’s accuracy must be tested by natives (in this case English). In particular, he proposed that when a more highly developed version of the software incorrectly recognized a word, students might view the computer’s error as an indication of a mispronunciation which needs correction. Since late 2000s, there has been a growing interest in CAPT by means of ASR because these systems provide learners automatic and individualized feedback in a private environment [13]. Two main tendencies in ASR for LL can be distinguished. First, computer-desktop applications for pronunciation training [11]. Second, the possibility of learning anytime and anywhere with a diverse range of smartphone applications makes possible learning not only individually but also collaboratively and competitively (i.e., Duolingo1, Babbel2, or ElsaSpeak3). This type of applications often turns into online courses [54], [55]. The core ASR systems integrated into the mentioned applications varies from open source frameworks with a great community support, such as CMU Sphinx, Kaldi, and Julius to commercial and very powerful ASR systems, such as HTK, Dragon Dictation, Google Now, Cortana, Siri, or Alexa, among others. Interested readers can find a detailed description about speech recognition fundamentals, characteristics and examples of ASR systems in Section 6.4.1. Choosing a proper ASR system is not trivial. The objective of the study and the subjects target must be taken into account: 1. ASR systems for native speakers. There are some characteristics to take into account, such as pronunciation variation, end-point detection, disfluencies, and background sounds, among others. Examples of this type of ASR systems are dictation software applications, personal assistants, or voice command systems. 2. ASR systems for non-native speakers. Systems more complex than the previous group which often have a degraded performance. They include atypical and pathological speech (i.e., a corpus populated with utterances of non-native speakers). The three main knowledge sources of ASR systems are clearly affected (grammar, set of words, and pronunciation deviations). Besides, there are problems in read speech (bad linking of words) and in spontaneous speech (more filled pauses). There are several inclusive strategies to improve non-native ASR systems’ performance. The first approach and more generalist one, is to optimize the acoustic models, lexicon, and language model to compensate for deviations. Another option is to restrict the search space. For instance, constraining learner’s output by elicitation strategies, such as reading aloud and repeating auditorily prompted sentences. 1http://duolingo.com 2http://babbel.com 3https://elsaspeak.com/en/
2.3. Speech Recognition in Language Learning 27 On the other hand, ASR systems performance can be improved by altering other external factors. First, analyzing what has been said, adjusting the grade of tolerance. Second, examining how has it been said (i.e., error detection and find deviations from typical speech). Finally, providing feedback to the learner. In summary, an ASR system must be optimized not only internally (system design) but also externally (technology, context, and train/test data). 2.3.1 ASR-based CAPT Systems in LL Experiments Empirical experiments in LL are scarce, and almost non-existent for CAPT systems in mobile devices even though ASR-based pronunciation training has several advantages, such as dynamic evaluation, individualized feedback, more intensive practice, anxiety-free context, and opportunities for repair [50], [56]. However, erroneous feedback could be given: false positives and false alarms [57]. Four categories of ASR errors that could be used as predictors of L2 learners’ difficulties are identified in [58]: homophones, minimal pairs, breached boundaries (in the context of two or more linked words), and negative cases. In phonology, a pair of words is said to be minimal (minimal pairs) when they differ in only one segment (sound) [59]. Learners carry out virtual risks of producing wrong word meanings when the correct phonemes are not properly uttered by simply altering a single segmental element. By simply altering a single phoneme, learners risk producing unwanted meanings. The distinction between both words is, a priori, a tough task for ASR systems due to the phonetic distance between words being very small. Finally, negative cases, such as can/can’t or legal/illegal are also important for ASR since their misrecognition can thoroughly change the meaning. There are some CAPT experiments which include theoretical lessons and activities related to them. In the experiment reported by [13], four training lessons were presented to the participants. In each of them an explanatory video was shown first, and then 25 sequential and guided exercises based on the video were proposed. Those exercises could be written questions to be answered orally by recording one of several possible answers or requiring the student to pronounce specific words for which example pronunciations are given. A similar training method is applied in [60]. First, a teacher explained the position of the tongue for the phonemes /r/ and /l/. Then, the training of prolonged /r/ and /l/ was administered by using spectrographic representations with overlaid formant-tracking results. The ASR system showed, in real time, the hidden Markov models (HMM) scores obtained in the production of the minimal pairs. The experiment described in [61] used 23 exercises of increasing difficulty. Each one of them emphasized the contexts: placement test, vowel/consonant, only one word or a sentence, and anticipation. It is important to note that this software offers six different hints in order to improve incorrect pronunciations: oneself and native exposure, instructions of how to pronounce the sound, image of side headcut with the position of the lips and listening to the word in a minimal pair or sentence. In [62], an HMM ASR-based CAPT system called PLASER presents 20 lessons, teaching two phonemes in each one. The exercises in each of the lessons are: (1) read-along: no assessment; (2) minimal pair listening: ear training; (3) minimal pair speaking: produce one of the words of the pair; and (4) word list speaking: produce one of the words of the list. In [63], [64], the Talk to me software provides six dialogue sequences (each one has thirty question–and–answer screens), where the program asks a question to which the user responds by uttering one of three answers presented on the screen. There is also optional explanatory feedback of sound articulations. The difficulty level of the
28 Chapter 2. CAPT Methodologies for Second Language Learning speech recognition can be adjusted to require a looser or tighter match to the underlying models. In [65], the PARLING system is reported. This is an HMM-CAPT system that sets learners the task of making up a word-based story through word game activities and the possibility of creating their own dictionary. The methodology of word production is simple: the user’s utterance of the word is accepted/rejected by the system as the answer. It also offers the possibility of listening to a native recording of the word. In [66], the Nuance Dragon Dictation software, a speakerindependent dictation system designed for continuous speech recognition installed on students’ mobile devices, gives immediate feedback to students when they read aloud the target words and phrases in French in 20–minute pronunciation activities. In [67], students interact with the Moby.Read application. In each test session learners read (1) a word list, (2) an easy practice passage, and finally (3) three grade–level passages. After that, they are asked to retell the passage in their own words, with all the details possible and then answer two short questions aloud. Scores are provided in real-time. Finally, there are other studies that promote training methodologies over the Internet, such as peer-reviewing of read sentences after tasks of reading and speaking with peers [68], or conversations with natives or other L2 learners after the training with minimal pairs and lists of words in activities of perception and native imitation—production [69]. 2.4 Text-To-Speech in Language Learning In the early 1980s, Texas Instruments’ Speak & Spell built the first single-chip voice synthesizer for a toy –called Spelling Bee– to teach children how to spell. The literature about the TTS benefits for pedagogical applications is very limited and almost non-existent for mobile devices contexts. It was only recently that some studies confirm TTS advances seem to be ready for use in LL activities [70], [71], [72]. The use of TTS systems as part of pedagogical tools has generated a great controversy and, there is still certain debate about their suitability for L2 pronunciation training (like ASR systems) and very few attempts to empirically measure their performance. However, recent research in speech synthesis has reported some benefits in terms of comprehensibility, naturalness, accuracy, and intelligibility [70], [72], [73]. TTS systems can raise learners’ awareness of certain language features in a personalized and learner-centered way [74]. In particular, their sound quality is adequate to be used in the generation of pronunciation models of phonemes, words or sentences for helping students to improve their discrimination and production skills. TTS systems are suitable in terms of promoting some of the ideal SLA conditions presented by [75], [76], such as learner fit, authenticity, potential for providing feedback, and learning strategy development. Interested readers can find an overview about speech synthesis and the features taken into account in this thesis to choose a TTS system in Section 6.4.2. 2.4.1 TTS-based CAPT Systems in LL Experiments The importance of listening to sounds before producing them is contrasted in [77], who said that the learner’s brain converts an unclear sound into the closest sound found in L1, and suggested that emphasis should be put on listening. This importance has also been accented by [37], [78], [79]. There are scarce experiments in LL which integrate speech synthesis due to controversy produced by the limitations of TTS systems [73]. However, the quality of
2.5. Summary 29 these systems has currently increased due to an even larger corpora and new statistical parametric and DNN to process both superficial [80] and HMMs [81] information from the mentioned corpora. Traditionally, the method called High Variability Phonetic Training (HVPT) consisted on exposing learners to multiple natural voices producing the target sounds, rather than a single voice (i.e., teacher’s voice in the classroom). Then, learners only had to choose the word that they have listened to (a single task). Studies with minimal pairs [82], [83], [84] and with single words [37], [85], [84] report learners’ perception and production improvement. Nowadays, this method currently can integrate different synthetic voices. There is even less research in speech synthesis for L2 CAPT systems, since most of them, to date, have used natural-speech as a model to listen to, imitate, and self– compare [86], [87], [88], [89], [90]. Imitation procedures have also been applied to pitch and intonation suprasegmental forms in sentence production [91], [92], [93], [94] or global speech characteristics [78], [95]. Some recent methods use manipulated natural-speech recordings in order to assist the identification and discrimination of individual phonemes, improving also production [83], [96]. Regarding the few experiments existing with speech synthesis in language learning, in [72], the majority of participants who listened to speech samples, alternately produced by TTS and a human, reported that TTS technology could and should be used as a tool for LL to perform perceptual activities with both natural and synthesis speech; whereas in [97], the NaturalReader TTS software system allowed students to complete weekly pronunciation tasks in a computer. They consisted of listen–and– rank, listen–and–categorize, and listen–and–repeat sentences. They were asked to fill in reports results manually after training. 2.5 Summary CALL has significantly contributed to new changes in SLA. In particular, CAPT systems are intended to be a useful resource in pronunciation training due to the emergence of new technologies and services for smart devices, and the unceasing growth in demand of both, off-line and online L2 courses. However, there are not yet enough empirical experiments on mobile CAPT systems that resort to ASR and TTS technologies nor reports on their effectiveness. In this thesis, off-the-shelf ASR and TTS systems are incorporated into different versions of mobile CAPT systems in a non-obstructive and user-friendly way, allowing designers and experts to personalize instructions for learners, and saving time and resource costs. They also offer the possibility of assessing different language level users. In this chapter, a general overview of pronunciation teaching in current trends, training activities, and findings has been described. The reasons why ASR and TTS technologies can be integrated into personalized mobile CAPT systems have been also analyzed. Then, the set of methodological decisions of the most relevant CAPT tool of the state-of-the-art have been reviewed. In particular, an exhaustive revision of CAPT experiments of the literature has been detailed. Firstly, the evolution of speech recognition in LL has been described, pointing out the most relevant experiment with ASR-based CAPT systems. Finally, the main features of speech synthesis systems in LL have been described. The experiments in the literature about TTS in LL contexts have also been pointed out.
31 Chapter 3 Assessment of Pronunciation with CAPT Systems Replicating a scientific experiment is crucial to ensure its validity. In particular, a clear statement of the assessment methods to process the obtained results leads to an increase in their significance and confidence level. Although the number of experiments with CAPT systems based on mobile speech technology is increasing, there is not yet a common protocol respecting the evaluation of the improvement in pronunciation when using them. In particular, there are very few contributions on assessment of CAPT’s pedagogical effectiveness with speech technology, and in some instances it remains unclear. On the one hand, the pre/post-test and pre/post-quest designs are the preferred subjective method to compare different groups of learners, gather their opinions, and measure the degree of change of their perception and production skills after specific training sessions. The main limitation of this approach is the possible inconsistency of scores provided by human raters, probably originated by the lack of common guidelines, and the fatigue that such an evaluation involves. On the other hand, several quantitative and very specific metrics obtained from ASR for pronunciation assessment are proposed in each one of the limited state-of-the-art experiments. In this thesis a mixed assessment approach for pronunciation training in CAPT systems is proposed. In the first stages of the experimentation, experts provide subjective measures of learner’s utterances. Then, these scores are correlated with objective and quantitative ones obtained automatically from the CAPT tool. The higher the correlation achieved, the greater the level of confidence scoring. The aim is to be able to automatically evaluate speakers in real time when practicing with the CAPT system. This assessment can also be applied to pre-test and post-test activities, in addition to a final CAPT system’s score. Thus, future experiments can rely on automatic and objective scores provided by the technology that can serve as support when it is not possible the human help, saving time and resources. In this chapter, the subjective techniques of pronunciation’s improvement assessment after using a CAPT system are described in the first place and a revision of the literature about experiments which apply these methods is carried out. Then, a description of the metrics used to assess pronunciation quality objectively with examples of the state-of-the-art is presented. Finally, experiments that combine both approaches are reported.
32 Chapter 3. Assessment of Pronunciation with CAPT Systems 3.1 Subjective Assessment Human ratings offer a subjective approach for pronunciation activities. A 79% of the experiments surveyed in the effectiveness of L2 pronunciation instruction review in [11] based their results in this technique. A preparatory session among the human raters for sharing common aspects for evaluation is recommended since scores could be inconsistent and unclear. These individuals can be native or specialists in the L2–target language. In particular, the most usual protocol reported in the review consists in comparing the pre-test scores with the post-test ones in different groups of participants (i.e., control and experimental). In some cases, a delayed post-test was also carried out to ascertain the long-term retention. These tests generally contained the same stimuli (i.e., the same words to discriminate in perception exercises or to utter in production ones). Different rating scales were also used (i.e., right/wrong/no response, scales from 0 to 3/10/100, among others). Some works report pronunciation improvement with perception and production activities of isolated phonemes (segmental level) with minimal pairs in a pre/posttest design [82], [83] or with isolated words [88]. As for the suprasegmental level, there are some studies that, instead of using minimal pairs, use lists of words read aloud to evaluate perception skills [85], production ones [37], [86], [90], [93], [94], or both of them [89], [96]. Furthermore, there are experiments that analyze either the perception or production of sentences with numerical scales. In particular, linguistic functions of prosody (elements of speech that are properties of syllables and larger units of speech), such as marking the location of pauses, the stressed words, and the direction for sentencefinal intonation are some activities for these perception tasks [78]. Phrases productions are analyzed in [68], [93]. Specific parts of phrases are assessed in [87]. Prosody is also evaluated with spontaneous speech production tasks in [91], [92]. There are other different approaches, such as the assessment of audio recordings by other classmates [98], oral presentations by experts [95], the analysis of spontaneous conversations, also assessed by experts [69], [87], [89], and the perceptual evaluation of the regenerated audio signal before and after transformation [99]. On the one hand, there are scarce studies about pronunciation improvement measurement with ASR-based CAPT systems. The production improvement achieved in [60], [62] follows a pre-test, training sessions and post-test design with an ASRbased CAPT tool and minimal pairs. Word and sentence-level perception and production activities improvement is numerically evaluated by human raters in [66], [100]. Production improvement of spoken words is analyzed in [13], [61], [65], [97]. On the other hand, there are even less studies about TTS implication in pronunciation improvement in CAPT tools, since they are beginning to appear. In particular, a combination of natural and synthetic speech is used to listen to minimal pairs in the experiment previously mentioned in [83]. Fully synthetic sentences are presented in [97], and the pronunciation improvement is assessed with a pre-test, post-test, and delayed post-test design. Finally, websites with TTS technology promote the self-study and are reported to be useful to improve pronunciation in a pre/post-test evaluation with human raters [101].
3.2. Objective Assessment 33 3.2 Objective Assessment Despite the accuracy and preciseness of human ratings, large CAPT experiments with a great number of participants become into a tough task of assessment in terms of time and resources. A powerful alternative is to automatically generate scores with technology. This second approach of assessment consists in providing objective measures and optionally, to correlate them with the scores of human raters [11]. On the one hand, the interaction with a CAPT system can be assessed in real-time during training. For instance, a numerical score can be provided after performing perception exercises; whereas in production tasks, a text with the recognized speech and a numerical score can be shown. On the other hand, pre-test, post-test, and delayed post-test tasks can be also assessed with technology, in a similar way as the previous case, but asynchronously. Even though there is a limited literature on objective assessment of CAPT systems, it can be categorized into three categories. First, objective measures can refer to intelligibility (acceptable/unacceptable scores); second, to quality (Goodness of Pronunciation, GOP-based scores [102], [103]) and finally, to nativeness-like (i.e., pitch contours, accent ratings, ASR-based confidence scores, among others). A better correlation between human ratings and pronunciation scores at a sentence(s) level based on both, L1 and L2 language characteristics of learners instead on only L2 ones, is reported in [104]. These automatic scores are obtained on Gaussian mixture models (GMMs) log-likelihood and HMMs confidence scores. A phonelevel comparison with a likelihood-based GOP is carried out in [13], [102], [105]. The production mistakes are assessed by comparing native speech to non-native one. Regarding studies with ASR-based CAPT systems, in some experiments an automatic right/wrong assessment is given after a user’s utterance by highlighting the wrong part (Dutch-CAPT system) [13] or showing the speech recognized (Nuance Dragon system) [66] without presenting to the learner any score or quality value. Another investigations show objective scores from an HMM-based ASR software, PhonePass [63]. In a posterior study, these scores are correlated to human rater ones [64]. A large correlation between human’s and ASR scores of orally produced words in sentences is also reported in [67]. Finally, diverse ASR system outputs are adopted for the assessment of basic English vocabulary in young children [106], [107]. Scores provided are based on phoneme-level language modeling and prove they can be used to obtain good classification results, even with a relatively small amount of acoustic training data. 3.3 Summary Pronunciation assessment in CAPT systems does not follow a common pattern. Despite the scarce number of studies about this topic, there are several evaluation techniques depending on the availability of human raters, the technology employed, and the scope of training. In this chapter, a detailed revision of the literature about assessment of CAPT’s pedagogical effectiveness has been conducted. One of the main contributions of this thesis is to provide an automatic scoring method for CAPT systems with speech technology, based on user’s results. This score can be obtained during and after training. It can also save time and resources to researchers and teachers when the number of learners is considerable.
35 Chapter 4 Corrective Feedback with CAPT Systems Corrective feedback (CF) in CAPT refers to the answers provided by the system when the learners make linguistic errors in their L2 pronunciation. Giving a proper CF is key to a CAPT system to be successful in its role of helping students to improve their pronunciation and increasing the effectiveness of the learning process. It also permits a more intensive and individualized practice within an immediate and anxiety-free context. CF can be explicit if the learner is informed of the corrected form of the error, or implicit, otherwise. Although there are several studies about CF in SLA, there is not a common framework with clear guidelines to follow in the field of CAPT. Results reported show different methods, variables and definitions that lead to mixed — and not always comparable— outcomes. Several techniques are carried out, from the easiest ones, such as reducing the result to correct or incorrect, to more complex methods, such as showing spectrograms and vocal tract videos. However, in most cases the feedback offered by these tools is insufficient or too difficult to understand by the users. Besides, teachers do not work systematically, being their corrections sometimes contradictory and ambiguous. Recent advances in speech technology have allowed to provide individualized and automatic CF in CAPT tools. Care must be taken to design and adapt CF to specific pronunciation training activities with this technology in order to avoid erroneous feedback as false alarms and false accepts. In this thesis, different types of CF are integrated into the experiments carried out to analyze which ones are the most suitable for building an effective CAPT tool. The aim is to automatically offer an appropriate training activity after a learner’s erroneous input and a set of advice to overcome the specific pronunciation problem. In this chapter, an overview about CF in SLA, from its essentials and different types to the possible integration into CAPT systems is presented. In particular, a revision of the literature about current CF techniques in CAPT experiments is carried out. First, the most significant features of relevant CAPT experiments with simple or isolated CF exercises are reviewed. Finally, other CAPT software systems with more advanced CF techniques are detailed.
42 Chapter 5. Game-based Learning with CAPT Systems a well-developed game goes beyond classroom boundaries and provides an incentive for social interaction [3]. Over the past decade, games have been included into higher education research tools [2], mostly due to their effective potential, including their entertainment value when teaching a certain skill [121]. Games also contribute to build productive social practices, helping individuals to take part in learning communities [3], [122]. As pointed out by [123], learning games can be categorized as individualistic or social, depending on how players organize their efforts. The former refers to users who play alone with the system, ensuring their own learning meets a preset criterion, independently from other participants. On the other hand, social games include not only the interaction with the machine but also with other players. Social practice can be classified into collaboration (cooperation among partners to accomplish shared learning goals and maximize their own and their teammates’ achievement), competition (among competitors, trying to perform faster and more accurately than other participants), or a combination of both of them. In any case, a learning game ought to have at least the next three main characteristics according to [124]: 1. A goal. The mixture of objectives and events needed to finish the game. It must be precisely and clearly established. It is the most relevant aspect in the game since its success depends on it. Achieving a certain amount of points or number of badges could be examples of game goals. 2. Obstacles. Challenges and adversities intended to complicate the game in order to avoid triviality. For instance, limiting the number of times an exercise can be performed or enhancing the activity’s difficulty along time. 3. Competition or collaboration. Players can compete against other users or try to beat the game itself. For instance, players can realize their individualistic game scores in a common leaderboard shared with other learners (individualistic efforts in an implicit competition). Other possibility is to promote challenges among users and divide the points depending on the results (explicit competition). Users can also form groups and try to reach a shared goal together (collaboration). Social learning structures must follow certain typical aspects in order to be successful. Even though both cooperative and competitive learning strategies have common features, they can be clearly differentiated. In the case of competitivelearning scenarios there must be present at least six characteristics [123]: 1. Negative goal interdependence. If a user wins, the others must lose. 2. Perceived scarcity. Only the best players can reach rewards and achievements since their quantity is limited. 3. Interaction with other parties. It can be direct (with oppositional actions), indirect (parallel or sequential actions, turns), or nonexistent (i.e., playing individually to reach a final score). In this thesis, the concept of explicit competition is related to a turn-based indirect interaction with other subjects via challenges; whereas implicit competition refers to a nonexistent interaction among users in a common competition. 4. Quantity of winners. It varies from one, to few or many winners.
5.1. Game-based Learning 43 5. Comparability among participants. User’s performances generate public information that can be optionally reviewed by the rest of the competitors. 6. Winning rules. The criteria for determining the winner must be clear. It can be objective or subjective, depending on the tasks. On the other hand, a cooperative structure must take into account other six particular characteristics [123]: 1. Positive goal interdependence. Players must realize they can attain their goals if and only if their teammates attain theirs. It can be enhanced with positive reward interdependence (i.e., group rewards). 2. Individual accountability. Related to the individual share of the work by each teammate. Players must know the results achieved by the rest any time. 3. Intergroup cooperation. Groups can help others to finish the task successfully or compare their strategies and results. 4. Desired behaviors. Related to the specific teamwork and taskwork skills to be learnt by the players. 5. Learning task. Two aspects must be clarified: what and how must be completed the assignment/goal. 6. Criteria for success. Both the learning task and the level of performance must be cleared settled. 5.1.1 Social Learning Games Nowadays the interest in social learning games is receiving a great attention from the literature. Social learning contexts for improving learning skills can be established with digital games, providing a learning scenario that offers learning contents, and a community that allows the condition for social learning [125]. Although there is not consensus in the literature about which approach is best for social learning games, the current trend in the state-of-the-art is mainly focused on collaboration over competition. While there are some studies that mention competition is related to violent and aggressive behaviors [126], there are also others that report similar effects in both approaches [127]. The effects derived from competitive learning games are influenced by other aspects of the game, such as the content, the rest of players, the complexity of the learning, or the game configuration, among others [128]. Actually, both alternatives are valid as long as they promote pro-social outcomes [129] and reduce as much as they can negative effects among players, such as aggression and aggression-related variables [130]. It is clear that more research is needed to determine under what conditions the competition can be most beneficial. Particularly, there are some studies about L2 teaching that report the benefits of collaboration over the negative consequences of competition in games [5], [131]. Others experiments explain the implications, differences and advantages of both of them, collaboration and competition, on player’s perception [132]. In the case of CAPT-based experiments, there are a few examples of collaborative learning, and to the best of our knowledge, there are no precedents in the case of competitive scenarios. For instance, in [133] three groups of students with individualistic and collaborative efforts are compared. It reports more qualitative gains and strategies
44 Chapter 5. Game-based Learning with CAPT Systems outcomes from the Collaborative CAPT Group. In [134], both the individualistic CALL and the collaborative computer-mediated technique approaches are reported positively by the students. However, motivation in games can be associated to the term “challenge”, with positive outcomes [135]. Competitiveness, in the context of game-based learning, also helps trainees to achieve their learning goals and perceive higher ability skills [4]. There are several studies that report motivation and engagement enhancement due to competition [136], [137], [138]. In [139], the comparison between students’ scores in a perception training controlled-competitive experiment tries to motivate them to improve their own results. In [140], an educational mathematics game shows the positive effect of competition in a collaborative learning situation for above–average students. Individualistic and competitive approaches are also directly compared in some experiments of the literature. A high enjoyment, future play motivation, and high physical intensity are reported in [141], thanks to competitive scenarios since parallel competition in separate physical spaces overcomes individual gaming. In [142], positive experiences when a competitive context is provided to competitive individuals are described; whereas detrimental effects are reported for the less competitive ones. In [143], it is reported that the competitive configuration in learning approaches requires further research since better results (but not statistically significant) were achieved by the users of the non-competitive condition in comparison to the competitive one. In [132], the players of a cooperative setting experienced greater enjoyment than those in a competitive configuration. 5.2 Gamification and L2 Pronunciation Training Gamification is defined as the use of game design elements in non-game contexts so as to enhance participant engagement and encourage desired behaviors with a product or service [144]. Gamification is also intended to reduce abandonment by designing attractive applications that generate pleasant and beneficial affection [145]. In the particular field of game-based methods and strategies for learning contexts, gamification uses game-based mechanics, aesthetics, and game thinking to engage individuals, promote learning, and solve problems [146]. Besides, educational gamification helps individuals to be immersed in learning, improves their motivation, and brings them playfulness [147]. The first commercial educational language learning tools, such as Sanako1and Rosetta Stone2appeared in the 1990s. It is not until the second decade of the twentieth century when the modern online applications appear (mobile, web and desktop), such as Duolingo3, Busuu4, or Babbel5. The former mentioned group of tools are pronunciation training services with methods and strategies based on content choice, focused on either self-training or academic institution solutions. However, 1http://www.sanako.com/ 2http://www.rosettastone.com 3http://www.duolingo.com/ 4https://www.busuu.com 5https://babbel.com
5.2. Gamification and L2 Pronunciation Training 45 the latter group of pronunciation training applications have changed the paradigm by including gamification elements with an evident intention to improve the user’s learning experience [148]. Although the number of applications with gamification elements intended for L2 pronunciation training is increasing, there are scarce experimental studies that validate their effectiveness. For instance, [149] describes a card-based game for L2 vocabulary acquisition with speech technology. In [65] a word game based on stories for children with an ASR system, PARLING, is presented. In [150], the Polish language is taught as a user role-based game, in which its complex grammar system and lexical problems are presented to learners as tasks and activities. In [118], a CAPT game for practicing Dutch oral and grammar skills based on speaking practice and feedback is reported. There are other innovative ways of pronunciation teaching with gamification, such as a multi-language karaoke application, SLIONS [151] or a recursive dialogue game for a personalized pronunciation training [152]. The previously mentioned studies and applications share game design elements which can be classified into [153]: 1. Points are the most extended and basic elements in language learning tools. They consist on a numerical representation of the result of an activity performed by a learner. Their scale can vary from complex ranges of numbers to simple binary outputs reporting user’s success or failure in the activity. In particular for pronunciation training, this important resource could be used to assess the goodness of the user interaction and to measure the user proficiency [150], [154], [155]. For instance, in [151], an overall score (0 to 100) from a user’s karaoke performance is shown to the learner. Besides, points can involve the accomplishment of user’s levels and badges and can lead to earning rewards in Duolingo and Babbel. 2. Badges display symbols or messages that represent user achievements. They are intended not only for informing user’s about their performance but also to motivate them to keep on training [149], [151], [152]. For instance, in [150], "The Lord of Memory" badge is given to a student who remembered most of new words of previous classes. In Busuu, individuals earns several different badges after concluding courses or talking tests. In Duolingo, users can earn badges after completing 10, 50, and 100 lessons, or 5, 10, and 30 skills, in addition to an extra incentive for making progress with the lessons, among others. 3. Leaderboards show a ranking that permits users to compare their relative success as regards the performance of other players [153]. This competitive indicator of progress allows users to contrast their own level regarding other learners, contributing to assemble a self conscience of level. It is also interesting because it also permits to configure competitions as a mean to promote social interaction between users. This game element is commonly used in social language learning applications, and almost non-existent in the state-of-the-art about pronunciation training studies with CAPT. For instance, in Duolingo and Babbel, progress is measured in terms of gaming-like elements by gaining experience (XP points), and leveling up, which affect to their social leaderboards.
46 Chapter 5. Game-based Learning with CAPT Systems 4. Performance graphs provide information about the player own progress over time. The difference with leaderboards is that, in this case, performance graphs do not compare the user’s performance to other players. Thus, it is an individual reference indicator instead of a social one. In particular, this resource is relevant for both students and teachers, since it can trace all right and wrong interactions with the system and can lead to personalize user’s training content based on the results achieved [152]. For instance, Sanako provides a complete dashboard for teachers to follow up the progress of the students in the language laboratory. In Duolingo and Babbel, graph statistics and historical records are available to users. In [118], a final report about all mistakes made by the learner is given after each conversation. 5. Meaningful stories stand for the narrative in which the gamified activities and characters are included in. They give a meaning far beyond the only purpose of achieving points and badges [146]. In particular, they could be used in language learning applications to connect different activities in the same flow of exercises. For instance, in [65], the word-based pronunciation training system, PARLING, displays well-known children’s stories. In [156], personalized activities in significant 3-D virtual environments are shown. Duolingo offers activities related to particular stories which can change with game progress, presented by an owl character, Duo. 6. Avatars represent players in the game and allow them to communicate themselves as well as objects within the environment. They vary from simply approaches as pictograms, to more complex ones, such as three-dimensional (3D) representations. They are intended to enrich the user’s experience during games. They are widely used in language learning games in virtual environments. For instance, in [150] a warrior-based avatar represents each group of students in the game. In [156], 3-D human-based avatars can interact with the whole environment. In [157], human-based avatars permits also more than two people to be involved in the conversations. 7. Teammates (teamplayers) are the rest of players in the game. They could be real or virtual ones, and can lead to conflict, competition or collaboration [146]. In particular, L2 learning applications with gamification elements must help users to overcome not only their own communication barriers but also to compete with other players [150]. For instance, Duolingo encourages users to compete with their friends to see who learns faster. In [150], each learner has a role in the game, and they form small groups, affecting other player’s actions. In [152], learners interact with virtual and simulated players in dialogues. The impact of these game elements on autonomy, psychological needs of competence, and social relatedness is analyzed in [158]. Some of them, such as points and leaderboards can be considered as extrinsic incentives for promoting performance on image tag task [159]. The design of effective leaderboards based on individuals preferences is analyzed in [160], which concludes that competition is a media rather than purpose when there are leaderboards in gamified applications.
5.3. Summary 47 5.3 Summary E-learning can be defined as any kind of learning or development content administered in a digital way. Recent technological advances have allowed researchers to integrate CAPT systems in online learning contexts, allowing users to learn anytime anywhere while keeping motivated. In this thesis several motivational elements are included in the CAPT prototypes of the experimentation. In particular, game approaches are carried out by means of prototypes of CAPT learning games for pronunciation training. Therefore, a large amount of data is automatically gathered susceptible to be in a speech corpus.
49 Part III Experimental Procedure
51 Chapter 6 Experimental Framework Developing an effective L2 CAPT tool for learners of a particular native L1 language requires the design of a proper set of training activities and corrective feedback techniques, along with a reproducible objective method of assessment, and an appropriate choice of speech technology. The purpose of this chapter is to present the dimensions of the experimental procedure, demonstrating an understanding and applicability of the specific concepts, theories, and technologies which are relevant and necessary for the experimentation. Firstly, the importance of minimal pairs in pronunciation training is described. A novel protocol for elaborating minimal pairs lists taking into account the speech technology integrated in the CAPT tool is also detailed. Secondly, the essentials of the training activities cycle included into the experimental prototypes are specified. Third, the strategies carried out to assess user’s pronunciation improvement with the CAPT tools are presented. Fourth, the fundamental principles and preferable characteristics of ASR and TTS systems are described. The selection of some state-of-the-art ASR and TTS technologies is also motivated, including a general outline for building a personalized ASR included in the CAPT tools developed. Then, the main feedback strategies adopted for each one of the experiments are commented. Sixth, the gamification elements included in the experimentation are mentioned. Finally, the aspects taken into account to select participants to carry out all the experiments presented in this work are mentioned and which will be discussed in Chapter 7. 6.1 Minimal Pairs After a review of the available methodological approaches to L2 teaching (see Chapter 2), the NCM has been partially followed as the basis for the design of learning activities in the CAPT systems we developed for the different experiments presented in this part of the thesis (see Section 2.1.1 for more details). As a reminder, the main aspect of this approach is the use of minimal pairs in a specific cycle of pronunciation activities, taking into account the learner’s L1 and L2. The concepts and methods of these approaches have been extrapolated to the experimental work to several languages, such as English and Spanish. Two different general approaches will be experimented for users’ guidance. First, a gamified playing methodology with free selection of activities in a learning game. Second, a guided and controlled pedagogical methodology with recommended activities of feedback based on user’s results that lead to a common end. Both approaches integrate current speech technology adapted and personalized to the CAPT systems.
58 Chapter 6. Experimental Framework prediction of the ASR system is composed of a list of n–best possible text hypotheses (nis adjusted in each experiment), ordered from highest to lowest confidence rates. In our case, these hypotheses are words and each possible word is followed by a numeric value (g–score), which is proportional to the reliability of the prediction (from 0% to 100% in a scale [0, 1]) although there is no documentation available on the specific meaning or interpretation of this score as a likelihood or similar. Thus, the utterance is considered correct as long as it is within the list of nelements returned by the ASR. For instance, an ideal utterance of the word mass, produced by a native American speaker would be associated to a g–score = 1.0 and a 5–list of strings as follows: "mass", "Mass", "masse", "masts", "mass.". The sequence of maximum wrong production attempts per word is adjusted in each experiment. Furthermore, synthesized models of the words are available to learners, who can play them as many times as they need while in production activities (requested listening) in order to improve their self-perception of the correct pronunciation to be obtained. The maximum number of listening attempts per word is adjusted in each experiment. 6.2.5 Mixed Activities This type of exercises consists of mixing up both, perception and production activities, being intrinsically more difficult. The strength of the relationship between them may vary according to the proficiency of the speaker-listener and the target (L2) sounds [36]. In particular, in the mixed activities both kind of activities (discrimination and production) are sequentially interleaved. While in each one of isolated discrimination and/or production activities users can fully concentrate on these tasks individually, the mixed activities represent a situation closer to real communication, where readiness both to understand and produce language must coalesce. From a methodological point of view, they convey an extra difficulty and are included it as a way to review and test both modes at once as well as the ability to shift between them, simulating a real conversation. 6.2.6 Selection of Activities A different set of activities is presented in each experiment depending on several factors: activities freedom of choice, activity goal, and progression along time. In terms of freedom to select activities, they can be restricted by the system (guidedbased) or freely selected by users (free will). The former case refers to activities that belong to a previously defined by experts controlled protocol that leads to a common end. In particular, the system recommends a specific type of activity based on user’s results. On the other hand, educational game-based prototypes of the experiments presented in this thesis give learners the freedom to choose the pronunciation activities. In some cases, the number of times the activities can be performed is limited. Activity goal includes either training or playing. An specific target sound and any type of activity can be chosen in the first case in order to let learners to train at their own peace. However, in playing activities in which there is a final reward and participates other subjects, these elements are not possible to be selected. The last aspect to take into account is progression over time. The content and difficulty of the activities change (locking/unlocking and decreasing/increasing, respectively) according to the evolution of user’s results and success rate values.
6.3. Assessment 59 6.3 Assessment The ultimate goal of pronunciation assessment is to generate a score for a nonnative speech utterance automatically and obtain results comparable to a human teacher/expert. As pointed out in Chapter 3, there is not a standard protocol for assessing user’s pronunciation improvement with mobile CAPT tools and speech technology. From a realistic and experimental point of view, two kinds of assessment strategies can be combined, which are usually classified into objective and subjective categories. The goal of a mixed approach is to help teachers and students to save time and resources by obtaining an objective score about learner’s performance with the system. The next subsections describe the potential data sources used in each experiment for the assessment of user’s pronunciation improvement (see a comparison in Table B.3 in Appendix B). 6.3.1 Subjective Assessment In the experiments carried out in this thesis, three different subjective approaches have been applied in which both students and educators take part (see Section 3.1 for more details and references): •Perceptual tests. Pronunciation at segmental level quality can be characterized by the perceptual parameters identified by human experts, while evaluation is the distance between the target and the reference phoneme characteristics. Giving scores can be either written manually or typed with a computer via a personalized web page. Raters must apply the same rules to score the utterances. These tests can be carried out at the beginning (pre-test), at the middle (middle-test), at the end (post-test), and some days or months after the experiment (delayed post-test). •Questionnaires. They provide quantitative data from learners during different stages of each experiment (i.e., pre-quest and post-quest). The users’ demographics, their opinion about the system’s experience (UX), the grade of motivation, and the reasons for participating or abandoning in the experiment are examples of topics included in these questionnaires. •Focus groups. This qualitative technique is carried out at the end of the experiment. A set of predetermined questions is asked to participants in a planned discussion, while the moderators take notes. Also, the whole session is recorded via audio or video with the participants’ consent. Individuals can be classified by common features into different focus group sessions. Apart from questionnaires’ information, focus groups allow to obtain extra relevant information in an alternative way. 6.3.2 Objective Assessment Giving an objective assessment about user’s performance helps learners and educators to keep track of the evolution of pronunciation improvement along time while saving time and resources. However, technology must be specifically adapted to the task since false alarms and false accepts can appear. Also, objective scores should be correlated to expert human scores in order to ensure their validity.
60 Chapter 6. Experimental Framework The CAPT systems developed in this thesis resort to ASR to obtain binary (right or wrong) ratings of word pronunciation. In this thesis, two instruments are employed to provide objective assessment (see Section 3.2 for more details and references): •Objective tests. User’s utterances of pre/middle/post/delayed post-tests can be evaluated not only by experts but also by speech recognition. It must be specified the objective of the evaluation (i.e., the whole utterance, a specific syllable, or the target phoneme). •Game scores. User’s interaction with the system is kept into log files. This information includes user’s activity results and several metrics can be defined. Learners realize about their performance and the system can personalize the activities content in function of this assessment. For instance, in discrimination activities, a right/wrong answer with numerical score is provided to the user and in the production ones is also shown a message with the most probable text spoken by the learner, evaluated by an ASR. Also, a game score at different stages can be correlated to the subjective and objective scores of the perceptual tests. 6.4 Speech and Software Technologies 6.4.1 Automatic Speech Recognition Automatic speech recognition (ASR) is the use of computer hardware and softwarebased techniques to identify and process human voice [165]. It is also known as automatic voice recognition (AVR), voice-to-text, speech-to-text, or simply speech recognition. ASR is not a simple task since there are different factors that affect it directly, such as the age and accent of the speaker, the codec used for the audio and compression artefacts, the sample rate, the background noise, the length of silences, the reverberation from varying the acoustic environment, and the artefacts from the hardware. As shown in Figure 6.2, ASR converts the user’s speech input (O) into a sequence of text hypotheses (W). ASR O: observable (unknown speech signal) W: text string (word sequence) FIGURE 6.2: Conceptual ASR system. The basic functionality of ASR is similar to any pattern recognition system, that is, some models are trained to subsequently recognize speech [166]. Two main model types can be distinguished:
6.4. Speech and Software Technologies 61 •Acoustic models represent the relationship between an audio signal and the phonemes (or other linguistic units). That is, they are the statistical representations of a phoneme’s acoustic information. In most cases they can be considered as task-independent models. •Language models assign probabilities to combining acoustic models in order to form sentences (list of words). Generally, they involve task-dependency. Alexicon (pronunciation model) serves as the link connecting the acoustic models and the language models. It is typically handcrafted by expert humans (a costly and time-consuming process). The uses of ASR are not limited to CALL. There are several practical applications of ASR systems such as to identify the words a person has spoken (i.e., dictations, voice commands, etc.), to provide information and to forward telephone calls, to help people with disabilities (i.e., fluidity or transmission of a conversation to a person with hearing problems) and to authenticate the identity of the person speaking into the system. In this thesis, ASR technology is included as an element of assessment and feedback for speech pronunciation. Choosing between an existing commercial off-the-shelf ASR system or a custom made and personalized one, is a crucial decision to obtain better or worse results in the specific domain of the problem to resolve. Sometimes, commercial ASR systems, such as Google ASR, Nuance Dragon, or Human VoiceBase are suitable for specific tasks, such as voice commands, telephone calls, or dictation assignments. However, a free and open-source ASR system as Kaldi could be better for specific educational purposes. In this thesis, both ASR types are incorporated in the experiments. Several features have been taken into account in order to select the ASR solution to be included in our CAPT tools [165]: •Accuracy: testing the ASR system with experts for its suitability before experimenting. •Confidence measures. The majority of ASR systems provide scores produced by extracting confidence features from the computation of hypotheses at the phonetic, word, and utterance level. Then, these features are processed using an accept/reject classifier for these hypotheses. They can be combined to linguistic scores and pragmatic constraints to offer to the speaker a corrective feedback. •Continuity: determines whether the system can recognize continuous speech or a pause between word and word must be forced. •Custom vocabulary. The accuracy will be higher if the target set of words is closed. •Documentation. Knowing the possibilities offered by the system would help to personalize and reach the objectives. •Environment robustness. Audio noise, stress, and sample rate are the most common factors that affect ASR systems performance. •Languages. Each language, accents, and dialect variants must be trained separately. It depends on the unit used to build the ASR models. For instance, languages with common phonemes can share some models if they are phonemebased.
62 Chapter 6. Experimental Framework •Learning curve: difficulty of understanding the ASR system design to integrate and personalize it (time and resources). •Price. The current trend in commercial ASR systems is to pay in transactions per unit of time instead of buying a complete ASR system. Other possibility is to use a reduced version of their capabilities for a limited period of time. Open-source systems are for free. •Word level timing: enables accurate linking to audio segments and helps enable comparison/merging of transcripts from multiple sources (i.e., taking punctuation from one transcript and applying it to another). Google ASR Technology Google’s speech recognition2is a general-purpose and commercial off-the-shelf service available for more than 120 languages and variants. It combines the power of cloud-based computing with the latest technology. Besides, thanks to the data gathered from millions of users using all software applications of the company, Google have improved the accuracy of their machine learning algorithms for achieving better results. To the best of our knowledge, the integration of the use of a generalpurpose ASR into a CAPT tool, such as Google ASR, constitutes a novelty in the field of pronunciation training. Initially, the first ASR launched by the company was the Google Voice Search3 application in 2008. Exceptional improvements on the accuracy levels of previous speech recognition technologies were reported. Then, Google introduced elements of personalization into its voice search results, and used this data to develop its Hummingbird algorithm for the Google Now application in 2013. It arrived at a much more nuanced understanding of language in use. From 2016, the company released Google Assistant, an artificial intelligence-powered virtual assistant able to purchase products, send money, identify objects and songs, search the Internet, schedule events and alarms, among others. Although originally Google ASR was conceived as a smartphone application, nowadays it is available in other fields, such as driving systems, home virtual assistants and security systems. All Android devices can incorporate this ASR system for free since Google is the proprietary of this smartphone operating system. The Google ASR system provides a hierarchically ordered n–best (nis defined by the user) list of probable sequence of words to match the input and a likelihood score (g–score) for each one. The main functionality is very simple: the user speaks and the system gives a string with the possible candidates in order, with the g–scores. Although its great performance and adaptability for developing purposes, the free version of the Google ASR system works as a black-box system. It does not allow to keep the recorded audio and the documentation is very limited. However, at the end of summer of 2017 Google launched the beta version of a non-free Google Cloud Speech-to-Text service (GCSTT) and released its stable version (1.0) at the beginning of the 2018. This online product consists in a speech API which allows researchers to customize the ASR capabilities to a particular domain of the problem. Nine main features of GCSTT can be pointed out2: 2https://cloud.google.com/speech-to-text 3https://play.google.com/store/apps/details?id=com.google.android. googlequicksearchbox
6.4. Speech and Software Technologies 63 1. Automatic punctuation: machine learning techniques grant to punctuate transcriptions accurately (i.e., periods, commas, and question marks). 2. Global vocabulary: a large words glossary of 120 languages and variants are supported. 3. Inappropriate content filtering: inappropriate text results can be filtered for some languages. 4. Model selection: four pre-built models are available: default, phone call, voice commands & search, and video transcription. 5. Multichannel recognition: audio recordings with two or more channels in which each speaker is in a channel (i.e., video conference or phone call) can be automatically separated and transcribed. 6. Noise robustness: audio recordings do not need to be pre-processed to handle noise problems. 7. Phrase hints: a set of words and phrases that are likely to be spoken can be given to the system to personalize and improve the results (i.e., custom words and names and voice-control use cases). 8. Real-time streaming or prerecorded audio support. Unlike the free ASR application, GCSTT allows users to keep the recorded file after recognition. It also provides a non-free platform for storing the audio files. Several audio encodings are supported, such as AMR, FLAC, and LINEAR16, among others. Speech recognition can be performed in three different ways: (a) Synchronous recognition: short (one minute length maximum) prerecorded audio samples can be processed in minimal time rates. (b) Asynchronous recognition: audio samples up to 180 minutes can be sent to be processed. Results can be periodically polled. (c) Streaming recognition: intended to process in real-time audio from a microphone, giving results while audio is being captured. It allows results to appear, for instance, while the user is still speaking. 9. Speaker diarization: automatic speaker identification is also possible. Although a likelihood score (g–score) is also given with each candidate sequence of words in an ordered n–best list, the official documentation warns researchers that ”This field is not guaranteed to be accurate and users should not rely on it to be always provided”4. Consequently, a rigorous process of adaptation of Google ASR and GCSTT service to the experiment prototypes has been required in order to maximize its scoring and diagnostic reliability (see Section 6.1.2). Kaldi Speech Recognition System In this thesis, the Kaldi framework is used for building a personalized ASR with different configurations and for analyzing comparative results with the generalpurpose Google ASR system (see a guide to elaborate an ASR system with Kaldi in Appendix E). Six main features of Kaldi can be pointed out [167]: 4https://cloud.google.com/speech-to-text/docs/reference/rest/v1/speech/recognize
64 Chapter 6. Experimental Framework 1. Complete recipes. They are intended to build speech recognition systems with broadly accessible databases, such as those supplied by the Linguistic Data Consortium (LDC)5: the Wall Street Journal Corpus, the Fisher-English Corpus, TIMIT, and more. Besides, they can serve as a template for training acoustic models on your own speech data. 2. Extensible design. Kaldi’s algorithms are generic. The majority of functionalities are based on interfaces that allow to customize the code operations. 3. Extensive linear algebra support: Kaldi supports standard BLAS6and LAPACK7routines with a custom-built matrix library. 4. Integration with FSTs. The OpenFST toolkit is included as a library. 5. Open license. The code is available in GitHub and licensed under the permissive free software license Apache v2.0. 6. Thorough testing. Detailed and careful test routines are included in the majority of the source code. 6.4.2 Text-to-speech Text-to-speech is a form of speech synthesis (artificial production of human speech) that converts text (input) into spoken voice (output) [168]. While voice response systems synthesize speech by concatenating sentences from a database of prerecorded words into fixed and invariable messages, TTS systems form sentences/phrases from scratch based on language’s phonemes and graphemes [169]. In fact, TTS systems are theoretically capable of "reading" any string of word sequence to form original sentences. As a general outline, three different stages can be distinguished in speech synthesis (see Figure 6.3). First, text-to-phoneme conversion, in which the text (rules, restrictions and dictionaries about sentences, words, phonemes, accents and stops, among others) is trained, analyzed, and processed. Second, the prosody modelling that includes intonation, rhythm, and intensity. A more natural and pleasant result for the user is achieved in this stage. Generally, this stage is diffusely shared between the text analysis module and signal generation one. Finally, the phoneme-to-speech conversion module, where the output signal speech is generated. It is based on the acoustic models and/or small units of pre-recorded wave-forms (i.e., concatenative synthesis, HMM-GMM based, DNN, hybrid approaches...). On the whole, the essence of TTS system design is to find a balance between flexibility, quality, and data. 5https://www.ldc.upenn.edu/ 6http://www.netlib.org/blas/ 7http://www.netlib.org/lapack/
6.4. Speech and Software Technologies 65 Text Text analysis Signal generation Prosody prediction Speech Rules, restrictions, dictionaries... Prosodic models Acoustic models, pre-recorded samples... FIGURE 6.3: Generic architecture of a TTS system. Giving more details about the design and building process of a TTS system exceeds the limits of this work. Interested readers can find an excellent revision of text-to-speech in [169] and details of speech prosody in speech synthesis in [170]. In this thesis, several features have been taken into account in order to select the TTS engine to be in our CAPT tools: 1. Customization: limitations about the number of words or languages and the possibility of changing speech characteristics, such as pitch and rate. 2. Flexibility: possibility of using vocabulary and phrases not employed for training during synthesis time. For instance, answering to unknown situations (i.e., a dialogue with the user or a very technical text). Also, it could refer to the adaptation to different voices or speech styles (i.e., reading a tale or a newspaper). 3. Intelligibility: quality of the audio generated. 4. Naturalness: human speech similarity. It is not required in all cases. 5. Price. Most of current off-the-shelf TTS systems are for free. Although it is possible to buy specific a whole TTS system, modern systems charge per transaction. 6. Quality: absence of noise and discontinuities, among others. 7. Similarity to the original voice. The TTS must capture the key features of human speakers successfully. Google TTS technology In the particular case of this thesis, the Google TTS8Android application has been employed for the experimentation. Seven main features of Google TTS motivate this election [171]: 1. Audio format flexibility. The audio can be generated in MP3, or LINEAR16, among others. 2. Audio profiles. The type of speaker from which the speech is intended to play can be selected (i.e., headphones or phone lines). 3. Multilingual. 180 voices and more than 30 languages and variants are supported. 8https://play.google.com/store/apps/details?id=com.google.android.tts
66 Chapter 6. Experimental Framework 4. Pitch and speaking rate tuning. Up to 20 semitones more than the default output and up to 4x faster speaking rates are available. 5. Text and Speech Synthesis Markup Language support. Pronunciation instructions can be specified, such as numbers, pauses, and date and time formatting. 6. Volume gain control. The output volume can be adjusted from -96db to 16db. 7. WaveNet voices. They provide sounds "more natural than other TTS systems" [7]. 6.4.3 Software Development In this thesis, an incremental and iterative development of the engineering methodology has been followed [172]. This methodology consists in developing initial versions of the CAPT tools, also called prototypes, in which successive improvements are applied, improving their quality of until the final version. The main phases in this methodology are planning, analysis and design, implementation, testing, deployment, and evaluation. Furthermore, a modular software design has also been implemented following a version control system (Git), in which the software code has been reused in the next versions of the CAPT tools (see Table B.1 in Appendix B for an estimation of the number of software development days in this thesis). FIGURE 6.4: Client–server model of the prototypes of this thesis, adapted from [31]. All the CAPT tools developed in this thesis have in common a client–server architecture (see Figure 6.4). The client is an Android device (version 4.4 or higher) in which the CAPT tool is installed; whereas the server part is divided into two components: external (Google) and personalized services (private own web server). The interaction results between the user and the CAPT tool in the client are saved as a JSON format in log files that compile all possible depersonalized data diachronically. They are sent to our web logger in the server part automatically. These log files contain the same structure of data fields for each experiment, including new fields when
6.5. Corrective Feedback Mechanisms 67 necessary. The audio files are also sent to the web server. The lists of minimal pairs elaborated by experts are defined in a file (JSONWordsDababase) which includes in each line the orthographic transcription, phonetic transcription, and the possible homophone words of the minimal pairs. The Android client must have access to the ASR and TTS technology integrated in the CAPT tool. Finally, the game versions of the tool use the Google Play Games platform which provides gaming service and software development kits of ready-to-use game features in the software applications. This platform is complemented by our CAPTManager component, specifically developed for our CAPT tools. Some of its functionalities are the user’s log-in and sign-up, and the scoring system, among others. For each one of the experiments, a CAPT tool for smart devices has been developed (see the download links in Section 9.3.7). These tools can be also run in desktop PCs by means of Android emulators. The iOS operating system was out of scope of this thesis. In particular, the CAPT tools have been developed with Android Studio version 3.0, the software development kit for Android version 26, and Java version 6. Six academic projects of the University of Valladolid have been related directly [173], [174], [175] or indirectly [176], [177], [178] with the experimental design of the tools, and other three more are currently being carried out. The particular characteristics of the technology employed during the experimentation taken into account have been (see a comparison in Table B.4 in Appendix B): 1. Smart devices that support Android version 4.4. or higher with full access to the Internet and a minimum system’s storage of 500 megabytes. The prototypes are installed in this instrument. 2. Device’s OS: Android version 4.4. or higher; or Windows 7 or higher (with NOX App Player 5.0.0 or higher9support). 3. Data server system: Linux standard distributions 2.6 or higher, derived from GNU/Linux; Windows Server platforms 2008 r2 or higher; or Mac OS X 10.6 or higher. Minimum system’s storage of 10 gigabytes. The server aims at gathering all statistical data of user’s interaction with the system, to provide some services through the Internet, and to keep audio files when needed. 6.5 Corrective Feedback Mechanisms In this section the main relation between the CF mechanisms included in the prototypes of the experiments presented in this thesis and those in the literature is pointed out (see a comparison in Table B.8 in Appendix B). The different CF strategies followed in each experiment of this thesis are enumerated at the end of this section (the definition of each one of them is available at Chapter 4). First, in all the experiments carried out in this thesis written visual feedback is provided to users, not only via the orthographic representation (like the majority of studies reported in Section 4.2) but also via the phonetic transcription of the words following the International Phonetic Alphabet (IPA) [179]. Explicit feedback is given to the learners in the discrimination exercises likewise [13], [37], [82], [85]. That is, the chosen word is highlighted in green color with a sound and a message of success. When the answer is wrong, it is highlighted in 9https://en.bignox.com/
74 Chapter 7. Experiments Exp. 1 Exp. 2 Exp. 3 Exp. 4 Minimal Pairs TipTopTalk! English Vowels Japañol COP Methodology Minimal pairs X X X X X Production activities X X X X X Training activities X X X X X Discrimination activities X X X X Exposure activities X X X X Mixed activities X X X X Guided protocol X X Theoretical-practical video X X Technology Log files X X X X X ASR X X X X X TTS X X X X X Audio recordings X X Game instruments Leaderboard (game points) X X Competition X X Badges (trophies) X X Challenges X Limited activity selection X Assessment Focus group X X Pre/Post-tests X X Questionnaires X X TABLE 7.1: Main elements included in each prototype of the experimentation classified by categories. Finally, a guided training protocol is the pedagogical approach followed in the English Vowels and Japañol prototypes. Audiovisual materials with theoretical explanations about the nature of the phonemes within each minimal pair and practical illustrations of the sounds of these phonemes are included, and a pre/post-test strategy for assessing user’s pronunciation improvement before and after training is followed. 7.2 Alpha Experiment Alpha is the name of the first experiment carried out, and a mobile learning application was developed from scratch as a means to provide L2–English pronunciation assessment for native Spanish speakers, named Minimal Pairs. It contained a minimal pairs set selected by a phonetics expert which users had to confront by producing their word utterances with general-purpose speech technology integrated into the learning application. This experiment was a starting point for discovering the weaknesses and limitations of current ASR and TTS technologies. Categorizing speakers by assessing their English pronunciation level with the help of ASR and TTS technology, from basic to native level, was also possible.
7.2. Alpha Experiment 75 7.2.1 Experimental Procedure The recruitment campaign lasted for 7 days. Learners who agreed to participate followed a one-session training protocol, as shown in Figure 7.2. All subjects (see Section 7.2.3 for more details about participants) were asked to perform the same pronunciation training activities with the system during a maximum time of 7 minutes, individually. The training session was carried out in a quiet testing room that contained a comfortable chair and a small table with a tablet in which the CAPT system software was installed. Before starting the session, a member of the research team gave instructions of use to participants. Then, each speaker had to perform the activities proposed by the system. All the interaction events, timestamped, and the results of the ASR engine for each attempted word were stored into log files for later analysis. Furthermore, a one-hour focus group session with some randomlyselected non-native participants of the experiment was conducted. Training Subjects Data Log files Group A Group B Analysis Report Group C Log files Log files Focus group FIGURE 7.2: Steps of the first experiment’s protocol. 7.2.2 Enrollment There were three different recruitment campaigns for this experiment. First, a group of American native learners of Spanish of the same course at the Language Center of Valladolid were asked to participate voluntarily by attending to their classroom. Their aim was to test the feasibility of the system. Second, students from the English philology degree of the University of Valladolid were invited to take part in the experiment via email. Finally, students from the Computer Engineering degree of the University of Valladolid were also asked to participate via invitation emails. These two groups of Spanish students were the main subjects for this study, while the American students were incorporated to test speech technology adequacy. Students filled in an agreement and a registration form with their demographic information. A specific time slot was reserved for each participant to perform the training activities individually. All participants were awarded with a diploma after
76 Chapter 7. Experiments completing the protocol, and a reward was given to those who also attended the focus group. 7.2.3 Participants Users were divided into three different groups, according to their English (en_US) pronunciation proficiency level: 1. Group A: 12 native American speakers between 18 and 26 years old. 5 were women and 7 were men. They all belong to the same L2 Spanish course at the Language Center of the University of Valladolid. 2. Group B: 21 undergraduate students of English Philology at the University of Valladolid between 18 and 26 years old. 11 were women and 10 were men. They all claim a C1–C2 English proficiency level as L2 and have passed the same advance English phonetic course. 3. Group C: 20 Computer Engineering students from the University of Valladolid between 18 and 26 years old. 6 were women and 14 were men. Their English proficiency level as L2 was lower than the rest of participants (B1–B2). All participants took part in the same testing activities of the experiment. It was expected group A achieved the best results, and group C the worst ones. Some randomly selected speakers of Group B and Group C took part in the focus group session. 7.2.4 Minimal Pairs CAPT System Description The training activities were performed by the speakers using a CAPT system developed from scratch, called Minimal Pairs. It is a software tool for smart devices which presents twelve American English minimal pairs of vowel and consonant contrasts. They are randomly chosen from a set of twenty difficult pairs for Spanish speakers selected by a phonetics expert. Participants had to produce the words correctly, so that the words were recognized by Google ASR. A production attempt was considered correct (right) when the orthographic transcription of the word (or some homophone) was included in one of the first five positions of the text hypotheses of the ASR result. Five attempts maximum were allowed. Participants could also listen to the synthesized form of the words (Google TTS) as feedback. Data related to user-system interaction was gathered via log files. Figure 7.3 shows a screenshot of the Graphical User Interface (GUI) of the CAPT system. For each minimal pair of the training session, both words are shown close to a representative image of each one. There is a button on the right side of each picture to listen to the synthesized word. The remaining time, the number of attempts per word, the number of remaining pairs, and the number of correct/wrong utterances are also displayed. There are three buttons at the top of the figure to go back to the previous pair, to go forward to the next one and to finish the training session. Instructions are written at the bottom of the screen. Figure 7.4 shows the result of a user interacting with the system with a minimal pair (the final state of the Figure 7.3). The speaker has correctly uttered the first word of the minimal pair and the interface has changed its main color to green, disabling the possibility of producing again the word. On the other hand, the second word
7.2. Alpha Experiment 77 FIGURE 7.3: Production activity GUI of the Minimal Pairs prototype (before). FIGURE 7.4: Production activity GUI of the Minimal Pairs prototype (after). has been disabled because the user has not been able to produce it correctly in five attempts. In this case, the main color of the word interface is changed to red. 7.2.5 Instruments and Metrics There were three different sources of data: •Registration forms: user’s demographic information, such as name, age, gender, L1, academic level, and final consent to analyze all gathered data. This information was carefully collected and saved into physical text documents. •User’s interaction log files. The CAPT tool gathered data associated with all low-level interaction events and monitored all user activities. This data was saved into local log files and automatically uploaded to a web server. From these, a set of experimental variables were identified and computed: 1. Training intensity. Which computed the amount of events tracked in the experiment. It was derived from the number of discrimination and production tasks; the number of times a particular phoneme was practiced; and the times a word was listened to.
78 Chapter 7. Experiments 2. Training performance. Which measured the number of successful answers obtained by the participant for each tracked activity during a specific time. The variable encompassed right and wrong discrimination tasks; right and wrong pronunciation tasks (with the n–best list of hypotheses and g–score values, see Section 6.2.4); success rates in discrimination and production tasks per phoneme; and time spent on performing training events. •Focus group session: The audio of the session was recorded via a camera and the most important opinions and requests of the participants are written by a member of the research team by taking notes. This meeting was carried out in a classroom with all participants face to face. One member of the research group conducted the session while other took notes and recorded the audio of the meeting. Firstly, an overview of the results obtained during the experiment was presented to the participants during 15 minutes. Then, subjects were asked about their perceptions and opinions about the experiment with the possibility of discussion with other participants (30 minutes). The last 15 minutes of the session were intended to suggest improvements and future work. 7.2.6 Results The first research question of this thesis about the inclusion of ASR and TTS systems in a CAPT tool (RQ1 and Issue 1.1) was tried to be answered with the results obtained in this experiment following the steps defined by the research objectives RO1, RO2, and RO3. These results are presented in next paragraphs according to their origin. That is, (1) results from the interaction between the students and the CAPT system during the training session and (2) results from the focus group session. The discussion of these results is included in Chapter 8. Results related to the training session have been partially published in [17]. User’s Performance Table 7.2 shows the total number of production attempts with the ASR system (#ASREvents), the total number of listenings with the TTS system (#TTSEvents) and the total time spent with the system (Time(s)). In terms of time to complete activities, group C was the slowest since its members used the TTS to listen to a word (requested listening) more times (606), and also, the number of attempts to produce a correct word with the ASR was higher (1094). Globally, the ASR system was used 2.42 times more than the TTS one, even though the TTS events were not restricted to a particular maximum number of events per word (as a reminder, students can produce a word with the ASR system up to five times). Statistically significant differences were found between the three groups and the three variables #ASREvents, #TTSEvents, and Time(s) of Table 7.2, as determined by one-way ANOVA test [180] (p< 0.001, 95% confidence level in the three cases). In particular, pairwise group comparisons were carried out to examine these differences in pairs. A t-test [181] at 95% confidence confirmed that there were statistically significant differences between all group pairs (p< 0.001) except for the #ASREvents column regarding Group B and Group C (p= 0.06). Regarding the results related to the ASR system, Table 7.3 shows the mean number of production events with the ASR system (#ASREvents column), the mean number of correct productions (SuccessASR column), the mean number of wrong
7.2. Alpha Experiment 79 Group # Speakers # ASREvents # TTSEvents Time (s) A 12 372 35 2431 B 21 1033 400 6677 C 20 1094 606 7492 Total 53 2499 1041 16600 TABLE 7.2: Descriptive data gathered with the Minimal Pairs CAPT system, adapted from [17] productions (FailASR column) and the mean percentage of success comparing the number of times the ASR system identifies as correct a word with the number of production attempts of such word (Recall column). Group A achieved the best results since its participants take -on averageless production attempts (31±7 of a total of 120 maximum attempts; five maximum attempts for each one of the 20 words of the 12 minimal pairs presented in the activity). Besides, Group A reached the highest SuccessASR rate, the lowest FailASR one and the best Recall rate. Thus, the higher the declared L2 level, the better production results with and without repetition. On the other hand, Group C reached a 73% of wrong attempts with the ASR and achieved the worst results in all rates. Statistically significant differences were found between the three groups and the four variables of Table 7.3, as determined by one-way ANOVA test (p< 0.001, 95% confidence level in the four cases). Pairwise comparisons with t-tests at 95% confidence were also run for all variables of the table, confirming that there were differences between all of them except for the #ASREvents column regarding Group B and Group C (p= 0.06). The differences between the SuccessASR rate of the participants of Group B and the rest of participants deserve attention. An ANOVA test and a t-test at 99% confidence confirmed that there were statistically significant differences in the same cases as 95%. except for the SuccessASR rate between Group A and Group B (p= 0.057). Group #ASREvents SuccessASR FailASR Recall(%) A 31±7 21±4 10±6 69.1±17 B 49±14 18±3 31±15 41.2±15 C 55±9 15±4 40±10 28.1±10 TABLE 7.3: ASR-related results gathered with the Minimal Pairs CAPT system, adapted from [17]. Values after the symbol ±represent the standard deviation. As explained in Section 6.2.4, an n–best list of predictions was provided by the off-the-shelf Google ASR system in each utterance (in this experiment a 5–best list). This list consists of pairs of a text and a numerical score, also called g–score, with values in a scale of [0, 1]. The next two tables report the results related to this issue. First, in Table 7.4 the mean g–score value of the right production attempts (Right column), the mean g–score value of the wrong production attempts (Wrong column), the mean g–score value of any production attempt (Total column) and the mean time spent with the tool each user (Time column) were represented, categorized by groups. Statistically significant differences were found between the three groups
80 Chapter 7. Experiments and the four variables of Table 7.4 (p< 0.001, 95% confidence level in the four cases). In particular, pairwise comparisons with t-tests found differences in all cases (p< 0.001, t-test, 95% confidence), except for Group A and Group B in the case of wrong attempts (p= 0.09, t-test, 95% confidence). That means native speakers -on averageneed less time to perform the activities than advanced learners and beginners, respectively, as intuited in Table 7.2. Group C speakers were the slowest ones since they need a high number of production attempts due to their wrong utterances (see Table 7.3) and the quality of the ASR response was not as much confidence as the other groups (0.55 vs. 0.59 vs 0.59). Besides, their utterances achieved better g–score values in the majority of cases, except when comparing wrong attempts values in Group A and Group B, which was similar without significant differences, as mentioned earlier. g–score Group Right Wrong Total Time (s) A 0.70±0.3 0.59±0.3 0.67±0.3 203±66 B 0.65±0.3 0.59±0.3 0.61±0.3 318±82 C 0.58±0.3 0.55±0.3 0.56±0.3 375±54 TABLE 7.4: ASR-related metrics gathered with the Minimal Pairs CAPT system, adapted from [17]. Values after the symbol represent the standard deviation. Second, Table 7.5 shows the distribution of the target word of an utterance when it was included in the ASR 5–list of results. The Group A production quality was higher than the rest of Groups since the target word was recognized in the first position of the results in 63.6% of the attempts, in contrast to Group B (51.8%) and Group C (47.3%). A Chi–square [182] test confirms the statistically significant differences between the three groups regarding the first position of the results (p= 0.0053, χ2= 21.7869, df = 8, at 95% level). Position (%) Group 1st 2nd 3rd 4th 5th A 63.6 18.8 8.4 6.8 2.4 B 51.8 21.8 13.7 10.6 2.1 C 47.3 23.8 13.8 9.7 5.4 TABLE 7.5: Mean distribution of the target word in each recognized utterance with the Minimal Pairs CAPT system, adapted from [17]. Table 7.6 sheds light about the reasons why native speakers also fail with the CAPT system when producing words in their L1, as shown in previous Tables. The list of 20 minimal pairs for this first experiment was designed by a phonetics expert in American English, based on his experience in the field after his teaching years. This list of words was directly integrated into the CAPT system. Results showed problems with infrequent in everyday English words, such as wreathe,luff, or wader which summed the 50% of the total wrong utterances of the experiment. Besides, the word wreathe was never identified by the CAPT system and the word luff was reported correct in only two events. The fifteen most frequently failed words in Groups B and C account for the 70% of the attempts. However, in the case of native
7.2. Alpha Experiment 81 speakers (Group A), this value was not reached after the word wader. Furthermore, this table shows words very frequently confused by Spanish speakers, such as peck, Dawn, or sue that were never confused by native speakers. In Chapter 8we will discuss the consequences of and actions to take to improve the selection of words and speech technology proper for a CAPT system. Group A Group B Group C Word % Word % Word % 1 wreathe 100 luff 100 wreathe 100 2 luff 94 wreathe 100 luff 98 3 wader 73 letch 97 letch 98 4 soot 64 loose 90 wader 96 5 sock 58 wader 88 sock 96 6 caber 56 peck 84 soot 96 7 letch 50 sue 84 Gwen 89 8 mass 38 sock 83 shun 88 9 don 33 dunce 81 sue 86 10 mess 33 dawn 80 dawn 85 11 Gwen 31 soot 79 were 83 12 shun 30 Gwen 76 peg 83 13 were 20 were 72 peck 82 14 dunce 12 don 71 loose 81 15 mat 11 zoo 70 dunce 81 TABLE 7.6: Most frequently unrecognized words by the ASR system in the Minimal Pairs prototype (in percentage), adapted from [17]. Finally, regarding the results related to the use of the TTS system as feedback for the production of words, Table 7.7 shows the average number of listenings to each word of the experiment with the TTS system (#TTSEvents column), the mean percentage of correct productions after listening (SuccessTTS column), the mean number of wrong productions after listening (FailTTS column), and the mean percentage of the number of times a learner used the TTS system with respect to the total number of listening and productions events (Rate column). Users listened to the synthesized models of the words when they had doubts about the way to produce them. The number of times they could synthesize the words was not limited. In particular, in Table 7.7 the results related to the words wreathe and luff were not included since these words were found the most problematic ones with the ASR system in this experiment as explained in Table 7.6. The use of the TTS by natives was negligible (2±2 on average). They only resorted to it when the system did not identify their utterance, being only a 5.5% of the production and synthesis events. Besides, their SuccessTTS rate was the worst and their FailTTS rate was the highest one, in comparison to the rest of participants. That means TTS feedback was not helping natives with the not-recognized words by the ASR. Although non-native speakers used the TTS more times than natives (27.6% and a 35.5%), the feedback provided seemed to be not sufficient enough since non-natives’ SuccessTTS rate values were low and similar to natives ones (30.3% vs.
82 Chapter 7. Experiments 29.4% vs 26.3%). In particular, Group C participants made use of the TTS system more than the rest of the groups (28.0% vs. 18.0% vs. 2.1%, respectively), including production and synthesis events (35.5% vs. 27.6% vs. 5.5%, respectively). Statistically significant differences were found between the three groups and the four variables TTSEvents,SuccessTTS,FailTTS, and Rate of Table 7.7, as determined by one-way ANOVA test (p < 0.001, 95% confidence level in the three cases). Pairwise comparisons with t-tests confirmed these results in all cases (p < 0.05) except for the FailTTS and the SuccessTTS rates between Group B and Group C (p= 0.08, at 95% confidence). Group #TTSEvents SuccessTTS(%) FailTTS(%) Rate(%) A 2.1±2 26.3 73.7 5.5 B 18.0±12 29.4 70.6 27.6 C 28.0±17 30.3 69.7 35.5 TABLE 7.7: TTS-related results gathered with the Minimal Pairs CAPT System, adapted from [17]. Focus Group Session This meeting was carried out with 10 non-native learners who participated in the test session. The notes extracted from the participants’ impressions, opinions and improvements about the ASR system were mainly positive. They supported the results presented in the previous Tables 7.2,7.3,7.4,7.5, and 7.6, in which non-native learners achieved the worst results in production activities. However, some comments, such as "I would like to hear and compare my utterances with the ones in TTS", "I felt frustrated after continuous failures", "I would like more training activities. For example, be able to listen to a word of a minimal pair and select which one is the correct answer", strengthened the idea of the necessity of designing and including new non-isolated training activities and corrective feedback techniques in further experiments to help users to overcome their production difficulties. • "I think this tool could be useful to improve my pronunciation with more sounds". • "I realized I cannot produce correctly similar words". • "The answer given by the tool in each utterance was very fast". • "I would like to produce sentences instead of isolated words". • "The TTS system helped me to produce better the sounds". • "I would like to hear and compare myself utterances with the system". • "I would appreciate an indicator about the percentage of production success per word". • "I felt frustrated after failing consecutively". A summary about user’s opinions gathered in the session about possible CAPT GUI’s improvements is: • "I would like to see the word’s phonetic transcription beside the orthographic one". • "I would like to know which words are the most difficult by assigning to them a representative color. For example, green, yellow and red, from the easiest to the most difficult ones".
7.3. Non-guided Learning Experiment 83 • "I felt frustrated when the timer continued counting while the ASR system was evaluating my utterance". • "I would like to pause the activity to take a break". Participants also reported several opinions about the training dynamics of this experiment and future work, from including more game elements and playing with other people outside the class, to combining more pedagogical and feedback resources: • "We think these pronunciation activities can be performed in a game for smart devices". • "A leaderboard would motivate myself to keep on training". • "I would prefer to train and play at home". • "I would like to challenge my friends in a mobile application either in the same room (local) or online". • "I felt myself frustrated to be under pressure in a class with an instructor. I would prefer to train in a stress-free place". • "I would appreciate a training tutorial and pronunciation instructions with mouth pictures". • "I would like more training activities. For example, to be able to listen to a word of a minimal pair and select which one is the correct answer". • "The level of difficulty should be adjustable to the necessities of each one". • "I would like to practice either isolated words and minimal pairs". Finally, the instructor of the session asked to the participants if they would prefer a tool for training, for playing or for both options. Gathered results confirmed that a 80% of learners would use the system for training, a 20% would use the system for playing, and the 100% would use the application for learning through playing. 7.3 Non-guided Learning Experiment The second practical approach of this thesis was the experiment called Non-guided Learning. Different pronunciation training activities were carried out when using general-purpose speech technology in the prototype developed for this experiment, TipTopTalk! In particular, this prototype included an implicit competition (see its definition at Section 5.1) in which users trained individually their L2 pronunciation with a gamified CAPT system for smart devices. Subjects were university students who participated voluntarily in the experiment. They practiced anytime anywhere, choosing the training activities at free will. The goal of this experiment aimed at assessing the possible learners’ pronunciation improvement along time while keeping users motivated at the same time they were training. User’s interaction data was automatically monitored to obtain pronunciation assessment results. Initially, this experiment was intended for native Spanish speakers who study American English as L2. However, due to the collaboration success achieved with different academic institutions and research groups, the prototype developed for this
90 Chapter 7. Experiments L2 en_EN L1 es_ES cn_ZH ANY Gender F M F M F M ANY #Subjects 18 21 1 1 19 22 41 Training mode EXP-T 94 97 - - 94 97 95.5 DIS-T 34 39 - 29 34 34 34 PRO-T 166 539 - - 166 539 352.5 Playing mode DIS-P 31 100 36 - 33.5 100 66.75 PRO-P 228 456 85 - 156.5 456 306.25 MIX-P 81 132 - - 81 132 106.5 L2 cn_ZH L1 es_ES cn_ZH ANY Gender F M F M F M ANY #Subjects 5 8 5 1 10 9 19 Training mode EXP-T 148 114 80 58 114 86 100 DIS-T 44 42 27 12 35.5 27 31.25 PRO-T 103 230 30 - 66.5 230 148.25 Playing mode DIS-P 40 32 26 36 33 34 33.5 PRO-P 223 252 107 87 165 169.5 167.25 MIX-P - - 119 108 119 108 113.5 TABLE 7.9: Average time (s) spent by users in each activity type of the TipTopTalk! prototype. The left tabular refers to American English as L2 subjects and the right one to Simplified Chinese as L2 ones. Fand Mmean ’female’ and ’male’, respectively. EXP-T,DIS-T, and PROTwere exposure, discrimination, and pronunciation activity types of the Training mode, respectively. DIS-P,PRO-P, and MIX-P stand for discrimination, pronunciation, and mixed activity types of the Playing mode, respectively. respectively, in en_EN as L2; and 148.25s and 149.75s for PRO-T and PRO-P, respectively, in cn_ZH as L2). Male native Spanish participants spent more time performing PRO-T and PRO-P activities than female ones, this difference being higher in PRO-T in both L2 cases (539s vs. 166s, en_EN as L2, respectively; and 230s vs. 103s, cn_ZH as L2, respectively); and in en_EN PRO-P activities (456s vs. 228s, respectively). Finally, PRO-T and PRO-P activities were performed faster in cn_ZH as L2 than in en_EN as L2 (352.5s vs. 148.25s, PRO-T, respectively; and 306.25s vs. 167.25s, PRO-P, respectively). FIGURE 7.8: Distribution of users by number of days with active participation in the competition of the TipTopTalk! prototype. Concerning the distribution of players activity throughout the 24 competition
7.3. Non-guided Learning Experiment 91 days, Figure 7.8 shows the accumulative number of days in which a player performed at least one activity type. A high number of users only played with the game occasionally (41% participated one day). Regarding the rest of the users, the plot describes that there was not a single user who participated more than 11 days. User’s Performance In this experiment a database of 87,918 entries was stored containing all user’s interaction data with the CAPT system. In particular, almost a 40% of these entries were related to the two different event types which lead to achieve points for the competition: perception and production activities. The former events were performed by learners in discrimination and mixed activity types of the Playing mode and in discrimination activities of the Training mode. Production events were completed in production and mixed activity types of the Playing mode and in production activities of the Training one. Table 7.10 shows the average number of these two activity events performed by each user in the experiment. Discrimination events were performed by more users than production ones in both the Training (86% vs. 62%) and in the Playing mode (100% vs. 64%). Although the selection of activity types was left to user’s will, results reveal a balanced choice between perception and production activities since there were no statistically significant differences between them in the average number of events performed in each mode (405.2 vs. 349.9 in the Playing mode; and 37.0 vs. 24.3 in the Training mode). The higher average number of Playing activities than Training ones performed by each user leads to statistically significant differences (U = 442.0, p< 0.001, Mann– Whitney Utest [184]). #Events #Participants|#Total Training mode Discrimination 37.0 (60.4%) 25/29 (86%) Production 24.3 (39.6%) 18/29 (62%) Playing mode Discrimination 405.2 (53.7%) 36/36 (100%) Production 349.9 (46.3%) 23/36 (64%) TABLE 7.10: Average number of discrimination and production events per participant of the TipTopTalk! prototype. The third column (#Participants|#Total) refers to the number of subjects who perform these activities (first value) and the total number of participants who perform an activity of the same mode (second value). Figure 7.9 represents the number of discrimination and production activities performed by learners per competition day. The highest value of activity was reached during the middle days of the experiment (42% and 50% in discrimination and production activities, respectively). A large number of discrimination activities were carried out during the first ten days of competition (65%). In the case of production activities, a 70% of the total events were registered from the 10th day of competition. Finally, a peak of discrimination activities was observed on the penultimate day of competition (7.6%). These perception and production events performed in the competition are represented in Figure 7.10. It shows the evolution of quality functions fDand fPalong a
92 Chapter 7. Experiments FIGURE 7.9: Distribution of discrimination and production activities per day in the TipTopTalk! prototype. chronologically ordered attempts sequence, s, varying uand kwith a window size of w=6. (see Section 7.3.5 for their definition details). Three groups of users are displayed depending on the value of the quality functions achieved in the introductory window s=6, which represents the initial competence of each user before using the CAPT system for the first time. FIGURE 7.10: Evolution along time of the pronunciation quality functions of the TipTopTalk! prototype, adapted from [18]. The first diagram refers to perception activities and the second one to production activities. The ordered sequence of minimal pairs attempts was represented in the abscissa and the quality function in the ordinate axis. Production activities registered different results. First, a positive tendency was displayed for the lowest initial value group of users (0.460 to 0.485) and the intermediate ones (0.585 to 0.645) according to their fP, until s=14. Then, these values varied, reaching a higher value (0.495) than the initial one (0.460) in the worst group and a lower value (0.570) respecting the initial one (0.585) in the intermediate group. Finally, subjects with the highest fP(6) =0.730 gravitated around this value until s=20, in which fPfell to 0.655.
7.3. Non-guided Learning Experiment 93 Table 7.11 displays the users’ success rate average in each activity type, categorized by their gender, L1 and L2. In relation to Figure 7.10, learners achieved the worst success rate values in all production activity types (32.5% and 37.7% in PRO-T and PRO-P, en_EN as L2, respectively; and 34.5% and 41.9%, in PRO-T and PRO-P, cn_ZH as L2, respectively). Discrimination results in both L2 targets confirms the improvement detected in the sequence of events along time represented in Figure 7.10, reaching higher success rate values than 60% in all cases except in DIS-T of en_EN as L2 (59.2%). Besides, users obtained better success rate values in all activity types of the Playing mode than in the Training mode. Female subjects achieved better results than male ones in all cases (Fand Mcolumns of the L1-ANY row). These results conformed to those reported in Table 7.9, in which female participants spent less time to perform the activities than male ones. Furthermore, native Chinese speakers had difficulties with perception activities since they did not reach success values higher than 50%. L2 en_EN L1 es_ES cn_ZH ANY Gender F M F M F M ANY #Subjects 18 21 3 1 21 22 43 Training mode DIS-T (%) 67.4 65.9 - 44.4 67.4 55.2 59.2 PRO-T (%) 42.7 22.2 - - 42.7 22.2 32.5 Playing mode DIS-P (%) 76.8 76.4 82.3 - 79.6 76.4 78.5 PRO-P (%) 36.8 35.0 41.2 - 39.0 35.0 37.7 MIX-P (%) 62.6 62.5 - - 62.6 62.5 62.6 L2 cn_ZH L1 es_ES cn_ZH ANY Gender F M F M F M ANY #Subjects 5 8 5 1 10 9 19 Training mode DIS-T (%) 60.0 78.3 70.8 50.0 65.4 64.1 64.8 PRO-T (%) 18.9 - 50.0 - 34.5 - 34.5 Playing mode DIS-P (%) 68.8 70.4 91.8 33.3 80.3 51.9 66.1 PRO-P (%) 24.3 39.3 62.1 - 43.2 39.3 41.9 MIX-P (%) - - 84.2 66.7 84.2 66.7 75.4 TABLE 7.11: Success rate (%) in each activity type of the TipTopTalk! prototype. The left tabular refers to American English as L2 subjects and the right one to Simplified Chinese as L2 ones. Fand Mrefer to ’female’ and ’male’, respectively. DIS-T and PRO-T are discrimination and pronunciation activity types of the Training mode, respectively. DIS-P, PRO-P, and MIX-P stand for discrimination, pronunciation and mixed activity types of the Playing mode, respectively. Terminal Characteristics User’s preferences about the device where the CAPT system was installed were also analyzed for future GUI improvements. Data gathered from participants who gave permission to share their smart device’s technical specifications led to three main groups of Android devices in this experiment. In particular, there were 45 smart devices with a screen size lower or equal than 5.5 inches (70.4%), 8 devices with a screen between 5.5 and 7 inches (12.4%), and 11 devices (17.2%) with a size equal or higher than seven inches.
94 Chapter 7. Experiments Questionnaire In this section we present the results obtained from the voluntarily-answered questionnaire about the CAPT system UX. Regarding the Likert-type scale questions (Figure 7.11), a 90% of the questionnaire respondents did not find difficulties or consider them insignificant when interacting with the system (question 1); while the same percentage thought the guide texts and tips helped them properly to understand the proposed activities (question 2). These results were in tune with the 80% of users who disagreed or fully disagreed in feeling lost to continue (question 3). In this case, they agreed and fully agreed that the support system, its controls, and commands were adequate and useful (80% and 90%, respectively, (question 4 and question 5). A 80% of the questionnaire respondents felt confident using the system (question 6); whereas a 50% reported frustration in at least one occasion and the other 50% did not, which confirmed the lower production success rate values shown in Table 7.11 (question 7). These results were reinforced with the fact that the 95% of users agreed and fully agreed in founding the mechanics of the system easy to understand (question 9). Finally, a 30% and a 45% of the questionnaire respondents agreed and fully agreed with finding fun the game activities, respectively (question 8). FIGURE 7.11: Likert scale questions of the TipTopTalk! prototype. Figure 7.12 displays the answers to the last five questions of the questionnaire about selection. Over half of the answers claim Discrimination as the favorite training activity type (question 1). However, the Mixed activities were preferred in the Playing mode confirming the results presented in Table 7.8 (question 2). Besides, the 85% of the questionnaire respondents claim that the words’ difficulty was adequate (question 3). Finally, up to a 90% of the users would like to challenge other people with the activities of the system and elaborate their own list of words (question 4 and question 5). Finally, the answers provided to the optional open-ended question, "Please, tell us any other suggestion, complaint, or future improvement", were:
7.4. Guided Learning Experiment 95 FIGURE 7.12: Selection-type questions of the TipTopTalk! prototype. • "I would like to have my own lists of words to practice at any time". • "I would like some type of daily rewards such as points or lives". • "I would appreciate the translation into my native language Chinese words". • "I felt frustrated when I saw a lot of wrong production attempts". • "Whenever I rotate my phone’s screen, the ’success’ sound is played again". • "I would prefer not to scroll in the menus of the application". • "Whenever I rotate my phone’s screen, the activity tip disappears". • "At the end of an activity my smartphone’s back button is disabled". Some of the answers agreed with the close-ended questions (i.e., elaborating own lists of words and feeling frustration after several wrong production attempts); whereas others recommended proposal and future improvements (i.e., extra rewards and word translations). The rest of the answers were related to usability aspects, such as GUI improvements and software bugs detected. 7.4 Guided Learning Experiment In the third experiment conducted in this dissertation, named Guided Learning, a controlled training protocol with a pre/post test strategy was followed to ascertain the pronunciation level improvement of the participants from different training groups. During the days between the tests, some participants trained with a pedagogical evolved version of the CAPT system from the previous experiments, which followed a specific training protocol. Two prototypes were developed for this experiment, corresponding to different L1 and L2. Native Spanish learners of English
96 Chapter 7. Experiments participated in the first one, English Vowels, while native Japanese learners of Spanish were involved in the second prototype, Japañol. Next subsections describe their specific details. 7.4.1 Experimental Procedure A four-week protocol was defined for this experiment. It included a pre-test, three training sessions, and a post-test, as shown in Figure 7.13. At the beginning, the subjects took part in the pre-test session individually in a quiet testing room while the sound of the session was recorded with a microphone and an audio recorder. All the students took the pre-test under the sole supervision of a member of the research team. In the case of the English Vowels prototype, learners were asked to read aloud the 25 minimal pairs contrasts administered via a sheet of paper with no time limitation. They were free to repeat each contrast as many times as they wanted if they thought they might have mispronounced them. In particular, the test included contrasts of the English pure vowels /A:/, /2/, and /æ/, that are usually reduced to Spanish vowel {a} [185]; vowel /e/, that is usually realized as a closer Spanish {e} [185], and vowels /i:/ and /I/, often reduced by Spanish speakers to {i} [74], [186]. Readers can find more details about this test in Table C.1 in Appendix C. Pre-test Post-test Training Subjects Data Test results Test results Log & audio files In-classroom group Experimental group Analysis Report FIGURE 7.13: Steps of the English Vowels prototype’s protocol, adapted from [23]. A total of three training sessions were carried out one week after running the pre-test. There was a time gap of at least 72 hours between them in order to avoid fatigue and promote memory consolidation [187]. Subjects were divided into two different groups after taking the pre-test (see Section 7.4.3 for more details). The training sessions of both groups were conducted at the same time in different locations (classroom and laboratory). A maximum time of 60 minutes was established in each one of them. The session’s learning content was divided into lessons. Two lessons were presented in each session and a minimal pair contrast was practiced in each lesson (block distribution). However, most phonemes were retaken in later sessions (spaced distribution). In particular, in the first session, phonemes /A:/–/æ/ and /æ/–/2/ were contrasted. In the second one, /A:/–/2/ and /e/–/æ/. The last session involved the phonemes /I/–/i:/ and /I/–/e/. Only /i:/, a vowel that is almost interchangeable with the Spanish /i/, was left out of a repeated practice scheme. On the one hand, students of the experimental group exclusively used the CAPT system (see Section 7.4.6 for more details). The software application was installed on an Android emulator (NOX App player) in the laboratory computers. Subjects
7.4. Guided Learning Experiment 97 were inside a cubicle, separated to each other by glass dividers, and used a headset with microphone. Before starting the first session, users in the CAPT-condition were instructed in place on how to use the software. During the rest of the session, each student worked individually (see Section 7.4.5 for activity details) and did not have any interaction with either classmates or instructors. Along the three experimental sessions, a total of 72 minimal pairs were presented to participants (12 in each lesson, 2 lessons per session)3. On the other hand, the in-classroom group participants had interaction with their classmates and the instructor (see Section 7.4.4 for more details). Finally, one week after the last training session, a post-test was carried out by all participants. The post-test contents and conditions were identical to those of the pretest. Both tests were assessed by three L2–English raters experts in phonetics, in random order and independently, some days after collecting all results. The raters were specialized EFL teachers who had worked extensively on L2–English pronunciation. They had no contact with the participants. They did not know if the utterances they were evaluating belonged either to pre-test or post-test realizations. The same number of training sessions, duration and spacing between them and the places for the English Vowels prototype (see Section 7.4.1) were repeated for the Japañol prototype. That is, a four-week protocol which included a pre-test, three training sessions, and a post-test was followed, as shown in Figure 7.14. In particular, the tests included 28 contrasts of the most difficult to produce Spanish consonant sounds by native Japanese speakers [188], [189], [190]: [T, s] sounds are usually reduced to Japanese consonant [s]; the Spanish sounds [T, f] are confused by Japanese speakers in perception activities since they do not exist in Japanese and the nearest Japanese sound is [F]; the consonant sound [f] is often confused with [x], especially when these sounds are followed by [u]; the Spanish phoneme /r/ is usually realized as [R] or [l]; and finally, the combination of fricative /f/ with /l/ or /r/ in onset, also triggers mispronunciation. Readers can find more details about this test in Table C.2 in Appendix C. Pre-test Post-test Training Subjects Data Test results Test results Log & audio files Placebo group In-classroom group Analysis Report Experimental group Test results FIGURE 7.14: Steps of the Japañol prototype’s protocol. 3https://github.com/eca-simm/minimal-pairs-enus-eses
98 Chapter 7. Experiments A total of 84 minimal pairs were presented to participants (12 in each lesson, 2 lessons per session, except for the last session that included 3 lessons)4. A blocked and spaced practice schedule was also followed within the sessions. Regarding the sounds practiced in each session, in the first one, sounds [fu]–[xu] and [l]–[R] were contrasted. In the second one, [l]–[r] and [R]–[rr]. The last session involved the sounds [fl]–[fR],[T]–[f], and [T]–[s]. Finally, subjects of the placebo group did not participate in the training sessions. They were supposed to take the pre-test and post-test and obtain results without significant differences. 7.4.2 Enrollment The recruitment campaign for the English Vowels prototype consisted in a call for volunteers from the same course of EFL for B1–B2 level of the Language Center of the University of Valladolid. The target subjects of the Japañol prototype were native Japanese speakers learners of Spanish from two different locations, the Language Center of the University of Valladolid and the University of Seisen (Japan). Students who gave consent, filled in a registration form with some personal information and signed an authorization. The training protocol sessions were carried out during their course’s classes. All participants were awarded with a diploma and a reward after completing all stages of the experiment. 7.4.3 Participants A total of 20 native Spanish students who qualified and registered for the same EFL course of the Language Center of the University of Valladolid were initially selected for taking part in the English Vowels prototype. This institution distributes its students along its different courses by means of an accurate level test. Two participants left the training sessions for personal reasons during the early stages, and were consequently discarded. Before being allowed to take this course at the University, the participants took a placement test. In this case, all of them had an intermediate B1–B2 level of English with a very little or no previous training in English phonetics (see Table B.7 for specific details). In this way we ensured that the experiment realistically reproduced the diversity of students that attend the same course of the Language Center of the University of Valladolid; and that all students had the same initial level of English. Furthermore, it was explicitly requested to participants not to do any extra work in English (extra lessons, conversation exchanges with natives, etc.) while the experiment was still active. Students were offered, through the mediation of the instructor, and with the reluctant agreement of the institution’s authorities, to cover a small part of the EFL course program (in particular, the teaching of a few English phonemes) by using a CAPT system. Participants were divided into two homogeneous groups since all of them obtained low pre-test scores (see Section 7.4.12 for more details about results): 1. Experimental group. 10 students who trained their English pronunciation with the CAPT system developed, during three sessions of 60 minutes. 2 were women and 8 were men. 4https://github.com/eca-simm/minimal-pairs-japanol-eses-jpjp
7.4. Guided Learning Experiment 99 2. In-classroom group. 10 students who attended to three pronunciation teaching sessions of 60 minutes within the EFL course, with their usual instructor, making no use of any computer-assisted interactive tools. However, two of them were discarded since they left the experiment before finishing all the stages. 5 were women and 3 were men. A total of 33 native Japanese speakers from 18 to 26 years old participated voluntarily for the Japañol prototype. All of them declared a low or intermediate level of Spanish as L2 with a very little or no previous training in Spanish phonetics (see Table B.7 for specific details). Besides, they were requested not to do any extra work in Spanish (extra phonetics research, conversation exchanges with natives, etc.) while the experiment was still active. They came from two different locations: 1. Language Center of the University of Valladolid. 8 students of the Spanish philology degree of the same University course who recently arrived to Spain from Japan in order to start an L2 Spanish course. 5 were women and 3 were men. 2. University of Seisen. 25 female students of the Spanish philology degree from Seisen, Japan. Participants were divided into three homogeneous groups since all of them obtained low pre-test scores (see Section 7.4.13 for more details about results): 1. Experimental group. 18 students who trained their Spanish pronunciation with our CAPT system, during three sessions of 60 minutes. 15 are women and 3 are men. 2. In-classroom group. 8 female students who attended to three pronunciation teaching sessions of 60 minutes within the L2–Spanish course, with their usual instructor, making no use of any computer-assisted interactive tools. 3. Placebo group. 7 female students who only took the pre-test and post-test. They did not attend neither the classroom nor the laboratory for Spanish phonetics instruction. Finally, a group of 10 native Spanish speakers from the Teatro Pie Izquierdo of Valladolid5(5 women and 5 men) participated in the recording of a total of 41,000 words included in the pre/post-tests and the CAPT tool developed (see details in Appendix D). The recording sessions were carried out in an anechoic chamber of the University of Valladolid. The dataset was intended to be part of an own ASR system for assessing the pre/post-test utterances gathered in the experimentation. 7.4.4 In-classroom Group Training Activities In both prototypes, in-classroom group participants were guided by a non-native L2–English or L2–Spanish teacher, respectively, with a vast experience in phonetics of the target L2. The teaching program included the same phonemes covered in the experimental group. Each 60–minutes session began with around 10 minutes of explicit articulatory instructions and auditory descriptions of the sounds, with ample exposure to contrasting examples. These examples were both produced by the instructor and extracted from the audio materials of an English or Spanish as L2 handbook, respectively. After exposure, students were asked to practice perception 5http://www.pieizquierdo.es/
106 Chapter 7. Experiments the University of Valladolid. They were informed that the first element of the pair is always the Spanish word, and the second one the English word. In each rating unit, one of the pairs was randomly tagged as A, and the other as B. The rater’s task consists in confronting each rating unit and answering two questions by selecting one of the four possible answers (see Table 7.13). Regarding the second prototype of this experiment, Japañol, since the number of audio samples and subjects was so much higher than the previous one, five expert phoneticians and native speakers assigned a correct/incorrect value to each preand post-test word of the participants of the Language Center of the University of Valladolid (see 7.4.3); whereas the rest of utterances of the participants from the University of Seisen were objectively assessed by two different ASR systems (see section 7.4.10). In the first case, all data was presented to human raters randomly and without user association via a web page. They were asked to focus on the specific sound of each word which should be generated correctly, ignoring the bad pronunciations of the rest of sounds. During the process, raters neither interacted amongst themselves nor with the subjects. Experts scores were computed by summing up the number of correct words per speaker and normalizing the result to the range [0, 10]. 7.4.10 Scoring Procedures An objective value of user’s performance was defined for both prototypes of this experiment, called game score,G, which consisted of the average value of the success scores obtained by a user in each of the training lessons using the CAPT system: G=1 N N ∑ i=1 Ls,i;G∈[0,10]. (7.4) where sis the speaker, iis the lesson, and Nrefers to the number of lessons attempted (N=6 in the English Vowels prototype and N=7 in the Japañol prototype). In particular, each lesson can be rated by a score, Ls,i, based on the user’s performance in the Discrimination (D), Pronunciation (P), and Mixed (M) modes (see Section 7.4.5 for specific details of each one). It can be expressed as: Ls,i=1 3(Ds,i+Ps,i+Ms,i); Ls,i∈[0, 10]. (7.5) The score in the Discrimination mode, Ds,i, is based on the number of discrimination tasks (DT) successfully attempted (see Table 7.12): Ds,i= 10 ∑ j=1 DTj;Ds,i∈[0,10]. (7.6) where DTjis the discrimination task’s value (1 if right, 0 if wrong), sis the speaker and iis the lesson. The score in the Pronunciation mode, Ps,i, is based on the number of production tasks (PT) successfully carried out (see Table 7.12): Ps,i= 10 ∑ j=1 PTj;Ps,i∈[0,10]. (7.7)
7.4. Guided Learning Experiment 107 where PTjis a production task (1 if right, 0 if wrong), sis the speaker and iis the lesson. The score in the Mixed mode is based on the number of Mixed-mode tasks (MT) successfully attempted (see Table 7.12): Ms,i=10 9 9 ∑ j=1 MTj;Ms,i∈[0,10]. (7.8) where MTjis a mixed task (1 if right, 0 if wrong), sis the speaker and iis the lesson. In the case of the Japañol prototype, the second objective method for assessing user’s utterances consisted on processing all utterances with GCSTT and Kaldi ASR engines, obtaining an n–best list of hypotheses and their confidence values. The ASR score was computed by summing up the number of correct words per speaker and normalizing the result to the range [0, 10]. 7.4.11 Statistical Tests Since most data gathered did not pass the Kolmogorov–Smirnoff nor Levene’s standard tests for assuming normality and homogeneity of variances, respectively, several non-parametric tests for non-normally distributed data were carried out to detect statistically significant differences: 1. Expecting that there might be a certain degree of variability in the scoring process by human agents, a consistency check based on Kendall’s coefficient [193] analysis was carried out. 2. Interand intra-group pairs comparisons between the pre/post-scores were carried out using a Mann–Whitney Utest [184] and Wilcoxon signed-rank test [194], respectively. 3. A Mann–Whitney Utest was also performed to analyze significant differences between following or not the recommended feedback. 4. A Kappa Fleiss index [195] computed the inter-rater consistency in the ABX and perceptual tests. 5. A Chi–square test [182] was run to measure the statistically significant differences between the two-way contingency table of the ABX results of the experimental and in-classroom groups. 6. A Pearson correlation [196] was carried out to compare the scores assigned by the software and the human raters scores of the post-test. 7.4.12 English Vowels Prototype Results This prototype was the first one to focus on a guided and pedagogical protocol (see Section 7.1). It mainly tried to give answers to the research question RQ2 (and its Issues) about the effects on user’s pronunciation improvement following a specific pedagogical training methodology, and also to reinforce the answers to RQ1 (and Issue 1.1) about the inclusion of current ASR and TTS systems into CAPT systems in a non-obstructive way for any L2 pronunciation proficiency level. This prototype also followed the steps defined in the research objectives RO1, RO2, RO3, and RO4. The most relevant results obtained in this prototype are presented in next subsections according to: results gathered from the interaction between the learners and
108 Chapter 7. Experiments the CAPT system during the training sessions related to performance (1) and behavior (2), and results extracted from the pre-test and post-test (3). These results have been partially published in [23], [24], and a discussion of them is addressed in Chapter 8. User’s Performance Table 7.14 reports the intensity of use of the CAPT tool by the speakers. n,m, and Mare the mean, minimum and maximum values, respectively. Time (min) row stands for the time spent on minutes per learner in each mode in the three sessions of the experiment. #Tries represents the number of times a mode was executed by each user. The symbol - stands for ’not applicable’. Mand. and Req. mean mandatory and requested (listening). The TTS system was employed in both listening types; whereas the ASR was only used in the #Productions row. Theory Exposure Discrimination Pronunciation Mixed n m M n m M n m M n m M n m M Time (min) 31.32 20.1 39.2 16.93 11.1 29.6 5.48 3.7 7 41.47 19.2 65.1 19.03 3.7 34.1 #Tries 6.4 6 8 11.9 7 17 7.2 6 9 12.6 6 21 9 6 18 #Mand.List. - - - 347 210 510 69.5 60 82 - - - 26.8 15 54 #Req.List. - - - 146.9 64 292 29.9 0 75 147.9 25 426 63.2 20 178 #Discriminations - - - - - - 69.5 60 82 - - - 26.8 15 54 #Productions - - - - - - - - - 441.5 166 806 174.1 87 382 #Recordings - - - 90.2 56 134 - - - - - - - - - TABLE 7.14: User’s performance with the CAPT system of the English Vowels prototype, adapted from [23]. A high rate of active user-time invested in interactive tasks is shown in Table 7.14 (114.0 minutes out of the total 180.0 minutes in three sessions). The activity registered with speech technology systems reached high values since as an average term, each subject listens to synthesized words for 831.2 times (calculated as the sum of the values of #Mand.List. of Exposure, Discrimination, Pronunciation, and Mixed modes, and #Req.List. values of column n) and uses the ASR system 615.6 times (calculated as the sum of the value nof the column #Productions of Pronunciation, and Mixed modes), reaching a rate of 8.04 uses of the TTS/ASR per minute. On the other hand, the variation in the minimal and maximum values of the variables shown in Table 7.14 clearly illustrates the differences between users. For instance, the fastest user in performing the Mixed mode’s activities spent 3.7 minutes; whereas the slowest one took 34.1 minutes. This contrast can also be observed in the time spent on the rest of the training modes and in the number of times learners practiced each one of them (row #Tries). These inter-user differences also affected both the number of times the users made use of the TTS (109 min. vs. 971 max.) and the number of times they used the ASR (253 min. vs. 1188 max.). The two main interactive training tasks —discrimination and production— included in Discrimination, Pronunciation, and Mixed modes, motivated the variety in the use of the tool regarding the differences between users, since a high quantity of attempts which involved them was registered. This affirmation is further illustrated in Table 7.15 where the number of correct and incorrect interactions per tested phoneme is shown. The final column (Total) indicates that production activities resulted tougher than discrimination ones: 53.5% vs. 81.2% of successful events,
7.4. Guided Learning Experiment 109 Successful (S) and Failing (F) Events Task A: æ 2 e I i: Total S (%) F S (%) F S (%) F S (%) F S (%) F S (%) F S (%) F Discrimination 143 (75.7%) 46 198 (81.1%) 46 114 (77.0%) 34 144 (86.2%) 23 105 (78.9%) 28 78 (95.1%) 4 782 (81.2%) 181 Production 151 (36.2%) 266 261 (53.6%) 226 127 (42.1%) 175 195 (76.5%) 60 115 (58.4%) 82 103 (85.1%) 18 952 (53.5%) 827 All productions 151 (8.9%) 1543 261 (15.5%) 1424 127 (10.6%) 1066 195 (31.1%) 433 115 (17.0%) 563 103 (37.1%) 175 952 (15.5%) 5204 Mandatory (M) and User-Requested (R) Listening Events A: æ 2 e I i: Total M R M R M R M R M R M R M R Discrimination 189 86 244 89 148 62 167 75 133 74 82 24 963 410 Production - 562 - 552 - 374 - 218 - 241 - 53 - 2000 TABLE 7.15: Right, wrong, and listening events categorized by phoneme of the English Vowels prototype, adapted from [23]. The symbol - stands for ’not applicable’. respectively. This difference was accentuated when the All productions events rate was compared: 15.5% vs. 81.2%. As a reader’s reminder, a maximum discrimination and production wrong sequence consisted in one and up to five wrong attempts, respectively. The most difficult phonemes for the learners can also be revealed from Table 7.15, since there were important differences that affect both discrimination and production activities. For instance, the phoneme /A:/ seemed to be the most difficult one since it showed the lowest success rate values in both cases (only a 8.9% in Production) and a 36.2% in All productions success rates in production and a 75.7% in discrimination). On the other hand, the phoneme /i:/ appeared to be the easiest one since it reached a 37.1% All productions and a 85.1% Production success rate values in production and a 95.1% success rate in discrimination. The number of times the learners requested the use of the TTS was also influenced by these differences: 648 for /A:/ vs. 77 for /i:/. Discrimination tasks A: æ 2 e I i: TPR (%) A: 143 34 12 - - - 75.7 æ19 198 11 16 - - 81.1 220 14 114 - - - 77.0 e11 - 144 12 - 86.2 I- - - 17 105 11 78.9 i: - - - - 4 78 95.1 PPV (%) 78.6 77.0 83.2 81.4 86.8 87.6 Production tasks A: æ 2 e I i: TPR (%) A: 151 143 123 - - - 36.2 æ78 261 35 113 - - 53.6 2121 54 127 - - - 42.1 e36 - 195 24 - 76.5 I- - - 33 115 49 58.4 i: - - - - 18 103 85.1 PPV (%) 43.1 52.8 44.6 57.2 73.2 67.8 TABLE 7.16: Confusion matrices of the English Vowels prototype, adapted from [23]. Left table: confusion matrix of discrimination tasks (diagonal: right discrimination tasks). Right table: confusion matrix of production tasks (diagonal: right production tasks). The CAPT system also revealed what the real difficulties of the users were in terms of the most difficult phonemes. The confusion matrices of discrimination and production tasks between the phonemes contrasted in each lesson are displayed in Table 7.16. In both tables the rows are the expected phonemes and the columns are the phonemes selected/produced by the user. TPR (true positive rate or recall) and PPV (positive predictive value or precision) are quality indicators. The symbol - stands for ’not applicable’.
110 Chapter 7. Experiments In particular, the phoneme /A:/ was the hardest to predict in discrimination tasks (TPR = 75.7%) since it had the lowest recall; whereas /æ/ was the most commonly confused (PPV = 77.0%) since it had the lowest precision. The easiest phoneme in this type of tasks was /i:/ since it had the highest precision and recall (PPV = 87.6% and TPR = 95.1%). On the other hand, in production tasks, the phoneme /A:/ obtained the lowest precision, and also recall values (PPV = 43.1% and TPR = 36.2%); whereas /i:/ obtained the highest recall (TPR = 85.1%), and /I/ the highest precision (PPV = 73.2%). User’s Behavior The recommendation of specific training modes and activities was a part of the feedback in the training protocol which can also be analyzed. First, Table 7.17 represents the number of times that each training mode which affects the game score G was practiced (Discrimination, Pronunciation, and Mixed modes, see Section 7.4.10 for more details about this score). Three different scenarios can be described: (1) a training mode was passed (grade 60% or higher) at the first attempt, (2) a training mode was passed after repetition (because in previous attempts the user did not reach a 60% grade) following the recommended feedback or not, and (3) a training mode was not passed (with or without feedback). Proposed modes Completed modes at first-attempt Mode repetitions Mode repetitions with feedback Mode repetitions without feedback Discrimination 60 51 10 6 [6,0] 4 [3,1] Pronunciation 60 35 61 40 [23,17] 21 [2,19] Mixed 60 43 34 12 [12,0] 22 [5,17] TABLE 7.17: Comparison between following recommended feedback or not of the English Vowels prototype, adapted from [24]. Numbers between square brackets correspond to [passed, failed]. Clear differences between the three training modes are stated in Table 7.17. The Discrimination mode was the easiest one since it was passed 51 out of 60 times at the first attempt (83.33%). However, the Pronunciation training mode was the most difficult one, with 61 repetitions and a 58.33% success at the first attempt. When repeated, only in two occasions it was passed without the help of the provided feedback. Besides, the experiment showed significant differences (U= 46.0, p< 0.001, Mann–Whitney Utest at 99% confidence level) between following or not the corrective feedback recommendations provided by the tool. In particular, these differences were higher in the case of the Pronunciation mode: without feedback, only a 10% of success was achieved; whereas the Mixed and Discrimination training modes achieved a 100% of success rate when the recommended feedback was followed. Second, the effect of listening to a synthetic model of a misproduced word as an explicit corrective feedback in production events of Pronunciation and Mixed modes was also measured. Table 7.18 shows the number of production sequences which led to a positive or negative improvement. Each production sequence was given by: 1
7.4. Guided Learning Experiment 111 A: æ 2 e I i: +93 (32%) [2.8] 85 (32.4%) [3.2] 70 (35.4%) [2.9] 24 (30.4%) [3.1] 51 (47.2%) [2.9] 18 (58.1%) [4.2] 341 (35.2%) [3.2] =168 (57.7%) [0] 148 (56.5%) [0] 101 (51%) [0] 41 (51.9%) [0] 44 (40.7%) [0] 7 (22.6%) [0] 509 (52.5%) [0] -30 (10.3%) [-2.4] 29 (11.1%) [-2.7] 27 (13.6%) [-2.3] 14 (17.7%) [-2.9] 13 (12%) [-2.3] 6 (19.4%) [-1.7] 119 (12.3%) [-2.4] 291 [0.1] 262 [0.2] 198 [0.2] 79 [0.1] 108 [0.2] 31 [0.8] 969 [0.3] TABLE 7.18: Sequences of wrong production, listen, and repeat of the English Vowels prototype. wrong production attempt followed by a number n(n>0) of requested listenings, and ending in a new production attempt (from 1..4), always from the same word practiced. 969 sequences that comply with these requirements were registered of a total of 1779 (952+827, as showed in Table 7.15). There are three numbers in each cell of Table 7.18. The first one refers to the number of sequences. The second one is the percentage of sequences with respect to the phoneme (column). The last number is an indicator of positive (+), negative (-), or none of them (=) improvement in a scale of [−4, 5] according to the target word’s recognition position in the n–best list of hypotheses (1 to 5 positions, and 6 if not recognized). All right production events are included in the +row; whereas wrong production events can be included in any of three rows. When the subtraction of the n–best list position (in this case 5-best list) of the last recognition word attempt minus the first one is positive, the result is included in the +row, when it is negative in the -row, and when it is 0 in the =row. When the last attempt was better than the first one, 3.2 points of difference were achieved. A slightly positive tendency of production improvement in the described sequences (0.3 out of 5) was confirmed as determined by a Wilcoxon signed-rank test at 95% confidence level since there were statistically significant differences between the (+) and (-) rows in Table 7.18 (Z= -10.362, p< 0.001). Results in Table 7.18 also corroborate results reported in Table 7.16 about the most difficult (/A:/) and easiest phonemes in the production activities (/i:, I/). Pre-test and Post-test Scores As explained in Section 7.4.1, the pre-test and post-test content was identical (each student produced the same words in both tests), but the tests were performed at the beginning and at the end of the experiment. In order to assess the consistency of raters scores of both tests, a Kendall’s coefficient analysis was carried out. A relevant inter-rater agreement was found (Kendall’s coefficient W = 0.493; items = 900, raters = 3, p= 3.1e-19). A high correlation between the scores assigned by the raters to the speakers was also reported. In particular, the Pearson correlations between the mean scores assigned to the speakers in the pre-test are: r= 0.87, p< 0.001 between Rater1 and Rater2, r= 0.73, p< 0.001 between Rater1 and Rater3, and r= 0.79, p< 0.001 between Rater2 and Rater3. In the case of the post-test are: r= 0.97, p< 0.001 between Rater1 and Rater2, r= 0.94, p< 0.001 between Rater1 and Rater3, and r= 0.95, p< 0.001 between Rater2 and Rater3. Table 7.19 shows the average scores assigned by the raters to the pre/post-test utterances. There were a total of 1200 scores for the in-classroom group (8 participants x 25 minimal pairs x 3 raters x 2 tests) and 1500 scores for the experimental group (10 participants x 25 minimal pairs x 3 raters x 2 tests). A comparison of pretest and post-test scores, granted by the three human raters (column Rater: 1,2,3),
112 Chapter 7. Experiments Group Rater Pre-test Post-test Difference (Wilcoxon signed-rank test) mean N mean N mean N Z r p Experimental 1 0.82 250 2.53 250 1.71 250 -7.864 0.50 <0.001 Experimental 2 0.99 250 2.45 250 1.46 250 -8.148 0.52 <0.001 Experimental 3 0.55 250 2.38 250 1.83 250 -7.422 0.47 <0.001 Experimental 1,2,3 0.85 750 2.59 750 1.74 750 -13.551 0.50 <0.001 In-classroom 1 0.41 200 0.68 200 0.27 200 -2.281 0.16 0.023 In-classroom 2 0.63 200 0.86 200 0.23 200 -3.056 0.22 0.002 In-classroom 3 0.27 200 0.61 200 0.34 200 -2.597 0.19 0.009 In-classroom 1,2,3 0.41 600 0.75 600 0.34 600 -4.566 0.20 <0.001 TABLE 7.19: Pre-test and post-test mean production scores of the English Vowels prototype, adapted from [23]. Mean is the average score assigned by a rater in a [0, 10] scale. The p-value is 2-tailed. shows that there was improvement in both groups: from 0.85 to 2.59 in the experimental group, and from 0.41 to 0.75 in the in-classroom group. Since the content of pre-test and post-test was identical, a word–by–word comparison could be carried out between pre/post-test utterances of the same items by each student. In particular, a Wilcoxon signed-rank test found statistically significant differences between pronunciation improvement in both groups. The CAPT-group obtained an improvement of 1.74 points (Z= -13.551, p< 0.001), with a large effect size (r= 0.50); and the in-classroom group obtained a 0.34 improvement (Z= -4.566, p< 0.001) with a small effect size (r= 0.20). In the case of the pre-test, results report that their scores were homogeneous at the beginning of the experiment since there were no statistically significant differences between groups scores (U = 18.0, p= 0.055, Mann–Whitney Utest) with a moderate effect size (r= 0.46). However, there were significant differences between both groups in the scores of the post-test (U = 9.0, p= 0.001, Mann–Whitney Utest), with a large effect size (r= 0.65). That means students who trained with the CAPT system achieved better pronunciation improvement values than learners in the in-classroom group regarding scores at the end of the experiment, both in absolute (1.74 vs. 0.34 of improvement) and in relative terms (205% vs. 82% of improvement). Group Post-test selected Pre-test selected Indifferent NA Experimental 73 (36.5%) 35 (17.5%) 79 (39.5%) 13 (6.5%) 13% 47% 28% 9% 2% 48% 34% 14% 1% 2% 26% 79% In-classroom 25 (15.5%) 18 (11.3%) 109 (68.1%) 8 (5.0%) 12% 28% 36% 24% 5% 27% 50% 16% 0% 2% 27% 73% TABLE 7.20: ABX test results of the English Vowels prototype. Indifferent indicates how often the rater shows no preference for any of the two stimuli. NA means ’no answer’. The second row of each group contains the Likert values (in percentage) assigned to each preferred option: from left to right, the left column means native like pronunciation and the right column means absolutely non-native pronunciation. Table 7.20 shows the results of the ABX test (described in Section 7.4.10). Although six raters answered this test, only four of them showed a reasonable interrater reliability index (Fleiss Kappa index fair agreement, k= 0.393, Z = 13.8 , p
7.4. Guided Learning Experiment 113 < 0.001) and were thus, kept to compute results in Table 7.20. A 60% of the experimental group samples received the highest marks in the post-test, against the 40% for the in-classroom group. A Pearson Chi–square test found statistically significant differences between the experimental and in-classroom group preferences (χ2(2) =30.461, p< 0.001). Raters tended to prefer the post-test performances of the experimental group students: 73 vs. 35, (p< 0.001, Binomial test), contrasting with the in-classroom group results: 25 vs. 18 (with no statistically significant differences). ID Rater1 Rater2 Rater3 Raters score Game score score score (mean) score (G) 1 5.52 5.94 5.92 5.79 8.30 2 3.76 3.92 4.52 4.07 7.80 3 3.62 3.90 2.54 3.35 8.10 4 2.40 2.62 2.56 2.53 8.10 5 2.24 2.54 2.52 2.43 7.40 6 2.16 3.04 2.32 2.51 7.80 7 2.12 1.58 0.90 1.53 7.40 8 2.08 2.28 2.04 2.13 7.40 9 1.28 1.96 0.32 1.19 7.30 10 0.44 0.70 0.18 0.44 7.10 TABLE 7.21: Correlation between the software and human raters post-test scores of the English Vowels prototype, adapted from [23]. Learners belong to the experimental group. Score’s scale is [0, 10]. ID is the user identifier. Subjective scores coming from the preand post-test correction by experts and objective ones were compared in Table 7.21. The average scores assigned by each rater to the subjects and the game score obtained by each one of them training with the CAPT system are displayed. Individually, the Pearson correlation between the game score and each rater was r= 0.84 for Rater1 (p= 0.002), r= 0.86 for Rater2 (p= 0.001) and r= 0.79 for Rater3 (p= 0.007). Besides, the correlation of Rater1, Rater2, and Rater3 together was r= 0.84 (p= 0.002). Finally, a potential rater score, Rater’, can be obtained from the CAPT tool score (G), with an average error of ±5.5% using a linear regression model (see Figure 7.18): Rater0=−21.724 +3.171 ∗G(7.9) In order to further validate the results about improvement between preand posttests, a correlation study was carried out between the time spent by students to fulfill the test and the results of this test. Each participant took an average of 79.72 seconds to complete the pre-test (59 seconds min. and 107 seconds max.) and an average of 95.61 seconds to complete the post-test (62 and 140 seconds min. and max.). A moderate correlation was found (r= 0.506, p= 0.032; and r= 0.459, p= 0.055, respectively) between the time that students spent on the post-test, on the one hand, and their performance (human raters score in the post-test) and achieved learning (difference between post-test and pre-test human raters score), on the other. These values suggest a certain impact of the time spent on the post-test over user’s performance and learning, but although post-test time was generally higher than pre-test time, the correlation between pre-test and post-test time, calculated as r= 0.75, p< 0.001, also suggests a dependence on the speaker. These results suggest that a proper
114 Chapter 7. Experiments FIGURE 7.18: Correlation between the game and human raters posttest scores of the English Vowels prototype, adapted from [23]. Circles, squares, and rhombuses represent the Rater1, Rater2, and Rater3 average scores, respectively. evaluation of the incidence of test completion time requires further and rigorous experimentation to get undeniable conclusions. 7.4.13 Japañol Prototype Results Results extracted from this prototype tried to reinforce the same answers to the research questions of the previous prototype, English Vowels: RQ2 and RQ1 (and their Issues), following the steps defined in the research objectives RO1, RO2, RO3, and RO4. Results related to participants from the Language Center of the University of Valladolid have been partially published in [25], [26], [27]. Other contributions related to this prototype have been partially published in [28], [29]. User’s Performance Results related to the interaction with the CAPT system by users of the experimental group (18) are displayed in Table 7.22.n,m, and Mare the mean, minimum and maximum values, respectively. Time (min) row stands for the time spent on minutes per learner in each mode in the three sessions of the experiment. #Tries represents the number of times a mode was executed by each user. The symbol - stands for ’not applicable’. Mand. and Req. mean mandatory and requested (listening). The TTS system was employed in both listening types; whereas the ASR was only used in the #Productions row. Japañol learners spent an average of 100.56 minutes performing the proposed activities in the three training sessions. A 85.25% of this time was consumed by carrying out interactive training modes (Exposure, Discrimination, Pronunciation, and Mixed modes). As a mean term, users listened to the TTS system 612.6 times and produced 291.8 times with the ASR system, reaching a rate of 9.0 uses of the TTS/ASR per minute. Important differences in the level of use of the tool depending on the user are also illustrated in Table 7.22. For instance, the fastest learner performing pronunciation
7.4. Guided Learning Experiment 115 Theory Exposure Discrimination Pronunciation Mixed n m M n m M n m M n m M n m M Time (min) 14.80 8.7 20.8 19.7 12.8 21.9 7.1 4.1 13.8 42.6 22.4 72.9 16.4 7.6 30.5 #Tries 7.8 6 10 10.6 7 16 8.5 7 15 9.8 7 14 6.7 3 10 #Mand.List. - - - 287.8 210 390 91.7 70 134 - - - 20.2 9 30 #Req.List. - - - 99.3 53 157 33.0 0 153 54.9 0 127 25.6 6 60 #Discriminations - - - - - - 91.7 70 134 - - - 20.2 9 30 #Productions - - - - - - - - - 208.8 116 356 82.9 38 181 #Recordings - - - 62.4 42 81 - - - - - - - - - TABLE 7.22: User’s performance with the CAPT system of the Japañol prototype. activities spent 22.43 minutes; whereas the slowest one took 72.85 minutes. This contrast can also be observed in the time spent on the rest of the training modes and in the number of times learners practice each one of them (row #Tries). The interuser differences affect both the number of times the users made use of the ASR (154 minimum vs. 537 maximum) and the number of times they requested the use of TTS (59 vs. 497 times). Task Event [fl] [fR] [l] [R] [rr] [s] [T] [f] [fu] [xu] Total Disc. Right t-s 123 (65.8%) 115 (62.5%) 239 (76.1%) 217 (71.4%) 215 (85.7%) 95 (74.8%) 214 (89.2%) 104 (96.3%) 115 (77.2%) 111 (74.0%) 1548 (76.9%) Disc. Wrong t-s 64 (34.2%) 69 (37.5%) 75 (23.9%) 87 (28.6%) 36 (14.3%) 32 (25.2%) 26 (10.8%) 4 (3.7%) 34 (22.8%) 39 (26.0%) 466 (23.1%) Disc. Mand.List. 187 184 314 304 251 127 240 108 149 150 2014 Disc. Req.List. 65 52 139 115 51 45 45 16 89 103 720 Prod. Right t-e 170 (31.8%) 137 (25.1%) 253 (45.1%) 289 (52.6%) 252 (51.5%) 134 (24.6%) 226 (28.4%) 116 (51.8%) 138 (19.2%) 140 (21.4%) 1855 (33.0%) Prod. Wrong t-e 364 (68.2%) 408 (74.9%) 308 (54.9%) 260 (47.4%) 237 (48.5%) 410 (75.4%) 571 (71.6%) 108 (48.2%) 580 (80.8%) 513 (78.6%) 3759 (67.0%) Prod. Right t-s 170 (78.3%) 137 (68.2%) 253 (84.9%) 289 (89.2%) 252 (89.0%) 134 (66.7%) 226 (67.7%) 116 (85.9%) 138 (56.6%) 140 (60.6%) 1855 (75.2%) Prod. Wrong t-s 47 (21.7%) 64 (31.8%) 45 (15.1%) 35 (10.8%) 31 (11.0%) 67 (33.3%) 108 (32.3%) 19 (14.1%) 106 (43.4%) 91 (39.4%) 613 (24.8%) Prod. Mand.List. - - - - - - - - - - - Prod. Req.List. 128 125 105 103 70 146 202 29 240 186 1334 TABLE 7.23: Right, wrong, and listening events categorized by sounds of the Japañol prototype. In Table 7.23 Disc. and Prod. correspond to discrimination and production tasktypes, respectively. Right t-s refers to a correct task sequence. Wrong t-s means incorrect task sequence: in Disc. it refers to a wrong attempt of discrimination; in Prod. it refers to five misproduction task-events. Right/Wrong t-e are correct/incorrect single production events. A wrong t-e occurs when the ASR does not include the produced word (or a homophone) in the first position of the n–best list of hypotheses. Mand. and Req. mean mandatory and requested (listening). The symbol - stands for ’not applicable’. Analyzing all perception and production events registered in Discrimination, Pronunciation, and Mixed modes it was evidenced that although their sequence success rate values, Right t-s, were similar (76.9% and 75.2%, respectively), the difference increased notably when compared to the single events rate, Right t-e (76.9% vs. 33.0%, discrimination and production, respectively). As a reader’s reminder, while a discrimination task conveyed a single attempt, up to five attempts could be conducted before a production task was passed. Table 7.23 also reveals which sounds were the most difficult ones for the learners since there were important differences that affect both discrimination and production activities. For instance, discrimination success rates related to words with sounds [fl] and [fR] reached the lowest values (65.8% and 62.5%, respectively); whereas sound [f] (when contrasting to [T]) seemed to be the easiest one (96.3% success rate). In the case of production activities, sounds [fu] and [xu] showed the lowest production success rates values, both in sequences (56.6% and 60.6%, respectively) and single events (19.2% and 21.4%, respectively). The number of times the TTS was
122 Chapter 7. Experiments Pre-quest Post-quest Competition Subjects Data Test results Log & audio files Experimental group Group of natives Analysis Report Group of staff members Log & audio files Log & audio files Focus group FIGURE 7.21: Steps of the COP prototype’s protocol. to participants who completed at least 60 challenges. A prize was also awarded to the first 15 classified (the better the position of the leaderboard the higher the prize). Finally, during the whole competition, the research team was available to answer emails from users asking for technical help about the installation and execution of the games. 7.5.3 Participants Initially, 354 users signed up online, completed the pre-quest, and started to compete; being 165 the final number of users who performed every stage of the experiment and neither belonged to the research staff (5) nor were native (2): registration, pre-quest, competition, and post-quest stages. There were some subjects who took part in this experiment as support (see Table B.7 for specific details). The rest of participants were native Spanish students at University level. They studied EFL for several years during primary and secondary school. In summary, subjects related to this experiment can be classified as follows: 1. Native Spanish speakers who fully participated (165). They participated during the whole competition. They were awarded with a diploma and an academic certification if they completed 60 or more challenges. The best 15 users of the leaderboard received a reward. 64 speakers took part in the four focus group sessions (16 learners in each one). They were 111 women and 54 men (average age = 21.44, SD = 1.82). 2. Native Spanish speakers who abandoned before the end (182). They were subjects who completed the registration form and the pre-quest, participated some days in the competition without achieving the academic certification, and completed the post-quest about reasons of early abandonment. They were 100 women and 82 men (average age = 21.01, SD = 1.42). 3. Natives (2). They participated only during the first seven days of the competition. They were intended to serve as bait for the rest of the users. They were
7.5. Competitive Learning Experiment 123 awarded with a reward. They were 1 woman and 1 man from USA (average age = 22.05, SD = 0.5). 4. Staff members (5). They took part only during the first seven days of the competition. They were intended to motivate users by submitting and accepting a great quantity of challenges to avoid a possible initial stagnation of the dynamics of the competition. They were 5 men from the ECA-SIMM group. 7.5.4 COP CAPT System Description The competition was run via a gamified CAPT tool for smart devices, called Clash of Pronunciations, COP. This new version of the CAPT system presented in the Non-guided Learning experiment (see Section 7.3.4) consisted in a turn-based social game in which users challenged each other and their results were reflected on a leaderboard. The main goal of a user playing with the developed CAPT tool was to achieve points by performing some pronunciation activities based on the NCM (see Section 2.1.1) in matches of challenges, trying to reach the best position possible in aleaderboard. A match could be played in two modes: Playing and Training. In the Playing mode, users got points by participating in challenges against other subjects (main difference with the Non-guided Learning experiment’s CAPT system). Each challenge involved a minimum of two and a maximum of five participants who performed the same activities in their respective matches. They also included twelve discrimination and production activities, in rows of two. In the Training mode, users played matches individually, which included exposure, discrimination or production activities. However, in this mode users did not get rewards nor points. In this experiment was included the Google ASR, GCSTT, and TTS systems. FIGURE 7.22: COP CAPT system screenshots of discrimination (first picture) and production (second picture) activities in a match. Leaderboard example (third picture). Adapted from [31]. The CAPT tool was populated with a database of 329 American English minimal pairs of vowel and consonant contrasts (English words and their phonetic transcription)6. These words were organized into lists of ten or more pairs of words, which 6https://github.com/eca-simm/minimal-pairs-cop
124 Chapter 7. Experiments each one of them corresponded to a pair of phonemes to contrast. The minimal pairs list to be used in the activities of a match was randomly selected by the system to try to keep the same variety of the difficulty level along the competition. In the discrimination activities of the Playing mode (see the first screenshot of Figure 7.22), users could listen to the sound as many times as they want, although the final score was penalized after the second listening. In the production activities of the Playing mode (see the second screenshot of Figure 7.22), there was a maximum number of three pronunciation attempts per word of the pair with a penalization after the second attempt. Additionally, the system invited users to listen to the correct pronunciation without penalization in the scoring. A production attempt was considered correct (right) when the orthographic transcription of the word (or some homophone) was included in the three first positions of the text hypotheses of the ASR result. There was a time limit of 100 and 10 seconds per production and discrimination activity, respectively. In the Training mode, in both games, users could freely select the list of minimal pairs to train and the activity type (exposure, discrimination, or production). In particular, there was no time limit and users cannot obtain points. Finally, users were rewarded with digital trophies (badges) and motivated with inspirational push messages sent to the CAPT system from the web server to keep users playing and training. In order to establish a competitive game configuration, the guidelines stated by [123] were followed (see Section 5.1 for more details about competitive scenarios in learning). That is, a player to participate in the Playing mode must create a new challenge or accept the invitation of a challenge sent by another users, trying to beat them (interaction with other parties). A challenge starts with the match issued by the player who creates the challenge against a set of selected users. All subjects of a given challenge perform the same discrimination and production activities included in the match in their respective match (nine activities, six for discrimination and three for production, interspersed). In each activity, the points obtained by the users depend directly on the quality of their performance (two points per first-time right attempt and one point in other right attempt. There are no negative points). Players also receive extra points when they beat players with a higher position on the leaderboard. These points are valid when the challenge is finished, that is, when the last user of the challenge performed her/his match. Then, the winner(s) and loser(s) of the challenge (negative goal interdependence) are declared; updating the leaderboard of the competition with their final scores (comparability among participants). See an example of the leaderboard in the last screenshot of Figure 7.22). The winner of the competition is the player who achieves more points during all the competition days, reaching the first position of the leaderboard. The winner of a challenge is the player who achieves the highest MatchScore. Ties can occur in challenges involving more than two players (winning rules). In order to guarantee a similar game level, the possible available opponents for a challenge belong to a range of ten positions above and below the creator’s current leaderboard position. There is a limit of 30 matches per user per day in order to avoid counterproductive extra working load. Only players who complete at least 60 challenges obtain an academic certification (perceived scarcity). Also, subjects who reach one of the fifteen first position on the leaderboard at the end of the competition obtain a reward (quantity of winners). To summarize, a player in this competition has the following options: •Submit (create) challenges. Each user can challenge up to four other learners (from those who are 10 positions below or above her/him of the leaderboard).
7.5. Competitive Learning Experiment 125 The user who initiates the challenge plays its first match and wait for the answer of the rest of users. •Play matches in challenges. Users must perform different activities with minimal pairs and obtain a score. When a player finishes the match, a message with the points and right/wrong attempts is displayed. Then, the player must wait for the other participants’ results in order to declare winners and losers of the challenge and add the points achieved to the leaderboard. •Train. In addition to playing matches, users can choose training activities as an unlimited option. They can choose exposure, discrimination, or production activities of the minimal pair contrasts they want. They do not add points to the leaderboard in this mode. •Respond to the received challenges. The user receives a message about the playmate who is challenging her/him. The incoming challenge can be accepted to perform the match or ignored without counting into the restriction of a maximum of 30 challenges per day (although one is deducted from the creator of the challenge). 7.5.5 Instruments Different sources of data for this experiment can be discerned: •Registration forms: user’s demographic information, such as name, age, gender, L1, academic level, and final consent to analyze all gathered data. This information was carefully collected and saved into digital text documents. •Pre-quest: an online questionnaire about three different categories of questions. The first category included Likert-scale type questions to evaluate the degree of competitiveness of the users. Second, the scale of scholar motivation for (EME-E), subdivided into extrinsic and intrinsic motivation adapted and validated in Spain [197]. Finally, a specific questionnaire for evaluating the self-concept in what concerns to pronunciation and discrimination of sounds level in the context of SLA [198]. This data was gathered into a secure web server. •User’s interaction log files. The CAPT tool gathered data associated with all low-level interaction events and monitors all user activities. This data was saved into local log files and automatically uploaded to a web server (see Section 7.5.6 for more details). •Audio recordings. In this experiment the use of the basic–free Google ASR system for Android and the online speech API called GCSTT was alternated. The latter system allowed us to keep the audio files sent to their speech recognition system. This data was gathered into a secure web server. •Post-quest: an online questionnaire for users who finished the corresponding competition about the usability of the tool (adapted from [199]). Three different questionnaires about reasons for playing, attitude toward competition, and information from users who abandoned the game before completing all the stages were also included. This data was collected and saved into a secure web server.
126 Chapter 7. Experiments •Focus group sessions. The audio of the session was recorded via a camera and the most important quotes and requests of the participants were written by a member of the research team by taking notes. This data was carefully collected and saved into digital text documents. 7.5.6 Metrics Game Intensity and Motivation User’s game intensity was characterized in terms of declared reasons for participating in the competition (motivation), and quantity and regularity in matches’ participation. Data related to user’s motivation was gathered in log files: •Number of active days: amount of days in which a user participates in Training or Playing matches. •Number of attempts: amount of discrimination and production activities performed by a user in Training or Playing matches. •Degree of motivation: subjective answers to the questionnaire at the end of the competition (motivation for participating, feelings during the competition, and reasons for abandonment). Performance Related to the amount of events tracked from each participant. Different indicators of the CAPT tool characterize user’s performance: •Production attempt: every attempt of producing correctly the proposed word of a pair. Binary value (true, false) indicating whether the orthographic transcription of the word matches to the user’s utterance result of the n–best list of hypotheses of the ASR. •Production success rate: percentage of right production attempts according to the total number of attempts of the user in the competition. •Discrimination attempt: every attempt of selecting correctly the word of a pair synthesized by the system. Binary value (true, false) indicating whether the user chooses the word of the minimal pair that the system synthesizes in the activity. •Discrimination success rate: percentage of right discrimination attempts according to the total number of attempts of the user in the competition. •Number of matches (Playing or Training mode) in which the user participates in (either launched or answered matches). •Match duration: time a user spends on performing the activities of a match. •Challenge win rate: number of challenges won by a player divided by the total number of challenges in which the user participated in. •Leaderboard position, rank: place on the competition’s leaderboard that a player occupies during a challenge. •Number of points obtained from a finished challenge in the Playing mode. The final amount of points achieved by a player in a challenge depends on the
7.5. Competitive Learning Experiment 127 Condition ExtraScore IsCreator MaxBaseScore BetterRank n=2n∈ {3, 4,5} True True True 0 0 True True False rank1−rank2 3 ∗(n−1) True False True −(rank1−rank2) −(n−1) True False False 0 0 False True True 0 0 False True False rank1−rank2 (n−1) False False True 0 0 False False False 0 0 TABLE 7.29: Extra points scoring system of COP (ExtraScore value). Table adapted from [31]. performance in her/his corresponding match and the rest of the players. It can be defined as: MatchScore =BaseScore +ExtraScore (7.10) The BaseScore is computed from the performance results of discrimination and pronunciation activities in each player’s match: BaseScore = 6 ∑ D=1 uD+ 6 ∑ P=1 vP;uD,vP∈ {α,β,γ}(7.11) where uDand vPare the weight values assigned to the activity performance value according to the result: αis the value assigned to a wrong attempt (0), β is the value referred to a right attempt with some help, such as a request for a word listening or performing more than one production attempt (1), and γis the value assigned to a right attempt without help (2). The ExtraScore is added after all players in the challenge finish their matches. As shown in Table 7.29, this value depends on the number of players, the player who launches the challenge (IsCreator) and the leaderboard position difference between the player and the opponents (BetterRank) and the BaseScore. IsCreator value is True when the player launched the challenge, MaxBaseScore value is True when the BaseScore achieved by the player is the highest one of the challenge, and BetterRank indicates if the leaderboard position of the player is higher than the position of the opponent(s). rank1 and rank2 are the leaderboard position of the player with the higher and lower position of the leaderboard, respectively. nis the number of players in the challenge. In particular, the extra points scoring system had the premise of rewarding courageous players. Those who challenged players above in the leaderboard and won, obtained more extra points (no penalties in case of losing the challenge). However, top players were penalized when they challenged worse ones in terms of leaderboard position and lost the challenge. Proficiency Improvement The learning improvement analyzed was related to the perception and production skills involved in the activities of the competition. Inter and intra-group success
128 Chapter 7. Experiments rates were compared of the same quantity of activities at the beginning (two first days) and at the end (two last days) of the competition. User Grouping The total number of Playing matches performed by each user is used to classify COP participants by three statistical tertiles in terms of their quantitative level of activity (performance): T1 (Constant), T2 (Habitual), and T3 (Casual), that is, high, medium, and low participation in the competition, respectively. 7.5.7 Results Participants of the TipTopTalk! prototype gradually lost interest, most probably due to habituation and lack of new motivational factors. This led us to analyze the effects of a challenge-based competition on user’s motivation, performance, and learning (RQ3). In comparison to our previous challenge-free version of the game, TipTopTalk!, the COP challenge-based competition ensured a higher and more stable level of motivation, while also providing a measurable increase in correct pronunciation of the phonemes addressed in the game. Both prototypes shared main gamification elements, such as leaderboards, points, profile avatar, badges, and performance graphs. The results obtained in this experiment, along those derived from the three previous ones, also reinforced the answers to the research questions about the inclusion of current speech technology in a CAPT system (RQ1) and the user’s pronunciation performance and learning with a specific pronunciation training methodology (RQ2); by following the research objectives RO1, RO2, RO3, and RO4. The most relevant results can be classified into four different categories. First, results related to user’s behavior with the system (i.e., days of active participation and preferred type of activity). Second, learner’s performance interacting with the CAPT system during the competition. Third, the results of the post-quest questionnaires. Finally, results from the focus group sessions. Some of these results have been published in [31]. Their discussion is included in Chapter 8. User’s Behavior Concerning the distribution of any player’s activity throughout the 24 competition days, Figure 7.23 shows the accumulative number of days in which some players’ activity was traced (as it was shown in Figure 7.8 for the TipTopTalk! prototype of the second experiment). There was a 17% of the total 354 subjects of the experiment who only participated one day; whereas a 25% played more than 11 days and only a 5% of learners were active during the whole 24 day competition period. In Figure 7.24 the total distribution of challenges in which each participant on average was involved during the competition is represented in grey bars, whereas the distribution of average number of challenges performed by a user per day is displayed in the horizontal black line. Although the global activity intensity registered falls along the days of competition as the number of participants, the average of challenges per active player remains constant (≈4.5%). User’s Performance Most of user’s data gathered was related to discrimination and production events on the Playing and Training modes. Table 7.30 represents the average number of
7.5. Competitive Learning Experiment 129 FIGURE 7.23: Distribution of users by number of days with active participation in the COP competition, adapted from [31]. FIGURE 7.24: Distribution of activity registered each day of the COP competition. these two activity events performed by each user in the experiment. In the Playing mode, an average user performed about six times more production activities than an average player of the TTT prototype (2168.0 vs. 349.9, see Table 7.10) and three times more discrimination activities (1413.2 vs. 405.2). A Mann–Whitney Utest shows statistically significant differences in both cases (U= 186.0, p< 0.001 for productions and U= 907.0, p< 0.001 for discriminations). In the Training mode, an average user of COP prototype performed almost four times more production activities (81.3 vs. 24.3) and two times more discrimination activities (72.1 vs. 37.0) than an average TTT user. However, a Mann–Whitney Utest indicates there were statistically significant differences only in production training activities (U= 606.5, p= 0.022). Table 7.31 shows the average values of user’s performance indicators interacting with the CAPT system (described in Section 7.5.6). Individuals were categorized by their English level declared in the pre-quest (native or non-native: A1–A2, B1–B2, C1–C2). Since data in Table 7.31 did not pass the Kolmogorov–Smirnoff nor Levene’s standard tests, several non-parametric tests for non-normally distributed data were carried out, in order to detect statistically significant differences. In particular, in Table 7.32 the results of a Kruskal–Wallis test [200] conducted to determine the possible statistically significant differences among the three non-native groups (C1– C2, B1–B2, and A1–A2); whereas in Table 7.33, statistical pairwise comparisons were
130 Chapter 7. Experiments #Events #Participants|#Total Training mode Discrimination 72.1 (47.0%) 108/128 (84%) Production 81.3 (53.0%) 102/128 (80%) Playing mode Discrimination 1413.2 (39.5%) 165/165 (100%) Production 2168.0 (60.5%) 165/165 (100%) TABLE 7.30: Average number of discrimination and production events per participant of the COP prototype. The third column (#Participants|#Total) refers to the number of subjects who perform these activities (first value) and the total number of participants who perform an activity of the same mode (second value). Table adapted from [31]. C1–C2 (48) B1–B2 (250) A1–A2 (49) Non-native (347) Native (2) Mean SD Mean SD Mean SD Mean SD Mean SD Playing mode Production success rate 66.8% 11.6 55.6% 13.4 45.0% 11.6 55.7% 14.1 89.4% 9.4 Discrimination success rate 79.8% 7.0 74.2% 9.7 69.8% 7.9 74.3% 9.5 95.3% 9.4 Number of matches 205.6 229.4 130.4 168.1 80.4 82.3 133 172.0 198 9.4 Match mean duration (s) 73.0 17.3 90.2 39.4 102.4 26.0 89.5 36.3 43.8 9.4 Leaderboard-related Challenge win rate 53.0% 22.2 44.4% 30.0 22.4% 16.4 43.6% 23.9 88.1% 9.4 Mean position 141.5 108.6 175.1 101.8 213.0 80.7 175.79 101.483 157.0 9.4 Mean number of points 13.5 5.5 11.5 5.1 9.5 6.3 11.4 5.4 21.51 9.4 Training mode Production success rate 29.1% 36.0 32.9% 35.9 32.9% 31.3 32.4% 35.7 - - Discrimination success rate 42.1% 41.8 41.5% 36.5 39.5% 38.6 41.3% 37.4 - - Number of matches 13.5 22.9 19.0 11.9 8.8 16.0 11.7 21.1 - - Match mean duration (s) 37.7 34.4 38.1 31.1 49.6 34.6 40.0 32.2 - - TABLE 7.31: Indicators of activity per declared level of English (CEFR) of the COP prototype. SD is the standard deviation. The symbol - stands for not applicable. carried out determined by Mann–Whitney Utests. Table 7.32 reports statistical differences for all indicators (p< 0.05) except for the Training mode’s production success rate (H(2) = 2.592, p= 0.274), discrimination success rate (H(2) = 0.677, p= 0.713), and number of matches (H(2) = 28.507, p= 0.395). In particular, the A1–A2 students were the least implicated ones in the competition with an average of 80.4 matches (Table 7.31). This quantity value was almost three and two times less than the activity registered by the C1–C2 and B1–B2 players (205.6 and 130.4 matches on average). The C1–C2 players were the most efficient ones since they achieved the highest rates of wins. These values were statistically significant different comparing group by group (see Table 7.33). Regarding the success rate in production and discrimination activities of the Playing mode, the C1–C2 players were also, on average, the most skilled ones, being the differences statistically significant higher than the other groups (see Table 7.33). These differences in the performance of each player have an
7.5. Competitive Learning Experiment 131 H df p Playing mode Production success rate 29.997 2 < 0.001* Discrimination success rate 27.826 2 < 0.001* Number of matches 12.329 2 0.002* Match mean duration (s) 21.400 2 < 0.001* Leaderboard-related Challenge win rate 23.530 2 < 0.001* Mean position 21.754 2 < 0.001* Mean number of points 24.439 2 < 0.001* Training mode Production success rate 2.592 2 0.274 Discrimination success rate 0.677 2 0.713 Number of matches 28.507 2 0.395 Match mean duration (s) 7.183 2 0.028* TABLE 7.32: Kruskal–Wallis test results of indicators of activity per declared level of English of Table 7.31 of the COP prototype. The * symbol means that there were statistically significant differences. C1–C2 vs. B1–B2 C1–C2 vs. A1–A2 B1–B2 vs. A1–A2 Playing mode Production success rate ( 1062.5, p= 0.005) ( 59.5, p< 0.001) ( 474.5, p< 0.001) Discrimination success rate ( 971.5, p= 0.001) ( 56.0, p< 0.001) ( 596.5, p< 0.001) Number of matches – ( 140.5, p= 0.002) ( 746.0, p= 0.005) Match mean duration (s) ( 1200.5, p= 0.033) ( 89.5, p< 0.001) ( 573.5, p< 0.001) Leaderboard-related Challenge win rate (1204.0, p= 0.034) (86.0, p< 0.001) ( 521.5, p< 0.001) Mean position ( 1219.0, p= 0.041) ( 98.0, p< 0.001) ( 543.0, p< 0.001) Mean number of points ( 1164.5, p< 0.020) ( 76.5, p< 0.001) ( 531.0, p< 0.001) Training mode Production success rate – – – Discrimination success rate – – – Number of matches – – – Match mean duration (s) – ( 182.5, p= 0.024) ( 853.0, p= 0.029) TABLE 7.33: Mann–Whitney Utest results (U,p) by declared level of English of Table 7.31 in the COP prototype. impact on the final positions of the leaderboard: C1–C2 players occupied, on average, higher leaderboard positions than B1–B2, and B1–B2 positions were, at the same time, higher than A1–A2 ones: 141.5, 175.1, and 213.0, mean positions, respectively (see Table 7.31). These differences were statistically significant in the three cases (see Table 7.33). However, only eight C1–C2 players reached one of the first 25 positions on the leaderboard, being the 17 positions remaining occupied by B1–B2 learners. These leaderboard positions aligned with the mean number of points obtained per match: 13.5 vs. 11.5 vs. 9.5, respectively (see Table 7.31). There were significant differences in all groups (see Table 7.33). Finally, there was also a last indicator which also evidences a higher proficiency of C1–C2 players than the rest: the average time spent on Playing matches (73.0s, 90.2s, and 102.4s, respectively, Table 7.31). C1–C2 players spent less time than the other group with statistically significant differences in all cases (see Table 7.33). Regarding training activities, although B1–B2 players trained, on average, more than the other groups, there were not statistically significant differences in any Training indicator, except for the Training match duration 37.7s, 38.1s, and 49.6s, C1–C2, B1–B2, and A1–A2, respectively, being statistically significant differences between the A1–A2 group and the other two groups (see Table 7.33). In order to report consistent results, the following results analyzed in this section only include data related to the native Spanish speakers who fully participated in