scieee AI-readable full text Open interactive document viewer

Online adaptation to altered auditory feedback is predicted by auditory acuity and not by domain-general executive control resources

Martin, Clara D.,Niziolek, Caroline A.,Dunabeitia, Jon Andoni,Pérez, Alejandro,Hernández Barros, Doris,Carreiras, Manuel,Houde, John

Full text

This is an electronic reprint of the original article. This reprint may differ from the original in pagination and typographic detail. Author(s): Title: Year: Version: Please cite the original version: All material supplied via JYX is protected by copyright and other intellectual property rights, and duplication or sale of all or part of any of the repository collections is not permitted, except that material may be duplicated by you for your research use or educational purposes in electronic or print form. You must obtain permission for any other use. Electronic or print copies may not be offered, whether for sale or otherwise to anyone who is not an authorised user. Online adaptation to altered auditory feedback is predicted by auditory acuity and not by domain-general executive control resources Martin, Clara D.; Niziolek, Caroline A.; Dunabeitia, Jon Andoni; Pérez, Alejandro; Hernández Barros, Doris; Carreiras, Manuel; Houde, John Martin, C. D., Niziolek, C. A., Dunabeitia, J. A., Pérez, A., Hernández Barros, D., Carreiras, M., & Houde, J. (2018). Online adaptation to altered auditory feedback is predicted by auditory acuity and not by domain-general executive control resources. Frontiers in Human Neuroscience, 12, 91. https://doi.org/10.3389/fnhum.2018.00091 2018 ORIGINAL RESEARCH published: 12 March 2018 doi: 10.3389/fnhum.2018.00091 Frontiers in Human Neuroscience | www.frontiersin.org 1March 2018 | Volume 12 | Article 91 Edited by: Xiaolin Zhou, Peking University, China Reviewed by: Xing Tian, New York University Shanghai, China Matthias Franken, Ghent University, Belgium *Correspondence: Clara D. Martin [email protected] Received: 22 November 2017 Accepted: 23 February 2018 Published: 12 March 2018 Citation: Martin CD, Niziolek CA, Duñabeitia JA, Perez A, Hernandez D, Carreiras M and Houde JF (2018) Online Adaptation to Altered Auditory Feedback Is Predicted by Auditory Acuity and Not by Domain-General Executive Control Resources. Front. Hum. Neurosci. 12:91. doi: 10.3389/fnhum.2018.00091 Online Adaptation to Altered Auditory Feedback Is Predicted by Auditory Acuity and Not by Domain-General Executive Control Resources Clara D. Martin1,2*, Caroline A. Niziolek3, Jon A. Duñabeitia1,4, Alejandro Perez1, Doris Hernandez5, Manuel Carreiras1,2,6 and John F. Houde7 1Basque Center on Cognition, Brain and Language, San Sebastian, Spain, 2IKERBASQUE, Basque Foundation for Science, Bilbao, Spain, 3Department of Communication Sciences and Disorders, University of Wisconsin–Madison, Madison, WI, United States, 4Facultad de Lenguas y Educación, Universidad Nebrija, Madrid, Spain, 5Department of Psychology, Center for Interdisciplinary Brain Research, University of Jyväskylä, Jyväskylä, Finland, 6Basque Language and Communication Department, University of the Basque Country, San Sebastian, Spain, 7Department of Otolaryngology, University of California, San Francisco, San Francisco, CA, United States When a speaker’s auditory feedback is altered, he adapts for the perturbation by altering his own production, which demonstrates the role of auditory feedback in speech motor control. In the present study, we explored the role of auditory acuity and executive control in this process. Based on the DIVA model and the major cognitive control models, we expected that higher auditory acuity, and better executive control skills would predict larger adaptation to the alteration. Thirty-six Spanish native speakers performed an altered auditory feedback experiment, executive control (numerical Stroop, Simon and Flanker) tasks, and auditory acuity tasks (loudness, pitch, and melody pattern discrimination). In the altered feedback experiment, participants had to produce the pseudoword “pep” (/pep/) while perceiving their auditory feedback in real time through earphones. The auditory feedback was first unaltered and then progressively altered in F1 and F2 dimensions until maximal alteration (F1 −150 Hz; F2 +300 Hz). The normalized distance of maximal adaptation ranged from 4 to 137 Hz (median of 75 ±36). The different measures of auditory acuity were significant predictors of adaptation, while individual measures of cognitive function skills (obtained from the executive control tasks) were not. Better auditory discriminators adapted more to the alteration. We conclude that adaptation to altered auditory feedback is very well-predicted by general auditory acuity, as suggested by the DIVA model. In line with the framework of motor-control models, no specific claim on the implication of executive resources in speech motor control can be made. Keywords: speech production, altered feedback, adaptation, auditory acuity, executive control INTRODUCTION Sensorimotor control has been studied for decades by exploring the role of visual feedback in reaching. Many studies have shown that when participants reach for a target, they initially miss it when the visual feedback of their hand position is shifted. This visual feedback alteration, when it is consistent, induces participants to adapt, and adjust their reaches to oppose the feedback shift. Martin et al. Predictors of Adaptation in Speech This learned adaptation, developed gradually in response to consistently altered feedback, is called sensorimotor adaptation (von Helmholtz, 1962; Welch, 1978; Rossetti et al., 1993; Redding et al., 2005). Drawing a parallel between visual and auditory feedback, several authors have shown that speakers adapt for alteration of auditory feedback during speech production. Both phenomena are explained by the existence of an internal forward model that enables human subjects to adjust their motor act on-line to any perturbation (Houde and Jordan, 1998, 2002; Purcell and Munhall, 2006a). This sensorimotor adaptation (SA) is highly variable across individuals. The goal of the present study is to identify some critical factors involved in SA using a speech production task. Sensorimotor Adaptation in Speech It has long been known that auditory feedback has an impact on speech: speakers respond to an increase in surrounding noise level by concomitantly increasing the volume of their speech (Lombard, 1911; Lane and Tranel, 1971), and disruptions of speech are consistently observed when auditory feedback is delayed (Lee, 1950; Yates, 1963). Speakers have also been shown to adapt for unpredictable shifts in pitch (Elman, 1981; Burnett et al., 1998), loudness (Heinks-Maldonado and Houde, 2005; Bauer et al., 2006), and formant frequencies (Purcell and Munhall, 2006a; Tourville et al., 2008; Cai et al., 2012; Niziolek and Guenther, 2013). In addition to these rapid adaptation responses to online feedback shifts, SA in vowel production was observed when speakers listened to their auditory feedback altered by a consistent, learnable shift in formants (Houde and Jordan, 1998, 2002; Purcell and Munhall, 2006b; Villacorta et al., 2007), or fundamental frequency (Jones and Munhall, 2000, 2005; Xu et al., 2004), showing an aftereffect which persists after the removal of altered feedback and which is thus separable from online adaptation. Sensorimotor adaptation can be explained in the following way: the perceived feedback during the execution of an action is compared with predicted feedback, and any inconsistency between the actual and predicted feedback lead to changes (adaptation) in the motor command parameters (Held and Hein, 1958; Hein and Held, 1962; Welch, 1986). The predicted feedback is generated internally from the conjunction of an efference copy of the motor command parameters with an internal model of both the motor system and the environment (“forward models,” see Jordan and Rumelhart, 1992; Guenther, 1995; Wolpert et al., 1995; Perkell et al., 1997; Tremblay et al., 2003; e.g., DIVA model, Guenther et al., 1998, 2006; Tourville and Guenther, 2011; State Feedback Control model of speech motor control, Houde and Nagarajan, 2011). Regarding the role of auditory feedback in speech production, the forward model would function as follows: When a speaker produces a speech sound, the motor command parameters are sent to the articulatory system while an efference copy of the descending motor command parameters is created. Based on the efference copy, the forward model generates a prediction of what should be the feedback of this oral production. Then, the actual sensory feedback perceived during speech sound production is compared to the predicted feedback. Whenever there is a mismatch between the actual and predicted feedback, there is online adaptation. For example, when a formant alteration is artificially introduced in the actual feedback, a mismatch will be detected, and adaptive motor command parameters will result in the production of a shifted version of the speech sound (Houde and Jordan, 1998, 2002; Purcell and Munhall, 2006a). The online change in motor commands after repeated mismatches leads to a long-lasting change in the motor command parameters. In fact, once the alteration in auditory feedback is removed, there is persistence in producing the shifted version of the speech sound, as shown by the adaptation aftereffect (Houde and Jordan, 1998, 2002; Purcell and Munhall, 2006a). Thus, while adaptation reflects online changes in motor commands, aftereffects can be attributed to persistent recalibration of the target sound which washes out when the alteration is removed. Variability in Adaptation and Aftereffect Adaptation and aftereffect are usually explored through formant alteration in auditory feedback during syllable production. Participants have to pronounce a syllable each time a cue appears on a computer screen. They listen to their auditory feedback through earphones in real time. Formants are not shifted during the initial trials (baseline trials). Then, the first and/or second vowel formant frequencies (F1 and/or F2) are progressively shifted during “ramp” trials to reach a maximal alteration of F1 ±x Hz / F2 ±y Hz. The participant is typically unaware of the alteration and continues to produce the same syllable while unconsciously modulating the verbal production as a consequence of the external alteration (hold trials). The adaptation to alteration is measured as the difference in vowel formants produced during hold vs. baseline trials. At the end of the experiment, the alteration in feedback is removed for the last block of trials. The aftereffect to alteration is measured as the difference in vowel formants produced during end vs. baseline trials. Several studies using similar designs consistently report significant adaptation and aftereffect at the group level: Participants generally alter their production so as to oppose the shift of the feedback alteration (adaptation) and this alteration of production is maintained even after removing the alteration in feedback (aftereffect; Houde and Jordan, 1998, 2002; Purcell and Munhall, 2006a). Nevertheless, when looking at individual adaptation and aftereffect values, a large variability is observed. For instance, Purcell and Munhall (2006b), despite the main significant adaptation, report a large variability in individual adaptation (from 9 to 126 Hz for a 200 Hz alteration). Houde and Jordan (2002) report that some participants adapt almost completely while others do not show any significant adaptation (see Figures 5A,C in Houde and Jordan, 2002; Figure 3 in Houde and Jordan, 1998; see also Villacorta et al., 2007). Cai et al. (2010) observed a significant adaptation to F1 alteration in only 60% of their participants. Note that large individual variability in adaptation and aftereffect was also reported in pitch-shift studies (e.g., Burnett et al., 1998; Sivasankar et al., 2005). One tentative explanation of variability in adaptation and aftereffect can be found in the literature on executive control. According to the major theoretical accounts of cognitive control (Botvinick et al., 2001; Holroyd et al., 2005), Frontiers in Human Neuroscience | www.frontiersin.org 2March 2018 | Volume 12 | Article 91 Martin et al. Predictors of Adaptation in Speech performance is adjusted when a conflict is encountered, and this “conflict monitoring” skill varies among people. During speech production, two sources of information have to be integrated: top-down feedforward information and bottom-up auditory feedback (mandatorily but not exclusively, since other sources of information have to be integrated such as somatosensory information for instance; Lametti et al., 2012). When an alteration is applied to the bottom-up auditory feedback, a conflict between those two sources of information is introduced. If such conflict monitoring requires executive functioning, one would expect people with lowest executive control skills to be poor in adaptation (i.e., poor in conflict monitoring and thus poor in resolving conflict between the two sources of information). Secondly, according to some theoretical postulates (e.g., DIVA model), auditory acuity should be another factor explaining variability in adaptation and aftereffect (Guenther et al., 1998, 2006; Tourville and Guenther, 2011). In the model, speakers’ targets for production are defined as sensory goals in a state space, and more acute speakers have smaller target regions. This assumption is supported by previous work showing that more acute speakers have reduced production variability (Perkell et al., 2004, 2008; Franken et al., 2017) and by models that equate acuity with higher resolution in auditory space, leading to smaller (more precise) target regions (Perkell, 2007, 2012). Here, we can infer that the same feedback alteration will push productions farther outside the target region for individuals with higher auditory acuity (i.e., with smaller target regions). Thus, higher auditory acuity should lead to finer detection of subtle alterations (errors) in auditory feedback, leading to larger amount of adaptation. In the present study, we investigated the relationship between adaptation and aftereffect and various factors that could play a relevant role in SA. In line with cognitive control models and the DIVA model, we specifically targeted two main factors: executive control and auditory acuity. First, if good adapters are better in conflict detection/monitoring they should outperform poor adapters in executive control. Second, if good adapters are better in detecting subtle alterations in feedback perception (i.e., smaller target regions) they should outperform poor adapters in auditory acuity. In other words, we tested the hypotheses that poor adapters are less sensitive to feedback alterations because (1) they perceive their feedback but they do not efficiently monitor errors in it (variability in adaptation would then be mainly explained by executive control skills) or because (2) they cannot perceive their feedback as accurately (variability in adaptation would then be mainly explained by auditory acuity). Here we explore these two hypotheses in greater depth. SA and Executive Control Since one possible explanation for variability in adaptation and aftereffect is differing efficiency of conflict monitoring, we investigated the relationship between executive control and adaptation/aftereffect to altered feedback. The underlying hypothesis is that the same networks that mediate domaingeneral conflict monitoring (cognitive control) also allow better internal self-speech monitoring (Schiffer et al., 2015). We hypothesized that the better the participants are in detecting, monitoring, and resolving conflict between two sources of information, the more they adapt to altered feedback. Such interplay between domain-general and speech-specific monitoring is still debated, and previous literature reveals contradictory results on this topic. On one hand, control of auditory feedback during selfproduced speech perception is thought to be involuntary in nature, independent of general cognitive executive resources (see Hu et al., 2015). In fact, several studies have shown that participants were unable to voluntarily suppress the adaptive response induced by F0 alteration in their feedback (Munhall et al., 2009; Keough et al., 2013; Patel et al., 2014), even when explicitly instructed to ignore the altered feedback (Zarate and Zatorre, 2008; Hu et al., 2015). Moreover, internal forward models do not make claims about general executive resources in auditory feedback processing ( e.g., Hickok et al., 2011; Houde and Nagarajan, 2011). On the other hand, recent studies revealed that feedback control can be modulated by experimental task (e.g., speaking vs. singing; Natke et al., 2003), learning experience (tone language experience; Chen et al., 2012; singing experience; Zarate and Zatorre, 2008), and auditory attentional load (Tumber et al., 2014). Zarate and Zatorre (2008), for instance, showed that adaptive responses to altered feedback can be suppressed when subjects are instructed to ignore their feedback, but this inhibition capacity was observed only in musicians. Tumber et al. (2014) showed that when less attention was available for auditory feedback monitoring (dual-task condition producing a high attentional load) participants adapted less to pitch alterations in auditory feedback (see also Liu et al., 2015 and Scheerer et al., 2016, for evidence of the influence of attention on adaptation). Thus, recent studies suggest that the level of reliance on auditory feedback during self-produced speech listening can vary, even if involuntary in nature. Because auditory attentional load modulates responses to altered auditory feedback (Tumber et al., 2014; Liu et al., 2015; Scheerer et al., 2016), it may be hypothesized that executive control resources play a role in processing competing feedforward and feedback sources of information. Here, we go a step further on the implication of general executive skills in altered feedback processing. We make the hypothesis that the level of adaptation to altered feedback also depends on general executive control skills. We test the prediction that the higher the general executive control skills, the higher the capacity of detecting and monitoring conflict, and consequently the higher the adaptation to the altered feedback. In order to investigate the role of executive control in altered auditory feedback processing, we measured the correlation between general executive control skills, and adaptation to altered auditory feedback. Participants had to perform a CVC production task in which auditory feedback was progressively altered. In a second experimental session, we independently estimated the executive control skills of each participant. In order to estimate conflict monitoring skills, we used classical executive control tasks including confliction resolution (Simon task, Simon and Rudell, 1967; Flanker task, Eriksen and Eriksen, 1974; numerical Stroop task, Tzelgov et al., 1992). Resolution of conflict between two sources of information is assessed by interference effects in those tasks: the larger the cognitive control skills, the Frontiers in Human Neuroscience | www.frontiersin.org 3March 2018 | Volume 12 | Article 91 Martin et al. Predictors of Adaptation in Speech smaller the interference effect (see Methods section for further details). Those tasks are the most commonly used in the literature to measure domain-general executive control capacities (using non-verbal stimuli) in a modality-independent manner (Roberts et al., 2006; Spagna et al., 2015). Three different tasks targeting executive control have been included since they measure different sub-types of conflict monitoring related to various forms of inhibitory processes (Miyake and Friedman, 2012; Duñabeitia et al., 2014). If domain-general executive control skills influence altered auditory feedback processing, we should expect those with the strongest executive control skills to exhibit the largest adaptation to altered feedback. If, as suggested by internal forward models, general executive control resources do not play a relevant role in auditory feedback processing (e.g., Hickok et al., 2011; Houde and Nagarajan, 2011), no correlation between executive control skills and adaptation should be observed. SA and Auditory Acuity One of the main predictions of the DIVA model is that auditory perception affects development of speech motor commands, so that better auditory acuity goes hand-in-hand with better-tuned speech production. In fact, the DIVA model learns the targets for speech sounds by taking an acoustic signal as input, which it uses to learn which motor commands match that signal. Having a more accurate representation of that input (i.e., higher auditory acuity) would result in a smaller region of acoustic space that produces a “match” with the target. Therefore, speech productions at the periphery of the target region would be recognized as an error (i.e., “mismatch”) only for the individuals with high enough acuity to distinguish the peripheral production from the target signal (see Figure 5 in Perkell, 2007)1. In other words, when auditory feedback is altered, the speaker adapts until formant values of her auditory feedback move into the target region. The extent of the target region being smaller for acuity speakers (Perkell, 2007, 2012), those speakers should better adapt their speech production to altered auditory feedback. This prediction (greater adaptation to alteration for better auditory acuity) was explored by Villacorta et al. (2007). In this previous study, participants were exposed to a classical altered feedback experiment. Their feedback was altered online (F1 shift) and the authors observed significant adaptation (online effect during alteration). The authors also measured auditory acuity to vowel formant differences, and showed that auditory acuity significantly correlated with adaptation: the better the auditory acuity, the greater the adaptation to feedback alteration (Villacorta et al., 2007). Note, however, that the link between adaptation and auditory acuity still has to be explored since several further studies did not observe any evidence for crossparticipant correlations between adaptation and auditory acuity for F1 (Feng et al., 2011; Cai et al., 2012). Furthermore, the evidence is scarce regarding the link between auditory acuity 1Note that other factors are known to influence which sounds are considered errors in production. The lexical status of the production, for instance, influences adaptation to alteration, participants tending to avoid crossing word/non-word boundaries (see for instance Bourguignon et al., 2014). Such influence of lexicality was avoided in the present study, since all sounds being produced and/or perceived were Spanish non-words (i.e., /pep/, /pip/, /pap/; see Method section). and aftereffect (when alteration removed). In a study comparing singers and non-singers, Jones and Keough (2008) observed that both groups adapted for F0 altered feedback, but only the singer group showed significant aftereffect. The results of this study suggest that auditory acuity (presumed to be larger in singers than non-singers) might affect aftereffects of adaptation more than adaptation itself. In the present study, participants were tested in an altered auditory feedback paradigm similar to the one used by Villacorta et al. (2007), except that F1 and F2 were simultaneously shifted. We explored the correlation between adaptation and/or aftereffect and auditory acuity at the individual level. To extend previous results obtained by Villacorta et al. (2007), we estimated auditory acuity in several dimensions by using a series of four auditory tests: pitch, loudness, melody, and transposed melody discrimination. First, to extend previous results to general (and not speech-related) auditory acuity, we included tones instead of speech sounds in the tasks. Basic match/mismatch detection at low levels of acoustic processing was assessed through pitch and loudness discrimination tasks. Additionally, two melody discrimination tasks assessed participants’ ability to detect changes in an incoming melody relative to a remembered target melody. We hypothesized that adaptation to altered feedback correlates with performance on these tasks, which parallel the comparison of incoming speech feedback with an internal representation of target speech sounds. Thus, we expected poor adapters to suffer from poor auditory acuity and vice-versa, for the 4 different auditory acuity sub-components tested here. MATERIALS AND METHODS Participants Thirty-six native speakers of Spanish (18 females) took part in the experiment. All subjects were right handed, had normal, or corrected to normal vision and self-reported normal audition. They had no previous history of psychological or neurological disorders. Five participants were removed from analyses because of a large number of failed formant tracks (more than 50% of the trials; see below for explanation). Analyses were then performed on 31 participants. Their mean age was 23.8 ±4.2 years old (range: 19–39). Despite the small sample size usually used in altered feedback experiments (around 10 participants; see for instance Houde and Jordan, 1998, 2002; Purcell and Munhall, 2006b; Villacorta et al., 2007), we tested three times more participants in order to achieve reasonable power in regression analyses. All participants were naïve to the purpose of the study. Participants received a payment of 10eper h for their collaboration. This study was carried out in accordance with the recommendations of the BCBL ethics committee with written informed consent from all subjects. All subjects gave written informed consent in accordance with the Declaration of Helsinki. The protocol was approved by the BCBL ethics committee. Altered Auditory Feedback Paradigm The adaptation paradigm used a feedback alteration device (FAD) that was designed by the last author using digital speech processing methods. The FAD induced real-time alterations Frontiers in Human Neuroscience | www.frontiersin.org 4March 2018 | Volume 12 | Article 91 Martin et al. Predictors of Adaptation in Speech in the subject’s own voice during vocalization. Participants were seated in front of a PC video monitor wearing a headmounted microphone and Sennheiser Koss earphones. They were instructed to pronounce a bilabial consonant-vowelconsonant (CVC) non-word (/pep/). Speech was transduced by the microphone through a Delta 44 sound card and into a computer. Speech was analyzed and re-synthesized in real time by the FAD. The FAD implemented a formant-shifting acoustic transformation and returned the altered feedback via earphones (in place of normal auditory feedback) with a delay of 30 ms (below the 50 ms threshold where speakers begin to be noticeably affected by feedback delays; Kalinowski and Stuart, 1996). It also recorded the formants of each participant’s utterance. Within the FAD, an analysis-synthesis process repeatedly captured from the microphone 3 ms frames of the participant’s speech (32 time samples at an 11.025 kHz sampling rate). Each frame was shifted into a 400 sample buffer, which was analyzed by computing a narrow-band magnitude frequency spectrum. Formants were isolated from the spectrum, modified, and recombined together with pitch and temporal envelope to create a new narrow-band magnitude spectrum. This way, frames were analyzed, modified and re-synthesized into new frames making up the altered speech output (for further details on the signal-processing device, see Katseff et al., 2012). Note that the formant tracking approach was based on linear prediction coding (LPC) analysis. LPC has proven to be a very successful approach to formant tracking for a majority of speakers and speech sounds, but there are always speakers for whom LPC analysis proves to be inaccurate and unstable (five among 36 participants in the present study). LPC analysis is generally good for voiced non-nasal vowels, but there are always a certain number of subjects whose productions will deviate from ideal “one tube” articulations, possibly nasalizing their productions or positioning their tongue so that side paths for air in the vocal tract are introduced. In such case, LPC analysis would be inaccurate resulting in a large amount of failed formant tracks. Furthermore, note that shifting formant peaks was accomplished by shifting poles of the LPC analysis of the input speech, and such pole movement (especially the F1 pole) would noticeably change the overall spectral amplitude if it was not adapted for in the output speech. Experimental Task Participants were asked to pronounce the non-word “PEP” (/pep/) 100 times, once per trial, each time a pink triangle appeared on the computer screen (with a break after every 15 trials). This non-word was chosen because the /e/ target vowel was centered in the vocal space, allowing alteration toward an existing vowel (/i/), and adaptation also toward an existing vowel (/a/). The consonant /p/ was chosen because it is labial, which creates less interfering coarticulation (Recasens, 1999). Participants were unaware that there were four phases in the experiment: (1) Baseline: No alteration was applied to trials 1–20. (2) Ramp: Trials 21–40 were progressively and linearly altered until maximal alteration (from vowel /e/ toward /i/). Progressively altering formant frequencies from /e/ toward /i/ involved decreasing F1 frequency and increasing F2 frequency (i.e., decreasing/increasing peaks of spectral power). For F1, the magnitude values of spectral power progressively decreased until a minimum of −150 Hz in 7.5 Hz steps. For F2, the spectral power increased in 15 Hz step sizes until a maximum of 300 Hz. Alteration was gradually introduced to minimize the participant’s awareness of the alteration. (3) Hold: Trials 41–80 were kept maximally altered. (4) End: Finally, alteration was completely removed for productions 81 until 100. Participants’ utterances were recorded and formants were tracked. During baseline trials, when the participant produced the vowel /e/, his auditory feedback was of this same vowel sound /e/. During hold trials, if the participant produced the vowel /e/ with no adaptation, his auditory feedback would be close to /i/. By shifting the formants of his production toward /a/, a participant could shift his auditory feedback closer to the formants of his original unaltered auditory feedback of vowel sound /e/. It is worth mentioning that in the present experiment as in previous ones using the same paradigm, participants were unaware of the manipulation and did not consciously detect the feedback alteration during ramp and hold trials. They detected the sudden change back to baseline, but none of them realized that this change restored unaltered feedback. Auditory Altered Feedback Data Processing F1/F2 formant values of each participant’s utterance (trial) were measured online by the apparatus, and used for data analyses. Formant values for each utterance were measured as the mean formant frequencies over the vowel interval of the utterance. The vowel interval was determined by thresholding the amplitude envelope of the utterance. The first 5 (baseline) trials were discarded from analyses as considered a period of signal equilibration (Purcell and Munhall, 2006b; Villacorta et al., 2007). Then, missed trial values (until four consecutive) due to a program failure in formant tracking, were created by interpolation of boundary data (Code available at http://www.mathworks.com/matlabcentral/ fileexchange/loadFile.do?objectId=4551&objectType=file). This procedure was applied to 0.9 ±1.4% of the trials on average (range of failed formant tracking from 0 to 5 over 95 trials). Secondly, formant data was smoothed by a robust version of local regression using weighted linear least squares fitting and a 1st degree polynomial model (span of the moving average = 5). This robust version assigns lower weight to outliers in the regression and zero weight to data outside six mean absolute deviations. Finally, formant values recorded from participant’s pronunciations were averaged for three different time frames: Baseline trials (trials 6–20), Hold trials (trials 61–80), and End trials (trials 81–100). Adaptive changes in formant frequency were calculated as the scalar projection of formant change in the direction opposite to the alteration, measured using the following equation (see also Niziolek and Guenther (2013),Figure 2A): C=F1x−F1b F2x−F2b•150 −300 335.41 Frontiers in Human Neuroscience | www.frontiersin.org 5March 2018 | Volume 12 | Article 91 Martin et al. Predictors of Adaptation in Speech where F1xand F2xare the formant values of production x, and F1band F2bare the formant values of the average baseline production. The vector (150 300)is the inverse of the alteration, representing perfect adaptation, and the denominator 335.41 is the magnitude of the alteration in Hz (√[1502+3002]). This equation computes the scalar projection of the difference vector (i.e., how a given trial’s F1 and F2 values differed from those of the mean baseline production) onto the alteration vector (i.e., how much the feedback was shifted: −150 Hz in F1 and 300 Hz in F2). Adaptation was defined as adaptive changes to formant frequencies measured at the end of the Hold phase, i.e., trials 61– 80. To get a single measure of adaptation per subject, we used the median value from these trials. Aftereffect was defined as adaptive changes to formant frequencies measured during the End phase, i.e., trials 81–100. To get a single measure of aftereffect per subject, we used the median value from these trials. Executive Control Tasks Since we hypothesized that poor adapters might have poor domain-general conflict monitoring skills, and so would adapt less for the alteration, we tested participants individually on their executive control skills. Classical psychological measures of executive control include tasks where participants face conflicting information and have to resolve such conflict in order to correctly perform the task. The most widely used psychological measures of executive control are the Simon (Simon and Rudell, 1967), Flanker (Eriksen and Eriksen, 1974), and numerical Stroop (Tzelgov et al., 1992) tasks. Each of the three tasks involves a slightly different type of conflict: the conflict in the numerical Stroop task concerns abstract (numerical) values, the conflict in the Simon task relies on the irrelevant spatial information of the stimulus, and the conflict in the Flanker task comes from surrounding distracters (Paap and Greenberg, 2013). In order to test executive control skills broadly, participants were tested in each of the three tasks. Numerical Stroop Task In the numerical Stroop paradigm (Tzelgov et al., 1992; see also Duñabeitia et al., 2014; Antón et al., 2016), participants have to decide which of two digits simultaneously displayed on a screen is physically larger than the other. A correct execution of the task requires ignoring the numerical values of the digits that could facilitate (congruent trials) or interfere (incongruent trials) with the task. In congruent trials, the physical size of the digits and their numerical magnitude align (e.g., physically larger digits correspond to those with greater numerical values). In neutral trials, both digits have the same numerical value, and there is no matching or mismatching information about the numerical magnitude that could modulate the response. In incongruent trials, the physical magnitude of the digits and their location in the mental number line provide mismatching pieces of information (e.g., a big 3 and a small 7). Forty-eight pairs of digits were presented, with one digit on the left and one digit on the right part of the screen. All digits were displayed in Courier New black font on a white background, small digits in size 32 and large digits in size 48. Sixteen pairs were congruent, 16 were incongruent, and 16 were neutral. Participants had to decide as fast and accurately as possible which digit was physically larger, pressing the left button of the keyboard when choosing the digit on the left, and the right button when choosing the digit on the right. Flanker Task In the Flanker task (Eriksen and Eriksen, 1974), participants have to determine the direction of the arrow displayed at the center of the computer screen. A correct execution of the task requires ignoring the direction of the surrounding arrows (distracters). Distracters are 4 arrows (2 immediately to the left and 2 to the right of the central target arrow) that can facilitate (congruent trials) or interfere (incongruent trials) with the task. In congruent trials, the direction of distracters is identical to the direction of the target arrow. In neutral trials, distracters are horizontal lines with no left or right directionality. In incongruent trials, the direction of distracters is opposite to the direction of the target central arrow. Forty-eight trials (sets of 5 arrows) were displayed at the center of the computer screen. All symbols were displayed in Courier New black font (size 48) on a white background (a dash mark for horizontal lines; the “less than” symbol for left arrow; the “greater than” symbol for right arrow). 16 trials were congruent, 16 were incongruent, and 16 were neutral. Participants had to decide as fast and accurately as possible which direction the central arrow pointed, pressing the left button of the keyboard when the arrow was pointing to the left, and the right button when the arrow was pointing to the right. Simon Task In the Simon task (Simon and Rudell, 1967), participants have to press a left button when a square is displayed on the screen and a right button when a circle is displayed. A correct execution of the task requires ignoring the spatial location of the shape that could facilitate (congruent trials) or interfere (incongruent trials) with the task. In congruent trials, the shape requiring a left button press is displayed on the left of screen and vice versa. In neutral trials, the shapes are displayed at the center of the screen. In incongruent trials, the shape requiring a left button press is displayed on the right of the screen and vice versa. Forty-eight trials (shapes) were presented one by one on the computer screen. Sixteen trials were congruent, 16 were incongruent, and 16 were neutral. Participants had to decide as fast and accurately as possible whether a circle or a square was displayed. Button press mappings were counterbalanced across participants. Executive Control Tasks Processing The processing procedure was similar for the three executive control tasks. Incorrect responses (<2% of the data) and reaction times below or above 2.5 standard deviations from the mean in each condition for each participant (<2.5% of the data) were excluded from the latency analysis. For each task, cognitive control skills were measured as the Frontiers in Human Neuroscience | www.frontiersin.org 6March 2018 | Volume 12 | Article 91 Martin et al. Predictors of Adaptation in Speech difference in mean reaction time between congruent and incongruent trials (hereafter called “interference effect”), and then submitted to regression analyses. Given the very high accuracy in each participant and each task, accuracy scores were not included in the correlation analyses. For the sake of completeness, two other measures were calculated: The congruency effect was measured as the difference in mean reaction time between neutral and congruent trials. The incongruity effect was measured as the difference in mean reaction time between incongruent and neutral trials. Results were similar when submitting those measures to regression analyses. It could be argued that auditory executive control tasks would have been better suited than visual ones. However, these tasks are claimed to reflect general executive control capacities in a modality-independent way (see for instance Roberts et al., 2006; Spagna et al., 2015). Moreover, it should be kept in mind that one of the aims of the current study was to tease apart effects attributable to executive control and those due to auditory acuity. For those two reasons, visual tasks were better suited for our study. Auditory Acuity Tasks Four different auditory acuity tasks were used in this experiment (taken from Foster and Zatorre, 2010a,b; Voss and Zatorre, 2011). Two of them measured low-level auditory processes (pitch and loudness discrimination), and the other two evaluated higher-level auditory processing (simple melody discrimination and transposed melody discrimination). The two low-level tasks contained four blocks each while the high-level tasks consisted of two blocks each. Task order was counterbalanced. Pitch Discrimination Task Participants had to decide which of 2 tones was higher in pitch (ISI =1,500 ms). They had to press the left mouse-button when the first sound was higher in pitch, and the right button when the second sound was higher. The reference tone was a 500 Hz pure tone (500 ms), and the task followed a 2-down/1-up staircase procedure (Levitt, 1970; the initial difference between the 2 tones was stepped down after two sequential correct responses and stepped up after a single incorrect response). This procedure produces runs of increasing and decreasing the difference between stimuli whose endpoints (reversal points) bracket the 71% discrimination threshold. The initial difference in frequency was 7%, and the initial step factor was 2 (to converge rapidly onto the subject’s approximate threshold). After two reversals, the step factor was reduced to 1.25 to determine the threshold with greater precision. One staircase run was completed after 15 reversals, and the geometric mean of the value of the last eight reversals was taken as the threshold. The threshold was therefore unaffected by the choice of starting difference because the first 7 endpoints were not entered into the calculation. Four separate runs were conducted for each subject and averaged to produce the final discrimination threshold. The median of the geometric mean obtained for each of the four blocks was calculated and included in the final regression analysis. Loudness Discrimination Task Participants had to decide which of 2 tones was louder (ISI =1,500 ms). The procedure was identical to that of the pitch discrimination task except that the tones differed in loudness, not pitch. The standard loudness reference of the tones was set at 65 dB sound pressure level and the initial difference was set at 10 dB. Participants had to press the left mouse-button when the first sound was louder, and the right button when the second sound was louder. As in the pitch discrimination task, the median of the geometric mean obtained for each of the four blocks was calculated and included in the final regression analysis. High-Level Auditory Discrimination Tasks For the two higher-level auditory tasks, subjects had to decide whether two successive sequences were identical or not via a button press on a 2-button mouse. The sequences were composed of 5–13 notes, and each note was 320 ms in duration, equivalent to eighth notes at a tempo of 93.75 beats per min. In the simple melody discrimination task, participants had to decide whether 2 sequential melodies were identical or different. All stimuli were unfamiliar melodies in the Western major scale. In half of the trials (“different” trials), the pitch of a single note was changed, anywhere in the melody, by up to ±5 semitones (median of 2 semitones). The change maintained the key of the melody as well as the melodic contour. The number of notes in a melody was progressively increased during the task. The task was divided into two blocks. The transposed melody discrimination task was identical to the previous task with 2 exceptions. First, all the notes of the second stimulus pattern were transposed 4 semitones higher in pitch (both in the “same” and “different” trials). Second, in “different” trials, one note was altered by 1 semitone to a pitch outside the pattern’s new key, maintaining the melodic contour. This task therefore required the listener to compare the pattern of pitch intervals (frequency ratios) between each successive tone, and not the absolute pitches of each tone, since those were always different in the 2 melodies of the pair due to the transposition. The number of notes in a trial was progressively increased during the task. This task required a more abstract relational processing, as opposed to the simple melody task, which could be accomplished by direct comparison of the individual pitch values. For the two high level auditory discrimination tasks, percentages of correct responses were calculated to be included in the regression analysis. Correct responses were “same” responses when the stimuli were identical and “different” responses when the stimuli were different. Total scores were divided by 120 (total number of trials). RESULTS Auditory Altered Feedback Median trial-by-trial adaptation for the 31 participants is shown in Figure 1. As a group, participants adapted progressively during ramp trials and reached a plateau during hold Frontiers in Human Neuroscience | www.frontiersin.org 7March 2018 | Volume 12 | Article 91 Martin et al. Predictors of Adaptation in Speech FIGURE 1 | Median trial-by-trial adaptation and aftereffect values in Hz (for a 335.41 Hz maximal shift). Baseline trials =5–20; Ramp trials =21–40; Hold trials = 41–80; End trials =81–100. Error bars reflect standard errors. trials (41–80). Adaptation progressively decreased after abrupt removal of alteration (trials 81–100). Median adaptation and aftereffect values obtained for each participant in the altered auditory feedback paradigm are shown in Figure 2. The majority of the participants adapted during full alteration (hold trials) but a large individual variability was observed. Adaptation values ranged from 4 to 137 Hz (median = 78 ±35), aftereffect values from −26 to 114 Hz (median =26 ±39). Note that the magnitude of the shift was the same for all participants (335.41 Hz), meaning that adaptation ranged from 1.0 to 40.8% (median =23.3) of maximal alteration, aftereffect ranged from −7.8 to 33.9% (median =7.7). Executive Control The interference effects obtained in the three executive control tasks also varied across participants. Interference effects ranged from −120 to 30 ms (median = −46 ±37) in the Flanker task, from −123 to 4 ms (median =−40 ±31) in the Simon task and from −117 to 27 ms (median =−20 ±37) in the Stroop task (see Appendix A for individual values). Auditory Acuity Median discrimination thresholds ranged from 0.2 to 3.1 (median =1.2 ±0.8) for pitch and from 0.3 to 2.1 (median = 0.8 ±0.4) for loudness. Percentages of correct responses in the melody discrimination tasks ranged from 67 to 90% (median = 77 ±0.6; see Appendix A). Regression Analyses Regression analyses were performed on the data across subjects, comparing adaptation and aftereffect to the series of measures obtained in auditory acuity and executive control tasks. We first examined correlations between the following variables: (1) Interference effects in the three executive control tasks (numerical Stroop, Simon and Flanker); (2) Loudness and Pitch discrimination thresholds; (3) Percentages of correct FIGURE 2 | Individual median adaptation and aftereffect values in Hz (for a 335.41 Hz maximal shift). Adaptation =Median shift in production in trials 61–80. Aftereffect =Median shift in production in trials 81–100. Each line links online adaptation and aftereffect values for one participant. responses (%CR) in the Simple Melody and Transposed Melody discrimination tasks. Not surprisingly, and as it has been already shown in earlier studies (e.g., Duñabeitia et al., 2014), the three interference effects Frontiers in Human Neuroscience | www.frontiersin.org 8March 2018 | Volume 12 | Article 91