scieee AI-readable full text Open interactive document viewer

Seeing a talking face matters: The relationship between cortical tracking of continuous auditory ‐visual speech and gaze behaviour in infants, children and adults

Tan, S.H. Jessica,Kalashnikova, Marina,Di Liberto, Giovanni M.,Crosse, Michael J.,Burnham, Denis

Abstract

Available online 15 April 2022

Full text

NeuroImage 256 (2022) 119217 Contents lists available at ScienceDirect NeuroImage journal homepage: www.elsevier.com/locate/neuroimage Seeing a talking face matters: The relationship between cortical tracking of continuous auditory ‐visual speech and gaze behaviour in infants, children and adults ✩ S.H. Jessica Tan a , ∗ , Marina Kalashnikova b , c , Giovanni M. Di Liberto d , Michael J. Crosse e , Denis Burnham a a The MARCS Institute of Brain, Behaviour and Development, Western Sydney University, Australia b The Basque Center on Cognition, Brain and Language, Australia c IKERBASQUE, Basque Foundation for Science, Australia d School of Computer Science and Statistics, Trinity College Dublin, Dublin, Ireland e Department of Mechanical, Trinity Center for Biomedical Engineering, Manufacturing AND Biomedical Engineering, Trinity College Dublin, Dublin, Ireland a r t i c l e i n f o Keywords: Auditory-visual speech benefit Cortical tracking Gaze behaviour Auditory-visual speech perception Infants Children Adults a b s t r a c t An auditory-visual speech benefit, the benefit that visual speech cues bring to auditory speech perception, is experienced from early on in infancy and continues to be experienced to an increasing degree with age. While there is both behavioural and neurophysiological evidence for children and adults, only behavioural evidence exists for infants –as no neurophysiological study has provided a comprehensive examination of the auditoryvisual speech benefit in infants. It is also surprising that most studies on auditory-visual speech benefit do not concurrently report looking behaviour especially since the auditory-visual speech benefit rests on the assumption that listeners attend to a speaker’s talking face and that there are meaningful individual differences in looking behaviour. To address these gaps, we simultaneously recorded electroencephalographic (EEG) and eye-tracking data of 5-month-olds, 4-year-olds and adults as they were presented with a speaker in auditory-only (AO), visualonly (VO), and auditory-visual (AV) modes. Cortical tracking analyses that involved forward encoding models of the speech envelope revealed that there was an auditory-visual speech benefit [i.e., AV > ( A + V )], evident in 5-month-olds and adults but not 4-year-olds. Examination of cortical tracking accuracy in relation to looking behaviour, showed that infants’ relative attention to the speaker’s mouth (vs. eyes) was positively correlated with cortical tracking accuracy of VO speech, whereas adults’ attention to the display overall was negatively correlated with cortical tracking accuracy of VO speech. This study provides the first neurophysiological evidence of auditory-visual speech benefit in infants and our results suggest ways in which current models of speech processing can be fine-tuned. 1. Introduction When listening to a speaker talk face-to-face, we process visual speech cues as well as the predominant auditory signal. These visual speech cues come from facial movements that occur in tandem with acoustic speech and can provide additional information that augments speech perception both in quiet (e.g., Fort et al., 2013 ; Navarra and SotoFaraco, 2007 ) and in noise (e.g., Moradi et al., 2013 ; Rudmann et al., 2003 ; Schwartz et al., 2004 ; Sumby and Pollack, 1954 ). The augmen- ✩ This research was funded by a doctoral scholarship to the first author funded by the MARCS Institute at Western Sydney University and the HEARing Cooperative Research Centre (CRC), and by HEARingCRC funding to the last author. The second author’s work is supported by the Basque Government through the BERC 2018–2021 program, and PIBA PI-2019–0054, and by the Spanish Ministry of Science and Innovation through the Ramon y Cajal Research Fellowship, PID2019– 105528GA-I00. ∗ Corresponding author. E-mail address: [email protected] (S.H. Jessica Tan). tation of speech perception by visual speech cues, or the auditory-visual speech benefit , has been widely studied. Most of these studies have been conducted with adults, but findings from studies with children and infants suggest that they too benefit from visual speech information, even though the degree of auditory-visual speech benefit increases with age. The studies reported here concern the auditory-visual speech benefit in 5-month-old infants, 4-year-old children and adults. Behavioural studies provide evidence of an auditory-visual speech benefit across ages. For instance, 7.5-month-olds successfully segmented words from a fluent speech stream that was blended with a backhttps://doi.org/10.1016/j.neuroimage.2022.119217 . Received 7 November 2021; Received in revised form 9 April 2022; Accepted 14 April 2022 Available online 15 April 2022. 1053-8119/© 2022 Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/by-nc-nd/4.0/ ) S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 ground voice when the auditory stimuli were paired with videos of a speaker’s talking face, but not when they were paired with a still image of the speaker’s face ( Hollich et al., 2005 ). Studies with children and adults found that children identified phonemes and words better in the auditory-visual modality compared to the auditory-only modality, and that this benefit is evident both in quiet ( Lalonde and Holt, 2015 ) and in noise ( Lalonde and Holt, 2016 ; Maidment et al., 2015 ; Ross et al., 2011 ). Additionally, comparisons between children and adults revealed that adults experienced greater auditory-visual speech benefit ( Maidment et al., 2015 ; Ross et al., 2011 ). The same developmental trend has been found in neurophysiological studies with children and adults. Knowland et al. (2014) presented 6to 11-year-olds and adults with auditory-visual words and with auditoryonly words. Both the children and the adults showed attenuated amplitude and shorter latencies of the auditory P2 event-related potential (ERP) component to auditory-visual compared to auditory-only words, but the adults additionally showed the attenuated amplitude and shorter latencies for N1 for auditory-visual compared to auditory-only stimuli. Together these results suggest that visual speech modulation of auditory ERP components is present, yet not fully developed in children ( Knowland et al., 2014 ). Other ERP studies have measured speech perception in auditoryvisual vs auditory-only and visual-only speech in terms of integration rather than enhancement. The criterion for auditory-visual integration is based on the relative magnitude of neural responses to auditoryvisual (AV) stimuli compared with the summation of neural responses to auditory-only (A) and visual-only (V) stimuli [i.e., by testing whether AV = ( A + V ) no integration, or whether AV > A + V , integration]. Using this method, Kaganovich and Schumaker (2014) revealed that peak amplitudes of N1 and P2, and the latency of P2 were attenuated in auditory-visual compared to the algebraic sum of ERP responses to auditory-only and visual-only /ba/, /da/, and /ga/ syllables in 7–8-yearolds, 10–11-year-olds, and adults, thereby indicating auditory-visual integration at all three ages. In a separate study, adult participants showed a significantly shorter latency of the auditory N1/P2 response peak when presented with /ka/, /pa/, and /ta/ in auditory-visual syllables than in auditory-only or visual-only syllables ( van Wassenhove et al., 2005 ). The same integration approach has not been used with infants; rather, the majority of the electrophysiological studies of auditory-visual speech perception in infants have involved the comparison of neural responses (in the form of ERPs) to congruent versus incongruent auditoryvisual syllables ( Bristow et al., 2009 ; Kushnerenko et al., 2008 , 2013 ) and short phrases ( Hyde et al., 2011 ; Reynolds et al., 2013 ). For example, Kushnerenko et al. (2008) examined 5-month-olds’ neural processing of conflicting auditory-visual syllables that typically result in the McGurk effect. Congruent stimuli consisted of auditory-visual /ba/ and auditory-visual /ga/ while incongruent stimuli consisted of the McGurk effect stimuli (auditory /ba/ dubbed onto a visual /ga/ which usually results in a “da ”or “𝛿a ” response) and a conflicting stimulus (auditory /ga/ dubbed onto a visual /ba/ which usually results in a combination, “bga ”, response). The ERPs in response to the conflicting stimulus were more positive over frontal areas and more negative over temporal areas compared to ERPs in response to the other stimulus types, suggesting that 5-month-olds detected the mismatch between the auditory /ga/ and visual /ba/ but integrated the auditory /ba/ and visual /ga/, treating it the same as they did for the integration of congruent auditoryvisual stimuli. Similar findings were reported in a study that used short phrases. Hyde et al. (2011) presented 5-month-olds with an auditory recording of the phrase, “Oh, hi baby ”, that was either paired with a matched video of a face saying the same phrase or a mismatched video of a face saying a different phrase. Mean amplitude of visual N1 and attentional Nc components were more negative in the asynchronous than the synchronous condition, while mean amplitude of auditory P2 component was more positive in the synchronous than the asynchronous condition. Although these infant ERP studies provide some neural level evidence for auditory-visual integration by comparing neural responses to congruent versus incongruent auditory-visual stimuli, they did not include auditory-only and visual-only conditions and so do not truly quantify auditory-visual integration and, in addition, do not afford comparison with the modulating effect of visual information found in children and adults. Beyond electrophysiological studies, the hemodynamic (fNIRS) approach has been used to investigate infants’ processing of auditoryvisual speech ( Altvater-Mackensen and Grossman, 2016 ; 2018 ). The neural responses of six-month-old German-learning infants were enhanced in the left inferior frontal regions when they were presented with matched auditory-visual speech as compared to when they were presented with mismatched auditory-visual speech ( Alvater-Mackensen and Grossman, 2016 ). A separate study compared infants’ processing of unimodal auditory, visual, and multimodal auditory-visual speech at the neural level by presenting six-month-old German-learning infants with unimodal and multimodal speech stimuli /a/, /e/, and /o/ ( AltvaterMackensen and Grossman, 2018 ). This study revealed that the infant participants did not show differential responses to unimodal and multimodal speech within the frontal regions and between hemispheres. Taken together, ERP studies with adults and children illustrate that auditory-visual integration occurs at a neural level and suggest that visual speech information is beneficial for speech perception. In contrast, infant ERP studies demonstrate only the detection of a mismatch between auditory and visual stimuli, and do not show whether visual speech information augments infants’ speech perception, i.e., whether there is an auditory-visual speech benefit. The fNIRS approach used with infants did not find any difference in neural responses to unimodal or multimodal speech within frontal regions In addition to the paucity of studies investigating auditory-visual speech benefit in infants, a major drawback of these studies in general is that in order to evoke brain responses they require presenting participants with multiple repetitions of identical short stimuli which are averaged and then compared between conditions. In the case of auditory-visual speech perception, this comprises the use of syllables or short phrases, stimuli that are not entirely representative of natural, conversational speech. A recent approach addresses this drawback by assessing cortical tracking, or the mathematical relationship between the speech dynamics and the corresponding brain responses (e.g., Ding and Simon, 2012 ; Fiedler et al., 2019; Golumbic et al., 2013 ; Gross et al., 2013; J. O’Sullivan et al., 2014). This approach has greater ecological validity than ERP approaches, as it allows the use of continuous stimuli rather than discrete, repeated stimuli, e.g., rather than single words, passages that more closely resemble natural speech, such as audiobooks or podcasts. Accordingly, this method has been increasingly used to examine auditory-only speech perception in adults (e.g., Ding and Simon, 2013 ; Ding et al., 2016 ), children ( Di Liberto, Peter, et al., 2018 ; Vander Ghinst et al., 2019 ), and infants (e.g., Jessen et al., 2019 ; Kalashnikova et al., 2018 ). Even so, the few studies conducted with adults so far suggest that cortical tracking is augmented when visual speech information from a speaker’s talking face is provided (e.g., Crosse et al., 2015 ; Crosse et al., 2016 ; O’Sullivan et al., 2019 ). Importantly, although there is evidence that cortical tracking of speech can be reliably measured in children and infants, whether cortical tracking of auditory-visual speech is enhanced in children and infants remains an open question, one that this paper will address. The auditory-visual speech benefit effect rests upon the assumption that listeners attend to a speaker’s facial movements. It is thus somewhat surprising that most auditory-visual speech perception studies do not concurrently examine participants’ looking behaviour to the speaker’s face (although see Foxe et al., 2015 ). It has been shown that while the eyes convey emotional and social information, the mouth translates information closely related to the temporal and acoustic properties of speech ( Yehia et al., 1998 ). Face viewing studies indicate that humans are cognisant of the various types of information that different facial features provide and will shift their gaze from one facial 2 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 region to another accordingly (e.g., Buchan et al., 2008 ; Lansing and McConkie, 1999 ). This attentional shift is observed even in infants as young as 6 months ( Tenenbaum et al., 2013 ). In addition, idiosyncratic differences between individuals in facial scanning patterns of the eye and mouth regions are related to perceptual performance ( Gurler et al., 2015 ; Mehoudar et al., 2014 ; Peterson and Eckstein, 2012 ). For instance, Gurler et al. (2015) found that individuals who report experiencing the McGurk effect more frequently also spend a larger proportion of time fixating on the speaker’s mouth. This finding points toward the strong likelihood that individuals’ idiosyncratic preferences in their fixation of the speaker’s mouth or eyes will influence the extent to which visual speech information augments their speech perception. Interindividual variations in looking behaviour to the speaker’s face may result in subtle but significant differences in speech perception. For example, the opening and closing of the mouth corresponds to the syllabic timescale of auditory speech ( Chandrasekaran et al., 2009 ), thus providing the richness of redundant cues relating to the start and end points of syllables that may augment speech perception, especially for listeners who fixate on the speaker’s mouth region. This pertains particularly to young infants in normal listening conditions because they are just beginning to acquire a language system. In this regard, Lewkowicz and Hansen-Tift (2012) provided evidence of a developmental trend in looking behaviour: infants move away from preferential attention to the speaker’s eye region to attending more to the speaker’s mouth region sometime between 4 and 8 months, and then back to attending more to the speaker’s eye region by 12 months of age. As this pattern coincides with the developmental timeline of speech production ( Imafuku et al., 2019 ), the researchers propose that the initial eye-tomouth attentional shift reflects infants’ attempt to extract the redundant cues present in auditory-visual speech while the second attentional shift converges with adults’ looking behaviour to a talking face and suggests some level of language expertise that reduces the need to focus specifically on the speaker’s mouth ( Lewkowicz and Hansen-Tift, 2012 ). Notably, relative attention to a talker’s mouth at 6 months is positively related to expressive language skills both then ( Tsang et al., 2018 ) and at 18 months ( Young et al., 2009 ), and to receptive vocabulary at 12 months ( Imafuku and Myowa, 2016 ). Failure to attend to the speaker’s mouth is associated with later language learning disorders ( Pons et al., 2019 ). Adults, by comparison, are proficient language users and instead focus more on the talker’s eye region under optimal listening conditions but will increasingly direct their attention to the talker’s mouth as listening situations become more challenging, such as when there is background noise (e.g., Buchan et al., 2008 ; Stacey et al., 2020 ; VatikiotisBateson et al., 1998 ). These findings raise the possibility that individuals’ idiosyncratic differences in looking patterns to a talking face will influence the degree of auditory-visual speech benefit experienced. Investigating whether this is indeed the case forms the second aim of this study. 1.1. This study and the hypotheses To examine whether cortical tracking of auditory-visual speech is enhanced in infants and children, and whether gaze behaviour modulates the extent of auditory-visual speech benefit, EEG and gaze data were simultaneously recorded as 5-month-old and 4-year-old participants watched short clips of a speaker in auditory-only (AO), visualonly (VO), and auditory-visual (AV) presentation modes. AO presentations consisted of still photos of the speaker’s face paired with auditory recordings, VO presentations consisted of silent videos of the speaker talking, and AV presentations consisted of both the videos and the auditory recordings. As this paradigm has been used previously with adult participants ( Crosse et al., 2015 ; Crosse et al., 2016 a, 2016 b), a group of adults was tested as a control. Behavioural studies illustrate that the auditory-visual speech benefit is evident across development. Neurophysiological studies show the same for children (using ERPs) and adults (using ERPs and cortical tracking), while none have yet directly examined the auditory-visual speech benefit in infants. Even so, ERP studies with infants that investigated their detection of auditory-visual asynchrony coupled with behavioural findings suggest that the auditory-visual speech benefit may also be evident at the neurophysiological level in infants. With these considerations in mind, we hypothesise that, across the three age groups, (1) cortical tracking of the speech envelope will be most accurate during AV presentations, followed by AO then VO presentations, and (2) auditory-visual speech benefit will be evident as indexed by the additive criterion [i.e., AV > ( A + V )]. Next, facial scanning and speech perception findings suggest that gaze behaviour may modulate cortical tracking accuracy differently for infants compared to children and adults. At five months, infants are likely to be in the process of shifting their attentional focus from the speaker’s eyes to the speaker’s mouth region ( Lewkowicz and Hansen-Tift, 2012 ; Pons et al., 2015 ). Furthermore, 5-month-olds are in the process of acquiring language and may benefit from any additional information that can be extracted from visual speech cues. Accordingly, we hypothesise that the proportion of time that infants spend attending to the speaker’s mouth will be positively correlated with cortical tracking accuracy when visual speech information is available, i.e., during VO and AV presentations. On the other hand, the same positive correlation is not expected for 4-year-olds and adults, given previous findings that older children and adults focus more on the speaker’s eyes when the auditory speech signal is clear (e.g., Lewkowicz and Hansen-Tift, 2012 ), presumably because the acoustic properties from the auditory signal are sufficient for speech perception and they turn to the eyes to seek out emotional and social information that may not be conveyed as clearly by auditory speech. 2. Methods 2.1. Participants Five-month-olds : A final sample of eighteen 5-month-old infants from Australian English monolingual backgrounds were included (M age = 5.49 months, SD = 0.30 months, 8 females). This sample size was decided upon by drawing on previous neurophysiological studies that investigated infant neural processing of AV asynchrony (e.g., Hyde et al., 2011 ; Kushnerenko et al., 2008 ; Reynolds et al., 2013 ) and compared children’s and adults’ neural processing of AV speech (e.g., Kaganovich and Schumaker, 2014 ; Knowland et al., 2014 ). An additional 20 babies were tested but excluded because of fussiness ( n = 6), excessively noisy EEG recordings ( n = 11), or insufficient gaze data ( n = 3). The attrition rate in this study is not uncommon for infant EEG studies (e.g., deBoer et al., 2007 ; Hyde et al., 2011 ; Reynolds et al., 2013 ). All infants came from a monolingual Australian English-speaking background. Four-year-olds : A final sample of 19 Australian English monolingual 4-year-olds were included (M age = 4.16 years, SD = 0.14 years, 12 females). An additional 14 children were tested but excluded because they were very fidgety and did not complete the experiment ( n = 5), had excessively noisy EEG recordings ( n = 3), or had insufficient gaze data ( n = 7). Adults : A final sample of 18 Australian English monolingual adults aged between 18 and 56 years were included (M age = 23.42 years, SD = 8.75 years, 15 females). An additional eight adults were tested but excluded because seven had insufficient gaze data, and one experienced technical failure. All infants and children were born full-term, not at-risk for any cognitive or language delay, with normal hearing and vision, and no history of ear infections. Prior to the study, their parents provided written informed consent, were briefed about the procedure and told that the session would terminate immediately if they wished so, or if their child showed any signs of distress during the session. All adult participants had self-reported normal hearing and normal or corrected-to-normal vision, were free of neurological diseases, and provided written informed 3 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 consent. Adult participants took part in this study as part of a Psychology course requirement and received research participation points. This study was approved by the Human Research Ethics Committee at Western Sydney University (approval number H11517). The approved protocol regarding participant recruitment, data collection and data management was adhered to. For all groups of participants, noisy EEG recordings were defined as datasets that contain more than 20 bad channels as in previous infant studies (e.g., Kalashnikova et al., 2018 ). Additionally, for analysis purposes, participants were required to have at least 10 out of 30 common trials across the three conditions (auditory-only, visual-only, and auditory-visual) with a minimum of 15% attention (as calculated by 𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 = 𝑡𝑜𝑡𝑎𝑙 𝑓𝑖𝑥𝑎𝑡𝑖𝑜𝑛 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 𝑡𝑜 𝑠𝑐𝑟𝑒𝑒𝑛 𝑑𝑢𝑟𝑖𝑛𝑔 𝑡𝑟𝑖𝑎𝑙 𝑡𝑟𝑖𝑎𝑙 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 ) to be included in the final sample. The exclusion criterion for attention (at least 15% attention in a minimum of 10 common trials) was decided upon because previous eye-tracking studies with young infants have used similar exclusion criterion (e.g., 15% in LoBue et al., 2016 ; 20% in Taylor & Herbert, 2013). As infant EEG studies have a typical attrition rate of 50–75% ( deBoer et al., 2007 ), the lower bound of 15% attention was chosen to reduce further data loss. The mean number of trials (per condition) included in the analyses are 15.83 for infants, 21.26 for 4-year-olds, and 25.61 for adults. The mean levels of attention across conditions are 56.24% for infants, 62.66% for 4-year-olds, and 79.95% for adults. 2.2. Stimuli Audiovisual recordings of 30 short speech passages were made by a female native speaker of Australian English experienced in producing infant-directed speech (see Appendix A for transcripts). To allow for infants’ limited attention span these passages were relatively short, but long enough to ensure an amount of EEG recording that was sufficient for analyses ( Crosse et al., 2021 ). These speech passages were adapted from Richoz et al. (2017) or from recordings of infant-directed speech between mothers and their babies and varied in durations from 8.44 s to 16.35 s ( M = 11.35 s, SD = 1.76 s). The recordings consisted of a close-up of the speaker’s face and shoulders against a white background. There were three presentation modes, auditory-only (AO), visual-only (VO) and auditory-visual (AV) with the unimodal auditory and visual recordings extracted separately from the auditory-visual recordings. In the auditory-only condition, a still image of the speaker’s resting face was shown on the screen as the auditory track was played. In the visualonly condition, the dynamic video of the speaker’s talking face was presented in silence. In the auditory-visual condition, both the dynamic video and its soundtrack were played together. The auditory recordings have a sampling rate of 44.1 kHz and a 16-bit resolution. The 30 speech passages were presented in three blocks. Each block consisted of 10 speech passages that were presented once in each modality. Presentation order was randomised across modalities in such a manner that the same sentence did not appear in two modalities on consecutive trials. Attention-getter stimuli were used throughout the experiment to maintain participants’ attention. The type and frequency differed between age groups. For 5-month-olds, attention-getters consisted of 2-s animations (often used in the infant calibration routine in Tobii Studio) that appeared after each trial. For 4-year-olds and adults, attentiongetters consisted of different pictures of ‘Minions’ that appeared in a random order after either two or three trials, with their frequency randomly determined. In addition, a different 3-s cartoon animation was played to mark the end of the block and to re-engage participants. 2.3. Procedure 2.3.1. Five ‐month ‐olds Infants sat on their caregiver’ laps approximately 70 cm away from the centre of an LCD screen. Continuous EEG data were recorded with a 128-channel Hydrocel Geodesic Sensor Net (HCGSN), NetAmps 300 amplifier, and NetStation 4.5.7 software (EGI Inc) at a sampling rate of 1000 Hz, with the reference electrode placed at Cz. Electrode impedances were kept below 50 k Ω. The EEG recordings were saved for offline analyses. Stimulus presentation was controlled using Presentation software (Neurobehavioural Systems). Triggers indicating the start and end of each trial were recorded along with the EEG. Eye-tracking recordings were co-registered with EEG recordings for two purposes: (i) to ensure that infants were attending to the visual stimuli and (ii) to examine whether gaze behaviour to the mouth region modulates cortical tracking of the speech envelope. To this end, a Tobii X120 eye tracker was placed below the screen to gather gaze fixation data. As the entire duration of the session was quite long for an infant study (approximately 25 min), the stimuli continued to play until infants showed signs of fussiness or until completion, whichever came first. 2.3.2. Four ‐year ‐olds and adults The procedure for 4-year-olds was identical to that for 5-month-olds with two exceptions. First, 4-year-olds were seated on their own. Second, the session was framed as a game; in order to motivate children to focus on the screen, children were required to press a button on a response pad whenever a picture of a Minion appeared on the screen ( Kaganovich and Schumaker, 2014 ). Adult participants were informed prior to the start of the experiment that they are part of a control group for an infant and child study. The procedure for adults was similar to 4-year-olds, except that adults also participated in a second EEG task which used similar stimuli but in adultdirected speech (ADS). Its order of presentation (immediately before or after the first task) was counterbalanced between participants (the results of this ADS session are not reported here). 2.3.3. EEG measure 2.3.3.1. Pre ‐processing. EEG data were pre-processed using EEGLAB ( Delorme and Makeig, 2004 ), FieldTrip ( Oostenveld et al., 2011 ), NoiseTools ( http://audition.ens.fr/adc/NoiseTools/ ), the mTRF Toolbox ( Crosse et al., 2016 ) and custom scripts in MATLAB R2019a (The Mathworks, Inc). First, EEG data from the three outer rings of the net were removed because these channels have been found to be very noisy in infants and children ( Di Liberto et al., 2018 ; Folland et al., 2015 ; Kalashnikova et al., 2018 ). EEG data from the remaining 92 channels were high-pass filtered at 0.1 Hz, low-pass filtered at 12 Hz with Butterworth 8th order filters. As infant and child EEG recordings are noisy due to movements, artefact subspace reconstruction (ASR; Kothe and Jung, 2014 ) was applied to remove noise. ASR uses a sliding window technique whereby each EEG window is decomposed via principal component analysis. Each EEG window is then statistically compared with reference EEG data obtained from clean portions of the EEG recording. Within each window, the ASR algorithm searches for principal subspaces that significantly deviate from the reference EEG data. These subspaces are rejected and then reconstructed using a mixing matrix computed from the reference EEG data ( Chang et al., 2019 ). As in Kalashnikova et al. (2018) , this study used a sliding window of 500 ms and a threshold of 20 standard deviations to identify corrupted subspaces. Noisy channels that were removed during ASR were replaced with an estimate of neighbouring clean channels using spherical interpolation. Finally, EEG data were re-referenced to the average of all channels (e.g., Kalashnikova et al., 2018 ) and later downsampled to 100 Hz to reduce processing time. To investigate the impact of visual speech cues on the cortical tracking of auditory speech, the speech stimuli were pre-processed in a manner following Jessen et al. (2019) . The auditory soundtracks of each video were extracted, downsampled to 100 Hz to match the sampling rate of the EEG data and characterised using the broadband speech envelope of the acoustic signal through the NSL toolbox that models the auditory peripherical and subcortical processing stages ( Ru, 2001 ). A spectrogram representation of each stimulus contained band-specific envelopes of 128 logarithmically-spaced frequency bands between 0.1 and 4 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 4 kHz. The broadband temporal envelope of each soundtrack was obtained by summing up the band-specific envelopes across all frequencies. 2.3.3.2. Data analysis. Cortical tracking of the speech envelope was measured by mathematically modelling response functions that describe the linear mapping between the stimulus speech envelopes and the corresponding neural responses. For this study, the stimulusresponse mapping function is modelled in the forward direction (see Crosse et al. (2016) for details), i.e., the resulting model describes an optimal linear transformation from the stimulus domain to the neuralsignal domain. Such a model is fit by conducting a lagged ridge regression between the envelope and the EEG data while accounting for likely time-delays between the acoustic input and the corresponding EEG response. The regression weights obtained with this procedure estimate the temporal response function (TRF) between envelope and EEG at each EEG channel. Significant non-zero weights reflect EEG channels where cortical activity is related to stimulus encoding ( Haufe et al., 2014 ). TRFs are similar to event-related potentials (ERPs) in that they allow for an examination of the amplitude, latency, and scalp topography of the stimulus-EEG relationship. Specifically, the distribution of TRF weights can be examined across the scalp at different latencies, or different relative time lags between the ongoing speech and EEG signals. For example, a time lag of 100 ms refers to the impact that a change in the speech stimulus at time t has on the EEG at time t + 100 ms. To investigate neural tracking of continuous stimuli, adult studies commonly compute response functions based on a subset (e.g., n − 1 trials) of the available data from each participant (e.g., Crosse et al., 2015 ), resulting in TRFs that are then used to model responses for the n th trial for each participant. This approach —subject-dependant modelling —requires lengthy datasets for each participant that are typically unattainable for the infant population. To account for the limited amount of available data from the infant sample, the subject-independent approach ( Di Liberto and Lalor, 2017 ) was used for this study. Instead of computing an individual response function for each participant, this approach involves computing an average response function over n − 1 participants that is then used to predict the EEG signal of the n th participant via leave-one-out cross-validation. The subject-dependant modelling approach has been shown to yield better results than the subjectdependant modelling approach when used with 5-min EEG recordings from 7-month-olds and adults ( Jessen et al., 2019 ). Subject-dependant modelling was used for each age group. In other words, an average response function was computed for each age group to predict the EEG signal of the nth participant from that age group. Initially, TRFs were calculated for each stimulus at time lags between − 200 and 1000 ms before selecting a temporal region of the TRF (0–600 ms) that included all relevant components to map the stimulus to the EEG signal with no visible response outside of this range. Leave-oneout cross-validation using Tikhonov regularization was conducted to assess how well the unseen EEG data could be predicted based on the TRF. The regularisation parameter of the ridge regression was set to 𝜆= 100 for all participants. The lambda parameter value was chosen to mitigate the potential failure of lambda tuning due to the limited amount of data available (for a discussion, see Crosse et al., 2021 ). Prediction accuracy was quantified by calculating the Pearson correlation coefficient between the predicted and original EEG responses at each electrode. If EEG data is indeed reflecting the encoding of the speech envelope, then the correlation values would be significantly greater than zero. To investigate auditory-visual speech benefit, ( A + V ) TRFs were computed and compared to AV TRFs in accordance with the additive criterion. The additive criterion was chosen to investigate auditory-visual speech benefit because this was used in previous studies with similar paradigms (e.g., Crosse et al., 2015 , 2016 ). The AV speech benefit was quantified as the difference in prediction accuracy for AV TRFs relative to A + V TRFs. Table 1 Means (and Standard Deviations) of spatial offsets (Measured in Pixels) in gaze data for each age group. 5-month-olds 4-year-olds Adults X-coordinate 39.91 (519.75) 72.85 (278.33) 33.26 (159.45) y-coordinate 25.37 (225.46) 98.80 (315.86) 164.78 (130.44) Fig. 1. Areas of interest (AOIs) defined for the speaker’s eye and mouth regions. 2.3.4. Gaze measures Means and standard deviations of the spatial offsets (xand ycoordinates) for each age group are reported in Table 1 . As 5-month-olds and 4-year-olds were more fidgety than adults during the study, there was a considerable amount of data loss from the eye-tracker for those groups. To circumvent the cumulative effect of data loss due to gaze as measured by the eye-tracker and to noisy EEG data, videos of participants who met the EEG data inclusion criterion ( ≤ 20 noisy channels) but had eye-tracking issues (i.e., participants were looking at the screen but their gaze was not detected by the eye-tracker) were coded frameby-frame manually using ELAN software (version 5.9) for whether or not they were looking at the screen. This resulted in hand-coded videos for 11 four-year-olds, and 3 five-month-olds. Areas of interest (AOIs) covering the top half and bottom half of the speaker’s face demarcated the speaker’s eye and mouth regions ( Fig. 1 ). These AOIs were of equal dimensions (640 ×340 pixels) and were adjusted using the derived mean spatial offsets of each age group. The proportion of total looks (PTLs) to these AOIs, in addition to attention, were computed for each trial: 1 Attention = [𝑡𝑜𝑡𝑎𝑙 𝑓𝑖𝑥𝑎𝑡𝑖𝑜𝑛 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 𝑡𝑟𝑖𝑎𝑙 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 ], (hereafter referred to as Attention) and 2 Proportion looking to the speaker’s mouth region (hereafter referred to as PTL Mouth) = [𝑡𝑜𝑡𝑎𝑙 𝑓𝑖𝑥𝑎𝑡𝑖𝑜𝑛 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 𝑡𝑜 𝑚𝑜𝑢𝑡ℎ 𝑡𝑜𝑡𝑎𝑙 𝑓𝑖𝑥𝑎𝑡𝑖𝑜𝑛 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 𝑡𝑜 𝑚𝑜𝑢𝑡ℎ + 𝑡𝑜𝑡𝑎𝑙 𝑓𝑖𝑥𝑎𝑡𝑖𝑜𝑛 𝑑𝑢𝑟𝑎𝑡𝑖𝑜𝑛 𝑡𝑜 𝑒𝑦𝑒𝑠 ]. Note that PTL Mouth is a relative measure of attention to the mouth compared to eyes, so chance is 0.5, scores > 0.5 show greater fixation to mouth than eyes and scores < 0.5 show greater fixation to eyes than mouth. All statistical analyses on these two gaze measures were conducted using custom scripts in MATLAB R2019a (The MathWorks, Inc). The 11 four-year-olds and 3 five-month-olds whose gaze data were manually coded were only included for analyses that examined attention to screen —they were excluded from analyses that involved PTL Mouth. 2.4. Statistical analyses Estimates of global field power were computed and topographic maps of TRF weights plotted to inspect the scalp regions where responses 5 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 Fig. 2. Global field power measured at each time lag for all ages. to the speech envelope were greatest. Mean TRFs were then computed for those scalp locations identified as regions of interest (ROIs) for each condition. To evaluate model performance, mean prediction accuracies were obtained by averaging across all electrodes belonging to the ROIs and then tested against zero. Additionally, these mean prediction accuracies were compared between conditions to investigate the auditory-visual speech benefit and any age differences in model performance. As agerelated anatomical differences may influence cortical tracking between groups independent of effects due to speech modality, TRF components and their respective prediction accuracy were not directly compared statistically between age groups. To examine gaze behaviour, ANOVAs were conducted for each age group to examine the differences in attention and proportion looking at speaker’s mouth between conditions (see Eqs. (1) and (2)). To examine the relationship between gaze behaviour and cortical tracking, Pearson’s correlations were conducted for each condition between (1) cortical tracking and attention, and (2) cortical tracking and looking preference for each age group, where cortical tracking is quantified by TRF prediction accuracy. 3. Results 3.1. Prediction accuracies First, as a preliminary step, global field power (GFP) —a referenceindependent measure of response strength across the entire scalp at each time lag ( Murray et al., 2008 ) — was estimated by calculating the TRF variance across all channels. The temporal profile of GFP for each age group showed clear TRF components at ∼200–400 ms for AO, AV and ( A + V ), but not VO ( Fig. 2 ). Topographies of TRF weights ( Figs. 3–5 ) revealed that the observed components were mainly located over the frontal, occipital and temporal scalp regions. To avoid diluting the effects of interest, subsequent analyses of TRFs were therefore focused on the frontal, occipital, and temporal groups of electrodes. These groupings were used in previous infant (e.g., Folland et al., 2015 ; Table 2 Mean prediction accuracies (and Standard Deviations), quantified by pearson’s r, of TRFs from frontal, temporal and occipital scalp ROIs for each condition and age group. AO VO AV A + V 5-month-olds .021 (0.018) .001 (0.008) .035 (0.019) .032 (0.018) 4-year-olds .020 (0.018) − 0.005 (0.011) .018 (0.020) .014(0.015) Adults .009 (0.011) .0004 (0.011) .022 (0.015) .007 (0.012) Peter et al., 2016 ) and child (e.g., Corrigall and Trainor, 2014 ) EEG studies to examine the average responses across scalp regions ( Fig. 6 ). To examine the presence of envelope tracking, TRF prediction accuracies at the three scalp ROIs were tested against zero. To assess the difference in the extent of envelope tracking, these prediction accuracies were then compared between conditions. Of interest are (1) the differences between cortical tracking of AO, VO and AV speech, and (2) the presence of an auditory-visual speech benefit as quantified by the additive criterion [i.e., AV vs. ( A + V )]. One-sample t -tests were first conducted to test prediction accuracies against zero. Next, one-way ANOVAs were conducted for each age group with their respective prediction accuracies as the dependant variable to examine whether prediction accuracies differed between conditions. Subsequent post-hoc comparisons were conducted using two-tailed paired-sample t -tests with Bonferroni-adjusted alpha levels where multiple comparisons were made. The same analyses were conducted with 15 randomly selected trials per condition for 4-year-olds and adults to examine whether different amounts of data from each age group influenced the results. Fifteen trials were chosen because infant data had the least number of trials included with approximately 15 trials per condition. 3.1.1. Evidence of cortical tracking All means and standard deviations of prediction accuracy for each condition and age group are set out in Table 2 . Five-month-olds : One-sample t -tests indicated that prediction accuracy of AO, AV, and ( A + V ) TRFs were significantly greater than zero 6 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 Fig. 3. (A) Topographies and TRFs of frontal, occipital and temporal locations, and (b) prediction accuracy of TRFs from 5-month-olds’ data. Fig. 4. (A) Topographies and TRFs of frontal, occipital and temporal locations, and (b) prediction accuracy of TRFs from 4-year-olds’ data. (AO: t (17) = 5.15, p = < 0.001, Hedges’ g = 1.16; AV: t (17) = 7.47, p < .001, Hedges’ g = 1.68,; A + V: t (17) = 7.42, p < .001, Hedges’ g = 1.67), but prediction accuracy of VO TRFs was not significantly greater than zero, t (17) = 0.75, p = .23, Hedges’ g = 0.17. Four-year-olds : Prediction accuracies of AO, AV, and ( A + V ) TRFs were significantly greater than zero (AO: t (18) = 4.93, p = < 0.001, Hedges’ g = 1.08; AV: t (18) = 3.86, p < .001, Hedges’ g = 0.85; A + V: t (18) = 3.96, p < .001, Hedges’ g = 0.87), whereas prediction accuracy of VO TRFs was not significantly greater than zero ( t (18) = − 2.13, p = .98, Hedges’ g = − 0.47). The analyses with 15 trials revealed only one difference: prediction accuracy of VO TRFs was significantly lower than zero ( t (18) = − 2.38, p = .03, Hedges’ g = 0.48. Adults : Prediction accuracies of AO, AV, and ( A + V ) TRFs were significantly greater than zero (AO: t (17) = 3.49, p = .001, Hedges’ g = 0.79; AV: t (17) = 6.11, p < .001, Hedges’ g = 1.38; A + V: t (17) = 2.48, p = .012, Hedges’ g = 0.56), whereas prediction accuracy of VO TRFs was not significantly greater than zero, t (17) = 0.17, p = .44, Hedges’ g = 0.04. Results from the analyses with 15 trials were not different. 3.1.2. Difference in strength of cortical tracking between conditions The one-way ANOVAs testing between conditions (AO, VO, AV, A + V ) were significant for all age groups (5-month-olds: F (3, 68) = 14.95, p < .001, 𝜂p 2 = 0.40; 4-year-olds: F (3, 72) = 9.63, p < .001, 𝜂p 2 = 0.29; adults: F (3, 68) = 9.22, p < .001, 𝜂p 2 = 0.29). To in7 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 Fig. 5. (A) Topographies and TRFs of frontal, occipital and temporal locations, and (b) prediction accuracy of TRFs from adults’ data. Fig. 6. Electrode groupings used for analyses. (A) frontal electrodes, (B) occipital electrodes, (C) temporal electrodes. spect the differences between conditions and to identify whether there was auditory-visual speech benefit [i.e., AV > ( A + V )], post hoc comparisons were subsequently performed using paired-sample t -tests with Bonferroni-adjusted alpha level of 0.013 (0.05/4). Five-month-olds : When prediction accuracies of AO, VO, and AV TRFs were compared, paired-sample t -tests indicated that prediction accuracy of AV TRFs was greatest, followed by AO, then VO TRFs (AO vs. VO: t (17) = 5.13, p < .001, Hedges’ g = 1.42; AO vs. AV: t (17) = − 4.07, p < .001, Hedges’ g = − 0.69; AV vs. VO: t (17) = 7.73, p < .001, Hedges’ g = 2.15). Prediction accuracy of AV TRFs was also significantly greater than ( A + V ) TRFs, t (17) = 2.82, p = 0.001, Hedges’ g = 0.16, suggesting that auditory-visual speech benefit was present at the scalp ROIs. Four-year-olds : Paired-sample t -tests revealed that the prediction accuracy of AO TRFs was significantly greater than that of VO TRFs 8 S.H. Jessica Tan, M. Kalashnikova, G.M. Di Liberto et al. NeuroImage 256 (2022) 119217 Fig. 7. Scatterplots of Attention (A) and proportion of total looking time to the mouth vs. Eyes (PTL Mouth) (B) for all conditions and age groups and their corresponding bar graphs (C: Attention; D: PTL Mouth). error bars represent standard errors of mean (SEM). With respect to attention, across age groups, greater attention was captured in the AV condition. With respect to the speaker’s mouth, adults fixated the speaker’s mouth to a greater extent in the AV condition than in AO and VO. ( t (18) = 5.66, p < .001, Hedges’ g = 1.68) but not significantly different from the prediction accuracy of AV TRFs ( t (18) = 0.58, p = 0.57, Hedges’ g = 0.14). The prediction accuracy of AV TRFs was significantly greater than that of VO TRFs ( t (18) = 4.75, p < .001, Hedges’ g = 1.39), but was not significantly greater than that of ( A + V ) TRFs ( t (18) = 1.06, p = 0.30, Hedges’ g = 0.21). The analyses with 15 trials had similar findings. Adults : Paired-sample t -tests showed that the prediction accuracy of AV TRFs was greatest, followed by AO, then VO TRFs (AO vs. VO: t (17) = 4.10, p < .001, Hedges’ g = 0.78; AO vs. AV: t (17) = − 3.85, p = .001, Hedges’ g = − 0.88; AV vs. VO: t (17) = 7.36, p < .001, Hedges’ g = 1.57). Prediction accuracy of AV TRFs was also significantly greater than ( A + V ) TRFs ( t (17) = 5.01, p < .001, Hedges’ g = 1.06), suggesting that auditory-visual speech benefit was present at the scalp ROIs. The analyses with 15 trials revealed only one difference: prediction accuracy of AO TRFs is not significantly different from that of VO TRFs, t (17) = 1.97, p = .07, Hedges’ g = 0.54. 3.2. Gaze behaviour 3.2.1. Attention Separate one-way within-subjects ANOVAs were conducted for each age group with Attention as the dependant variable (see Eq. (1) in Statistical Analyses) and Condition as the independent variable. The ANOVAs revealed a significant main effect of Condition for all age groups (5month-olds: F (2, 34) = 3.58, p = .04, 𝜂p 2 = 0.17; 4-year-olds: F (1.44, 25.89) = 26.67 with Greenhouse-Geisser correction, p < .001, 𝜂p 2 = 0.60; adults: F (2, 34) = 7.16, p = .002, 𝜂p 2 = 0.30). Subsequent post-hoc comparisons between conditions were made using paired-sample ttests with Bonferroni-adjusted alpha level of 0.017 (0.05/3). Fig. 7 contains scatterplots and bar graphs of Attention and PTL Mouth for all conditions and age groups. Five-month-olds : Attention was significantly greater in the AV than the VO condition ( t (17) = 2.93, p = .009, Hedges’ g = 0.50), but the differences between AO and VO and between AO and AV conditions were not significant (AO vs. VO: t (17) = 1.49, p = .15, Hedges’ g = 0.34; AO vs. AV: t (17) = − 0.94, p = .36, Hedges’ g = 0.14). Four-year-olds : Attention was significantly greater in the AV than in the AO condition ( t (18) = 6.10, p < .001, Hedges’ g = 1.54) and in the VO condition ( t (18) = 9.19, p < .001, Hedges’ g = 1.43), whereas the difference in attention between AO and VO conditions was not significant ( t (18) = − 1.00, p = .33, Hedges’ g = − 0.26). Adults : Attention was significantly greater in the VO than the AO condition ( t (17) = 3.58, p = .002, Hedges’ g = 0.38) and in the AV than the AO condition ( t (17) = 3.06, p = .007, Hedges’ g = 0.40). The difference in attention between VO and AV conditions was not significant ( t (17) = 0.11, p = .91, Hedges’ g = 0.01). Age comparisons: An Age x Condition mixed-design ANOVA was conducted with Attention as the dependant variable. The main effects of Condition and Age, and the Age x Condition interaction were significant (Condition: F (1.68, 87.50) = 26.00 with Greenhouse-Geisser correction, p < .001, 𝜂p 2 = 0 . 33 ; Age: F (2, 52) = 25.21, p < .001, 𝜂p 2 = 0.49 ; Age x Condition: F (3.37, 87.50) = 12.47 with Greenhouse-Geisser correction, p < .001, 𝜂p 2 = 0.32). To examine the Age x Condition interaction, we conducted independent-samples t -tests for each condition. Five-montholds attended less to the screen than 4-year-olds only in the AV condition ( t (35) = − 4.45, p < .001), whereas they attended to the screen similarly during AO and VO presentations (AO: t (35) = 0.45, p = .65; VO: t (35) = − 1.53, p = .13). Five-month-olds attended less to the screen than adults in all conditions (AO: t (34) = − 4.94, p < .001; VO: t (34) = − 7.43, p < .001: AV: t (34) = − 6.10, p < .001). Four-year-olds attended less to the screen in AO and VO conditions than adults but not during AV presentations (AO: t (35) = − 4.99, p < .001; VO: t (35) = − 5.64, p < .001; AV: t (35) = − 1.88, p = .07). 3.2.2. PTL to the speaker’s mouth Separate one-way within-subjects ANOVAs were conducted for each age group (DV: PTL Mouth, IV: Condition). The ANOVAs were significant for 5-month-olds and adults (5-month-olds: F (2, 26) = 4.98, p = .01, 𝜂p 2 = 0.28; adults: F (1.35, 23.00) = 13.40 with Greenhouse-Geisser correction, p < .001, 𝜂p 2 = 0.44), but not for 4-year-olds ( F (2, 14) = 1.82, p = .20, 𝜂p 2 = 0.21). Subsequent analyses involved one-sample t -tests to assess whether PTL Mouth was significantly greater than chance and paired-sample t -tests with Bonferroni-adjusted alpha level of 0.017 9