scieee AI-readable full text Open interactive document viewer

The nature of first and second language processing: The role of cognitive control and L2 proficiency during text-level comprehension

Pérez Muñoz, Ana Isabel,Hansen, Laura,Bajo Molina, María Teresa

Abstract

Text comprehension relies on high-level cognitive processes as it is the ability to revise an erroneous inference. Recent models of language processing hold that native language processing is proactive in nature (highly predictive), whereas processing seems to be weaker in the second language. However, if a prediction fails because unexpected information is encountered, reactive processing is needed to revise previous information. Twenty-four highly proficient late bilinguals were presented with narratives in L1-English and L2-Spanish. Each text demanded the revision of an initial predictive inference. Reading times and N400 amplitude suggested inferential revision is less efficient in the L2 compared to the L1. Importantly, these effects were modulated by individual differences in cognitive control and L2 proficiency. More efficient L1 comprehension was related to a balance between proactive and reactive control and lower L2 proficiency, whereas more native-like L2 comprehension was associated with a strong proactive control and higher L2 proficiency.

Full text

1 Running Head: FIRST AND SECOND LANGUAGE DURING TEXT-LEVEL COMPREHENSION The nature of first and second language processing: The role of cognitive control and L2 proficiency during text-level comprehension* Ana Pérez, Laura Hansen 1 and Teresa Bajo University of Granada (Spain) *ACKNOWLEDGEMENTS This research was supported by a doctoral research grant from the Spanish Ministry of Education, Culture and Sports (FPU AP2010–3434) to Laura Hansen, by a postdoctoral contract funded by Spanish Ministry of Science and Innovation (PSI2012–33625) to Ana Pérez, and by grants of MINECO (PSI201233625; PCIN-2015-132; PSI2015-65502-C2-1P) and Junta de Andalucía (Excelencia2012-CTS 2369) to Teresa Bajo. Address for correspondence: Ana Pérez 2 and Laura Birke Hansen 3 Mind, Brain and Behavioral Research Centre (CIMCYC) Department of Experimental Psychology, University of Granada. C/ Profesor Clavera s/n, CIMCYC, 18011, Granada, Spain. E-mails: [email protected].uk; [email protected] Phone: +44 (0)7706619219 2 Abstract Text comprehension relies on high-level cognitive processes as it is the ability to revise an erroneous inference. Recent models of language processing hold that native language processing is proactive in nature (highly predictive), whereas processing seems to be weaker in the second language. However, if a prediction fails because unexpected information is encountered, reactive processing is needed to revise previous information. Twenty-four highly proficient late bilinguals were presented with narratives in L1-English and L2-Spanish. Each text demanded the revision of an initial predictive inference. Reading times and N400 amplitude suggested inferential revision is less efficient in the L2 compared to the L1. Importantly, these effects were modulated by individual differences in cognitive control and L2 proficiency. More efficient L1 comprehension was related to a balance between proactive and reactive control and lower L2 proficiency, whereas more native-like L2 comprehension was associated with a strong proactive control and higher L2 proficiency. Keywords: late bilinguals; reading comprehension; inferential revision; cognitive control; L2 proficiency 3 1. Introduction Mastering a non-native language can be very challenging, especially if the second language (L2) has been acquired relatively late in life. Adult learners can reach high levels of proficiency in their L2, and with increasing proficiency many aspects of L2 processing can become more and more native-like (see Birdsong & Molis, 2001). However, some studies suggest that L2 processing remains different from native language (L1) processing. Empirical evidence has demonstrated that between-language differences are present in both the syntactic and semantic domains (Dufour & Kroll, 1995; Duyck & De Houwer, 2008; Wartenburger, Heekeren, Abutalebi, Cappa, Villringer, & Perani, 2003; for reviews, see Clahsen & Felser, 2006; Slabakova, 2006), although these differences tend to be quantitative rather than qualitative in the semantic domain (see Slabakova, 2006; Wartenburger et al., 2003). Studies looking at lexical and semantic processing in L1 and L2 have used procedures involving words or sentences (see e.g., Foucart, Martin, Moreno, & Costa, 2014; Martin et al., 2013), while relatively fewer studies have involved texts and higher order discourse processes. Importantly, different from lexical and sentence processing, text processing requires the construction of a mental representation (i.e., situation model, van Dijk & Kintsch, 1983) by constantly integrating new information, which depends on multiple high-level comprehension processes (derived from the integration of linguistic and pragmatic information) not present at lower-level processing (exclusively based on linguistic properties). For instance, a crucial aspect of text comprehension is the ability to generate information that has not been explicitly described, referred to as inference making (Cain & Oakhill, 1999). To this end, readers must connect several pieces of information presented in the text and/or activate prior knowledge from long-term memory (see McNamara & Magliano, 2009). An important type of inference in text comprehension is prediction. Predictive inferences help to anticipate upcoming concepts in the story (Beeman, Bowden & Gernsbacher, 2000; see also Pérez, Paolieri, Macizo & Bajo, 2014), they are inherently proactive and tend to be 4 automatically encoded in proficient L1 readers when a) information is quickly and easily available in memory, or b) they are necessary to provide text coherence (McKoon & Ratcliff, 1980). Moreover, as the text unfolds, initial predictions can become outdated, in which case they have to be replaced with new inferences, a process known as inferential revision (see Pérez, Cain, Castellanos & Bajo, 2015). Although research into these high-level text comprehension processes in the L2 is scarce, preliminary evidence suggests that readers with advanced L2 proficiency are able to generate inferences in their L1 and L2 (Horiba, 1996), and with increased proficiency, L2 readers seem to show more efficient integration compared to less proficient readers (Yang, 2002). However, no previous study has directly compared inferential revision in the L1 and the L2. Thus, the main goal of the present study was to understand if the ability to revise inferential information during text comprehension is less efficient in the L2 compared to the L1. A current theoretical hypothesis states that whereas native language processing tends to be highly anticipatory, with comprehenders continuously predicting upcoming information, non-native speakers show a “Reduced Ability to Generate Expectations” (RAGE hypothesis, Grüter & Rohde, 2013; Grüter, Rohde & Schafer, 2014). According to this hypothesis, rather than relying on predictive processing as occurs in L1 comprehension, L2 comprehension primarily relies on a posteriori integration and thus, tends to be relatively less proactive than L1 processing. Studies looking at brain activity during L1 and L2 sentence have provided evidence for this hypothesis. A sensitive marker related to integration and prediction is the N400 component (Kutas & Hillyard, 1980). The N400 reflects the ease with which the meaning of a word can be integrated into the current mental representation, with larger amplitude for words that are unexpected (e.g., “It was raining so he grabbed his… coat”) compared to expected words (e.g., “umbrella”). Interestingly, this effect has also been interpreted as indicating the costs of 5 revising an active prediction (scalar inference) in underinformative sentences like “some people have lungs” (Nieuwland, Ditman, & Kuperberg, 2010; Nieuwland & Kuperberg, 2008). Evidence using highly constrained sentences indicates that L2 readers are less likely to make lexical predictions (i.e., whether the gender of an article is expected with respect to the predicted word) than L1 readers (Martin et al., 2013; but see Foucart et al., 2014), suggesting L2 comprehension is performed by passive integration of encountered words, rather than by active lexical prediction (see also Lau, Holcomb, & Kuperberg, 2013). Furthermore, a recent study investigating causal inferences during text comprehension (Foucart, Romero-Rivas, Gort, & Costa, 2016) has shown that, in contrast to native speakers, L2 speakers did not manifest significant N400 differences in texts that were causally unrelated compared to causally related texts (instead, they found an early and late positivity). Moreover, when comparing intermediately related and causally unrelated texts, there were no significant differences either in native or L2 speakers, and only a tendency to significance was generally found, showing more negativity in intermediately related than causally related texts. Although the authors argued that this marginally significant result, along with a late negative component (620-750ms), is evidence for the ability of L2 speakers to generate causal inferences during online comprehension, we believe their results are, at least, ambiguous. Overall, then, the previously discussed findings indicate that quantitative differences in the N400 might, in fact, reflect qualitatively different comprehension processes. The less predictive nature of L2 processing has been interpreted as being due to limited availability of working memory (WM) capacity (Hopp, 2013). L2 processing requires more WM resources compared to L1 processing (Dussias & Piñar, 2010; Ransdell, Arecco, & Levy, 2001), and even simple linguistic processes such as lexical access increase activation in brain areas associated with cognitive control when executed in the L2 (Ma et al., 2014). As more cognitive resources need to be allocated to these lower-level linguistic processes, less resources are available for higher-level semantic-pragmatic processes, where less proficient 6 L2 comprehenders in particular often experience difficulties (Horiba, 1996; Horiba & Fukaya, 2015; Yang, 2002). Similarly, the depletion of WM resources by lower-level processes can explain why L2 comprehenders might revert from a proactive (active prediction) processing to a more reactive (passive integration) one. However, L1 processing relies heavily on domain-general cognitive control as well, especially when it comes to linguistic processes beyond the single word or sentence level (Arrington, Kulesz, Francis, Fletcher, & Barnes, 2014; Borella, Caretti, & Pelegrina, 2010; Pérez et al., 2015). An alternative approach to assess the role of cognitive control within-language is to quantify how cognitive control is implemented during task performance. According to the general dual-mechanisms framework within the executive control field (Braver, 2012), individuals may employ two different control modes (proactive or reactive) to exert attentional control during ongoing task performance. Which control mode is implemented depends on individual tendencies as well as situational demands. Proactive control is implemented pre-emptively, through sustained goal maintenance and anticipatory monitoring throughout performance of a cognitive task. This type of control is highly dependent on WM capacity (Braver, 2012). Reactive control, on the other hand, consists of the momentary and transient activation of the task goal in the light of conflict or interference, which has been associated with inhibition, both empirically and theoretically (Morales, Gómez-Ariza, & Bajo, 2013). Since reactive control is resource-economic, it is usually employed when cognitive resources are limited either due to individual capacity limits or when the task demands are particularly high. From this perspective, rather than relying on domain-general control to some degree, it might be the case that cognitive control is implemented in a different way when bilinguals are processing in their L2 vs. their L1. For instance, the RAGE hypothesis suggests L1 comprehenders mostly remain in a proactive control mode, but shift towards the less demanding reactive control mode when processing in the L2. 7 Although many bilingual studies have explored the relationship between language and cognitive control, most of them have focused on how control processes are engaged during language selection (Levy, McVeigh, Marful, & Anderson, 2007; Martín, Macizo, & Bajo, 2010; Morales, Paolieri, & Bajo, 2011; Blumenfeld & Marian, 2013), or on how repeated practice at language selection may enhance attentional control (Bialystok, Craik, & Luk, 2008; Costa, Hernández, & Sebastián-Gallés, 2008; Morales, Yudes, Gómez-Ariza & Bajo, 2015). However, very few studies have directly addressed how cognitive control mechanisms are recruited for the resolution of semantic difficulties during L1 and L2 sentence and/or text comprehension (Moreno, Bialystok, Wodniecka, & Alain, 2010), and none has investigated the role of cognitive control in inferential revision in both languages. Crucially, given that exclusive reliance on active predictions based on previous information (proactive control) would counteract the revision of inferences, inferential revision also requires the flexibility of passive integration (reactive control) to accommodate unexpected upcoming information. Thus, both proactive control, which facilitates the generation of predictive inferences, and reactive control, which enables revision in case prior predictions are not fulfilled, are requisite for successful text comprehension. Accordingly, an additional goal of our study was to explore whether the ability to revise inferential information in L1 and L2 was related to differences in proactive/reactive cognitive control. 1.1. The present study A paradigm developed to investigate inferential revision is the situation model revision task (Pérez et al., 2015). In this task, participants are presented with short narrative texts (see Table 1). The first three sentences of each text present a constrained Context that facilitates a predictive inference (“guitar”). In the following sentence, readers are presented with one of three conditions: neutral, where the sentence does not refer back to the inference (“…at the prestigious national concert hall.”); non-update, consistent with the inference 8 primed by the context (“…with a beautiful curved body.”); and update, mismatching the inference primed in the context and facilitating the generation of a new inference (“...with a matching bow.”). This latter condition primes the replacement of the previous inference with a new one (revision) and the integration of this new information into the situation model. Reading times are measured in the fourth sentence (RT sentence). The final sentence presents the disambiguating word (“violin”), that is always inconsistent with the inference primed by the context (“guitar”), but consistent with the inferential information facilitated in the update condition (“matching bow”). Event-related potentials are recorded during presentation of this disambiguating word (ERP word). Importantly, the status of the ERP word depends on the condition presented in the RT sentence, being a) “expected” when coming from the update condition (“matching bow”  “violin”), because previous inferential information is coherent; b) “unexpected” when coming from the non-update condition (“curved body” related to the idea of guitar  “violin”), because previous inferential information is improbable but still plausible; and c) “uncertain” when coming from the neutral condition (“concert hall”  “violin”), because previous information is not related to the critical words. <Insert Table 1 about here> In line with previous literature, reading times for RT sentences allow us to draw conclusions about inference making when reading the three sentences presented as context. Coming from a specific situation (idea of “guitar”), longer RTs for the update (“matching bow”) compared to the non-update (“curved body”) and neutral (“concert hall”) conditions would indicate that readers are able to generate the predictive inference facilitated in the context, and subsequently detect a mismatch when new information is presented. In addition, the N400 elicited by the ERP word reflects the processing cost of revising and integrating this word (“violin”). Therefore, a reduced N400 elicited in the expected condition (coming from “matching bow”) compared to the unexpected (coming from “curved body”) and uncertain 9 (coming from “concert hall”) conditions, would indicate that comprehenders have been able to take advantage of the prior update information to successfully revise their initial inference, and integrate the new prediction into their situation model. Based on previous literature, we anticipate inference making and revision-integration to be less efficient in the L2 compared to the L1. In addition, given that the only previous study on inferential revision found WM differences to be associated with the revision-integration processes but not with inference making (Pérez et al., 2015), we expect between-language processing differences to be more pronounced for inferential revision than for inference making. Furthermore, RT and N400 data is expected to be modulated by individual differences in both cognitive control and L2 proficiency. Moreover, to understand whether differences in cognitive control entail processing differences in L1 and L2 text comprehension, we evaluated proactive/reactive control by means of the Behavioural Shift Index (BSI) of the AX-CPT task (see Method). This index reflects the individual tendency towards a strong proactive control (BSI near 1) or a balance between proactive and reactive control (BSI near 0). Although previous studies have not explored the effect of cognitive control on high-level comprehension processes in terms of proactive/reactive control, we tentatively propose that efficient inferential revision processes might be supported by a good balance between proactive control (necessary for predictions) and reactive control (required in revision), in both the L1 and L2. Finally, because processing differences due to language status can often be explained in terms of linguistic proficiency alone (Horiba & Fukaya, 2015; Kaan, 2014; Newman, Tremblay, Nichols, Neville & Ullman, 2012; Yang, 2002), we also expect text comprehension processes will be more native-like in more L2 proficient readers. 2. Method 2.1. Participants 16 11.70, 11.46, and 11.81, for the neutral, non-update and update conditions, respectively), or Spanish, F(2, 178) = 1.79, p = .17 (Ms = 12.00, 12.13, and 12.49, respectively). The fifth sentence ended with the ERP word (“violin”) which was always inconsistent with the predictive inference biased by the context (“guitar”) and consistent with the prediction supported by the RT sentence in the update condition. Consequently, the ERP word was inconsistent in the uncertain and unexpected conditions, prompting inferential revision, and consistent in the expected condition, assuming revision-integration to have taken place already. EEG was recorded at the onset of this word. At the end of each text, a comprehension sentence requiring a true or false judgment was presented to ensure that participants read for understanding. The two critical words (“guitar” and “violin”) were controlled in character and syllable length, neighbourhood size, age of acquisition, concreteness, frequency, familiarity and imageability for both languages (see Table 3). The words did not differ in most of these measures in the L1-English: number of characters, t(89) = -0.74, p = .47; number of syllables, t(89) = -0.47, p = .64; neighbourhood size, t(89) = -0.48, p = .63; age of acquisition, t(89) = - 2.14, p < .05; concreteness, t(89) = 2.06, p < .04; frequency, t(89) = 0.58, p = .56; familiarity, t(89) = 0.73, p = .47; and imageability, t(89) = 0.81, p = .42. In addition, none of these measures differed in the L2-Spanish: number of characters, t(89) = -0.16, p = .87; number of syllables, t(89) = 0.41, p = .69; neighbourhood size, t(89) = 0.62, p = .54; age of acquisition, t(89) = -0.60, p = .55; concreteness, t(89) = 0.90, p = .37; frequency, t(89) = 1.19, p = .24; familiarity, t(89) = 0.67, p = .51; and imageability, t(89) = 1.34, p = .19. <Insert Table 3 about here> A prior norming study suggested that native English speakers activated the non-update concept after reading the context in their L1, but not the update concept (see Pérez et al., 17 2015). Accordingly, we run a second norming study to provide evidence of concept preferences in native Spanish speakers. Twenty participants (M = 30.6 years old; range: 2633) read the context of each text (sentences 1-3) and were then presented with a single word. Their task was to rate from 1 (Improbable) to 5 (Very probable) how probable was the word in the context of the story. The word was either the non-update concept, which was most strongly supported by the context (“guitar”), or the update concept, which was a plausible but not probable alternative (“violin”). Two versions of the same questionnaire were created to ensure participants saw only one of the two concepts for each text. A linear mixed model with Participants and Items as random factors and Concept type as fixed factor, was performed on the mean rate. Results demonstrated a main effect of concept type, F(1) = 199.29, p < .001, dv = .37, where the non-update concept was highly probable (M = 4.40, SE = 0.09) compared to the update concept (M = 2.65, SE = 0.13), which was considered almost “neutral” (score of 3). As intended, this difference suggests that, after reading the context, native Spanish speakers were significantly more likely to activate the non-update concept than the update concept. 2.3. Procedure There were two experimental sessions. Participants who volunteered for participation were contacted by e-mail or telephone and asked a few screening questions to ensure whether they met L2 proficiency requirements. Only those participants who had lived in a Spanishspeaking country for at least one year, had learned their L2 after the age of 10 and considered themselves to have a high level of proficiency were invited to session 1. In this session, participants completed both language questionnaires (LHQ and BSWQ) and all behavioral tasks (vocabulary, verbal fluency, WM and AX-CPT). Only participants who reached an accuracy cut-off value of 60% on the vocabulary test were invited to session 2. In session 2, EEG recordings took place while participants completed the situation model revision task (approx. 90 minutes). This task was administered in two blocks, one in 18 English and the other in Spanish. Each trial started with a fixation cross (“+”) that remained on the screen until the participant pressed the “Yes” or “No” key on the keyboard to start reading. The first four sentences were presented one sentence at a time, and participants were asked to read each sentence at their own pace, pressing the same key to display the next sentence. RTs for the fourth sentence were registered. Subsequently, the fifth sentence was presented word by word with a fixed stimulus-onset asynchrony (SOA) of 300ms per word. In addition, there was a delay of 700ms after the ERP word to ensure the recording of activity during a sufficiently long time window (SOA = 1000ms). Participants were instructed to try not to blink during this final sentence, in order to prevent excessive noise in the EEG data. Finally, a comprehension sentence was presented, and participants were instructed to press the “Yes” key if they thought the sentence was true, or “No” if they thought it was false. Each of the 90 experimental texts was presented to each participant only once, in one of the two languages and the three conditions (six cross conditions). The assignment of language and condition to text was counterbalanced across participants, so that each participant read 15 texts within each factor level combination of condition and language. The order of language block was also counterbalanced. A practice of three trials ensured that instructions were understood. In addition, at the end of the second session, participants were also asked to complete the second vocabulary test (5-10 minutes). 2.4. Apparatus Tasks were presented by the E-prime software (Schneider, Eschman, & Zuccolotto, 2002), administered on a 19” inch. CRT video monitor (refresh rate = 75 Hz). For the situation model revision task, we recorded scalp voltages using a SynAmps2 64 channels Quik-Cap, plugged into a Neuroscan SynAmps RT amplifier with a continuous sample rate of 250 Hz. The ground (FCZ) and reference (FPZ) electrodes worked as referential signals. The vertical and horizontal electrooculogram (VEOG and HEOG, respectively) was registered 19 supraand infraorbitally to the left eye and at the outer canthi. Impedances were kept below 5 kΩ. Subsequently, the electrical signal was amplified with a 1-30 Hz band-pass filter. Blinks and ocular movements were corrected by using singular value decomposition. Trials with artifacts (2.94%) were rejected, and recordings from electrodes with high level of artifacts (>1%) were substituted by the average value of the group of nearest electrodes. Epochs from - 200 and 800ms with respect to the presentation of the ERP word were averaged and analyzed. We applied a baseline correction, using the average EEG activity in the 200ms previous to target onset as a reference. ERPs were averaged for each factor level combination by participant, text, and region of interest. Individual averages were re-referenced off-line to the average of left and right mastoids (M1 and M2). 2.5. Data analysis Reading times index. RTs (in milliseconds) were measured for the RT sentence of the situation model revision task. To factor out differences between the L1 and the L2 in baseline reading speed, we divided the RT sentence (fourth sentence) by averaged RTs for the first three sentences (context) of each text. In addition, this also helped to control for differences in text content due to the linguistic properties of each language (e.g., passive vs. active voice in English and Spanish respectively), although it is still possible that cultural aspects may have influenced how our bilinguals were comprehending in both languages. Event-related potentials. To analyze ERPs, we used the same six regions of interest (ROI) referenced by Pérez et al. (2015): left frontal (F1, F3, F5, FC3, and FC5), right frontal (F2, F4, F6, FC4, and FC6), central (C1, C2, CZ, FCZ, and CPZ), left parietal (P1, P3, P5, CP3, and CP5), right parietal (P2, P4, P6, CP4, and CP6), and occipital (O1, O2, POZ, PO3, and PO4). The N400 component was measured as the mean amplitude (in microvolts) in the time window from 300 to 500ms, averaged for each ROI, and ROI was included as a predictor variable in the N400 analysis. Outliers, defined as amplitude values 2.5 standard deviations 20 above or below the mean by language, condition, and ROI (0.86%), were replaced with corresponding mean values (see Pérez et al., 2015). Linear mixed models. LMMs were conducted using the lmer function of the lme4 R package, version 1.1-7 (Bates, Maechler, Bolker, & Walker, 2015), with participants and items as random factors, and language, condition, ROI, and centered values for both the BSI of cognitive control and L2 proficiency as fixed factors (see Schielzeth, 2010). Separate models were conducted for each dependent variable (RTs and N400). Thus, the full fixed structure run with RTs contained two three-way interactions (language x condition x BSI + language x condition x L2 proficiency), whereas the full fixed structure of the N400 model contained two four-way interactions (including ROI) as well as all their lower level interactions and main effects. Texts containing target words that participants did not know (11%) and texts to which the comprehension sentence was answered incorrectly (8%), were eliminated from analyses. The mean final number of trials used per participant was 37.79 texts in the L1-English (neutral = 12.79, non-update = 12.33 and update = 12.67), and 35.67 texts in the L2-Spanish (neutral = 11.50, non-update = 11.83 and update = 12.33). Accordingly, the relatively small amount of items used for the electrophysiological analyses entails the need to interpret our N400 results with caution. First, keeping the full fixed structure, we fitted each model with the maximal random effects structure by participants and items using restricted Maximum Likelihood (Barr, Levy, Scheepers, & Tily, 2013). Convergence problems were solved by removing one by one the effects for which less variance was observed when the summary function was applied to the partially converged solution (for participants or items), until the model converged 4 . Secondly, keeping the maximal random effects structure possible, we conducted stepwise model comparisons starting from the most complex model using Maximum Likelihood (ML) and removing effects that did not account for significant variance in the data, as determined by χ² 21 Log-likelihood tests. Finally, for models with significant fixed effects, p values were provided by the anova function of the lmerTest R package (Kuznetsova, Brockhoff, Christensen, 2015), using ML. Explained deviance was calculated using the pamer.fnc function of the LMERConvenienceFunctions R package (Tremblay & Ransijn, 2015). This statistic assesses the overall goodness of fit and serves as a generalization of R² by measuring the marginal improvement or reduction in unexplained variability in the fixed component after accounting for a given predictor effect (see Pérez, Joseph, Bajo & Nation, 2016). To follow-up on threeway interactions, we divided the data into subsets according to the levels of language and/or condition and fitted adjusted LMM for these subsets. To qualify two-way interactions, we ran pairwise comparisons within each factor level combination by using the test Interactions function of the phia R package (De Rosario-Martínez, 2013). 3. Results Our results are organized into two sections. We first analyzed the RT index extracted from the RT sentence, and then we examined the N400 amplitude recorded in response to the ERP word (see Table 4 for means and standard errors of RTs and ERP amplitude). Taking into account the large number of results presented in this study, we focused on the fixed effects of each LMM. Summary details (lmerTest package) regarding model fit and random effects of each model are provided in the Appendix. <Insert Table 4 about here> 3.1. RT sentence: Inference making To address the question whether comprehenders had previously generated the predictive inference and were able to detect a mismatch in the update condition (see Introduction for hypotheses), we performed a LMM with language (L1 vs. L2), condition (neutral vs. non-update vs. update), and both individual differences indices (BSI and L2 22 proficiency) as fixed factors, and RT index (RT sentence/context, in milliseconds) as the dependent variable. Main factors. The final model (Model 1, Appendix) demonstrated significant main effects of language, F(1) = 5.34, p <.05, dv = .01, where readers took longer in the L1 (M = 0.929, SE = 0.01) compared to the L2 (M = 0.924, SE = 0.02); and condition, F(2) = 5.52, p <.01, dv = 1.38, where RTs were longer in the update (M = 1.00, SE = 0.02) compared to the non-update (M = 0.94, SE = 0.02), t(89) = 5.93, p <.001, and the neutral condition (M = 0.92, SE = 0.01), t(85) = 8.04, p <.001; and the non-update was marginally longer than the neutral condition, t(84) = 2.26, p =.07. In addition, there was a significant two-way interaction of language x condition, F(2) = 5.21, p <.01, dv = .20 (see Figure 1). According to the interaction, although the effect of condition was significant in both languages [L1-English: χ² (2) = 64.98, p <.001, and L2-Spanish: χ² (2) = 22.96, p <.001], pairwise comparisons within language revealed different patterns. In the L1, comprehenders took longer to read the update compared to the non-update, t(289) = 6.59, p <.001, and the neutral condition, t(251) = 7.21, p < .001; with no differences between the non-update and the neutral condition, t(237) = 0.67, p =.98. These effects indicate that reading in the L1, comprehenders had generated the predictive inference prompted by the context, and then detected new inconsistent information (longer RTs in the update condition). In the L2, on the other hand, the pattern was somewhat different. That is, although comprehenders took longer in the update than in the neutral condition, t(272) = 4.79, p <.001, the difference between the update and the non-update condition was not significant, t(290) = 2.11, p =.28; and there was a marginal difference between the non-update and the neutral condition, t(259) = 2.69, p =.08. These results suggest that reading in the L2, comprehenders performed some level of inferential processing (longer RTs in the update compared to the neutral), but they had difficulties to fully generate the predictive inference prompted by the context (lack of difference between the update and non- 23 update), and therefore, required more information to make it (marginally longer RTs in the non-update compared to the neutral). <Insert Figure 1 about here> BSI of cognitive control. The main effect of BSI and its interaction terms with the other variables were dropped from the final model during the backwards stepwise procedure, as neither of them made a significant contribution to the model (all ps >.05). Thus, cognitive control did not explain language differences in the ability to predict and subsequently detect a mismatch with that prediction. L2 proficiency. Participants’ L2 proficiency significantly interacted with language, F(1) = 5.00, p <.05, dv = .19, and condition, F(2) = 6.96, p <.001, dv = .27. No other effect was significant, (ps >.05). Both two-way interactions with L2 proficiency resulted from opposite regression slopes: a) for language, χ² (1) = 5.00, p <.05, lower L2 proficiency was related to faster RTs in the L1 compared to the L2 (see Figure 2a); and b) for condition, χ² (2) = 13.92, p <.001, higher L2 proficiency was associated with longer RTs in the update compared to the non-update and neutral conditions (larger differences between conditions, see Figure 2b). No other pairwise comparison was significant (all ps >.36). Therefore, linguistic proficiency signaled a differential tendency between languages, where comprehenders with lower L2 proficiency were faster in their L1, and between conditions, where more L2 proficient comprehenders approached a more native-like inference making. <Insert Figure 2 about here> 3.2. ERP word: Revision-Integration To assess the question whether comprehenders had revised their previous prediction and integrated the newly inferred concept into their situation model (see Introduction for 24 hypotheses), we conducted an LMM with language, (L1 vs. L2), condition (uncertain vs. unexpected vs. expected), ROI (left frontal vs. right frontal vs. central vs. left parietal vs. right parietal vs. occipital) and both individual differences indices (BSI and L2 proficiency) as fixed factors, and N400 component (mean amplitude, in microvolts) for the ERP word as the dependent variable. Main factors. The final model (Model 2, Appendix) showed a main effect of condition, F(2) = 8.97, p <.001, dv = .07, where as predicted, the expected condition manifested less negativity than the unexpected, t(89) = 3.81, p <.001, and the uncertain, t(90) = 3.50, p <.01; no differences were observed between the unexpected and the uncertain condition, t(88) = 0.12, p =.99. Importantly, the two-way interaction between language and condition was also significant, F(2) = 18.41, p < .001, dv = 0.08. No other effects reached significance (all ps >.05). Pairwise comparisons within language in the interaction demonstrated that, although the effect of condition was significant in both languages [L1English: χ² (2) = 29.26, p <.001, and L2-Spanish: χ² (2) = 7.73, p <.05], again they revealed different patterns (see Figure 3). In the L1, as hypothesized, the expected condition showed less negativity compared to the unexpected, t(121) = 4.25, p <.001, and the uncertain condition, t(118) = 4.96, p <.001; with no differences between the unexpected and uncertain conditions, t(116) = 0.88, p =.95. This pattern suggests that when reading in their L1, comprehenders revised/integrated their situation model by replacing a misleading predictive inference (“guitar”) with a more plausible one (“violin”) in the previous update condition (less negative amplitude in the expected compared to the unexpected and uncertain conditions). In the L2, by contrast, the expected condition was only marginally less negative than the unexpected, t(125) = 2.77, p =.07; and no differences were found between the expected and the uncertain, t(124) = 1.54, p =.64, or between the unexpected and the uncertain condition, t(122) = 1.09, p =.89. Thus, although comprehenders carried out some 25 level of revision when reading in their L2, this process seemed quantitatively less efficient than in the L1 (only marginal tendency of less negativity in the expected than in the unexpected condition). Interestingly, pairwise comparisons within condition demonstrated significant differences between languages only in the expected condition, χ² (1) = 9.30, p <.01, with more negativity in the L2 compared to the L1. Language differences were not significant in either the unexpected, χ² (1) = 2.19, p =.42, or the uncertain condition, χ² (1) = 0.53, p =1.00. Thus, when reading in the L2, comprehenders did not fully replace their initial prediction with the new inference (larger negativity in the expected condition for the L2 compared to the L1), confirming greater difficulties to revise their situation model in L2 processing. <Insert Figure 3 about here> BSI of cognitive control. There was a significant three-way interaction between language, condition, and the BSI, F(2) = 9.17, p <.001, dv = .07 (see Figure 4). No other effects reached significance (all ps >.05). To follow up on this three-way interaction, we first divided the data by language. Significant interactions between condition and BSI were found in both languages [L1-English, F(2) = 13.57, p <.001, and L2-Spanish, F(2) = 3.61, p <.03], but once more, analyses within language signaled different patterns. In the L1, a smaller BSI (more balanced reliance on proactive and reactive control) predicted less negativity in the expected condition, χ² (1) = 13.32, p <.001; and no differences in the unexpected, χ² (1) = 0.22, p =1.00, and the uncertain condition, χ² (1) = 0.24, p =1.00. Therefore, when reading in the L1, comprehenders with a more proactive-reactive balance were better at revising the no longer relevant predictive inference and integrating the new inferential information into their situation model. In the L2, on the other hand, the BSI did not manifest differences in any of the three conditions: expected, χ² (1) = 0.66, p =1.00, unexpected, χ² (1) = 0.04, p =1.00, and uncertain, χ² (1) = 2.13, p =.43. The interaction between condition and BSI in the L2 came 32 comprehenders to reduce the costs of lower-level linguistic processes in the L2. Overall, the L2 pattern suggests that the ability to generate a prediction and then, replace it with new information is costlier in the L2, where lower-level linguistic processes tend to be more resource-consuming (e.g. Horiba, 1996; Yang, 2002). These results contribute to our understanding of the nature of L1 and L2 processing. Recent theoretical proposals hold that L2 processing tends to be less proactive than L1 processing, and that this factor may account for many differences observed between the L1 and the L2 across linguistic domains (RAGE hypothesis, Grüter & Rohde, 2013; Grüter et al., 2014; Grüter, Lew-Williams & Fernald, 2012; Hopp, 2013; Martin et al., 2013). Underlying this notion is the belief that typical L1 comprehension is highly proactive, in that good comprehenders continuously predict upcoming information on the basis of incrementing lexical, semantic and morphosyntactic cues. Our observations support the view that when comprehension requires high-level cognitive processes like the replacement of a previous misleading prediction with a new one, a balance between proactive and reactive control (that is, a more flexible cognitive system) predicts better understanding. At least this was true for L1 comprehension. In the L2, on the other hand, the ability to revise inferential information was not predicted by more flexible cognitive control, but by strong proactive control, suggesting the processing cost of lower-level linguistic processes had been minimized. Interestingly, it has been suggested that the larger N400 effect found in L1 comprehenders could be reflecting the combination of both active prediction of the upcoming information and passive integration of the encountered word, whereas a smaller N400 effect in L2 comprehenders would indicate less semantic processing, with the exclusive use of posteriori integration (Lau et al., 2013, Martin, et al., 2013). Therefore, in relation to our results, the larger N400 effect associated with a balance between proactive and reactive control showed by good L1 comprehenders points at the use of both active prediction and passive integration mechanisms for efficient comprehension in the L1. In contrast, the larger N400 effect 33 associated with proactive control in good L2 comprehenders suggests that an active prediction process was what distinguished them from poor comprehenders. 4.4. L2 proficiency in L1 and L2 text processing Our data also align with those of others in suggesting that L1 vs. L2 differences might ultimately be due to proficiency asymmetry between the two languages. In regards to the inference making process, higher L2 proficiency was generally associated with longer RTs in the update compared to the non-update and neutral conditions, suggesting a more native-like inference making process. Similarly, the revision-integration process in the L2 was more native-like with higher L2 proficiency, as indicated by less negativity in the expected condition (“violin” coming from “matching bow”). Note also that to some extent, higher L2 proficiency and proactive control played a similar role in L2 comprehension (both associated with a greater difference between conditions), indicating that to some degree, cognitive and linguistic abilities can compensate each other. Finally, some attention should be dedicated to findings demonstrating that L2 proficiency predicted performance not just in the L2 but also in the L1. Lower L2 proficiency was related to faster RTs in the L1 compared to the L2, signaling comprehenders with less proficiency in the L2 were faster comprehending in the L1 than those with more proficiency in the L2. As suggested by the N400 results, contrary to L2 comprehension, lower (rather than higher) L2 proficiency was associated with better discrimination between conditions in the L1, indicating a better inferential revision-integration process. Importantly, this difference between conditions in the L1 was attenuated with higher L2 proficiency, indicating less efficient comprehension in the L1 compared to the L2. This means that L1 vs. L2 processing differences were less marked in participants with higher L2 proficiency not just due to enhanced processing efficiency in the L2, but also to reduced efficiency in the L1. Although surprising, these findings cohere with some recent evidence suggesting that the acquisition of 34 a second language later in life can modulate an already established L1 (Baus, Costa, & Carreiras, 2013; Chang, 2012; Linck, Kroll, & Sunderman, 2009; Malt, Li, Pavlenko, Zhu & Ameel, 2015), an effect called attrition. The observation of a smaller N400 effect in the L1 for the most proficient bilinguals also mirrors the findings of one previous study where semantic effects were generally reduced in bilinguals’ compared to monolinguals’ L1 (Ardal, Donald, Meuter, Muldrew & Luce, 1990). These data suggest that there is a trade-off between L1 and L2 processing efficiency in active bilinguals: to reach very high levels of proficiency in their L2, late bilinguals might “sacrifice” processing efficiency in their L1, or alternatively, it could be the case that a permeable L1 system that is susceptible to change is requisite for reaching native-like proficiency in a late L2 (see Kroll, Bobb, & Hoshino, 2014). Although further study is needed to explore the mechanisms and temporal dynamics underlying this relationship, these findings speak to an evolving and reciprocal relationship between language systems. 4.5. Conclusions To sum up, the present study extends previous research into sentence processing by showing that the efficiency of high-level text comprehension processes such as inference making and inferential revision-integration, is reduced in an L2 acquired in adulthood, compared to the L1. Modulatory effects suggest that these processing differences may be ultimately rooted in reduced linguistic proficiency and consequentially, limited availability of cognitive resources to engage control processes in the L2. Thus, to some extent, individual differences in cognitive control can compensate limitations in linguistic proficiency, whereas very high proficiency in the L2 can, in principle, compensate non-native language status. Modulatory effects of L2 proficiency on the native language bear witness of a bidirectional and dynamic relationship between a bilingual’s language systems. Further study is needed to fully understand the dynamic interaction between L1 and L2 processing, cognitive control and 35 linguistic proficiency, during online text comprehension, as well as possible differences of this interaction by comparing monolinguals and bilinguals. 5. References Alonso, M. A., Fernández, A., & Díez, E. (2015). Subjective age-of-acquisition norms for 7,039 Spanish words. Behavior Research Methods, 47, 268-274. Ardal, S., Donald, M. W., Meuter, R., Muldrew, S., & Luce, M. (1990). Brain responses to semantic incongruity in bilinguals. Brain and Language, 39, 187-205. Arrington, C. N., Kulesz, P. A., Francis, D. J., Fletcher, J. M., & Barnes, M. A. (2014). The contribution of attentional control and working memory to reading comprehension and decoding. Scientific Studies of Reading, 18, 325-346. Barr, D. J., Levy, R., Scheepers, C., & Tily, H. J. (2013). Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language, 68, 255-278. Bates, D., Maechler, M., Bolker, B., & Walker, S. (2015). lme4 package. Retrieved from https://cran.r-project.org/ Baus, C., Costa, A., & Carreiras, M. (2013). On the effects of second language immersion on first language production. Acta Psychologica, 142, 402-409. Beeman, M. J., Bowden, E. M., & Gernsbacher, M. A. (2000). Right and left hemisphere cooperation for drawing predictive and coherence inferences during normal story comprehension. Brain and language, 71, 310-336. Bialystok, E., Craik, F. I. M., & Luk, G. (2012). Bilingualism: Consequences for mind and brain. Trends in Cognitive Sciences, 16, 240-250. 36 Birdsong, D., & Molis, M. (2001). On the evidence for maturational constraints in secondlanguage acquisition. Journal of Memory and Language, 44, 235-249. Blumenfeld, H. K., & Marian, V. (2013). Parallel language activation and cognitive control during spoken word recognition in bilinguals. Journal of Cognitive Psychology, 25, 547567. Borella, E., Carretti, B., & Pelegrina, S. (2010). The specific role of inhibition in reading comprehension in good and poor comprehenders. Journal of Learning Disabilities, 43, 541-552. Braver, T. S., Paxton, J. L., Locke, H. S., & Barch, D. M. (2009). Flexible neural mechanisms of cognitive control within human prefrontal cortex. Proceedings of the National Academy of Sciences, 106, 7351-7356. Braver, T. S. (2012). The variable nature of cognitive control: a dual mechanisms framework. Trends in Cognitive Sciences, 16, 106-113. Brysbaert, M., Warriner, A. B., & Kuperman, V. (2014). Concreteness ratings for 40 thousand generally known English word lemmas. Behavior research methods, 46, 904-911. Cain, K., & Oakhill, J. V. (1999). Inference making ability and its relation to comprehension failure in young children. Reading and Writing, 11, 489-503. Chang, C. B. (2012). Rapid and multifaceted effects of second-language learning on firstlanguage speech production. Journal of Phonetics, 40, 249-268. Chiew, K. S., & Braver, T. S. (2014). Dissociable influences of reward motivation and positive emotion on cognitive control. Cognitive, Affective, & Behavioral Neuroscience, 14, 509529. 37 Clahsen, H., & Felser, C. (2006). How native-like is non-native language processing? Trends in Cognitive Sciences, 10, 564-570. Costa, A., Hernández, M., & Sebastián-Gallés, N. (2008). Bilingualism aids conflict resolution: Evidence from the ANT task. Cognition, 106, 59-86. Davis, C. J. (2005). N-Watch: A program for deriving neighborhood size and other psycholinguistic statistics. Behavior research methods, 37, 65-70. Davis, C. J., & Perea, M. (2005). BuscaPalabras: A program for deriving orthographic and phonological neighborhood statistics and other psycholinguistic indices in Spanish. Behavior Research Methods, 37, 665-671. De Rosario-Martínez, H. (2015). Phia package. Retrieved from https://cran.r-project.org/ Dufour, R., & Kroll, J. F. (1995). Matching words to concepts in two languages: A test of the concept mediation model of bilingual representation. Memory & Cognition, 23, 166180. Dussias, P. E., & Piñar, P. (2010). Effects of reading span and plausibility in the reanalysis of wh-gaps by Chinese-English second language speakers. Second Language Research, 26, 443-472. Duyck, W., & De Houwer, J. (2008). Semantic access in second-language visual word processing: Evidence from the semantic Simon paradigm. Psychonomic Bulletin & Review, 15, 961-966. Foucart, A., Martín, C. D., Moreno, E. M., & Costa, A. (2014). Can bilinguals see it coming? Word anticipation in L2 sentence reading. Journal of Experimental Psychology: Learning, Memory, and Cognition, 40, 1461. 38 Foucart, A., Romero-Rivas, C., Gort, B. L., & Costa, A. (2016). Discourse comprehension in L2: Making sense of what is not explicitly said. Brain and language, 163, 32-41. Frey, L. (2005, January). The nature of the suppression mechanism in reading: Insights from an L1-L2 comparison. In Proceedings of the Annual Meeting of the Cognitive Science Society, 27, 714-719. Grüter, T., Lew-Williams, C., & Fernald, A. (2012). Grammatical gender in L2: A production or a real-time processing problem? Second Language Research, 28, 191-215. Grüter, T., & Rohde, H. (2013). L2 processing is affected by RAGE: Evidence from reference resolution. In the 12th conference on Generative Approaches to Second Language Acquisition (GASLA). Grüter, T., Rohde, H., & Schafer, A. (2014, May). The role of discourse-level expectations in non-native speakers’ referential choices. In Proceedings of the annual Boston university conference on Language Development. Hopp, H. (2013). Grammatical gender in adult L2 acquisition: Relations between lexical and syntactic variability. Second Language Research, 29, 33-56. Horiba, Y. (1996). Comprehension processes in L2 reading: Language competence, textual coherence, and inferences. Studies in Second Language Acquisition, 18, 433-473. Horiba, Y., & Fukaya, K. (2015). Reading and learning from L2 text: Effects of reading goal, topic familiarity, and language proficiency. Reading in a Foreign Language, 27, 2246. Kaan, E. (2014). Predictive sentence processing in L2 and L1: What is different? Linguistic Approaches to Bilingualism, 4, 257-282. 39 Kroll, J. F., Bobb, S. C., Misra, M. M., & Guo, T. (2008). Language selection in bilingual speech: Evidence for inhibitory processes. Acta Psychologica, 128, 416-430. Kuperman, V., Stadthagen-Gonzalez, H., & Brysbaert, M. (2012). Age-of-acquisition ratings for 30,000 English words. Behavior Research Methods, 44, 978-990. Kutas, M., & Hillyard, S. A. (1980). Reading senseless sentences: Brain potentials reflect semantic incongruity. Science, 207, 203-205. Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B. (2015). lmerTest package. Retrieved from https://cran.r-project.org/ Lau, E. F., Holcomb, P. J., & Kuperberg, G. R. (2013). Dissociating N400 effects of prediction from association in single-word contexts. Journal of cognitive neuroscience, 25, 484-502. Levy, B.J., McVeigh, N.D., Marful, A., & Anderson, M.C. (2007). Inhibiting your native language: The role of retrieval-induced forgetting during second-language acquisition. Psychological Science, 18, 29-34. Li, P., Sepanski, S., & ZhAo, X. (2006). Language history questionnaire: A web-based interface for bilingual research. Behavior Research Methods, 38, 202-210. Linck, J. A., Kroll, J. F., & Sunderman, G. (2009). Losing access to the native language while immersed in a second language: Evidence for the role of inhibition in second-language learning. Psychological Science, 20, 1507-1515. Ma, H., Hu, J., Xi, J., Shen, W., Ge, J., Geng, F., & Yao, D. (2014). Bilingual cognitive control in language switching: An fMRI study of English-Chinese late bilinguals. PloS one, 9, e106468. 40 Malt, B. C., Li, P., Pavlenko, A., Zhu, H., & Ameel, E. (2015). Bidirectional lexical interaction in late immersed Mandarin-English bilinguals. Journal of Memory and Language, 82, 86-104. Martín, M.C., Macizo, P., & Bajo, M.T. (2010). Time course of inhibitory processes in bilingual language processing. British Journal of Psychology, 101, 679-693. Martin, C. D., Thierry, G., Kuipers, J. R., Boutonnet, B., Foucart, A., & Costa, A. (2013). Bilinguals reading in their second language do not predict upcoming words as native readers do. Journal of Memory and Language, 69, 574-588. McKoon, G., & Ratcliff, R. (1980). The comprehension processes and memory structures involved in anaphoric reference. Journal of Verbal Learning and Verbal Behavior, 19, 668-682. McNamara, D. S., & Magliano, J. (2009). Toward a comprehensive model of comprehension. Psychology of Learning and Motivation, 51, 297-384. Morales, J., Yudes, C., Gómez-Ariza, C. J., & Bajo, M. T. (2015). Bilingualism modulates dual mechanisms of cognitive control: Evidence from ERPs. Neuropsychologia, 66, 157-169. Morales, J., Gómez-Ariza, C. J., & Bajo, M. T. (2013). Dual mechanisms of cognitive control in bilinguals and monolinguals. Journal of Cognitive Psychology, 25, 531-546. Morales, L., Paolieri, D., & Bajo, M.T. (2011). Grammatical gender inhibition in bilinguals. Frontiers in Psychology, 2, 284. Moreno, S., Bialystok, E., Wodniecka, Z Alain, C. (2010) Conflict resolution in sentence processing by bilinguals. Journal of Neurolinguistics, 23, 564-579. 41 Newman, A. J., Tremblay, A., Nichols, E. S., Neville, H. J., & Ullman, M. T. (2012). The influence of language proficiency on lexical semantic processing in native and late learners of English. Journal of Cognitive Neuroscience, 24, 1205-1223. Nieuwland, M. S., Ditman, T., & Kuperberg, G. R. (2010). On the incrementality of pragmatic processing: An ERP investigation of informativeness and pragmatic abilities. Journal of memory and language, 63, 324-346. Nieuwland, M. S., & Kuperberg, G. R. (2008). When the truth is not too hard to handle: An event-related potential study on the pragmatics of negation. Psychological Science, 19, 1213-1218. Pérez, A., Cain, K., Castellanos, M. C., & Bajo, T. (2015). Inferential revision in narrative texts: An ERP study. Memory & Cognition, 43, 1105-1135. Pérez, A., Joseph, H. S., Bajo, T., & Nation, K. (2016). Evaluation and revision of inferential comprehension in narrative texts: an eye movement study. Language, Cognition and Neuroscience, 31, 549-566. Pérez, A. I., Paolieri, D., Macizo, P., & Bajo, T. (2014). The role of working memory in inferential sentence comprehension. Cognitive processing, 15, 405-413. Ransdell, S., Arecco, M. R., & Levy, C. M. (2001). Bilingual long-term working memory: The effects of working memory loads on writing quality and fluency. Applied Psycholinguistics, 22, 113-128. Rodríguez-Fornells, A., Krämer, U. M., Lorenzo-Seva, U., Festman, J., & Münte, T. F. (2012). Self-assessment of individual differences in language switching. Frontiers in Psychology, 2, 1-15.