A longitudinal study of the effects of model texts on EFL children’s written production
Abstract
The financial support by the Spanish Ministry of Economy and Competitiveness (MINECO) under grant FFI 2016-74950-P (AEI/FEDER/UE) and by the Basque Government under grant IT904-16 are hereby gratefully acknowledged.
Full text
1 A longitudinal study of the effects of model texts on EFL children’s written production María Luquin and María del Pilar García Mayo Abstract As written corrective feedback tools, model texts have been claimed to improve language learners’ subsequent production, but almost exclusively in terms of lexical gains. However, scarce research has been carried out with EFL children, an underrepresented population in the literature, and much less from a longitudinal perspective. The main aim of this study was to determine the extent to which sustained exposure to models can have an impact on the written production of child EFL learners. Thirty pairs of 11–12-year-old Spanish EFL children were randomly assigned to a control group, a treatment group and a long-term treatment group, who engaged in two four-stage collaborative writing cycles of three weeks each. The children’s collaborative texts were transcribed and analyzed considering different measures (type of clause, syntactic complexity, lexical diversity, accuracy, fluency and holistic assessment). Our findings revealed that model texts led to a decrease in the number of pre-clauses and an increase in the syntactic complexity of the texts in the short run. A long exposure to models showed that the children were able to produce fewer proto-clauses and more clauses, feature higher lexical diversity in their texts and make fewer errors. Keywords: children, CAF, EFL, longitudinal design, model text, longitudinal design 1. Introduction This is the accepted manuscript of the article that appeared in final form in System 120 : (2024) // Article ID 103190, which has been published in final form at https://doi.org/10.1016/j.system.2023.103190. © 2023 Elsevier under CC BY-NC-ND license (http:// creativecommons.org/licenses/by-nc-nd/4.0/)
2 The worldwide introduction of a foreign language (FL) at an early age responds both to a very clear social demand and to the conviction that young children have greater auditory and oral plasticity, which will allow them to assimilate the FL better than at a more advanced age (Nicholas & Lightbown, 2008). However, there is no definitive evidence supporting the potential advantages of early-age FL learning (Copland & Garton, 2014). This lack of clarity emerges as a result of multiple factors that hinder the process of language acquisition. These factors encompass challenges like overcrowded classrooms, insufficient immersion in the target language, limited chances for interaction outside the classroom setting, and a constrained curriculum duration (García Mayo, 2018). It is, therefore, also important to acknowledge the necessity of enhancing the amount of time and active involvement with the FL (Manchón, 2009). It is true that young learners (YLs) enjoy great facility for the development of basic communication skills, such as oral comprehension and production, fundamentally linked to social skills (Kellogg, 2008). In the Spanish context, however, out of the four basic skills in the FL curriculum, the writing skill has been neglected in favor, almost exclusively, of the speaking skill (Cánovas Guirao, 2017). This lack of emphasis on writing clashes with research in the field of second language acquisition (SLA) oriented toward a writing-to-learn perspective (Manchón, 2011), which suggests that learners can benefit from receiving information about the accuracy and appropriateness of their texts, supporting the advancement of their L2 knowledge (Ferris, 2010). One way of achieving this goal is through feedback provision, which seems to activate cognitive processes such as hypothesis formation and testing, attention, metalinguistic reflection, noticing (Schmidt, 2001), or problem-solving strategies (García Hernández et al., 2017), all of them deemed crucial for language acquisition. As a result of its language acquisition potential, written corrective feedback (WCF) has been extensively researched in recent decades (Bitchener & Knoch, 2009). In particular, researchers have proposed model
3 texts as an effective WCF technique on the grounds that they constitute a more discursive feedback method which does not explicitly single out errors but treats the text as a whole. Studies using models as a form of WCF have been conducted both with adult/adolescent (e.g., Hanaoka & Izumi, 2012; Kang, 2020; Martínez Esteban & Roca de Larios, 2010; Montealegre Ramón, 2019; Wu et al., 2023; Yang & Zhang, 2010) and child (e.g., Cánovas Guirao et al., 2015; Coyle & Roca de Larios, 2020; Criado et al., 2022; Lázaro-Ibarrola, 2021; LázaroIbarrola & Villareal, 2021; Roothooft et al., 2022; Villarreal & Lázaro-Ibarrola, 2022) participants, but more work is needed to elucidate whether the benefits reported in these studies only lead to greater precision in L2 writing or to language development in the long term (Polio, 2012). Pedagogically, many EFL teachers seem not to be aware of the importance of WCF so researchers should try to transfer knowledge from empirical studies for it to play an important role in language learning curricula. Thus, the main aim of this study was to determine the extent to which sustained exposure to models can have an impact on the written production of child EFL learners. 2. Literature review 2.1. Child L2 acquisition in foreign language contexts Child L2 acquisition in FL contexts is a rapidly growing field of study, as more and more schools around the world are introducing English at the pre-primary and primary levels (García Mayo, 2018). The literature has presented several reasons for the perceived "hastiness" in FL learning. One primary argument is advocated by governments, parents, and society, emphasizing the need for learners to be proficient in the FL to access international education and employment opportunities in today's globalized world (Copland & Garton, 2014; Enever & Moon, 2009; García Mayo, 2018). Another rationale stems from the belief that younger is
4 better, supported by significant benefits observed in immersion and bilingual settings (Lyster, 2007). However, the potential advantages of early FL learning lack conclusive evidence (Copland & Garton, 2014; DeKeyser, 2013). Various factors impede language learning, such as large class sizes, limited exposure to the target language, restricted opportunities for interaction outside the classroom, and limited curriculum time (García Mayo, 2018; García Mayo & Hidalgo, 2017; Huang, 2015). Age, too, has been a significant variable in child L2 acquisition research, with findings showing that the notion of "the younger, the better" does not universally apply (García Mayo & García Lecumberri, 2003; Muñoz, 2006). Considering all aspects, the notion derived from bilingual or immersion contexts that YLs effortlessly absorb the FL does not apply when transferred to environments with limited exposure to language input, especially in large group settings (Copland & Garton, 2014). While introducing a FL in primary curriculum may have eventual benefits for children, there is a notable lack of research on how children actually approach specific learning tasks in such contexts. Furthermore, there is a need to explore effective teaching methods tailored to YLs' specific needs in these constrained settings compared to adults (García Mayo, 2018). One particular learning task that demands further exploration is writing, along with the feedback provided on the written output, which appears to engage cognitive processes such as formulating and testing hypotheses, focusing attention, engaging in metalinguistic reflection, recognizing linguistic patterns, and employing problem-solving strategies (Williams, 2012). In English as a Second Language (ESL) and English as a foreign language (EFL) settings, students might be accustomed to receiving and assessing feedback, particularly when it comes to writing
5 assignments. However, this pattern is more prevalent among older learners who typically display greater receptiveness to feedback concerning their linguistic mistakes. Conversely, the tendency with younger learners is to prioritize oral tasks over written tasks and feedback provision on their written work is often overlooked. This challenge underscores the reason why children frequently lack familiarity with composing in a second or foreign language (Cánovas Guirao, 2017). Some authors have linked this lack of engagement with writing to an underestimation of YLs cognitive or linguistic capacities (e.g., Lázaro-Ibarrola, 2023; Muñoz, 2017; Tellier & Roehr-Brackin, 2017; Villarreal & Munarriz-Ibarrola, 2021; Villarreal & Martínez-Sánchez, 2023). The potential impact of WCF on facilitating language acquisition has been extensively investigated in recent years (Bitchener & Knoch, 2009). It is within this context that exploring the contribution of WCF to the language learning of children emerges as a promising avenue for investigation. We also need to know more about the role that different types of feedback play in L2 learning (Ellis, 2008) and the extent to which oral interaction affects written production (Storch, 2021). Additionally, another big challenge for the SLA field is the design of longitudinal studies, since they may provide valuable insights into changes that children undergo over time (Long, 2014). 2.2. WCF in the L2 classroom: model texts as a written feedback alternative SLA research has paid attention to the potential of feedback in written production and its impact on the development of the learner's interlanguage through hypotheses testing (Polio, 2012). The argument is that feedback is related to the processes of attention to form, which is crucial for acquisition (Ellis, 2016). However, research on feedback has posed a series of limitations
6 as to its effectiveness when it is delivered in the form of correction of errors. Although positive results have been reported regarding the effects of feedback on a limited number of linguistic forms (e.g., the English article system or the simple past tense) (Bitchener, 2008), some researchers have found fault with the merits of error correction (EC) arguing that its implementation may (i) lack clarity, precision, and consistency (Ellis, 2009); (ii) lack sensitivity on the part of teachers to students’ needs and ability levels (Hyland, 1998); (iii) produce confusion, anxiety and stress among learners due to the overwhelming number of corrections, making it challenging to discern non-target-like elements (Hyland, 1998); (iv) result in passive reception and superficial hypotheses testing (Adams, 2003) and (v) lead to superficial and transient memorization of the explicit rules or corrections (Truscott, 1996). Consequently, the value of traditional EC for L2 development is still a vexed question (Martínez Esteban & Roca de Larios, 2010). Other feedback techniques should be deployed at least as partial options to traditional feedback methods. One of the alternatives that has drawn special attention among researchers are model texts, since there is evidence that they play an instrumental role in promoting noticing and metalinguistic awareness, and in engaging students in deeper processing than more traditional feedback methods such as EC (Hanaoka, 2007; Hanaoka & Izumi, 2012; Martínez Esteban & Roca de Larios, 2010). Models are a type of feedback which consists of providing learners with native or native-like texts that they compare with their original draft (Martínez Esteban & Roca de Larios, 2010). This more discursive feedback method does not explicitly single out errors but treats the text as a whole providing appropriate language, organization, mechanics, style and ideas for a given context, rather than offering lists of corrected errors, editing symbols or metalinguistic codes (Cánovas Guirao, 2017).Models also appear to stimulate cognitive conflict by presenting information that contradicts students' hypotheses about language mechanics, prompting them to contemplate alternative forms for expressing their intended meanings (Coyle & Roca de
7 Larios, 2020). Because errors are not explicitly indicated, modeling has also proven effective in encouraging learners to identify their own mistakes, potentially leading to more in-depth cognitive processing (Sachs & Polio, 2007). In fact, the primary advantage of models lies in creating an optimal environment for students to recognize the likenesses and disparities between their interlanguage and the target language, prompting a re-evaluation and affirmation of their L2 comprehension and consequently, a refinement of their output (Sachs & Polio, 2007). Furthermore, when interacting with the native-like model text, learners are not only capable of rectifying their limitations but also have the opportunity to encounter native-like vocabulary, grammar, ideational, and organizational structures. This exposure serves as an impetus for incorporating such elements into their subsequent texts, aiming to convey meaning more effectively. Consequently, it contributes to their writing proficiency, extending beyond mere error correction. For example, Coyle and Cánovas Guirao (2019) highlighted the use of longer sentences on the part of the students, while Villarreal and Lázaro-Ibarrola (2022) observed more syntactically complex structures to aptly encode the intended message. What is more, although Ellis (2016) notes that explicit WCF is more effective than implicit WCF in that the former is more likely to assure attention to form, he argues that implicit forms of WCF (such as models) have also proved to be beneficial and may even have a greater longterm effect. On the other hand, according to Villarreal and Lázaro-Ibarrola (2022, p. 2), model texts can be seen as bridging "the boundary between direct and indirect forms,”. So, we could argue that they might possess the advantages of both approaches. Taken together, models would seem to serve as a valuable feedback tool for EFL learners. Notwithstanding this great pedagogical appeal, there is a small body of literature that is concerned with the impact of models on the children’s drafts. 2.3. Empirical research on models
8 The typical sequence of research with models consists of engaging learners in a three-stage writing task involving (i) noticing of linguistic problems while writing a picture-based story (Stage 1 (S1)), (ii) comparison of their initial drafts with the native-speaker model, which may offer solutions to their problems, and noticing (or not) of those solutions (Stage 2 (S2)) (during this stage, learners also take note of the differences between their drafts and the model's solutions), and (iii) rewriting of their original text, which works as an immediate post-test to see whether learners incorporate the solutions into the revised texts (Stage 3 (S3)). Some studies have included a delayed post-test which was carried out a week (Luquin & García Mayo, 2021) or two months later (Hanaoka, 2006a, 2007). So far models have been examined in both individual and collaborative writing with EFL highschool students of Spanish (García Mayo & Loidi Labandibar, 2017; Martínez Esteban & Roca de Larios, 2010; Montealegre Ramón, 2019), and Korean (Kang, 2020); with primary school children (Luquin & García Mayo, 2020, 2021; Cánovas Guirao et al., 2015; Coyle & Roca de Larios, 2020; Coyle et al., 2018; Lázaro-Ibarrola, 2021; Criado et al., 2022; Lázaro-Ibarrola & Villarreal, 2021; Roothooft et al., 2022; Villarreal & Lázaro-Ibarrola, 2022); and with Japanese ESL (Abe, 2008) and EFL (Hanaoka, 2006a, 2007) university students. Models have also been compared to reformulations with two Japanese EFL learners (Hanaoka, 2006b), and with Japanese (Hanaoka & Izumi, 2012) and Chinese EFL college students (Yang & Zhang, 2010). Likewise, the use of models has been contrasted with EC with primary school EFL learners in Spain (Coyle & Roca de Larios, 2014). Out of all the studies on models carried out with children, to our knowledge, only a few have been conducted on how models affect the children’s final drafts. The first attempt was made by Coyle and Roca de Larios (2014), who studied the effect of EC and models on the noticing and revisions of 46 Spanish EFL children (aged 11-12). Participants were divided into four
9 groups and received different treatments of EC and models. The results showed that the group that received both EC and models had the highest scores in accuracy and fluency. The study suggests that the use of EC helps children notice errors, the use of models helps children produce more accurate and fluent language, and the combination of both can be more effective than either strategy alone. Later, Coyle et al. (2018) conducted a study based on Cánovas Guirao (2017), which examined the role of models in collaborative writing and feedback cycles over a five-month period including a period of instruction. The participants were 16 Spanish EFL children (aged 10-11) divided into a group that received instruction over a period of six weeks and a group that did not. The study focused on the process of writing by identifying the trajectories the children followed across tasks and added a long-term dimension to previous one-shot studies. The 14 trajectories were categorized based on the nature of challenges (or the absence thereof) encountered by the learners during the creation of their initial drafts. These encompass a single trajectory where no related output was initially present, three trajectory groups illustrating the development of learners' unresolved, unreported, and resolved problems (, and a concluding trajectory addressing the incorporation of new content during text revision. Each trajectory took a distinct course, shaped by the processing mechanisms initiated by the learners at various stages of the task. The authors found that the gains in acceptability and comprehensibility of the children's written output were achieved as a result of using trajectories with greater potential to promote language learning. The study certainly opened a new dimension toward the processing of WCF but the small number of participants and the lack of statistical analysis prevents the extrapolation of the results. Also, improvement on the part of the teaching group could also be attributed to a task-repetition effect and not only to prolonged training.
16 usage of verb tenses, and morphological aspects such as subject-verb agreement. Additionally, they addressed concerns related to the coherence, cohesion, and stylistic elements of the text. In the third stage, the rewriting stage, every pair received the same picture prompt again (no model text or initial drafts were provided this time). The children were then tasked with rewriting the story, aiming to recollect and integrate the aspects they had identified in the feedback from the prior week. The students were unaware of this task in advance, as this approach aimed to prevent the memorization of corrections. The final stage, or delayed posttest, involved the production of a text based on a new visual prompt (see Appendix C). This prompt narrated a story similar to the first one and contained the same number of cartoons. The aim was to allow students to incorporate some of the features present in the initial story. We deliberately chose a different visual stimulus that was not too dissimilar to ensure that the tasks remained isomorphic. The purpose behind this approach was twofold. Firstly, we wanted to assess how much of the feedback the students could remember and retain in the short term. Secondly, we aimed to distinguish potential task-repetition effects from actual learning from the model. This ensured that any new linguistic knowledge the children might have acquired and recalled was not solely due to performing the same task twice. Between February and May, both the LTG and the CG received a single picture prompt every month and undertook stages 1 (writing), 2 (comparison or self-correction), and 3 (rewriting) for each picture. These intermediate sessions were led by their corresponding teachers. 3.4. Data analysis The participants’ collaborative texts (a total of 180 texts: 30 from S1, 30 from S3, and 30 from S4 in Cycle 1, and 30 from S1, 30 from S3, and 30 from S4 in Cycle 2) produced throughout the stages in both cycles were transcribed and analyzed considering the following aspects: Type of clause, CAF and holistic measures. The incorporations of feedback into revised texts are
17 typically coded as either accurate or not, as documented in previous studies with adults (see, for example, Yang & Zhang, 2010). This methodological decision completely misses out all the partial gains that children may obtain from WCF, even more so when the ways in which YLs respond to feedback are much less regular and stable. Although we decided to replicate other studies (e.g., Cánovas Guirao, 2017; Coyle & Roca de Larios, 2014; Torras, 2005) in the analysis of the children’s written output, we found it important to use much more nuanced approaches and a variety of them to capture minor gains across texts (Coyle & Roca de Larios, 2014; García Hernández et al., 2017; Villarreal & Lázaro-Ibarrola, 2022). The texts, produced by learners with limited and insufficient knowledge that in many cases limits the ability to form propositions and articulate them consistently, forced the adoption of analytical criteria from the idiosyncratic perspective of the nature of the texts themselves. To assess the immediate effect of the models on the children's writing, the quality of the third draft (Stage 4) was compared to the initial draft (Stage 1) in both cycles (comparing draft 1 to draft 3 and draft 4 to draft 6). We opted to juxtapose the initial draft with the delayed post-test draft, eschewing the intermediate Stage 3 draft to mitigate potential task-repetition effects. Our objective was to scrutinize the evolution of students as writers, focusing on the manner and facets through which they refined their skills, rather than assessing their competence as mere replicators of the identical task. To study the long-term impact of feedback, the beginning and final drafts (comparing draft 1 to draft 6) were compared. A mixed ANOVA was employed for the analysis, incorporating both a between-groups variable (group) and a within-groups variable (time). To pinpoint differences, post-hoc tests using the Bonferroni correction were applied to adjust the alpha level for the number of comparisons, minimizing the risk of Type 1 error. For correlated variables, the Mauchly test was conducted to verify the assumption of sphericity. In cases where this assumption was not met, the Greenhouse-Geisser test was
18 employed to rectify the lack of sphericity. The significance threshold was established at α = 0.05. Effect sizes were also computed for each statistical procedure. Partial eta squared was used to gauge the effect size for both one-way and mixed ANOVA. Cohen's (1988) recommendations were followed, categorizing effect sizes between 0.06 and 0.1 as 'small,' 0.15 as 'medium,' and between 0.15 and 1 as 'large.' Post-hoc tests were evaluated using the d family of effect sizes. As outlined by Cohen (1992), an effect size of d = 0.2 corresponds to a 'small' effect, 0.5 signifies a 'medium' effect, and 0.8 indicates a 'large' effect. 3.4.1. Type of clause Building upon the work of Torras (2005) and Cánovas Guirao (2017), we segmented the participants' texts into clausal units based on their grammatical correctness. This coding scheme was chosen for its suitability in conducting a comprehensive analysis of written output from young EFL learners, characterized by simplicity and brevity in their texts. Three units were identified: pre-clause, proto-clause and clause. They are defined and exemplified as follows: Pre-clause: grammatically incorrect unit of language consisting of fragmented or distorted strings of words, at times incomplete, in which the meaning intention is not always apparent. Example: And tought with the plum
19 Proto-clause: Linguistic unit in which the children’s meaning intention is clear but which contains grammatical inaccuracies or gaps in the clausal unit. Example: Then go to the family of Helen Clause: grammatically accurate unit of language which may present a slight inaccuracy in spelling, lexis, grammar or concordance. Example: One day morning at six o’clock, Lucy was sliping In order to track the children’s writing development, the total number of units for each clause type was tallied and compared in their first and revised drafts in both cycles and across groups. 3.4.2. CAF measures and lexical diversity With the purpose of identifying potential progress in the linguistic acceptability and comprehensibility of the learners’ written texts from their original to their revised texts in both cycles, we followed Torras (2005), Torras et al. (2006), Cánovas Guirao (2017) and Coyle and Roca de Larios (2014), since they also analyzed the written output of child EFL learners. The measures were classified into four main areas: (i) accuracy, (ii) fluency, (iii) syntactic complexity and (iv) lexical diversity. (i) Accuracy: Following Cánovas Guirao’s (2017), an error ratio was used to measure overall accuracy: [number of linguistic errors/total number of words] × 10. Like Cánovas Guirao (2017), we decided to use a 10-word ratio rather than the 100-word ratio as the children’s texts were relatively short (i.e., less than 100 words). Error ratios were computed and compared as displayed in Table 1 below.
20 Table 1 Identification and revision of errors ERROR RATIOS Original text Revised tex t Today it is Monday and the sun is going up Ana is sliping at 6.00 a.m. the clock has started to ring. Ana don’t want to woke up but she has to go to the shool but her clock do things to Ana wake up. Now, Ana is brushing her head and washing her teets. Sudently, she takes her shool bag and goes to the shool. Pair 29, DTG, Stage 1, Cycle 1 Today is Monday morning and Martine is sleeping in her bed at six o´clock her clock´s alarm starts ringing but Martine isn´t want to wake up. Martine puts her feet on her pillow, but is sleep. Now, her clock takes a feather and starts touching her feet for she get up to go to school. Now Martine´s washing her teeth and combing her hair. She takes her school bag and goes to school. Pai r 29, DTG, Sta g e 3, C y cle 1 Nº words: 65 Nº errors: 20 Error ratio: (20/65)x10=3,08 Nº words: 66 Nº errors: 8 Error ratio: (8/66)x10=1,21 Today (1) it is Monday and the sun is going up (2). Ana is (3) slieeping (4) and at 6.00 a.m. the clock (5) has started starts to ring. Ana (6) don’t doesn’t want to (7) woake up but she has to go to (8) the (9) shool school (10). (11) but her clock (12) does things to (13) Ana wake up Ana. Now, Ana is (14) brushing combing her (15) head hair and washing brushing her (16) teets teeth. (17) Sudently suddenly, she takes her (18) shool school bag and goes to (19) the (20) shool school. Today is Monday morning and Martine is sleeping in her bed (1). at six o´clock her (2) clock´s alarm clock starts ringing but Martine (3) isn´t doesn’t want to wake up. Martine puts her (4) feet foot on her pillow, but (5) is sleep continues sleeping. Now, her clock takes a feather and starts touching her (6) feet foot (7) for she get up to go to school to wake her up (because she has to go to school). Now Martine´s (8) washing brushing her teeth and combing her hair. She takes he r school b a g an d g oes to school. (ii) Fluency: the total number of words per text was also considered. An online text analysis tool (https://textinspector.com/workflow) was used for word calculation. (iii) Syntactic complexity: Following Torras et al. (2006), syntactic complexity was measured as number of subordinate and coordinate clauses over total number of clauses. Table 2 features examples of how fluency and syntactic complexity were coded.
21 Table 2 Example of fluency and syntactic complexity codification One day a scientist was doing a potion while is dog was sleeping. When he finishes his potion, he was excited to taste it and he drinks the potion. When he drinks the potion he started feeling bad. Suddenly a bright light apier and a loud sound sounds and he turn into a cat. When the dog heard the sound it woke up and started fighting with the cat. ‘Miau’ were the last words of the scientist cat. Pai r 13, TG, Stage 3, Cycle 2 Subordinate clauses: - While his dog was sleeping - When he finishes his potion - When he drinks the potion - When the dog heard the sound Coordinate clauses: - and he drinks the potion. - and a loud sound sounds - and he turn into a cat. - and started fighting with the cat. Total word count: 78 Number of sentences1: 6 Number of clauses: 14 Number of subordinate clauses: 4 Number of coordinate clauses: 4 Clauses per sentence: [14/6]=2,33 (iv) Lexical diversity: some extensively used measures such as lexical complexity (number of verbs, adjectives, noun, etc. types) as used by Torras et al. (2006) or lexical density (what proportion of the text contains lexical words) were discarded. This methodological decision was motivated by the fact that the focus of the present analysis was not to obtain the total number of lexical words or the different types of lexical categories present in the children’s texts, but rather to explore how many different words appear in each text. Consequently, lexical diversity (or lexical richness) was used as a measurement for newly learned vocabulary, since we considered that it might better reflect the development of the children’s interlanguage. 1 A maximal clause, which conveys an independent meaning
22 Lexical diversity is usually calculated using a type-token ratio (TTR). A high TTR indicates a high degree of lexical variation. Although the TTR can be an extremely useful measurement for calculating the lexical diversity of a text, a common problem with this measure is that it does not work as effectively when dealing with texts of different length. Guiraud (1960) proposed a measure called ‘Root Type Token Ratio’ (RTTR) which is obtained by dividing the number of types by the square root of the number of tokens thus partially addressing the problem of TTR’s variance on text length. Each transcribed text was uploaded to a software tool (https://textinspector.com/workflow) which would automatically calculate the number of types and tokens. Prior to uploading the children’s written texts, spelling mistakes were corrected as the aim of lexical diversity is to analyze how diverse the range of words used is and therefore the software must recognize the words. Once this information was provided, we applied Guiraud’s (1960) formula to the data (types/√tokens), thus obtaining the RTTR for each text. 3.4.3. Holistic measures We also assessed the texts both quantitatively and qualitatively taking into account measures of adequacy, coherence, cohesion, grammatical accuracy, lexical range and mechanics. A three-point scoring rubric was used to evaluate the writings, 3 being good, 2 average and 1 poor. We opted for the rubric developed by Villarreal and Munarriz-Ibarrola’s (2021) (see also Hedgcock & Lefkowitz, 1992). This rubric was chosen due to its ability to effectively target precise learning objectives and skills, facilitating a thorough assessment of the desired outcomes. To ensure rating reliability, the participants’ written production was coded by one researcher and 20% of the data was independently coded by a second researcher. Inter-rater agreement resulted in 97%, and differences were discussed until total agreement was reached.
23 4. Results In order to examine the impact of the models on the children’s written production in the shortterm, the quality of the third draft (Stage 4) in relation to the first one (Stage 1) was measured in both cycles (draft 1 vs. draft 3 and draft 4 vs. draft 6). To examine the effect of the feedback in the long run, the first and last draft (draft 1 vs. draft 6) were compared (see Appendix D for a table containing descriptive statistics, including the type of clause and CAF measures, organized by group and draft). As for pre-clauses, within the global domain of type of clause, the Greenhouse-Geisser correction was used since the sphericity assumption was not met (X2 = 13.74; p = .017). Results revealed a main effect for Pre-clause (F(2.21,59.72) = 16.92; p = < .000; ηp2 = 0.38), Group (F(2,27) = 3.74; p = .037; ηp2 = 0.21), and for the Pre-clause x Group (F(4.42,59.72) = 8.03; p = .048; ηp2 = 0.07) interaction. Pairwise comparisons with Bonferroni adjustment showed some statistically significant intraand inter-subject differences. As for the former, the TG significantly reduced the number of pre-clauses from drafts 1 to 3 (p = .004, 95% CI [0.33, 2.21], d = 0.70), 4 to 6 (p = .045, 95% CI [-0.40, 1.13], d = 0.25) and 1 to 6 (p = .024, 95% CI [0.13, 2.60], d = 0.58). The LTG also improved significantly from their first to their third composition (p = .004, 95% CI [0.31, 2.29], d = 0.68), and from the first to the last one (p = .031, 95% CI [-0.25, 3.20], d = 0.48). The self-correction group showed no improvement in terms of pre-clauses. Group differences were only observed on the first delayed post-test (draft 3) where both treatment groups outperformed the CG (p = .007, 95% CI [0.28, 2.02], d = 0.61 for CG-TG; p = .002, 95% CI [0.44, 2.23], d = 0.70 for CG-LTG). In conclusion, both treatment groups exhibited a decrease in the number of pre-clauses from drafts 1 to 3 and from drafts 1 to 6. Furthermore, they surpassed the CG in the post-test of Cycle 1 (draft 3) by producing a lower count of preclauses.
24 Regarding proto-clauses, Mauchly’s test did not indicate any violation of sphericity (X2 = 4.07; p = .539). A main effect was found for Proto-clause (F(3,81) = 4.73; p = .004; ηp2 = 0.14) and for the interaction between Proto-clause and Group (F(6,81) = 2.82; p = .044; ηp2 = 0.14). Post-hoc analyses once more unveiled significant disparities between draft 1 and draft 6 within both model groups (p = .032, 95% CI [-1.61, 3.52], d = 0.37 for TG; p = .006, 95% CI [0.63, 4.97], d = 0.67 for LTG), but not for the CG, indicating that in the TG there was a notable reduction in proto-clauses from draft 1 to draft 6, while a significant decrease was observed for the LTG during this transition. This time, no between-group differences were observed for this aspect. Concerning clauses, sphericity was met (X2 = 8.34; p = .139), as indicated by Mauchly’s test. There is a significant main effect for Clause (F(3,81) = 14.18; p = < .000; ηp2 = 0.34), for Group (F(2,27) = 5.52; p = .043; ηp2 = 0.12), and for the Clause x Group (F(6,81) = 7.96; p = .027; ηp2 = 0.46) interaction. Further pairwise comparisons revealed statistically significant within-group and between-group differences. In this case, we found that the CG wrote statistically fewer clauses in draft 6 in contrast to draft 1 (p = .002, 95% CI [-0.48, 5.30], d = 0.41), and the LTG produced a significantly higher number of clauses in draft 6 in comparison with draft 4 (p = .012, 95% CI [-3.66, -0.34], d = -0.63). No differences across drafts were observed for the TG. Looking at between-group differences, post-hoc analyses revealed that the LTG outperformed both the CG (p = .001, 95% CI [1.30, 5.21], d = 0.78) and the TG (p = .014, 95% CI [0.39, 4.11], d = 0.56) in draft 6. In summary, we noted a decrease in clause production from the initial to the final version of the texts within the CG. Conversely, for the LTG group, a significant increase in the number of clauses was observed from drafts 4 to 6. Across all groups, the LTG group generated a significantly higher number of clauses in draft 6 compared to the other two feedback conditions.
25 Moving on to the area of complexity, and more specifically to subordinate clauses, we used the Greenhouse-Geisser correction since the sphericity assumption was rejected (X2 = 20.13; p = .001). Mixed ANOVA showed a main effect for Subordinate clause (F(2.07,56.10) = 11.49; p = < .000; ηp2 = 0.29), and for the Subordinate clause by Group (F(4.15,56.10) = 9.62; p = .037; ηp2 = 0.15) interaction effect. Further post-hoc tests located these differences for both experimental groups between drafts 1 and 3 (p = .006, 95% CI [0.45, 3.55], d = -0.67 for TG; p = .024, 95% CI [-2.29, 1.09], d = -0.18 for LTG) and 4 and 6 (p = .048, 95% CI [0.62, 3.16], d = 0.37 for TG; p = .001, 95% CI [1.34, 3.22], d = 0.13 for LTG), which means that following the receipt of the models in both cycles, specifically on both post-tests, the treatment groups integrated a greater number of subordinate clauses when compared to their initial drafts in Stage 1. No statistically significant differences were found for the CG or between groups. With reference to coordinate clauses, the assumption of sphericity was not rejected (X2 = 10.93; p = .053). The results obtained from the analysis of coordinate clauses only revealed a significant main effect for Coordinate clause (F(3,81) = 18.11; p = < .000; ηp2 = 0.40). In terms of lexical diversity, Mauchly’s tests indicated that sphericity was not violated (X2 = 6.27; p = .281). We observed a significant main effect for Lexical diversity (F(3,81) = 23.18; p = < .000; ηp2 = 0.46), Group (F(2,27) = 9.33; p = .028; ηp2 = 0.09), and for the interaction between Lexical diversity x Group (F(6,81) = 16.43; p = .021; ηp2 = 0.17). For the Bonferroni post-hoc test, there is a statistical difference between drafts 1 and 6 within the three groups. That is, the lexical repertoire of the children in both the CG (p = .013, 95% CI [10.53, 123.03], d = 0.62) and the TG (p = .002, 95% CI [23.21, 124.97], d = 0.76) appeared to be significantly richer in draft 1 as opposed to draft 6. On the contrary, the LTG (p = .004, 95% CI [18.84, 125.56], d = 0.70) appeared to increase the quantity of lexical words, and this rise was found to be
32 increase in clauses within such a short time frame may support the idea that the greater the exposure to models, the greater the benefits the children may obtain, at least in certain linguistic aspects. Although further study is warranted given the scant research existing to draw a firm conclusion, these positive findings hint at a possible relation between models and the linguistic acceptability of the children’s written output. Exposure to the model text within the first and second writing cycles was also found to have positive short-term effects on enhancing the syntactic complexity of the children’s third draft, which is visible in the higher number of subordinate clauses produced by the treatment groups on both post-tests. No differences across groups or cycles were found. No differences were observed either when analyzing the evolution of lexical diversity indicators from preto posttests in each group, but the results revealed that the LTG showed higher lexical diversity than the other two groups both when they initiated (draft 4) and finished (draft 6) Cycle 2. The fact that they outperformed the other two groups as soon as Cycle 2 began manifests once again a positive relationship between repeated exposure to models and the variety of words used in the children’s texts. As a matter of fact, long-term results revealed that the number of different words used by the learner pairs in the CG and the TG decreased significantly from draft 1 to draft 6. On the contrary, the group benefitting from models during four months showed a significant improvement, visible from their first to their last written text. If we consider other similar studies, we observe mixed results. For example, in their cross-sectional study with children, Villarreal and Lázaro-Ibarrola (2022) observed that the model group wrote more grammatically and lexically complex texts on the post-test. This finding, however, does not support that of Lázaro-Ibarrola (2021), who did not report any text improvements in terms of CAF with a similar sample and after having revised the exact same models as in Villarreal and Lázaro-Ibarrola’s (2022) study. For her part, Cánovas Guirao (2017) revealed that all children
33 wrote less grammatical texts than in the first writing cycle, whereas Roothooft et al. (2022) observed that the children’s drafts significantly improved in lexical diversity, but not in syntactic complexity, with models. Given the conflicting results, more textual analyses are needed to determine whether the benefits associated with models lead to gains in lexical diversity and syntactic complexity. With respect to accuracy, no meaningful short-term findings were observed. The only statistically significant result was for the LTG, who made significantly fewer mistakes from draft 1 to draft 6, and differences were also observed with respect to the CG and the TG in both drafts 4 and 6. The short-term findings correlate favorably with Villarreal and Lázaro-Ibarrola (2022) and Lázaro-Ibarrola (2021), who did not obtain any remarkable gains for any of the groups. The long-term results, however, somewhat mirror the findings obtained in the measures for type of clause, where we observed a significantly higher number of clauses or acceptable units after a long treatment with models. This lack of impact on accuracy in the short run may support the hypotheses of the children’s developing metalinguistic ability and the primacy of meaning over form (VanPatten, 2004). Furthermore, model texts have garnered praise primarily for their emphasis on lexis (Chandler, 2003; Coyle & Roca de Larios, 2014), while also being linked to syntactic complexity advantages as noted by Villarreal and Lázaro-Ibarrola (2022). Whatever the reason, one more time these obstacles were overcome by the children who had access to several models for the past four months, as they managed to produce more accurate and thus more acceptable and comprehensible texts. In line with previous research (Cánovas Guirao, 2017; Coyle & Roca de Larios, 2014; LázaroIbarrola, 2021; Roothooft et al., 2022; Villarreal & Lázaro-Ibarrola, 2022), fluency appears to be resilient to models. The similarity in the total number of words written before and after
34 exposure to the two feedback techniques indicates that neither form of feedback has much impact, at least in the short term, on the length of the texts. It seems that the need to convey meaning fosters qualitative changes, but not quantitative ones (Coyle & Roca de Larios, 2014; Lázaro-Ibarrola, 2021; Villarreal & Lázaro-Ibarrola, 2022). Coyle and Roca de Larios (2014) ascribe this stability of fluency across drafts to the briefness of the texts the children were asked to write. In a similar vein, Villarreal and Lázaro-Ibarrola (2022) attribute this outcome to task type, which clearly directs the children’s attention to content and inhibits more creative attempts. However, this finding does not coincide with Criado et al. (2022), whose participants in the model group demonstrated advancement in fluency and the self-correction group exhibited even greater fluency than the feedback group. It is possible that the previous penand-paper studies are at a disadvantage when compared to the digital writing used by Criado et al. (2022). This prompts speculation that digital writing might hold an upper hand over traditional paper-based writing, as suggested by Villarreal et al. (2021). Notwithstanding the above, we did find long-term changes in the number of words produced. When comparing the first and last texts, fluency underwent a pronounced decline in the three groups, regardless of the feedback condition. We interpret this sharp drop in the children’s textual fluency in relation to the end of academic year. The participants were clearly tired and eager to finish the tasks. The majority of participants exhibited a general lack of concentration and motivation, as previously observed in the study by Kopinska and Azkarai (2020), resulting in brief dialogues and shorter texts, which was also reported by Cánovas Guirao (2017). Finally, as for the global analysis of draft quality, the holistic scores showed an across-cycle improvement of the TG’s textual cohesion whereas, in line with the quantitative analysis, the LTG enhanced their lexical repertoire from draft 1 to draft 6. When the groups were contrasted,
35 the TG seemed to do significantly better than the CG in terms of coherence on the first posttest, while the LTG obtained a significantly higher score in cohesion than the CG and the TG on the second post-test. Lázaro-Ibarrola (2021) and Roothooft et al. (2022) also found a notable improvement in the overall evaluation of primary school students who examined their written pieces against model texts, while Villarreal and Lázaro-Ibarrola (2022) did not observe any significant differences. The combination of different measurements when analyzing written production is key to be able to grasp any improvements in draft quality. However, we believe that short-term upgrading may not be as detectable as long-term upgrading when it comes to holistic analysis, since any minor gain obtained in two weeks’ time may not be humanly noticeable. In addition, because the arbitrariness of this tool alongside the short range of values provided in the rubric used in this study made it difficult to guarantee a reliable assessment on the children’s written texts, the results from such analyses should thus be treated with the utmost caution and only be considered as an add-on analysis to the quantitative assessment. In sum, the short and long-term gains observed through type of clause, CAF and holistic analyses offer further empirical evidence that models can and do help L2 learners upgrade their written output and develop their emerging interlanguage. We can thus confirm that models help improve the overall written production of primary EFL students in both the short and long run. More specifically, in the short-term, models helped the children reduce the number of preclauses and proto-clauses as well as increase the syntactic complexity of their texts. After a long exposure to models, the children were able to (i) produce fewer pre-clauses, fewer protoclauses and more clauses, (ii) use a higher diversity of lexical words in their texts and (iii) make fewer errors. As was the case with incorporations, practice with model texts appeared to facilitate the feedback processing, decreasing its complexity and enhancing its noticing processes. This increased noticing then resulted in upgrading of their written texts and
36 consequently in L2 development. Nonetheless, we believe that practice is not enough when it comes to YLs. If we want children to make the most out of models, we need consciousness raising activities to help younger and weaker learners enhance their meta-awareness of language. When compared to previous research, our findings concerning the effect of model texts on text quality among children cast a new light on the controversy about which aspects of language are most benefitted but also help to reinforce the view that models do have a favorable effect on the children’s written output. This study also underscores the need to include a wide variety of fine-grained measures not to miss out any minor improvement and the exploration of different tools that corroborate whether small differences do or do not show a potential trend. 6. Conclusion The present study offers the first longitudinal assessment of models with a substantial sample size of participants, enabling robust statistical analyses to strengthen the findings regarding the enduring advantages of model texts on EFL children's written production. The results of this study highlight the need for young EFL learners to engage in writing activities and to receive and assimilate feedback on their written work in order to enhance their grasp of the L2 (Lázaro-Ibarrola, 2023). In EFL contexts, language teachers often rely on textbooks that ostensibly promote communicative language instruction, offering a range of exercises encompassing formal aspects, vocabulary, and reading, all organized thematically. However, writing frequently takes a back seat, often receiving less attention than oral tasks. Consequently, it becomes imperative for teachers to recognize the advantages of writing and
37 feedback processing in L2 learning, as well as the theoretical implications of writing-to-learn, as highlighted by Cánovas Guirao (2017) and Manchón (2009). Model texts represent just one form of feedback technique, rather than the sole option available. Instead of approaching alternative and traditional WCF as mutually exclusive, teachers can combine these two approaches to serve distinct purposes. For instance, they could offer unfocused and/or indirect feedback on initial drafts, directing attention to a wide array of aspects and facilitating self-correction. Subsequently, direct and/or focused feedback could be supplied for revised drafts, allowing children to rectify remaining errors and produce more precise texts. Prioritizing an understanding of the advantages inherent in diverse feedback methods and determining which technique aligns best with the needs of our learners should be a key consideration for language instructors. Our research has some shortcomings that need to be acknowledged, although we believe that these limitations could be a springboard for future work on the topic. The study is limited first and foremost because of the methodological decisions concerning the analysis of the children’s writing. The texts produced by children at initial stages of L2 learning forced the researchers to adopt analytical criteria from the perspective of the nature of the texts themselves. Therefore, the low quality and briefness of the written texts meant that the instruments and procedures normally used in research with older learners could not be used. Furthermore, the use of distinct measures hampers the possibility of making compelling comparisons. Although we decided to replicate other studies for the analysis of the children’s written output and used the clausal unit measure to capture minor gains across texts, we also wanted to include CAF measures to yield more detailed information. Even so, the present study may fall short when it comes to providing a full and accurate picture of the children’s writing development. Similarly, the arbitrariness of
38 the rubric used in the holistic analysis alongside the short range of values provided in it made it difficult to guarantee a reliable assessment on the children’s written output, and therefore the results obtained from such analyses should be treated with caution and only be considered as a complementary analysis to the quantitative assessment. Finally, it is important to recognize that in this study models served as a means to an ultimate goal. The noticeable gains across different contexts are evidently shown only in a specific picture narration task. However, this insight implies that the observed long-term improvements are indicative of the potential of children as prospective EFL writers. Sustained exposure to models (and alternative feedback techniques) over time may foster this growth. Consequently, conducting follow-up tasks to delve into the ramifications of these gains presents a promising direction for future research.
39 APPENDICES Appendix A: Picture story S1 (Taken from Lapkin et al., 2002)
40 Appendix B: Model text S2 Martine’s alarm clock (Taken from Lapkin, Swain & Smith, 2002)
41 Appendix C: Picture story S4
48 Long, M.H. (2014). Second language acquisition and task-based language teaching. London: Routledge. Luquin, M., & García Mayo, M. P. (2020). Collaborative writing and feedback: An exploratory study of the potential of models in primary EFL students’ writing performance. Language Teaching for Young Learners, 2(1), 73–100. Luquin, M., & García Mayo, M. P. (2021). Exploring the use of models as a written corrective feedback technique among EFL children. System, 98(1), 1–13. Lyster, R. (2007). Learning and Teaching Languages through Content. Amsterdam: John Benjamins. Manchón, R. M. (2009). Broadening the perspective of L2 writing scholarship: The contribution of research on foreign language writing. In R. M. Manchón (Ed.), Writing in foreign language contexts. Learning, teaching, and research (pp. 1-19). Clevedon: Multilingual Matters. Manchón, R. M. (Ed.) (2011). Learning-to-Write and Writing-to-Learn in an Additional Language. Amsterdam: John Benjamins. Martínez Esteban, N., & Roca de Larios, J. (2010). The use of models as a form of written feedback to secondary school pupils of English. International Journal of English Studies, 10, 143–170. Montealegre Ramón, F. (2019). Exploring the learning potential of models with secondary school EFL learners. Spanish Journal of Applied Linguistics, 32(1), 155-184. Muñoz, C. (Ed.). (2006). Age and the Rate of Foreign Language Learning. Clevedon: Multilingual Matters. Muñoz, C. (2017). Tracing trajectories of young learners: Ten years of school English learning. Annual Review of Applied Linguistics, 37, 168-184. https://doi.org/10.1017/S0267190517000095.
49 Nicholas, H., & Lightbown, P.M. (2008). Defining child second language acquisition, defining roles for L2 instruction. In J. Philp, R. Oliver, & A. Mackey (Eds.), Second language acquisition and the younger learner: Child’s play? (pp. 27–52). Amsterdam/Philadelphia: John Benjamins. Polio, C. (2012). The relevance of second language acquisition theory to the written error correction debate. Journal of Second Language Writing, 21, 375 – 389. Reinders, H. (2009). Learner uptake and acquisition in three grammar-oriented production activities. Language Teaching Research, 13(2), 201–222. Roothooft, H., Lázaro-Ibarrola, A., & Bulté, B. (2022). Task repetition and corrective feedback via models and direct corrections among young EFL writers: Draft quality and task motivation. Language Teaching Research. https://doi.org/jtfh Sachs, R., & Polio, C. (2007). Learners’ uses of two type of written feedback on a L2 writing revision task. Studies in Second Language Acquisition, 29(1), 67-100. Schmidt, R. (2001). Attention. In P. Robinson (Ed.), Cognition and second language instruction (pp. 3–32). Cambridge: Cambridge University Press. Storch, N. (2016). Collaborative writing. In R. M. Manchón & P. K. Matsuda (Eds.), Handbook of second and foreign language writing (pp. 387–406). Berlin: De Gruyter. Storch, N. (2021). Collaborative writing: Promoting languaging among language learners. In M. P. García Mayo (Ed.), Working Collaboratively in Second/Foreign Language Learning (pp. 13–34). Berlin: De Gruyter Mouton. Tellier, A. & Roehr-Brackin, K. (2017). Raising Children’s Metalinguistic Awareness to Enhance Classroom Second Language Learning. In M. García Mayo (Ed.), Learning Foreign Languages in Primary School: Research Insights (pp. 22-48). Bristol, Blue Ridge Summit: Multilingual Matters. https://doi.org/10.21832/9781783098118-004.
50 Torras, R. (2005). Procesos psicolingüísticos implicados en la adquisición del inglés en el contexto de la enseñanza primaria [Psycholinguistic processes involved in learning English in a primary school context]. Lenguaje y Textos, 23, 89–112. Torras, M.R., Navés, T., Celaya, M.L., & Pérez-Vidal, C. (2006). Age and IL development in writing. In C. Muñoz (Ed.), Age and the Rate of Foreign Language Learning (p. 156182). Bristol: Multilingual Matters. Truscott, J. (1996). The case against grammar correction in L2 writing classes. Language Learning, 46(2), 327-369. VanPatten, B. (2004). Input processing in SLA. In B. VanPatten (Ed.), Processing instruction: Theory, research and commentary (pp. 5–32). Mahwah, NJ: Erlbaum. Villarreal, I., Bueno-Alastuey, M. & Sáez-León, R. (2021). Computer-based collaborative writing with young learners: Effects on text quality. In M. García Mayo (Ed.), Working Collaboratively in Second/Foreign Language Learning (pp. 177-198). Berlin, Boston: De Gruyter Mouton. https://doi.org/10.1515/9781501511318-008 Villarreal, I., & Lázaro-Ibarrola, A. (2022). Models in collaborative writing among CLIL learners in Primary School: Linguistic outcomes and motivation matters. System, 110(4), 1-21. Villarreal, I., & Martínez-Sánchez, A. (2023). Exploring immediate and prolonged effects of collaborative writing on young learners' texts: L2 versus FL. IRAL International review of applied linguistics in language teaching, 61. https://doi.org/10.1515/iral-2022-0062. Villarreal, I., & Munarriz-Ibarrola, M. (2021). “Together we do better”: The effect of pair and group work on young EFL learners’ written texts and attitudes. In M. P. García Mayo (Ed.), Working Collaboratively in Second/Foreign Language Learning (pp. 89-116). Berlin, Boston: De Gruyter Mouton. https://doi.org/jtfj
51 Wigglesworth, G., & Storch, N. (2012). What role for collaboration in writing and writing feedback. Journal of Second Language Writing, 21(4), 364-374. Williams, J. (2012). The potential role(s) of writing in second language development. Journal of Second Language Writing, 21, 321–331. Wu, Z., Qie, J., & Wang, X. (2023) Using model texts as a type of feedback in EFL writing. Frontiers in Psychology, 14. doi: 10.3389/fpsyg.2023.1156553 Yang, L., & Zhang, L. (2010). Exploring the role of reformulations and a model text in EFL students’ writing performance. Language Teaching Research, 14, 464–484. Dr. Maria Luquin holds a PhD in Second Language Acquisition from the University of the Basque Country (UPV/EHU). Her dissertation, which was supervised by Prof Dr. María del Pilar García Mayo and funded with a predoctoral research grant (FPI) from the Ministry of Science, Innovation and Universities, was awarded the highest mark “cum laude”, and has been proposed for the Outstanding PhD dissertation award. She currently holds the positionof Assistant Professor at the Public University of Navarre (UPNA). Her main research interests lie in the use of collaborative writing tasks and model texts as a feedback technique in primary school EFL settings. Moreover, she also focuses on the impact of factors such as motivation or attitudes, as well as on focus on form. In 2021, María Luquin carried out a research stay at Concordia University (Montreal, Canada), under the supervision of Dr. Kim McDonough. https://orcid.org/0000-0002-3464-8396 Dr. María del Pilar García Mayo (https://laslab.org/staff/pilar) is Full Professor of English Language and Linguistics at the University of the Basque Country (UPV/ EHU). She has published widely on the L2/L3 acquisition of English morphosyntax and the study of conversational interaction in EFL. She has been an invited speaker to universities in Europe, Asia and North America and is an Honorary Consultant for the Shanghai Center for Research in English Language Education. Prof. Garcia Mayo is the director of the research group Language and Speech and the MA program Language Acquisition in Multilingual Settings. She is the editor of Language Teaching Research. She was a member of AILA Executive Board-International Committee and the international relations representative of the Spanish Society for Applied Linguistics and currently belongs to the Steering Committee of the Spanish State Research Agency. https://orcid.org/0000-0002-1987-4889