Phonetic imitation of t-glottaling by Czech speakers of English
Abstract
142
Full text
Phonetic imitation of t-glottaling by Czech speakers of English Pavel Šturm (Charles University, Prague)— Joanna Przedlacka (University of Oxford)— Arkadiusz Rojczyk (University of Silesia in Katowice) PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK ABSTRACT The paper focuses on the ability of Czech speakers to explicitly imitate native English realizations of the phoneme /t/ as [ʔ] (t-glottaling). In Czech, glottalization occurs as aboundary signal of wordinitial vocalic onsets. We hypothesize that this allows for abetter imitative performance in the intervocalic context as compared to non-prevocalic contexts. However, an alternative hypothesis based on language-external facts (frequency in the learners’ English input) predicts the opposite pattern. Our experiment involves 30 participants in ashadowing task. In addition to words with /t/, words with /k/ are examined to establish if speakers can generalize to aphonologically similar category to which they have not been exposed. Speakers adapted their pronunciation after exposure to t-glottaling to some degree. Our hypothesis was confirmed for the shadowing task, while the alternative language-external hypothesis was confirmed for the post-test task, suggesting adifferent pattern of performance in terms of imitation versus learning. KEYWORDS adaptation, glottalization, glottal stop, phonetic imitation, t-glottaling DOI https://doi.org/10.14712/18059635.2022.1.8 1 INTRODUCTION In this study we investigate the phenomenon of “t-glottaling”, i.e., the replacement of the English alveolar plosive [t] with aglottal stop [ʔ] in certain positions (often described as syllable-coda positions), resulting in pronunciations such as [ˈbeʔə] for better or [ˈhɒʔ] for hot. Arelated topic is that of “glottal reinforcement”, when the glottal gesture is produced along with the oral gesture rather than as its replacement (e.g. [ˈhɒʔt]). Considering that in many varieties we also find the alveolar flap [ɾ] in intervocalic positions, and most varieties have astrongly aspirated (and potentially affricated) plosive [tʰ] as well, the allophonic variation of the English /t/ phoneme is thus complex (cf. Skarnitzl and Rálišová 2022) and potentially challenging for learners. We present the results of aphonetic shadowing experiment with Czech speakers of English in which the participants were explicitly instructed to imitate native speakers’ productions. This language group was selected for structural reasons. In contrast to the English words presented above, Czech words involve neither glottal replacement nor reinforcement of [t]. As such, those variants are unfamiliar to the Czech speaker. However, although the Czech word bota /ˈbota/ ‘shoe’ would never be OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 143 pronounced as [ˈboʔa], the glottal stop is acommon sound in the Czech inventory, appearing as aboundary marker in vowel-initial words or morphemes (see Section 1.3). Therefore, Czech speakers are quite familiar with the glottal stop in the prevocalic position, including the intervocalic position (Potkal naopak Adama. [ˈpotkal ˈnaʔopak ˈʔadama] ‘On the contrary, he ran into Adam.’). This leads us to hypothesize that if Czech speakers are supposed to learn or phonetically imitate English words with t-glottaling, the task will be easier and the imitators more successful in the intervocalic context ([ˈbeʔə]) than in non-prevocalic contexts (before apause, hot [ˈhɒʔ], or aconsonant, hot weather [ˈhɒʔweðə]). However, an alternative hypothesis predicts that the intervocalic context should in fact be harder to imitate, since Czech learners encounter t-glottaling more frequently in preconsonantal or prepausal positions if we assume an extensive exposure to Standard Southern British English (see Jakšič and Šturm 2017) as opposed to varieties of English where t-glottaling occurs also intervocalically (see Sections 1.2 and 1.4). The experiment primarily aims to test these two conflicting hypotheses. 1.1 PHONETIC IMITATION AND LANGUAGE ADAPTATION Phonetic imitation appears to be afundamental human behaviour that plays acrucial role in language learning and acquisition. Infants imitate speech sounds in their ambient language to acquire new words (Kuhl and Meltzoff 1996). Later in life, phonetic imitation may be asource of language adaptation when speakers gradually pick up an accent of anew region (Chang 2012; Sancier and Fowler 1997). Finally, phonetic imitation may be adriving force in propagating sound changes when aspeaker imitates and adopts anew sound and passes it to the next interacting speaker (Labov 2001; Lin et al. 2021; Siegel 2010). Different research paths have attempted to investigate the nature of speech imitation and strived to identify the factors that shape its magnitude (Pardo et al. 2018). One approach concentrates on analyzing speech production in interacting partners in order to observe how the degree of imitation (in these studies referred to as convergence or accommodation) is modulated by social distance, the level of interaction, or the perception of adialogue partner (Babel 2010, 2012; Babel et al. 2014; Gasiorek et al. 2015; Giles et al. 1991). Other studies look into the process of phonetic imitation using speech shadowing tasks in alaboratory setting (Goldinger 1998; Kwon 2019, 2021; Mitterer and Ernestus 2008; Mitterer and Müsseler 2013; Namy et al. 2002; Nielsen 2011; Shockley et al. 2004). In atypical paradigm of ashadowing task, the participant first reads words presented in an orthographic form in order to elicit their baseline productions. Next, they hear and repeat the same words after amodel talker (shadowing) and finally re-read the words again in apost-test phase. The comparisons between the three conditions permit an insight into the degree of imitation changes and the level of post-exposure retention. Considering that successful learning of speech sounds in asecond language requires an effective interaction of perception and production, anumber of studies have also directed their attention to the degree of imitation after exposure to second language (L2) speech. In such studies, shadowing after amodel talker in L2 throws light on how the acoustic properties of an L1 sound inventory constrain the successful OPEN ACCESS
144 LINGUISTICA PRAGENSIA 1/2022 attainment of L2 speech sounds (de Jong et al. 2009; Flege and Eeefting 1988; Hao and de Jong 2016; Jia et al. 2006; Llompart and Reinisch 2018; Podlipský and Šimáčková 2015; Rojczyk 2013; Rojczyk et al. 2013; Schouten 1977; Zając and Rojczyk 2014). The results show that speakers are able to imitate phonetic features that are absent in their L1 relatively effectively. Compared to the baseline condition, imitated L2 productions after amodel speaker tend to be more native-like. This effect has been found for voice onset time (Flege and Eefting 1988), vowel duration (Podlipský and Šimáčková 2015; Zając and Rojczyk 2014), spectral properties of vowels (Jia et al. 2006; Llompart and Reinisch 2018; Rojczyk 2013), the lack of release in stop consonants (Rojczyk et al. 2013), or tones in tone languages (Hao and de Jong 2016). The results from these studies suggest that direct imitation may temporarily bypass phonological constraints emerging from cross-linguistic differences. For example, Rojczyk (2013) tested twenty-two Polish learners of English in how they imitated the quality of the non-native trap vowel /æ/.1 This vowel is especially problematic for Polish learners because it is subsumed by two Polish neighbouring vowels, /ɛ/ and /a/. The results showed that the imitated productions dissimilated successfully from both vowels compared to abaseline reading task. In another study, Llompart and Reinisch (2018) investigated the link between imitation and perception for two non-native contrasts differing in the level of difficulty: dress–trap (difficult) and fleece–kit (easy). The analysis of the productions by German learners of English revealed that both imitation and perception were more successful for the easy contrast than for the difficult one. Moreover, the ability to imitate the dress–trap opposition was closely related to the perception of this contrast, which implies astrong impact of phonological representations on the magnitude of imitation. The authors concluded that imitation is linked with the level of attainment of non-native contrasts but does not need to reflect the learners’ productive usage of such non-native distinctions. As suggested by one of the reviewers, it is important to differentiate between explicit and implicit imitation in reviewing prior research on the role of imitation in L2. Although we agree that this is an important distinction to be made, it is sometimes difficult or even impossible to achieve since the methodological descriptions in previous studies frequently lack such information. For example, in Podlipský and Šimáčková (2015: 2) we are only informed that “in shadowing, they repeated each word right after they heard it”. Llompart and Reinisch (2018: 603) do use the term “explicit”, but not referring to the type of imitation directly: “participants… were explicitly told that they had to wait until the native speaker finished talking before imitating”. In contrast, Zając and Rojczyk (2014: 502) overtly suggested that their method relied on aprecise distinction between implicit and explicit imitation by specifying that “twenty participants took part in the first session in which target-model words were presented without specific instructions inducing imitative behaviours: the participants were only instructed to wait until the recorded voice stopped producing the word and then read this from the screen. Another twenty participants took part in the second session in which they were instructed to imitate the words they heard as 1 The words in small capitals refer to John Wells’ 24 lexical sets (representative keywords) for English vowels (Wells 1982). OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 145 faithfully as they could”. Interestingly, the authors reported that providing speakers with explicit instructions to imitate did not have asignificant effect on the magnitude of convergence with anative model talker. To our knowledge, no previous studies have investigated the magnitude of phonetic imitation of English glottal articulations by non-native speakers. In the following two sections, we focus on glottalization in English and Czech and present the hypotheses that emerge from comparing the two language systems and other factors. 1.2 GLOTTALIZATION AND T-GLOTTALING IN ENGLISH The phenomenon of t-glottaling refers to the realization of the English phoneme /t/ as aglottal stop [ʔ] rather than avoiceless alveolar plosive or similar sounds in words like cat or city. T-glottaling occurs in many British English varieties, having spread surprisingly quickly in the latter part of the twentieth century, eventually losing its negative connotations (Fabricius 2002; Hughes, Trudgill and Watts 2013). T-glottalling, at least in some phonological contexts, is attested not only in current standard and non-standard varieties of British English (Gavaldà 2016; Schleef 2021), but also in American English (Eddington and Channer, 2010; Seyfarth and Garellek 2020). Importantly, non-prevocalic environments seem to be more favoured in terms of t-glottaling than intervocalic ones (Cruttenden 2014; Fabricius 2002). The term “glottal stop” is acover term for arange of glottal gestures that can be subsumed under abroader term, “glottalization”. The phonetic realization of these glottal events varies, from glottal plosives to various lenited variants, creaky voice or laryngealization (e.g. Ashby and Przedlacka 2014; Keating, Garellek and Kreiman 2015; Redi and Shattuck-Hufnagel 2001). The canonical glottal plosives are produced by acomplete closure of the vocal folds, obstructing the airflow into supralaryngeal cavities. As aresult, the subglottal pressure increases and is subsequently released by arapid parting of the folds. Cruttenden (2014: 182) comments on the auditory impressions of the glottal plosive as “its presence being perceived auditorily by the sudden cessation of the preceding sound or by the sudden onset (often with an accompanying strong breath effort) of the following sound”. However, other glottal variants seem to be more prevalent. According to Ashby and Przedlacka (2011: 50–51), “examples of ‘glottal stops’ with silent hold phases are hardly to be found in natural speech at all. Real glottal events are in fact themselves almost invariably ‘lenited’, consisting chiefly of aperiod of disturbed vocal fold vibration […] still serving as syllable margins”.2 T-glottaling is closely related to another phenomenon affecting the pronunciation of words like cat. T-glottaling is aglottal replacement of the alveolar segment, so that no trace of the oral articulation of [t] is left. However, an alternative strategy is aglottal reinforcement of [t], which can be seen in segmental terms as a(partially overlapping) sequence of aglottal and an alveolar stop [ʔt] (Cruttenden 2014: 184). The word cat could thus be pronounced as [kæʔ] or [kæʔt]. As these strategies are in some respect equivalent, we can describe both types of pronunciation as “glottal articula2 See Docherty and Foulkes (1999) and Ashby and Przedlacka (2014) for illustrations of glottal events in varieties of British English. For glottalization in American English, see for instance Seyfarth and Garellek (2020) or Kaźmierski (2020). OPEN ACCESS
146 LINGUISTICA PRAGENSIA 1/2022 tions”. Although glottal articulations of /t/ are acommon feature of British English, they are not represented in pronunciation dictionaries (Sturiale 2012), which typically contain only phonemic transcription. The phonological contexts where t-glottaling occurs are summarized under (1). In one environment, the /t/ phoneme is word-final (Vt#, right) or followed by aconsonant (VtC, football). Both represent asyllable-final, non-prevocalic position and thus can be treated together. The contexts (1b) are intervocalic (VtV), either within aword (city) or across aword boundary (alot of). They can be considered syllable-final only if we accept Wells’ syllabification, which assigns the segment to the syllable coda (Wells 1990). The location of stress is not relevant to the present study, as the /t/ can follow the stressed vowel immediately or at adistance (senator). Finally, the context of (1c) is aspecial case, as it is often not clear whether the following segment is agenuine syllabic consonant (=1a), or aschwa nucleus intervenes (=1b). However, contexts (1c) are not the focus of our study. (1) Examples of glottalization in English /t/-glottaling (a) Non-prevocalic contexts eat [ˈiːʔ], right [ˈraɪʔ], football [ˈfʊʔbɔːl], get down [ɡeʔˈdaʊn] (b) Intervocalic contexts city [ˈsɪʔi], letter [ˈleʔə], senator [ˈsenəʔə], alot of [ˈlɒʔəv] (c) Preceding asyllabic consonant bottle [ˈbɒʔl], button [ˈbʌʔn] Glottalization in English is also connected to the phenomenon of linking (liaison) in connected speech. Typically, there is asmooth transition between the word-final and word-initial segment, and no break is perceived between the words. Consonantto-vowel linking would be the norm in an hour [ən‿aʊə], vowel-to-vowel linking in two hours [tuː‿aʊz]. Transient glides are used after high vowels, whereas liaison /r/3 is used after non-high vowels (Cruttenden 2014: 315–317). However, initial vowels can be pronounced with glottalization when the prosodic context or the pragmatic situation necessitates it, as in emphasis: Ihaven’t seen [ʔ]ANYBODY. She’s [ʔ]AWFULLY good. Finally, the law [ʔ]ACTED (cf. Cruttenden 2014: 183). This can occur also within words, where glottalization may function as amorpheme boundary marker. The morpheme starts with avowel, following avowel or consonant, as in reaction [riːˈʔækʃən], co-operate [kəʊˈʔɒpəreɪt] or post-empiricism [pəʊstʔemˈpɪrɪsɪzm ]. Unlike in the examples under (1), the glottal stop here is not arealization of aparticular segment, but ameans of vowel hiatus resolution when the elements are not linked. 3 In their study of the speech of BBC newsreaders, Mompeán and Gómez (2011) report that laryngeal gestures were the most common hiatus breaking strategy in the potential r-liaison sites where no rhotic was used, with the prevailing realization being creaky voice and the canonical stops only occurring in asmall minority of cases. OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 147 1.3 GLOTTALIZATION IN CZECH In the Czech language, glottalization has ademarcative function. The glottal plosive or its lenited variants subsumed under glottalization (see previous section) appear at some lexical and morphemic boundaries, cuing the beginning of avowel-initial word or amorpheme (prefix or stem, but not suffix). In effect, the glottal gesture is amarker of vowel-initial “phonological words”, as exemplified under (2). For instance, mimoevropský ‘non-European’ consists of two phonological words: mimo ‘outside (of)’ and evropský ‘European’. Although words in (2a) are all stressed on the first syllable, words in (2b) demonstrate that the presence of glottal articulations does not depend on stress location. Example (2c) represents aspecial category, as the words are preceded by non-syllabic prepositions that form one phonetic syllable with the first syllable of the word itself. It is the only context in which the usage of glottalization is mandatory in standard Czech pronunciation (Hála, 1967; Palková, 1994). The pronunciation [ˈkaktsɪ] or [ˈɡaktsɪ] instead of [ˈkʔaktsɪ] is thus anon-standard variant. (2) Examples of glottalization in Czech (a) Word-initial contexts akce [ˈʔaktsɛ] ‘action’, Evropa [ˈʔɛvropa] ‘Europe’, obsah [ˈʔopsax] ‘content’, útok [ˈʔuːtok] ‘attack (noun)’ (b) Word-medial, prefixor stem-initial contexts protiakce [ˈprocɪʔaktsɛ] ‘counteraction’, mimoevropský [ˈmɪmoʔɛvropskiː] ‘nonEuropean’, bezobsažný [ˈbɛsʔopsaʒniː] ‘contentless, without content’, zaútočit [ˈzaʔuːtotʃɪt] ‘attack (verb)’ (c) Contexts with non-syllabic prepositions kakci [ˈkʔaktsɪ] ‘to action’, sEvropou [ˈsʔɛvropou] ‘with Europe’, vokně [ˈfʔokɲɛ] ‘in (the) window’ The contexts (2a) and (2b) provide achoice in standard Czech between apronunciation with or without glottalization. The likelihood of glottalization is affected strongly by the need for speech clarity, as the presence of glottalization cues the wordor morpheme-initial parses of the stream of speech. However, only 12% of Czech words begin with avowel (Šturm and Bičan 2021), so this function should not be overrated; on the other hand, many of these words have high frequency, for instance a‘and’, ale ‘but’, aby ‘so that’, už ‘yet, already’, or on ‘he’. The rate of occurrence of glottalization in Czech was studied by Volín (2012). He compared aformal speaking style (newsreading on the Czech radio) with spontaneous conversations. Overall, glottalization was present in 95% of potential contexts in the former, whereas it was only 65% of potential contexts in the latter. Women glottalized more often than men regardless of the style. Importantly, although glottalization in the sense above is anatural part of Czech utterances, t-glottaling does not occur in Czech. OPEN ACCESS
148 LINGUISTICA PRAGENSIA 1/2022 1.4 RESEARCH QUESTIONS AND HYPOTHESES There are several research questions considered in this study, reflected in the hypotheses under (3). We must distinguish two types of items— shadowed and nonshadowed— according to whether or not they are presented auditorily during the exposure phase. If shadowing triggers learning of the glottal gesture, the post-test reading task should show ahigher rate of glottalization not only in the shadowed items (3a), but also in non-shadowed items (3b). In other words, the expectation is that the process will be applied productively to words of similar structure, yet not previously heard. These are either from the same category (e.g., meet, corresponding to shadowed feet) or from anew but phonetically and distributionally similar category (/k/ in week). Furthermore, the strength of adaptation after exposure should depend on the type of reaction that is required in the task and on memory. According to the hypothesis in (3c), immediate explicit phonetic repetition will elicit higher rates of glottalization than the delayed post-test reading task (performed when the auditory memory of the stimulus has already faded away). The two tasks are otherwise comparable as they both include orthographic intervention on the screen. If the data do not support hypotheses (3a) and (3b) for the post-test performance but do support (3c) for the shadowing, it might mean that there is no actual learning involved, only explicit imitation without an attempt at learning the glottal articulations. The remaining two hypotheses are related to the comparison of the interacting languages. Hypothesis (3d) predicts that Czech speakers of English will adapt their speech towards glottalization more readily in the intervocalic rather than in the non-prevocalic contexts. Where Czech speakers might expect a[t] or [ɾ] (as in city), native speakers of some English varieties produce t-glottaling instead ([ˈsɪʔɪ], [ˈbeʔ], [ˈfʊʔbɔːl]). Importantly, only the first of these— intervocalic [ʔ]— has acorresponding segmental structure in the Czech language (e.g., [ˈnaʔopak]), whereas the other two contexts— pre-pausal and pre-consonantal— are unfamiliar in Czech. We argue that it is the familiar structure that should be more easily shadowed and retained. Alternatively, however, the non-prevocalic context might in fact prove easier to imitate, given the frequency of t-glottaling in various word positions in the learners’ input. As Fabricius (2002) showed, t-glottaling is aconsistent feature of Standard Southern British English only in pre-consonantal environments (and less consistently in pre-pausal environments); furthermore, her data from pre-vocalic environments suggest that the intervocalic position is prone to resist t-glottaling in this variety. As aresult, the Czech speakers would be more ready to accept VʔC or Vʔ# forms that they encounter quite frequently, as compared to the relatively rare VʔV forms that might thus be less familiar. The experiment should resolve which of the two aspects— language internal or language external— plays amore crucial role. (3) Hypotheses regarding the effect of phonetic shadowing (a) Speakers will glottalize more in both post-exposure tasks than in the baseline task. (Shadowing task triggers learning.) OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 149 (b) Non-shadowed items with similar characteristics will also show ahigher rate of glottalization in the post-test as compared to the baseline. (Shadowing task triggers learning and the process is generalized based on analogy.) (c) Speakers will glottalize more in the immediate repetition task than in the delayed post-exposure task. (Imitation is present, learning not necessarily or to alower degree.) (d) Speakers will glottalize more in VtV contexts than in VtC or Vt# contexts. (Language structure is akey factor.) (e) Speakers will glottalize less in VtV contexts than in VtC or Vt# contexts. (Language input frequency is akey factor.) 2 METHOD 2.1 MATERIAL The material includes several types of items. Some appeared in all three tasks (shadowed items, 4a-c), others only in the baseline and post-test tasks (non-shadowed items, 4d-e). In order to facilitate fluency in reading, difficult or less familiar words were not used. Crucially, the /t/ segments appeared either in an intervocalic (VtV) position, or in anon-prevocalic (VtC or Vt#) position, where # marks aword-final position before apause. There were 16 items in each category. Filler items comprised 16 additional words that did not involve a/t/ or /k/ segment at all, with the aim of concealing the object of investigation to some degree. Finally, the material contained 18 more words to test whether participants generalize beyond the auditorily presented words. Nine had a/t/ segment in the same positions, nine a/k/ segment in the non-prevocalic position (note that in fact, picture, and practice only the velar plosive is expected to be glottalized). These non-shadowed tokens will be referred to as from the same or different category, respectively. (4) Experimental material (a) VtV: beautiful, better, butter, cutting, daughter, eat up, energetic, forty, getting, hotter, it all, lot of, patriotic, relative, sort of, water. (b) VtC or Vt#: alot, bat, bit, cat, cut, eight, feet, fit, football, hot, hot weather, sit down, start, straight, what, white. (c) Filler items: ago, always, blue, brother, deal, eyebrows, floor, girls, home, children, money, mouse, nothing, roof, shoes, walls. (d) Non-shadowed tokens of /t/ (same category): city, later, letter, putting, alot more, meet, nightlife, nut, rot. (e) Non-shadowed tokens of /k/ (different category): background, book, fact, joke, picture, practice, sack, shock, week. Instead of selecting and recording asingle speaker, we opted for amulti-speaker approach that increases the variability of voices and idiolects. The recordings of words with the target segments were obtained from two sources. First, most of the reOPEN ACCESS
150 LINGUISTICA PRAGENSIA 1/2022 cordings come from acorpus of Southern British English collected for an earlier sociophonetic study (Przedlacka 2002), where realizations of /t/ intervocalically and non-prevocalically was one of the phonetic variables. The words or short phrases were responses to aspoken lexical questionnaire (e.g. What happens to water in 100° C?) in one-to-one interviews in the subjects’ schools (aged 14 to 16). The questionnaire was preceded by an informal chat about their interests and plans. The speakers appeared relaxed during the questionnaire and often made informal asides. Those factors indicate that the data represents arelatively casual speaking style in spite of the short, often single-word utterances (note especially that it was not aword-list reading task). The original analogue field recordings were digitized at asampling rate of 16 kHz. The accent of the informants was Estuary English and Cockney. Phonologically, their pronunciation was fairly close to Standard Southern British English (the accent variety described in standard textbooks on British English pronunciation). Avariety of /t/ tokens were produced. For the present experiments, words with t-glottaling, i.e., with glottal (stop, creak or other) realizations of /t/ (all transcribed as [ʔ]) were purposely selected. That is, glottally reinforced voiceless alveolar stops ([ʔt]) were excluded, so that the type of material could be restricted to asingle category. The second source of recordings are YouTube videos. Since the number of tokens with glottal replacement in the corpus mentioned above was not sufficient or included the target segments in other contexts, several more speakers and words were added. Care was taken to make the stimuli as similar to the previous recordings as possible. Namely, single-word utterances were selected. The speakers’ age or background was often unknown, but they appeared in their twenties or thirties. The quality of the recordings was comparable; we also converted the recordings to mono files with asampling rate of 16 kHz to fit in with the previous recordings. All stimuli were normalized to an RMS of 70 dB in Praat (Boersma and Weenink 2021) and saved as wav files, with an additional short silence at the beginning and end of the stimuli. 2.2 PARTICIPANTS In total, 30 participants were recorded for this study (15 male and 15 female). Their age ranged from 19 to 39 years, with amean of 25.3 (SD = 5.0). They were all native speakers of Czech without any reported hearing or speaking disorders. They were compensated financially for their time. Most of the participants studied at the Faculty of Arts, Charles University, Prague, aminority worked at the university library or had no connection to the university. In effect, there were three groups of participants: (i) students of English studies, i.e., expert users of the English language familiar with accents variation (2 participants), (ii) students of phonetics as afull programme (not an introductory class), i.e., expert listeners (5 participants), and (iii) other participants (n = 23). Given their expert knowledge, the seven participants from groups (i) and (ii) will therefore be analyzed separately. In addition, an attempt was made to take into account the approximate level of the speakers’ English. The participants were asked for their CEFR level (e.g., B1, C2 etc.) and the onset of learning English. There was one participant with level A(length of learning OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 157 shadowed /k/ (different category).6 However, the differences between the categories were not substantial; rather, they indicate interesting tendencies. Acomparison can also be made between the pre-test and post-test results. For all types of items, there was aminor increase of glottalized tokens in the post-test task. Task Set Glottalization Percent glottalized No Yes Pre-test Shadowed 941 19 2% [1.2%– 3.1%] Non-shadowed: /t/ 265 5 2% [0.6%– 4.3%] Non-shadowed: /k/ 268 2 1% [0.1%– 2.7%] Shadowing Shadowed 652 308 32% [29.1%– 35.1%] Post-test Shadowed 881 79 8% [6.6%– 10.2%] Non-shadowed: /t/ 257 13 5% [2.6%– 8.1%] Non-shadowed: /k/ 264 6 2% [0.8%– 4.8%] Table 2:The number of glottalized and non-glottalized tokens according to task and set of items. The 95% confidence intervals were computed from abinomial test. To evaluate these results, alogistic model was fitted to the preand post-test data. The effect of task was significant (χ2(1) = 7.89, p = 0.005), but none of the other fixed effects reached significance, including set (χ2(2) = 4.45, p = 0.108). The interaction between task and set was not significant either (χ2(2) = 1.59, p = 0.452), suggesting that there is ahigher probability of glottalization in the post-test phase regardless of the type of items. Therefore, when individual variation and the other control variables are considered in the model, the interesting differences apparent in Table 2 are not supported in ageneralized model. 3.4 REALIZATION OF GLOTTALIZATION It would be useful to differentiate between glottal replacement and glottal reinforcement in the data (it should be noted that the audio stimuli contained only glottal replacement). Table 3 shows that reinforcement occurred mainly in the non-prevocalic position, and that it was almost absent in the shadowing task. The speakers thus imitated the t-glottaling accurately, i.e., as aglottal gesture and not as acombination of aglottal and oral gesture. In the post-test phase, the contribution of reinforcement to the glottalized tokens was quite substantial. Moreover, Table 3 reveals that alveolar flaps [ɾ] were produced— quite expectedly— in the intervocalic position, unlike plosives with no audible release or deleted segments, which were associated mostly with the non-prevocalic position. 6 For /k/: words sack (3×), shock, joke, background, produced by three different speakers (M01, F03, F06). For /t/: words rot (5×), nut (2×), meet, alot more (2×), nightlife, letter (2×), produced by eight different speakers (F02, F03, F05, F06, F08, M10, M11, M13). Furthermore, there was another item with aword-final plosive, [p] (eat up), outside our analysis; it was replaced by aglottal stop once (M07, non-expert). OPEN ACCESS
158 LINGUISTICA PRAGENSIA 1/2022 Glottalized Non-glottalized Task Position ʔ ʔtt d ɾtd ø Pre-test Intervocalic 0 0 468 25 107 0 0 0 Non-prevocalic 5 19 543 27 0 18 16 2 Shadowing Intervocalic 168 1 176 6 77 0 0 52 Non-prevocalic 134 5 141 12 0 36 27 125 Post-test Intervocalic 19 3 403 9 155 1 0 10 Non-prevocalic 39 33 421 20 0 27 25 70 Table 3.The number of tokens from each category depending on task and position. 4 DISCUSSION 4.1 IMITATION AND LEARNING This study examined the process of explicit phonetic imitation of English words with t-glottaling by Czech participants. In the set of hypotheses assembled under (3), the first three concern the form and degree of learning, i.e., what an increase in glottal articulations of /t/ would reflect. Hypothesis (3a) was confirmed, as there was an increase in post-test glottalization as aresponse to the shadowing task. Hence, some degree of learning must have taken place, although it is open to question what its nature and extent is (see below) and whether it is ashort-term or along-term effect. The small size of the effect hinders any serious interpretation, apart from saying that long-term learning is probably not triggered, and the effect would disappear within several days. Importantly, there were higher rates of glottalization in the post-test not only in the shadowed items, but also in the non-shadowed items (3b). These are novel items in the sense that they have not been presented as sound stimuli to the participants. The Czech participants are expected to productively apply t-glottaling to words that are in some respect analogous to the shadowed items. Presumably, position of the target segment is acrucial conditioning of the process (e.g. /t/ in talk would never be glottalized). Since there were no instances of glottal /k/ or /p/ in the shadowing stimuli, two scenarios may follow. On the one hand, the glottal gesture learned from the shadowed items may be extended to all phonetically similar segments (voiceless plosives) in the relevant positions (intervocalic or non-prevocalic). In that case, both types of non-shadowed items (/t/ and /k/) should show increased glottalizations in the post-test, and to asimilar degree. On the other hand, if the basis for analogy is category membership, we would predict glottalization of the non-shadowed /t/ items to the exclusion of the /k/ items. Adifference between the two phonemes should ensue. The participants were able to generalize the imitation of t-glottaling to novel items of both types, although we must keep in mind that the number of glottalized tokens in the post-test was generally low. Statistically speaking, the interaction of task with item type was not significant, leading to apost-test increase in glottalization for all types of items. This seems to point to the latter conclusion: generalization OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 159 to [k] in week stems from the fact that [k] occupies the same position as in feet (and the similarity of [k] to [t] phonetically/phonologically is just aprerequisite). However, the patterns were more nuanced. The three sets of items yielded agradual decrease in the differences betwee the pre-test to post-test results: the departure from the control condition was highest for shadowed items (change from 2% to 8%), lower for non-shadowed items from the same category (=/t/, from 2% to 5%), and lowest for non-shadowed items from the new category (= /k/, from 1% to 2%). This suggests that (i)there are in fact some differences between /t/ and /k/, and (ii) the generalization ability is not so robust.7 Proceeding to the hypothesis in (3c), acomparison of the two post-exposure tasks again suggests that proper learning does not occur. Speakers glottalized much more extensively after immediate exposure in the shadowing task (32% of tokens glottalized) than in the delayed post-test reading task (8% glottalized). Taken together, the findings so far indicate that participants are quite responsive to glottal articulations in immediate phonetic shadowing, but do not retain these productions in adelayed task. In fact, the process at play seems to be imitation— which was explicitly called for in the instructions— rather than learning. Nevertheless, we cannot preclude the possibility of acombination with learning as there was at least some increase from pre-test to post-test.8 Moreover, if elided /t/sor plosives with no audible release (Table 3) are counted as (imperfect) imitation (an intermediate category between [t] and [ʔ]), the imitative behaviour of our speakers might actually be better than what we reported based on glottalized items only, and the effect of exposure to t-glottaling would stand out more. 4.2 LANGUAGE INTERNAL VS. EXTERNAL FACTORS Hypotheses (3d) and (3e) involved acomparison between two phonological environments. Structural differences between English and Czech led us to the prediction that imitating words with the glottal stop [ʔ] in intervocalic position (VtV) would be an easier task for the participants than imitating words with the target in non-prevocalic position (VtC or Vt#). The rationale was the presence of glottal stops in Czech intervocalically, but not before aconsonant or apause. Alternatively, ahigher frequency of t-glottaling in VtC/Vt# than in VtV in the English input, especially in standard varieties, would predict the opposite pattern of results. Facilitation of t-glottal7 For instance, Nielsen (2011) examined the amount of aspiration after exposing her participants to lengthened VOT tokens of /p/. The imitated production was generalized to new instances of the target phoneme /p/, and also to the novel phoneme /k/, albeit to alesser degree. Nevertheless, the generalization effect was much stronger in her experiment than in ours. The reason may of course relate to the different target feature (t-glottaling vs. VOT), or to adifference in method (also, see the following footnote). 8 The low retention of glottal articulations could also be linked to the following explanation. Three words in the material included afinal plosive that was not glottalized in the stimuli, which runs counter to the expectations based on /t/. The distracting presence of nonglottalized word-final [p] (eat up) and [k] (patriotic, energetic) could have lowered post-test generalization to /k/. OPEN ACCESS
160 LINGUISTICA PRAGENSIA 1/2022 ing in one of the two positions was expected for both tasks (immediate shadowing and the post-test reading). In the baseline condition— before exposure— it is easy to explain why there were no glottal tokens intervocalically and asomewhat higher incidence of glottalization word-finally or before aconsonant. Czech learners of English are likely to be targeting either Standard Southern British English or General American pronunciation (Jakšič and Šturm 2017). In the former, glottal reinforcement or replacement is widespread in the non-prevocalic position, whereas t-glottaling in the intervocalic position is acharacteristic of non-standard varieties and might be perceived as stigmatizing. In the latter, intervocalic /t/sare usually realized as alveolar flaps (this was indeed the pronunciation of many of our speakers). So Czech learners have no impetus to glottalize VtV in their own speech, and some might glottalize non-prevocalic /t/. In the shadowing task, there was indeed alarger number of glottalized tokens in the intervocalic condition (35% vs. 29%), but the difference was not significant. The predicted effect appeared only when the participants were split into experts, who produced significantly higher rates of glottalization in the VtV position, and nonexperts (without such adifference). This could not be extended simply to the English proficiency level, since advanced and intermediate speakers behaved in the two positions in asimilar way. The hypothesis (3d) was thus supported only for some participants. What seems to be the case considering all the data is that the higher rate of glottalization in the intervocalic position reflects aconscious copying of the glottal realization of /t/, which is especially salient intervocalically. Naturally, the expert participants are likely to be attentive to such details of articulation as t-glottaling in following the instruction to imitate the recordings “as closely as possible”. Moreover, in the post-test task, we found evidence of the opposite direction of the effect, namely, the non-prevocalic position being associated with higher rates of glottalization (12%) compared to VtV (4%). This supports hypothesis (3e). As in the pre-test, Czech learners do not have the motivation to glottalize intervocalic contexts due to the absence of t-glottaling in standard British and American accents, asort of anti-t-glottaling VtV constraint. Without the immediacy of phonetic shadowing, the minimal learning that might have occurred as aresponse to the exposure of tglottaling is thus reflected mainly in the non-prevocalic position. 4.3 PARTICIPANT VARIABLES Biological sex of the participants was not asignificant predictor. It is well known that women usually glottalize more than men in various speech corpora (e.g., Seyfarth and Garellek 2020; Volín, 2012), but it is not clear whether we should also predict adifference in terms of the imitation ability. More important participant aspects in our study were those already mentioned: the proficiency level in English and expertness. Since their effects were similar (for more proficient participants, the shadowing intervention led to somewhat higher rates of glottalization both immediately and with delay), one might expect astrong correlation between the two variables. However, avariance inflation factor (VIF) analysis did not reveal any substantial multicollinearity in the predictors. Moreover, there were 13 advanced participants, but only 7 expert participants (with substantial overlap). It would be necessary to examOPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 161 ine abalanced group of participants in afactorial design to evaluate the contribution of these effects reliably. Finally, afuture experiment should provide exclusively auditory stimuli (without orthography) during the shadowing task, as especially the non-expert participants may have been reluctant to replace the canonical alveolar [t] with aglottal stop when presented with the words in their orthographic form on the screen. There is unfortunately no space to discuss individual variation, although it is clear that various dispositions of the speaker or other speaker variables can affect the degree and conditions of phonetic imitation. Different speakers showed different patterns in terms of (i) the general rate of glottalization, (ii) the strength and direction of the position effect, (iii) the strength of retention of the imitated glottal gestures in the post-test, (iv) the particulars of the articulated target sounds (see Table 3). To illustrate some of these, Figure 5 is presented below with the data from the shadowing task. For instance, speaker M14 produced all glottalized tokens in the intervocalic position, whereas M01 or F11 in the non-prevocalic position. One of the positions could be favored by some speakers, or both positions could be perceived as more or less equivalent by others. Figure 5:The effect of position in the shadowing task for individual participants. OPEN ACCESS
162 LINGUISTICA PRAGENSIA 1/2022 5 CONCLUSIONS This paper has focused on the imitation and learning of glottal realizations of English /t/ by native speakers of Czech. The glottal variant, awell-established feature of both standard and non-standard varieties of British English, is rarely explicitly taught to users of English as L2. We have provided evidence that imitation is facilitated if the target sound appears in the same position (VtV) in the imitators’ target and source language. However, the structural similarities do not positively impact learning, as the successfully imitated native productions tend not to be retained (given the current level of experimental exposure). Other factors seem to play arole in long-term retention, such as previous exposure to glottal allophones of /t/ in classroom and informal learning. Furthermore, expert listening skills and expert knowledge of L2 seem to facilitate imitation. Acknowledgements Research supported by the National Science Centre Poland grant Phonetic imitation in anative and non-native language (UMO-2019/35/B/HS2/02767) and by the European Regional Development FundProject “Creativity and Adaptability as Conditions of the Success of Europe in an Interrelated World” (No. CZ.02.1.01/0.0/0.0/16_019/0000734). REFERENCES Ashby, M. and J.Przedlacka (2011) The stops that aren’t. English Phonetics 14–15. Festschrift commemorating the retirement and the 70th birthday of Professor John C.Wells. The English Phonetic Society of Japan. Ashby, M. and J.Przedlacka (2014) Measuring incompleteness: Acoustic correlates of glottal articulations. Journal of the International Phonetic Association 44, 283–296. Babel, M. (2010) Dialect convergence and divergence in New Zealand English. Language in Society 39, 437–456. Babel, M. (2012) Evidence for phonetic and social selectivity in spontaneous phonetic imitation. Journal of Phonetics 40, 177–189. Babel, M., G.McGuire, S.Walters and A.Nicholls (2014) Novelty and social preference in phonetic accommodation. Laboratory Phonology 5, 123–150. Bates, D., M.Mächler, B.Bolker and S.Walker (2015) Fitting linear mixed-effects models using lme4. Journal of Statistical Software 67, 1–48. Boersma, P. and D.Weenink (2021) Praat: Doing phonetics by computer (version 6.1.42). Retrieved from http://www.praat.org/. Chang, C.B. (2012) Rapid and multifaceted effects of second-language learning on first-language speech production. Journal of Phonetics 40/2, 249–268. Cruttenden, A. (2014) Gimson’s Pronunciation of English, 8th ed. London: Routledge. de Jong, K., Y.C.Hao and H.Park (2009) Evidence for featural units in the acquisition of speech production skills: Linguistic structure in foreign accent. Journal of Phonetics 37/4, 357–373. Docherty, G. and P.Foulkes (1999) Derby and Newcastle: Instrumental phonetics and variationist studies. In: Foulkes, P. and G.Docherty (eds) Urban voices, 47–71. London: Arnold. Eddington, D. and C.Channer (2010) American English has go? alo? of glottal stops: Social diffusion and linguistic motivation. American Speech 85, 338–351. OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 163 Fabricius, A. (2002) Ongoing change in modern RP: Evidence for the disappearing stigma of t-glottalling. English World-Wide 23/1, 115–136. Flege, J.E. and W.Eefting (1988) Imitation of aVOT continuum by native speakers of English and Spanish: Evidence for phonetic category formation. Journal of the Acoustical Society of America 83/2, 729–740. Forster, K. I. and J.C.Forster (2003) DMDX: AWindows display program with millisecond accuracy. Behavior Research Methods, Instruments, & Computers 35, 116–124. Gasiorek, J., H.Giles and J.Soliz (2015) Accommodating new vistas. Language and Communication 41, 1–5. Gavaldà, N. (2016) Individual variation in allophonic processes of /t/ in Standard Southern British English. The International Journal of Speech, Language and the Law 23/1, 43–69. Giles, H., J.Coupland and N.Coupland (1991) Contexts of Accommodation: Developments in Applied Sociolinguistics. Cambridge: Cambridge University Press. Goldinger, S.D. (1998) Echoes of echoes? An episodic theory of lexical access. Psychological Review 105/2, 251–279. Hála, B. (1967) Výslovnost češtiny I(výslovnost slov českých) [The pronunciation of Czech I(pronunciation of Czech words)]. Prague: Academia. Hao, Y.C. and K. de Jong (2016) Imitation of second language sounds in relation to L2 perception and production. Journal of Phonetics 54, 151–168. Hughes, A., P.Trudgill and D.Watt (2013) English Accents & Dialects: An Introduction to Social and Regional Varieties of English in the British Isles. London: Routledge. Jaeger, T.F. (2008) Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. Journal of Memory and Language 59, 434–46. Jakšič, J. and P.Šturm (2017) Accents of English at Czech schools: Students’ attitudes and recognition skills. Research in Language 15, 353–369. Jia, G., W.Strange, Y.Wu, J.Collado and Q.Guan (2006) Perception and production of English vowels by Mandarin speakers: Age-related differences vary with amount of L2 exposure. Journal of the Acoustical Society of America 119/2, 1118–1130. Kaźmierski, K. (2020) Prevocalic t-glottalling across word boundaries in Midland American English. Laboratory Phonology: Journal of the Association for Laboratory Phonology 11/1, Article 13. Keating P., M.Garellek and J.Kreiman (2015) Acoustic properties of different kinds of creaky voice. In: Proceedings of the 18th International Congress of Phonetic Sciences, paper 821. Glasgow: University of Glasgow. Kuhl, P.K. and A.N.Meltzoff (1996) Infant vocalizations in response to speech: Vocal imitation and developmental change. Journal of the Acoustical Society of America 100, 2425–2438. Kwon, H. (2019) The role of native phonology in spontaneous imitation: Evidence from Seoul Korean. Laboratory Phonology: Journal of the Association for Laboratory Phonology 10/1, Article 10. Kwon, H. (2021) Anon-contrastive cue in spontaneous imitation: Comparing monoand bilingual imitators. Journal of Phonetics 88, 101083. Labov, W. (2001) Principles of Linguistic Change, Volume 2: Social Factors. Oxford: Blackwell. Lenth, R. (2020) emmeans: Estimated Marginal Means, aka Least-Squares Means. (1.5.2–1) [Computer software]. Available at: https://CRAN.R-project.org/ package=emmeans. Lin, Y., Y.Yao and J.Luo (2021) Phonetic accommodation of tone: Reversing atone merger-in-progress via imitation. Journal of Phonetics 87, 101060. Llompart, M. and E.Reinisch (2018) Imitation in asecond language relies on phonological categories but does not reflect the productive usage of difficult sound contrasts. Language and Speech 62/3, 594–622. Mitterer, H. and M.Ernestus (2008) The link between speech perception and production is OPEN ACCESS
164 LINGUISTICA PRAGENSIA 1/2022 phonological and abstract: Evidence from the shadowing task. Cognition 109/1, 168–173. Mitterer, H. and J.Müsseler (2013) Regional accent variation in the shadowing task: Evidence for aloose perception-action coupling in speech. Attention, Perception & Psychophysics 75, 557–575. Mompeán, J.A. and F.A.Gómez (2011) Hiatus resolution strategies in non-rhotic English: The case of /r/-liaison. In: Proceedings of the 17th International Congress of Phonetic Sciences, 1414–1417. Namy, L.L., L.C.Nygaard and D.Sauerteig (2002) Gender differences in vocal accommodation: The role of perception. Journal of Language and Social Psychology 21, 422–432. Nielsen, K. (2011) Specificity and abstractness of VOT imitation. Journal of Phonetics 39, 132–140. Palková, Z. (1994) Fonetika afonologie češtiny. Praha: Karolinum. Pardo, J.S., A.Urmanche, S.Wilman, J.Wiener, N.Mason, K.Francis and M.Ward (2018) Acomparison of phonetic convergence in conversational interaction and speech shadowing. Journal of Phonetics 69, 1–11. Podlipský, V.J. and Š.Šimáčková (2015) Phonetic imitation is not conditioned by preservation of phonological contrast but by perceptual salience. In: Proceedings of the 18th International Congress of Phonetic Sciences, Glasgow. Przedlacka, J. (2002) Estuary English? ASociophonetic Study of Teenage Speech in the Home Counties. Frankfurt: Peter Lang. R Core Team. (2020) R: Alanguage and environment for statistical computing: Version 4.0.3 [software]. Vienna: R Foundation for Statistical Computing. Available at: https://www.r-project.org. Redi, L. and S.Shattuck-Hufnagel (2001) Variation in the realization of glottalization in normal speakers. Journal of Phonetics, 29/4, 407–429. Rojczyk, A. (2013) Phonetic imitation of L2 vowels in arapid shadowing task. In: Levis, J. and K.LeVelle (eds) Proceedings of the 4th Pronunciation in Second Language Learning and Teaching Conference, 66–76. Ames IA: Iowa State University. Rojczyk, A., A.Porzuczek and M.Bergier (2013) Immediate and distracted imitation in secondlanguage speech: Unreleased plosives in English. Research in Language 11, 3–18. Sancier, M.L. and C.A.Fowler (1997) Gestural drift in bilingual speaker of Brazilian Portuguese and English. Journal of Phonetics 25, 421–436. Schleef, E. (2021) Individual differences in intra-speaker variation: T-glottalling in England and Scotland. Linguistics Vanguard 7(s2), 20200033. Schouten, M.E.H. (1977) Imitation of synthetic vowels by bilinguals. Journal of Phonetics 5, 273–283. Seyfarth, S. and M.Garellek, M. (2020) Physical and phonological causes of coda /t/ glottalization in the mainstream American English of central Ohio. Laboratory Phonology: Journal of the Association for Laboratory Phonology 11/1, Article 24. Shockley, K., L.Sabadini and C.A.Fowler (2004) Imitation in shadowing words. Perception and Psychophysics 66, 422–429. Siegel, J. (2010) Second Dialect Acquisition. Cambridge: Cambridge University Press. Skarnitzl, R. and D.Rálišová (2022) Phonetic variation of Irish English /t/ in the syllabic coda. Journal of the International Phonetic Association, online first version. DOI: https:// doi.org/10.1017/S0025100321000347. Sturiale, M. (2012) No bot’le no party: T-glottalling and pronouncing dictionaries. Language and History 55/1, 63–74. Šturm, P. and A.Bičan (2021) Slabika ajejí hranice včeštině [The Syllable and its Boundaries in Czech]. Praha: Karolinum. Volín, J. (2012) Jak se vČechách “rázuje” [How the glottal stop is realized in Czechia]. Naše řeč 95, 51–54. Wells, J. (1982) Accents of English. Volume 1: An Introduction. Cambridge: Cambridge University Press. Wells, J. (1990) Syllabification and allophony. In: Ramsaran, S. (ed) Studies in the Pronunciation of English: ACommemorative OPEN ACCESS
PAVEL ŠTURM — JOANNA PRZEDLACKA — ARKADIUSZ ROJCZYK 165 Volume in Honour of A.C.Gimson, 77–86. London: Routledge. Wickham H. (2009) Ggplot2: Elegant Graphics for Data Analysis. New York: Springer. Winter, B. (2020) Statistics for Linguists: An Introduction Using R.London: Routledge. Zając, M. and A.Rojczyk (2014) Imitation of English vowel duration upon exposure to native and non-native speech. Poznan Studies in Contemporary Linguistics 50/4, 495–514. Pavel Šturm Institute of Phonetics Faculty of Arts, Charles University Address: nám. Jana Palacha 2, 116 38, Prague, Czech Republic ORCID ID: 0000-0001-5521-029X [email protected] Joanna Przedlacka Phonetics Laboratory Faculty of Linguistics, Philology and Phonetics, University of Oxford Address: Clarendon Institute, Walton Street, OX1 2HG, Oxford, UK ORCID ID: 0000-0001-5016-0043 [email protected] Arkadiusz Rojczyk Institute of Linguistics University of Silesia in Katowice ul. Grota-Roweckiego 5, 41-205 Sosnowiec, Poland ORCID ID: 0000-0002-7328-5911 [email protected] OPEN ACCESS