Language is embiggened by words that don’t exist: the case of acircumfix Aleš Klégr (Praha) The Simpsons’ Springfield town motto “Anoble spirit embiggens1 the smallest man.” ABSTRACT The paper deals with parasynthetic formations combining the prefix enand the suffix -en, which are sometimes regarded as an example of acircumfix in English. The aim is to find more instances of this pattern than the usual three or four mentioned in the literature (enlighten, embolden, enliven, and embiggen). After searching three corpora of several billion words without much success, an experiment was made to search the Web for hypothetical verb tokens constructed from monosyllabic adjectives on the pattern provided by the four initial verbs. The search confirmed that more than ahundred such verbs occur on the Web. The discovery of so many en-Adj-en verbs unacknowledged in standard reference books is attributed to the effect of big data on the Web; it is assumed that the en-Adj-en pattern is the type of process whose function is primarily pragmatic, occasion-specific and discourse-oriented, rather than lexical (i.e. concept labelling). As aresult, although the pattern is available for active use, these formations, after having served their purpose, rarely get beyond the nonce-word stage, let alone enter the lexicon. KEYWORDS parasynthesis, circumfix, en-Adj-en pattern, conjectural forms, Web search experiment 1. PRELIMINARIES The first part of the paper’s title quotes the headline of Bauer’s (2008) brief note on the Web. He addresses the question of the status of words we may have seen or heard but which are not listed in the standard reference books, and uses the word embiggen as an example. “What we probably want to say in such cases,” he concludes, “is that there is such aword, but that it is not found in the kind of English that many of us happen to speak or write, just as “morphosyntactic” is probably not in the kind of English most of the readers of this column will speak or write. At least we can find some traces of ‘embiggen’.” This seems to imply that the ‘existence’ of aword is proportional to the extent of its use. As the footnote below explains, the word embiggen has been around for quite some time. It is technically described as being formed by parasynthesis: aprefix-suffix combination. The em-en added to the adjectival base big 1 The word embiggen was coined by Dan Greaney in 1996 for “Lisa the Iconoclast”, an episode from Season 7 of The Simpsons TV Series— https://en.wikipedia.org/wiki/Lisa_the_Iconoclast. See also Peters (2005) The Simpsons: Embiggening Our Language with Cromulent Words. Verbatim: The Language Quarterly 30/2 (Summer 2005), 1–5.
54 LINGUISTICA PRAGENSIA 1/2018 is sometimes regarded as an example of acircumfix in English. The paper examines the distribution of this presumed circumfix in contemporary English, its frequency of incidence, or rather availability, together with aspects of its use, and attempts to place the findings in abroader context. 2. INTRODUCTION: PARASYNTHESIS AND CIRCUMFIXATION In English, as in other languages, we find words derived by the combination of aprefix and asuffix. Some of these combinations are due to aserial process (e.g. decarbonize), some appear to have been formed by adding prefix and suffix simultaneously (e.g. decaffeinate), which is called parasynthesis. “If neither of these affixes is used on its own”, says Bauer (1988b: 28, 325–326), “and the two seem to realize asingle morpheme, they are sometimes classed together as acircumfix”; if acircumfix “is taken to be asingle affix, it is adiscontinuous morph.” In amore recent book, Bauer, Lieber and Plag (2013: 500–503) show that the distinction in English between the three types of multiple hierarchical formations (the prefix is serially attached to an existing suffixed derivative, asuffix is attached to an existing prefixed word, and the possibility of either analysis) and parasynthetic formations (defined as instances of simultaneous prefixation and suffixation, in which neither aprefixed base nor asuffixed base has been attested before the appearance of the suffixed and prefixed form) is not always easy to make. Using the verb decaffeinate as an example, they point out that “there was no verb caffeinate at the time of creation to which the reversative prefix demight have attached. There was also no potential base decaffein for the attachment of the suffix –ate […]. This means that the meaning ‘remove X’ is expressed through both the prefix and the suffix, with the suffix contributing the verbal semantics and the prefix the privative meaning. It might seem that we would be forced to posit acircumfix de-ate with that meaning, […], but both deand -ate occur independently elsewhere with the relevant meanings, which is not typical of acircumfix.” They suggest two alternative explanations: the existence of aputative verb caffeinate ‘provide (with) caffeine’ subsequently prefixed with the reversative prefix de-, and the operation of analogy, i.e. the parasynthetic form is created on the pattern of the many existing derivatives with the same structure (deacylate, decapacitate, dechlorinate, dehyphenate). In asubsequent book, Bauer (in Lieber and Štekauer 2014: 127) upholds the rigorous criteria for the recognition of acircumfix, the “type of parasynthesis that seems to gain most attention”: “In convincing instances of parasynthesis, the two parts must make up asingle affix, which is usually taken to imply that if we have aword of the form X-Base-Y where X…Y is the circumfix, there is no semantically related form X-Base and not semantically related form Base-Y. Amore restrictive requirement would be that there must never be any words of form Y-Base or Base-Y which fulfill the same function as X-Base-Y.” We may ask why and whether it is necessary to insist that the two components X and Y of acircumfix must not appear on their own in X-Base or Base-Y and what warrants the claim that if they do it is not “typical of acircumfix” or aconvincing instance of parasynthesis. This not to say that the restrictive requirement is purely
ALEŠ KLÉGR 55 arbitrary, but even if some languages do have circumfixes that satisfy this strict requirement, may there not be other languages for which it does not apply in full? Given the tendency of affixes to be multifunctional and have overlapping uses, would not the existence of circumfixes made up only of “dedicated” components be something of aluxury in alanguage? As amatter of fact, in Czech morphological theory the X-Base-Y forms are recognized as circumfixes even when the Y is deprepositional or afree-standing reflexive particle. If anything, an example such as decaffeinate above evidently creates aproblem and shows that insistence on the restriction results in the need to invent alternative interpretations. It is not surprising then that authors are sometimes not consistent in the use of the label circumfix. Are there such instances in English, i.e. are there circumfixes in English? Lieber (1992: 155) echoes Bauer’s (1988b) complaint that little or no attention has been given to circumfixes (in the generative theory of morphology) but adduces no examples of circumfixation in English. Bauer (1988b: 315) himself tentatively allows that “Bepatched may illustrate arare circumfix in modern English”, as “[t]ere is no verb bepatch from which bepatched could have come.” In his later books quoted above, though, he sounds rather skeptical about the existence of circumfixes in English. Yet, we can read in Bauer and Valera (2015: 77) that “The sample contained one case of other word-formation processes, namely circumfixation, in the word enlighten.” This would suggest that the en/em-en combination need not be completely excluded from consideration as acircumfix. Finally, the issue of circumfixes has been recently raised by Stump (2017) in connection with his micromorphology hypothesis which he defends in his paper. He formulates the hypothesis at the affix level (an affix may be morphologically complex, i.e., acombination of other affixes) and the rule level (arule of affixation may be morphologically complex, i.e., the conflation of other rules of affixation). Referring to Bauer’s (1988a) postulated synaffixes (including both continuous conflated affixes such as -ic-al and -abil-ity and discontinuous affixes such as circumfixes or complex markings some or all of whose components are nonconcatenative), he poses aquestion whether acircumfix should be seen as acomplex affix whose two parts are discontinuous. And concludes (p.121) that “[if] so, then conflated affixes are only one kind of complex affix, circumfixes should be seen as complex affixes that result from the composition (rather than the conflation) of arule of prefixation and arule of suffixation.” Again, does the assumed composition of aprefix and asuffix and the respective rules resulting in acircumfix necessitate that the two affixes have to appear only as parts of the circumfix? 3. THE PATTERN EN/EM-BASE-EN In spite of, or because of, the unclear theoretical status of the en/em-en combination as acircumfix, it is occasionally described as such in some English and other sources (e.g. Byrd and Mintz 2010; Čermák 2008, 2011). At any rate, the pattern en/em-Baseen is probably the most quoted, yet understudied case of parasynthesis in English— Byrd and Mintz (2010: 2018) even claim the en-X-en is the only circumfix in Eng-
56 LINGUISTICA PRAGENSIA 1/2018 lish— and for this reason it was targeted in this study. The remarkable thing about this pattern is that in both the linguistic literature and dictionaries (e.g. The Concise Oxford English Dictionary, 2004), it is exemplified by only three verbs, enlighten, enliven and embolden. The pattern (or the concept of circumfixation, for that matter) is not mentioned in such standard accounts of English word-formation as Adams (1973, 2001), Bauer (1983), Bauer and Huddleston (2002) or Plag (2003). This would suggest that the pattern is in fact non-productive or unavailable (in the sense that it cannot “be used to produce new words as they become necessary”, Bauer 2004: 205). Indeed, Marchand (1969: 163) in his seminal monograph on word-formation mentions verbs derived from adjectives and nouns by the suffix –en to which the prefix en/emwas added between 1500 and 1650, but notes that most of these verbs have become obsolete. Similarly, Adams (2001: 42) claims that “Two prefixes, native beand foreign en-, cognate with by and in, no longer appear in new formations”. Bauer and Huddleston (2002: 1703) mention the prefix en/emon its own or only in combination with the suffix –ment. As Ihad some doubts about the unavailability of the en/em-en pattern in modern English— supported by the attention given to embiggen— Iconducted atwo-pronged, corpusand Web-based, search to find out whether the pattern is no longer productive. The form of the pattern on which the study focuses is derived from the three verbs, enlighten, enliven and embolden, which may be regarded as prototypical examples. It includes the prefix en- (changing to the allomorph embefore bilabial consonants b, m, p), the adjectival base (monosyllabic, with either aclosed or an open syllable), and the suffix –en; the pattern is class-changing, i.e. results in averb. 4. DATA SOURCES AND RETRIEVAL METHODS The data was collected from two sources, corpora and the Web. In the first step, three corpora (see Corpora in References) were searched, the BNC, Araneum Anglicum Maius and the NOW (News of the World) corpus. The corpus query had the form [lemma="e[n,m].+en"&tag="V.+"] for the BNC and the Araneum corpus accessed using the KonText interface of the Czech National Corpus; in the NOW corpus, the individual verb forms of the type en/em-Base-en+0/-s/-ed/-ing were looked for and manually processed. The results were disappointing: the BNC (100 million words) contained only the three dictionary-attested verbs, enlighten, enliven and embolden (in this order of frequency), the much larger corpus Araneum (1.2 billion words) contained only two more verbs of this type, enrichen and enhippen. Finally, NOW, the largest corpus of the three (over 4.4 billion words), added nine more deadjectival verbs (embiggen, endarken, enlargen, enwisen, etc.). All in all, the three corpora contain only 14 different deadjectival verb lemmas, including the predictable enlighten, enliven and embolden, and ahandful of nounand verb-based formations (enfilthen, enhearten, enscripten, entricken, etc.) that were excluded from the sample. Given the paucity of hits in the corpora, adifferent approach to data retrieval was adopted in the next stage: an experiment testing for the presence of hypotheti-
ALEŠ KLÉGR 57 cal verbs on the Web. The advantage of using the Web and Google Search, with all its pitfalls2, is that the Web represents the largest, most varied and up-to-date repository of contemporary language. The retrieval experiment proceeded in two steps. The initial idea was that if native speakers are familiar with and use the en-Adj-en pattern to create new verbs from (monosyllabic) adjectives at all, then the best way is to start with the most frequent adjectives and see whether hypothetical verbs formed from these adjectives can actually be found on the Web. Using the BNC Frequency List of Adjectives (compiled by Leech, Rayson and Wilson 2001), the first 100 most frequent adjectives were sifted for all monosyllabic ones. The list includes 35 monosyllabic adjectives from which conjectural verb forms on the en-Adj-en pattern were created. These hypothetical verbs, or rather their characteristic forms, the –ed past tense form/past participle, the 3rd person present tense form or the ing-form, were then searched for on the Web (in February 2014). With adjectives starting with b, m or p, both the prefix enand its allomorph emwere tested for. To make as sure as possible that the formations originated with native speakers, the search was restricted to the UK and the US domains (it was also assumed that it would be native, rather than nonnative, speakers who engage in such creative formation). The search targeted verbs in an utterance, which signals that the formation was used in actual text. As the aim of the search was to ascertain whether someone ever thought of creating such averb at all, even asingle occurrence was counted as asuccessful hit. By the same token, no count was kept of how many occurrences were found, as the experiment is not about the frequency of occurrence, but the availability of the pattern in current English as such. Instances with dubious context (non-sentential, incoherent, etc.) or not based on an adjective were excluded. Quite unexpectedly the strategy with the most frequent adjectives as bases proved hugely successful. So much so that the second part of the retrieval experiment went one step further and used amuch larger set of monosyllabic adjectives that were chosen randomly this time, i.e. regardless of their frequency ranking. The procedure was the same: the adjectives were turned into hypothetical verbs and googled (February to October, 2017). As with the first set, some of them appeared several times, many only once, but only those used in contemporary language were taken into account, dating typically from the end of the last century up until 2017. Again the frequency of individual formations on the Web was disregarded. 5. DATA ANALYSIS: CORPUS FINDINGS Although the corpus findings are rather meagre in terms of the number of different verbs of the type sought, they can provide the frequencies of their occurrence 2 The pros and cons of using the Web as corpus have been extensively debated for some time, see, for instance, the 2003 Special Issue on the Web as Corpus, Computational Linguistics 29/3, or Hundt, Nesselhauf and Biewer (2007). In spite of critical voices, esp. regarding the use of Google (Kilgarriff 2007), the Web offers invaluable information on rare phenomena (Jurkiewicz-Rohbacher, Kolaković and Hansen 2017).
58 LINGUISTICA PRAGENSIA 1/2018 within the corpus. The corpora included the following lemmas based on the en-Adjen pattern: the BNC (3)— enlighten, enliven, embolden; Araneum (5)— enlighten, enliven, embolden, enrichen, enhippen; NOW (News of the World; 13)— enlighten, embolden (variant enbolden), enliven, embiggen (with the variant enbiggen), emplumpen, enbrighten (no emvariant), enfatten, endarken, enlargen, enrichen, enrighten, enslicken, enwisen. All three corpora shared only the verbs enlighten, enliven, and embolden; two of them, Araneum and NOW, shared one more, enrichen. Five of the verbs appeared only once in the 5.7 billion-word corpora. Altogether these text corpora contain 14 lemmas: embiggen (enbiggen), embolden (enbolden), emplumpen, enbrighten, endarken, enfatten, enhippen, enlargen, enlighten, enliven, enrichen, enrighten, enslicken, and enwisen. The formal features of the formations, such as the presence of the alternation of initial ento embefore bilabials b, m, p, or its absence (enbolden, enbrighten, enbiggen) will be discussed later. Unlike the Web data, the corpora provide frequencies of these lemmas and so can give us an idea of how often they are used. The following overview shows both their total frequencies and their respective frequencies in the BNC, Araneum, and NOW corpora (misspelled instances were discounted): lemma Total BNC/Araneum/NOW enlighten 29677 242/4791/24644 embolden 11986 53/1122/10811 enliven 5010 182/1256/3572 embiggen 158 —/—/158 enrichen 25 —/6/19 enlargen 10 —/—/10 enbrighten 6 —/—/6 endarken 5 —/—/5 enbolden 4 —/—/4 enwisen 2 —/—/2 emplumpen 1 —/—/1 enfatten 1 —/—/1 enhippen 1 —/1/— enrighten 1 —/—/1 enslicken 1 —/—/1 Table 1.Summary of the corpus findings Table 1 makes it clear why only the first three of the fourteen en-Adj-en lemmas identified in the corpora are found in standard dictionaries and can be safely considered part of the lexicon (i.e. listemes). Given the total size of the corpora (over 5.7 billion words), the remaining lemmas are clearly marginal and ephemeral (with the possible exception of embiggen which may have enjoyed, or perhaps still does, something of afashion as Bauer’s note and Peters’ article suggest), and can be considered nonce words.
ALEŠ KLÉGR 59 All this would suggest that for all practical purposes the en-Adj-en pattern is indeed non-productive in current English as generally assumed. To make sure that this is really the case, another attempt was made using atwo-part Web experiment. 6. DATA ANALYSIS: WEB EXPERIMENT FINDINGS The first part of the experiment started with identifying all monosyllabic adjectives in the first 100 most frequent adjectives from the list compiled by Leech, Rayson, and Wilson (2001). These monosyllabic adjectives could be potential bases for en-Adj-en verbs. The batch of one hundred most frequent adjectives included the following 35 monosyllabic ‘seed’ adjectives (in order of frequency): good, new, old, great, high, small, large, young, right, big, late, full, far, low, bad, main, sure, clear, black, white, free, short, strong, true, hard, poor, wide, close, fine, wrong, nice, French, red, prime, dead Table 2.The list of 35 monosyllabic seed adjectives selected from the first 100 most frequent adjectives in the BNC Next, 35 hypothetical verbs were formed from these seed adjectives on the en-Adj-en pattern and searched for on the Web. The search revealed that 27 verb forms of the 35 hypothetical verbs were indeed found, that is, only 8 of them (marked in bold in the list above) were not attested on the Web. In other words, more than 3/4 (77 per cent) of the hypothetical verbs appeared on the Web at least once (sometimes more than once). As the adjectives forming unattested verbs are spread throughout the whole set of 35, presumably frequency played no part in this. Here are (unedited) examples of all the 27 verbs in context: It’s well-known that apicture of athousand words maketh abattle report engoodened. / Enoldened now, Sansa began to plan, to assimilate. / Today our insufficiently engreatened nation got its answer / Most pre-workout products are made for lifters. Their purposes wary between/are often combined of the following: enhighened mental focus (this is something many of them do), … / Our institutional branding is still there but centered and ensmallened. / Enlargened Top View of A1N Single Layer Defect. / In flashback, we see adigitally enyoungened Kurt Russell head over heels in love with Peter’s mom. / Ihope you’ll find this information worth your time to read and get enrightened in the process. / Well, another week full of comics is already dawned (even if slightly enlatened from atrip to grandma’s) / Iam enfullened with jealiosity. / Isuspect Harmless deliberately enlowened her posting level after the hiatus in order to attract voters / It was further embaddened by the fact that the defense counsel he was talking about hadn’t even been in the same courtroom. / The continuous efforts ensurened asignificant progress in linking broadcasters with the NDMA in Thailand, / Abgott, Xerath and Phyrexia are among the bands that will perform at the Enblackened 2010 festival held at the
60 LINGUISTICA PRAGENSIA 1/2018 Camden Underworld on Saturday May 15th 2010 / when Iwoke up this morning the world had been enwhitened. / The page seems to be enshortened / It became akind of religious myth enstrongened after the fall of communism / Yea, and they drink, for more enhardened joy, Man’s blood for wine, / there is akind of universal cultural impoverishment and their lives are empoorened by it. / we also enwidened our whole product range / The wall is viewed from so close-up, so it is almost as if one is being enclosened and smashed by the wall. / And Saje, dude, stop sitting so enwrongened. You’ll end up hurting yourself. / Whatever miraculous ennicening and ensmartening process has been utilized here should be applied to the entire world. / the pseudonym she uses in ‚No Future for You‘ is just aposher, slightly enFrenchened version of / I’m not convinced that „enreddened“ will catch on ... / Ikilled most of them, but Ican’t seem to figure out how to kill the kamakazi guys without becoming endeadened myself. Verb forms not confirmed on the Web come from the adjectives clear, far, fine, free, main, new, prime, and true, i.e. verbs like *enclearen, *enfaren, *enfreen, *ennewen, *entruen, etc. The possible reasons for their non-occurrence will be discussed later. In view of the encouraging results, the second part of the retrieval experiment sought to increase the sample of verbs to be searched for and at the same time avoid the potential effect of selecting adjectives by frequency. For this reason, two and ahalf times more seed adjectives were chosen, this time randomly without considering their (corpus) frequencies. This second set (see Table 3) thus included the following 91 seed adjectives from which hypothetical verbs were formed and looked for on the Web, using the same procedure as in the first part: bald, bare, bland, blank, bleak, blind, blue, blunt, brave, brief, broad, broke, brown, coarse, cold, cool, crass, crisp, crude, cute, damp, deaf, deep, dense, dim, drunk, dull, fair, fast, fierce, fit, flat, fresh, glad, gold, grim, gross, gruff, harsh, hot, huge, just, kind, lewd, limp, long, loud, mad, meek, mild, near, neat, odd, pale, posh, proud, queer, quick, rash, raw, round, rude, safe, scant, sharp, sick, sleek, slight, slim, smart, smooth, soft, sound, sparse, steep, stern, stiff, still, straight, strange, stuck, sweet, swift, tall, thick, tight, vague, warm, weird, wet, wild Table 3.The list of 91 monosyllabic seed adjectives randomly chosen This time the search confirmed the occurrence of 81 hypothetical verbs (89 per cent) based on the adjectives. It should be noted that as the choice of the adjectives was haphazard, the actual figures could easily be different with adifferent set of adjectives, which, however, does not detract from the fact that the number of attested verbs is conspicuously high. The verbs which were not found on the Web derive from these 10 adjectives: crass, just, near, odd, queer, rash, raw, scant, stern, still (again highlighted in the above list). In order to illustrate the variety of adjectives appearing in this kind of formation, asample of verbs based on different types of adjectives is presented here: Icertainly prefer to be smug and merry, but Iwas enbleakened by hormones and self-pity. / are you enblindened by your fury, only wishing to see someone who
ALEŠ KLÉGR 61 defeats you as having an unfair advantage, to ease the anger of losing? / The sky is enbluened with agobbet of beetle wings, and there are no owl scats to distract you from the incessant hobbling of tiny gnats; Ihave been embluened ! / Idon’t object to brevity, but the ideas being enbriefened have to have at least alittle merit to start with. / The microstructure will be encoarsened even at initiation of its solidification, / You fingering yourself in your dirty, spunk encrispened bed probably only burns like 50 calories. / Once again the observant techbeat watcher finds his or her lower-torso garments endampened by fear, as news emerges that heavyweight US military nerds / Just like asand-endensened flow, you swept into my world./ ever since his attack upon me in Waynesburg, where he was backed by big John King and his host of friends, he has been getting more enfiercened, ... / Oh here’s acouple of interview clips where they (including apartially re-enfittened Paul) talk about The Saturdays / They preferred to record their automatic writing in perfectly correct syntax; the world, not the sentence, was in need of an enfreshened vision. / Photos and videos with the hashtag ‚engoldened‘ on Instagram. / … the week engrimmens further. / For it is of the engrossened ether of which this ethereal planet is composed that earth is suffused. / As well as dumfounded, Iwas also instantly engruffened, but Iswallowed it down, / I’m enkindening and gentling in my decline. ;-). / Whether she is for lewd purposes or not, like the devils we are, she will be enlewdened. / Stratford hurries to the shell trolley, enmeekened. / Fla has been enposhened for de telly :-) / Safety is in danger, my friends, and only our freedom can keep it ensafened from further endangerment. / I’m ensickened of this, honestly... I’m going to fag all of that thing’s posts / my brain has become ensmoothened after reading this nonsense / Perhaps it portrays the enstrangened American society of the late 80s and early 90s. / Allies touched by pulses of Song of Celerity are enswiftened by her music. / Surrounding them were leafy bowers and enthickening moist vines that crept like reptiles along the ground / Here’s my situation (slightly envaguened to protect the innocent): / Enriched and Enwettened by Their Presence. / Endowed with psychological characteristics and the ability to experience emotions, strange and enwildened beasts with human heads wander the indefinite It is important to stress that although most, if not all, of these formations are very unusual and surprising to many native speakers when asked to comment on them, they are not difficult to understand and in context their use makes sense. The illustrative sentences are sometimes long, sometimes short, but they are generally sufficient to allow drawing inferences about the stylistic environment and the stylistic value of these verbs. 7. MERGING THE WEB AND CORPUS FINDINGS When the two Web samples, verbs from most frequent monosyllabic adjectives and verbs from randomly selected adjectives, are put together, the most impressive finding is the sheer quantity of en-Adj-en verb forms that can be found on the Web attest-
68 LINGUISTICA PRAGENSIA 1/2018 are broadly speaking two types of word-formation processes. The role of the first type of processes is to extend and innovate the lexicon by naming new concepts and by recategorizing existing word classes. They will tend to produce formations, typically rule-governed, that will eventually become permanent additions to the lexicon if the circumstances are right. The other type of processes, occupying the creative end of the spectrum, have adifferent purpose; their role is to contribute to discourse production by introducing pragmatic information, attitudes, evaluative aspects, by setting the stylistic tone of the utterance, or by contributing to the cohesiveness of the text (see Hohenhaus 2007). Unlike the first type they will typically produce evanescent formations as required by the particular situation for which there is no need to be permanently stored. Nonce formations with discourse and other functions (e.g. textual deixis, pronominalising, dummy-compounding, etc.) are discussed, for instance, by Hohenhaus (2007, 2015). The two categories of word-formation processes (lexicon-expanding and discourse-oriented) will not be mutually exclusive, and despite the tendency of patterns to belong to one rather than the other, the same pattern may produce formations with different kinds of function. On this interpretation, the en-Adj-en pattern belongs to the latter type and can be activated by speakers whenever the occasion requires it but the fanciful nonce words will not be circulated and become generally accepted (with rare exceptions, such as embiggen). To sum up, the paper starts by exploring the English circumfix en-Base-en (under more relaxed criteria) which has received some attention but was illustrated by only four examples. Then it proceeds from acorpus search for more instances of the pattern to web search queries that extracted over ahundred of these formations. To explain the existence of so many en-Adj-en verbs unknown to standard reference books it is hypothesised that en-Adj-en formations belong to the type that arises as afunction of text and disappears with the text in which they serve as stylistic markers; their role is to contribute to the discourse and not to the lexicon. As such they are unlikely to gain much currency, they will not become institutionalised (unless they enjoy ashort period of vogue, such as to embiggen popularized by aTV series) and will not make it into dictionaries (for asimilar position on nonce-formations and their lexicalizability see Hohenhaus 1998). Because of their passing nature they will be discovered only through the use of big data on the Web. On the whole, the results of Web search confirm that the description of many word-formation processes, their range and use in current English, may profit considerably from systematic exploitation of the available Web resources. REFERENCES Adams, V. (1973) An Introduction to English Word-Formation. London and New York: Longman. Adams, V. (2001) Complex Words in English. London and New York: Routledge. Bauer, L. (1983) English Word-Formation. Cambridge: Cambridge University Press. Bauer, L. (1988a) Adescriptive gap in morphology. In: Booij, G., van Marle, J. (eds) Yearbook of morphology 1988, 17–27. Dordrecht: Foris. Bauer, L. (1988b) Introducing Linguistic Morphology (2nd ed. 2003). Edinburgh: Edinburgh University Press.
ALEŠ KLÉGR 69 Bauer, L. (2004) Morphological productivity. Cambridge: Cambridge University Press. Bauer, L. (2014) Concatenative Derivation. In: Lieber, R. and Štekauer, P. (eds) The Oxford Handbook of Derivational Morphology, 118–135. Oxford: Oxford University Press. Bauer, L. and R.Huddleston (2002) Lexical Word-Formation. In: Huddleston, R. and G.K.Pullum. The Cambridge Grammar of the English Language, 1621–1721. Cambridge: Cambridge University Press. Bauer, L., R.Lieber and I.Plag (2013) The Oxford Reference Guide to English Morphology. Oxford: Oxford University Press. Bauer, L. and S.Valera (2015) Sense Inheritance in English Word-Formation. In: Bauer, L., L.Körtvélyessy and P.Štekauer (eds) Semantics of Complex Words, 67–84. Cham-HeidelbergNew York-Dordrecht-London: Springer. Byrd, D. and T.H.Mintz (2010) Discovering Speech, Words, and Mind. Malden-Oxford: Wiley-Blackwell. Čermák, F. (2008) Diskrétní jednotky vjazyce: případ cirkumfixů. Slovo aslovesnost 69/1–2, 78–98. Čermák, F. (2011) Morfematika aslovotvorba češtiny. Prague: Nakladatelství Lidové noviny. Dressler, W.U. (2000) ‘Extragrammatical vs. marginal morphology.’ In: Doleschal, U. and A.M.Thornton (eds) Extragrammatical and Marginal Morphology, 1–10. Munich: Lincom. Hohenhaus, P. (1998) Non-lexicalizability— as acharacteristic feature of nonce-formations in English and German. Lexicology 4/2, 237–280. Hohenhaus, P. (2007) How to do (even more) things with nonce words (other than naming). In: Munat, J. (ed.) Lexical Creativity, Texts and Contexts, 15–38. Amsterdam— Philadelphia: John Benjamins. Hohenhaus, P. (2015) Anti-naming through nonword-formation. SKASE Journal of Theoretical Linguistics 12/3, 272–291. Hundt, M., N.Nesselhauf and C.Biewer (eds) (2007) Corpus Linguistics and the Web. Amsterdam— New York: Rodopi. Jurkiewicz-Rohbacher, E., Z.Kolaković, and B.Hansen (2017) Web Corpora— the best possible solution for tracking rare phenomena in underresourced languages: clitics in Bosnian, Croatian and Serbian. In: Bański, P. et al. (eds) Proceedings of the Workshop on Challenges in the Management of Large Corpora and Big Data and Natural Language Processing, 49–55 (CMLC-5+BigNLP) 2017 incl. the papers from the Web-as-Corpus (WAC-XI) guest section. Birmingham, 24 July 2017. Mannheim: Institut für Deutsche Sprache, 2017. Kilgarriff, A. (2007) Googleology is bad science. Computational Linguistics 33/1, 147–151. Kilgarriff, A. and G.Grefenstette (2003) Introduction to the Special Issue on the Web as Corpus. Computational Linguistics 29/3, 333–347. Leech, G. (1990) Semantics. Harmondsworth: Penguin Books. Leech, G., P.Rayson and A.Wilson (2001) Word Frequencies in Written and Spoken English. Harlow: Pearson Education. Lieber, R. (1992) Deconstructing Morphology: Word Formation in Syntactic Theory. Chicago: The University of Chicago Press. Lieber, R. and P.Štekauer (eds) (2014) The Oxford Handbook of Derivational Morphology. Oxford: Oxford University Press. Marchand, H. (1969) The Categories and Types of Present-Day English WordFormation. München: C.H. Beck’sche Verlagsbuchhandlung (Oscar Beck). Peters, M. (2005) The Simpsons: Embiggening Our Language with Cromulent Words. Verbatim: The Language Quarterly 30.2 (Summer 2005), 1–5. Plag, I. (2003) Word-Formation in English. Cambridge: Cambridge University Press. Soanes, C. and A.Stevenson (eds) (2004) Concise Oxford English Dictionary, 11th Edition. Oxford: Oxford University Press. Stump, G. (2017) Rule conflation in an inferential-realizational theory of morphotactics. Acta Linguistica Academica 64/1, 79–124. Zwicky, A.M. and G.K.Pullum (1987) Plain Morphology and Expressive Morphology. Proceedings of the Thirteenth Annual Meeting of the Berkeley Linguistics Society, 330–340.
70 LINGUISTICA PRAGENSIA 1/2018 INTERNET SOURCES Bauer, L. (2008) Language is embiggened by words that don’t exist. Available at: http://www.stuff. co.nz/blogs/opinion/735053/i-Language-isembiggened-by-words-that-don-t-exist-i Facebook exchange: https://www.facebook.com/ permalink.php?story_fbid=301321023242571 &id=296963433678330 Leech, G., P.Rayson and A.Wilson (2001) Word Frequencies in Written and Spoken English: based on the British National Corpus. Companion Website. BNC Frequency lists. List 5.3: Frequency list of adjectives (by lemma). Available at: http://ucrel.lancs.ac.uk/bncfreq/ lists/5_3_all_rank_adjective.txt Urban Dictionary: ‘embiggen’. Available at: http://www.urbandictionary.com/define. php?term=embiggen CORPORA Benko, V. (2015) Araneum Anglicum Maius (Global English, 15.04) Praha: Ústav Českého národního korpusu FF UK. Available at http:// www.korpus.cz British National Corpus. Praha: Ústav Českého národního korpusu FF UK. Available at http:// www.korpus.cz Český národní korpus— SYN2000 (2000). Ústav Českého národního korpusu FF UK, Praha. Available at http://www.korpus.cz News on the Web (NOW) Available at: http:// corpus.byu.edu/now/ Aleš Klégr Department of English Language and ELT Methodology Faculty of Arts, Charles University nám. J.Palacha 2, 116 38 Praha 1, Czech Republic
[email protected]