Full text
1 Measuring semantic distance across time: An analysis of the collocational profiles of a set of near-synonyms in American English Daniela Pettersson Traba Over the last decades, several studies have analyzed the collocational preferences of particular sets of near-synonyms from a synchronic viewpoint, while their diachronic development has generally been disregarded. The aim of this paper is to partially fill this gap by examining the collocational behavior of the adjectives fragrant, perfumed, and scented, which denote the concept SWEET-SMELLING, over the time span 1810--2009. To this purpose, instances of the three near-synonyms and their L5-R5 collocates were extracted from the Corpus of Historical American English (COHA) and then submitted to statistical modelling. Results indicate that, at the beginning of the time span analyzed, the collocational preferences of scented and perfumed are very similar but, over time, scented becomes semantically closer to fragrant, while at the same time taking over some of its functions. Keywords: near-synonymy; collocation; semantic vector spaces; collocational networks; diachrony 1 Introduction It is a well known fact that most synonyms existing in language are actually nearsynonyms, that is, words or constructions which share the same core denotational meaning, but differ in peripheral aspects or in other dimensions of meaning such as connotation, style, and/or collocation (e.g. Cruse, 2000: 159--160; Liu, 2010). 1 In fact, languages tend to work against absolute synonymy, which leads to competition between semantically related words (e.g. Bolinger, 1977: ix--x, 9; Croft, 2000: 176). As argued by Samuels (1972: 62), ‘if […] two exact synonyms exist for a time in the spoken chain, either one of them will be less and less selected and eventually discarded, or a difference of meaning, connotation, nuance or register will arise to distinguish them.’ This has a clear diachronic implication, since synonyms are expected to become functionally more dissimilar over time. The idea that semantically related concepts compete is recurrent in research on language change, in which we can find several examples of synonymous expressions that have eventually become semantically less similar or cases in which one item has come to substitute the other. For instance, following the Norman Conquest, many Anglo Saxon words were duplicated by French loans with the same meaning. Over time, however, they underwent semantic change; this is the case, for example, of the pairs pigpork and cow-beef (Jackson, 1988: 66; Murphy, 2003: 161). However, recent research has argued that the competition theory could be an oversimplification since, besides differentiation, synonyms can also undergo a process of attraction in which they become semantically more similar (De Smet, D’hoedt, Fonteyn, and Goetham, 2018). In this scenario, synonymous words or constructions begin to mirror each other’s behavior and Postprint version of accepted manuscript. The paper was published as: Pettersson-Traba, Daniela. 2021. Measuring semantic distance across time: An analysis of the collocational profiles of a set of near synonyms in American English. Journal of Research Design and Statistics in Linguistics and Communication Science 6(2): 138-165. DOI https://doi.org/10.1558/jrds.40245
2 thus come to share more semantic space. In fact, a certain degree of attraction might very well be a precondition for replacement, as synonymous expressions probably need to share a great deal of their semantic features for them to be considered interchangeable by language users and for substitution to occur. Finally, yet another possibility is that of stability, to wit, when no changes in the functional profiles of the near-synonyms occur. One way in which one can measure semantic (dis)similarity between near-synonyms is by analyzing their collocational preferences. In particular, it is possible to examine the degree of collocate overlap between two or more near-synonymous expressions by considering the number of significant collocates they share as well as the number of significant collocates one synonym exhibits but not the other/s (Gries, 2001: 83). According to this criteria, two related words would be more similar the more significant collocates they have in common, and more dissimilar the more collocates they do not have in common. The importance of collocations, first emphasized by Firth (1957) and Sinclair (1966) and summarized in the quote ‘You shall know a word by the company it keeps’ (Firth, 1957: 11), has been one of the central tenets of lexical semantics. The term ‘collocation’ has received several slightly different interpretations over the years along the following dimensions identified by Gries (2013: 138--139): (i) the nature of the linguistic items analyzed, that is, words or more schematic categories such as parts of speech (POS) and constructions, (ii) the number of items constituting a collocation, ranging from strings of two words to longer sequences, (iii) the frequency threshold for an expression cooccurring with a node word in order to be considered a collocate, (iv) the distance between the items making up the collocation, i.e. whether they are directly adjacent, syntactically related, or within a context window of x words, (v) the specificity of the lexical items involved, to wit, word forms or lemmas, and (vi) the degree of compositionality and predictability of the collocation. For the purposes of the present study, we will follow Stubbs (2001: 24) and define collocation as ‘a lexical relation between two or more words which have a tendency to cooccur within a few words of each other in running text’, although we will establish in detail in Section 2 a specific frequency threshold and context window for items to be considered collocates in the present analysis. Over the last fifty years, advancements in corpus linguistics have led to the emergence of rigorous investigations into the role of individual collocates on various linguistic phenomena, including semantic ones such as polysemy and near-synonymy. Although several studies have analyzed the collocational preferences of particular sets of nearsynonyms (among other types of distributional patterns) from a synchronic viewpoint to quantify their semantic (dis)similarity (e.g. Kjellmer, 2003; Taylor, 2003; Divjak and Gries, 2006, 2008; Divjak, 2010; Liu and Espino, 2012; Liu, 2013; Desagulier, 2014), very few have paid attention to the diachronic evolution of the collocational behavior of specific groups of lexical near-synonyms (but see Primahadi-Vijaya-R. and Rajeg, 2014; Pettersson-Traba (2018)). Primahadi-Wijaya-R. and Rajeg (2014) conducted a corpus based analysis of the nominal collocational profiles of the near-synonymous adjectives hot and warm during the last one and a half centuries (i.e. 1860--2009) in American English. Their results uncovered various diachronic patterns that contributed to these two near-synonyms exhibiting different collocational behavior in Contemporary American English. For example, over time, warm comes to cooccur more frequently with nouns such as heart, welcome, and smile, therefore highlighting this this adjective’s prominence
3 in the metaphorical sense ‘of the heart, feelings, etc.: Full of love, gratitude, approbation, etc.; very cordial or tender’ (Oxford English Dictionary [OED] s.v. warm adj. 12a). In contrast, hot increasingly collocates with lexical items such as dog from the 1920s onwards, and with nouns referring to people (e.g. girl, guy, and woman) from the 2000s onwards with the meaning ‘[…] (originally a woman): sexually attractive; sexy’ (OED s.v. hot adj. 12i). Pettersson-Traba (2018) conducts a preliminary study on the diachronic development of the attributive uses of fragrant, perfumed, scented, and sweet-smelling, in the latter part of Late Modern and Present-day American English. She delineates the internal semantic structure of this set of synonyms by paying special attention to their noun collocates, which are grouped into semantic categories on the basis of the classification in the Historical Thesaurus of the Oxford English Dictionary. The results show that the four adjectives undergo major changes over the time span examined (1850- -2009), going from being used mostly to qualify entities which can exhibit a natural pleasant smell (e.g. flowers and trees) to modifying objects which are artificially sweetsmelling (e.g. oils and shampoos). This change is hypothesized to be a result of extralinguistic factors, to wit, socio-economic changes such as industrialization and mass production that took place during the period examined, in particular at the end of the 19th century, which have led to an ever-increasing need to allude to artificially scented lotions and candles rather than naturally fragrant plants. Moreover, fragrant and perfumed, which initially were the most frequent adjectives, are gradually replaced by scented, thus reflecting a change in the relation between the synonyms over time. The present paper further contributes to the line of research initiated in Pettersson-Traba (2018) by examining the individual collocates of three of the four adjectives examined, to wit, fragrant, perfumed, and scented, in the Corpus of Historical American English (COHA; Davies 2010--). Illustrative examples of the three lexical items are found in (1)- -(3). (1) The kitchen was warm from the slowly burning range and fragrant with the beans which had been cooking since yesterday. (COHA, 1934, Fiction, Folks) (2) She too, in a sense, is a portrait of captive refinement, but she is more captive than refined. Her tiny body perfumed and irritable, she rubs against the bars of her cage. (COHA, 1993, Fiction, CityManyDays) (3) She worked the cloth over a small sliver of scented soap she had scavenged and lathered herself liberally, reveling in the pungent fragrance. (COHA, 1972, Fiction, FlameFlower) The decision of analysing these three particular near-synonymous adjectives was made after careful examination of dictionaries and thesauri. While some additional adjectives are defined in the same way as the adjectives selected here (e.g. fragranced and sweetsmelling) or are listed as their synonyms (e.g. odorous and redolent) and could consequently have been included in the study, these adjectives have not been chosen for different reasons. The adjective fragranced is categorized as ‘rare’ in some of the sources and is not attested in COHA. Similarly, sweet-smelling displays a relatively low frequency in the corpus and thus the retrieval of significant collocates of this adjective yields few hits. The adjective odorous does not necessarily entail the positive connotation that the selected adjectives display and the meaning of redolent differs somewhat from fragrant, perfumed, and scented as it does not imply the trait ‘sweetness’ by definition. However,
4 an analysis with a wider range of near-synonyms, including the abovementioned adjectives, might prove valuable in future research. Despite the exclusion of sweet-smelling, the results of the present study are based on an almost four times larger dataset than that in Pettersson-Traba (2018), given that occurrences of the three adjectives in all their syntactic environments, including also predicative and postpositive uses, are considered (cf. (1)--(3)). Furthermore, the semantic analysis of the adjectives is considerably more fine-grained, since the focus here is on specific collocates and not on more abstract semantic classifications as in the case of Pettersson-Traba (2018). By zooming in on the individual collocates it is possible to identify differences between the adjectives which may be obscured if the collocates are grouped into broader classes such as semantic categories (e.g. CLEANING AND PERSONAL CARE, FARMING AND HORTICULTURE, PEOPLE, PLANTS, and WEATHER). In sum, the present paper aims at analyzing the idiosyncratic collocational preferences of the three nearsynonymous adjectives, which denote the concept SWEET-SMELLING, in 19th-and 20thcentury American English. This set of near-synonyms is particularly interesting to examine as made evident by a thorough examination of different types of reference material (e.g. dictionaries and thesauri): generally speaking, no clear and detailed information about the usage patterns and nuances of meaning of the three near-synonyms under investigation is provided. In fact, in many cases the adjectives are defined in terms of each other as the following definition of fragrant illustrates: A. Pleasant smelling, perfumed. (Newbury House Dictionary of American English s.v. fragrant) Moreover, the definitions and examples of usage, though offering valuable information about their similarities, do not provide a comprehensive picture of their differences, which prevents users from fully comprehending how to distinguish them. In what follows, an overview of the information provided in seven reference material is offered. These are the historical dictionary Oxford English Dictionary (OED; 2012--) and six present-day English dictionaries and thesauri, namely Lexico (2019), Cambridge Dictionary (CD; 2019), Collins online Unabridged English Dictionary (Collins; 2012--), Newbury House Dictionary of American English (NHDAE; 2019), Longman Dictionary of Contemporary English (LDOCE; 2015--), Merriam-Webster Dictionary and Thesaurus (2019), and MacMillan Dictionary (2009--). The main reasons for choosing these reference materials are that (i) with the exception of the OED, they are all PDE dictionaries, and should therefore reflect how the adjectives are currently used; (ii) most of them include an American English section, which is of great importance here, since the present analyses focus on this variety of English; and (iii) they provide definitions, examples of usage, and suggested synonyms of the adjectives at issue. Both in the OED and the PDE dictionaries and thesauri, the three adjectives seem to be entirely interchangeable as they share the same basic and central meaning: ‘having a sweet or pleasant smell or odor’. In fact, many of the dictionaries provide only this sense for each of the adjectives, thus making them seem monosemic and undistinguishable. Only a few of the dictionaries consulted offer information about their particular nuances of meaning and usage patterns. Some reference sources provide additional senses for
5 perfumed (Lexico, Collins, and OED) and scented (CD, Collins, and OED). According to these dictionaries, there seems to be a difference in nuance in the case of these two adjectives depending on whether the source of the pleasant smell or odor is natural or artificial. This difference is reflected in the examples of usage in the types of nouns that the two adjectives modify. When denoting a ‘naturally sweet or pleasant smell or odor’ the two adjectives modify nouns such as flower, bower, cherry, pine, and breeze. Contrariwise, when denoting an ‘artificial pleasant smell or odor’, they modify nouns which refer to manmade objects, as for instance, SOAP, GLOVE, CANDLE, PAPER, LAMP, and LEATHER. Although these apparent differences in nuances of meaning can be expected to bring about differences in usage patterns among the adjectives, especially when it comes to the types of nouns they typically modify, the information provided is too limited to know for which types of nouns the adjectives normally serve as modifiers. To illustrate this point, even though the ‘artificial’ sense is only provided for perfumed and scented, examples in which fragrant collocates with nouns referring to manmade objects, can easily be found, as in (4) and (5): (4) Inside are quirky old settees, painted chests and weathered wood hutches brimming with fragrant soaps and candles. (LDOCE, s.v. fragrant) (5) With soap still to be invented, the fragrant oils and waters were used in bathing and for perfuming hair. (Lexico, s.v. fragrant) Distinguishing the contexts in which fragrant, perfumed, and scented can be used interchangeably and those in which they cannot is, therefore, not a straightforward task. It is only by means of a careful examination of their collocational preferences that we can begin to understand the complex interrelations of this particular set of near-synonyms, and this is precisely the goal of the present contribution. Section 2 deals with the data and the methodology employed to achieve this goal and, in Section 3, the results of the study are presented and discussed. Finally, some concluding remarks and suggestions for future research are put forward in Section 4. 2 Data and methodology The data for the present study was extracted from COHA, which contains more than 400 million words from American English. It covers the time span 1810--2009, which, for the purposes of the present paper, was divided into four fifty year periods: (i) P1, from 1810 to 1859, (ii) from P2, from 1860 to 1909, (iii) P3, from 1910 to 1959, and (iv) P4, from 1960 to 2009. 2 Since the focus here lies on the collocational behavior of the nearsynonyms fragrant, perfumed, and scented, the lemmas of all noun collocates in a context window of 5 words to the left and 5 words to the right of the adjectives were also retrieved from COHA by making use of the COLLOCATE and POS tag options in the corpus. This L5-R5 context window was selected in this case since it has been shown that tighter windows such as 1 or 2 words often lead to data sparseness, especially if low frequency items are considered (Sahlgren, 2006). Moreover, such tight windows are often more appropriate to retrieve semantically (dis)similar terms such as synonyms or antonyms of the target word (Peirsman, Heylen, and Geeraerts, 2008: 40). In turn, if one is interested in typical collocates of the target, it is desirable to loosen the context window somewhat, for instance, to L5-R5. Another option would be to consider only collocates which are syntactically connected to the target, in the present case either syntactically —or
6 semantically— modified nouns of the adjectives (e.g. the fragrant flower, the flower is fragrant) or adverbs which modify the adjectives (e.g. the deliciously fragrant flower), among others. However, this would also lead to a lower number of retrieved collocates, and thus again to data sparseness, as in the case of tighter context windows. In addition, only the noun collocates of the adjectives were considered because, as argued by, for example, Geeraerts (1986) and Gries (2001; 2003), nouns are more informative than other word types when it comes to the semantics of adjectives. Geeraerts (1986) study on the Dutch adjective vers ‘fresh’ demonstrated that the fine-grained aspects of meaning of polysemous adjectives can be discovered by examining the nouns they modify. He provides evidence of vers having different meanings depending on the nouns it accompanies. For instance, when modifying wound, vers means ‘fresh, recent’, whereas when modifying the noun air, it takes on a slightly different meaning, namely ‘fresh, pure, untainted’, or ‘optimal.’ The retrieval process resulted in 4,990 tokens of the adjectives, which collocated with a total of 10,740 tokens and 2,682 types (lemmas) of context words. These were subsequently fed into two types of statistical analyses in order to identify (dis)similarities in the near-synonyms’ collocational preferences and, therefore, their semantic structure. The methodology employed here demonstrates how techniques of a more quantitative nature can complement and enhance qualitative inquiries into linguistic data. First, the data was analyzed by means of Semantic Vector Space (SVS) modelling. SVS is a technique that enables us to measure semantic (dis)similarity between related words on the basis of their collocational profiles. An SVS model is constructed following a series of steps (e.g. Levshina, 2015: 326; Hilpert and Correia Saavedra, 2017): 1. We create a table containing the raw cooccurrence frequencies of the target words (in the columns) and the context words (in the rows). In this case, a separate column was included for each near-synonym in each period, since we are interested in the evolution of their collocational preferences over time. 2. The raw cooccurrence frequencies are transformed into collocational strength values using an association measure such as Pointwise Mutual Information (PMI; e.g. Church and Hanks, 1990: 23). Even though a variety of association measures exist to compute collocational strength such as log likelihood or ΔP (e.g. Gries, 2013), PMI is used in the present analysis given that it is already provided in COHA when searching for the collocates of a specific word in the corpus. 3. We compute cosine similarity scores (Levshina, 2015: 328) between the target words (here fragrant, perfumed, and scented in each period) and transform them into distance values for an easier visualization of the results. Two different visualization techniques are employed in the present study in order to explore and interpret the semantic distances between the adjectives. These are cluster analysis (Levshina, 2015: 301), a method that helps identifying groups or clusters of objects on the basis of their (dis)similarities, and multidimensional scaling (MDS; Levshina, 2015: 336--337), an approach that depicts differences between two or more objects as distances in a twoor three-dimensional plot. Cluster analysis and MDS complement one another because the former is useful to identify coherent and delimited groups in the data, while the latter can be said to be semantically more realistic in that it
7 does not suppose a categorical either-or split between individual clusters, but instead displays a more precise and continuous cline of semantic (dis)similarity (Jansegers and Gries, 2017: 14). The second type of statistical procedure to which the data was submitted was collocational network analysis which supplements the findings of the SVS analysis (Brezina, McEnery, and Wattam, 2015; Baker, 2017: 95--101). 3 Put simply, this method consists in selecting the most prominent collocates of the target words (in this case fragrant, perfumed, and scented) and visualizing them in a collocational network. 4 The idea is that one can observe, on the one hand, increases or decreases in the frequency of the target words and, on the other, the competition between them: if one target word is drawing away collocates from another, this means that it is taking over some of its semantic space. Four collocational networks were built, one per period, which enabled us to determine which collocates the near-synonyms share at different points in time and to identify potential variations in the internal semantic structure of the set. Following the methodology in Baker (2017: 98--100) for low frequency words, two thresholds were established for a collocate to be considered significant and thus be included in the networks: a minimum frequency of 5 with the target words and a PMI value of 3 or higher in each of the four periods. These thresholds drastically decreased the number of collocates of the nearsynonyms, if compared to the SVS analysis, thus allowing us to conduct a more qualitative and careful examination of the collocational profiles of the adjectives. Higher thresholds would lead again to data sparseness, particularly in the case of perfumed and scented, which display a lower frequency in COHA (cf. Table 1). Therefore, the threshold proposed by Baker (2017: 96) for words of a higher frequency, namely a minimum raw frequency of 10 and a PMI of 6 was here discarded. On the contrary, lower thresholds, that is a minimum frequency of lower than 5 and a PMI lower than 3, would include collocates which might not be particularly typical of the adjectives. Additionally, a PMI of 3 or higher is often considered to indicate that two items co-occur significantly more often than expected by chance (e.g. Church and Hanks 1990; Church et al. 1991; Church et al. 1994; Liu 2010). 3 Results A total of 4,990 instances of the three near-synonyms were retrieved from COHA. As shown in Table 1, fragrant is the most frequent adjective of the set, with a total frequency of 3,395 (68.04 %), followed by perfumed (808 instances, 16.19 %) and then scented (787 instances, 15.77 %). If we examine their relative frequencies in each of the four periods, an interesting trend can be observed. Fragrant, despite being the most common adjective in all four periods, decreases in frequency over time, from 75 % in P1 (1810--1859) to 59.33 % in P4 (1960--2009). Scented, on the other hand, goes from being the least frequent adjective of the set in P1, with a relative frequency of 10.17 %, to occupying the second position (22.60 %) in P2 (1860--1909). The most dramatic changes, however, take place from P2 to P3 (1910--1959), that is, during the transition from the 20th to the 21st centuries, since fragrant decreases more than 10 percentual points in frequency while scented increases by almost 10 %. Finally, perfumed does not change much in the time span examined but its relative frequency increases slightly, from 14.86 % in P1 to 18.06 % in P4. However, if the normalized frequencies of the adjectives per million words are considered, it is clear that all three synonyms become less frequent over time, although
8 this downward tendency is much more pronounced in the case of fragrant and perfumed. Scented, on the other hand, remains relatively stable with only a minor decrease in P4. The distribution shown in Table 1 is statistically significant according to a chi-square test of independence (χ2 = 125.27, df = 6, p < 0.001): fragrant is significantly more frequent than expected in P1 and P2, but significantly less frequent than expected in P3 and P4, while the opposite tendency is true for scented. The distribution of perfumed, in turn, is not statistically significant. Table 1: Frequency of the synonyms per period Synonym P1 (1810-- 1859) P2 (1860-- 1909) P3 (1910-- 1959) P4 (1960-- 2009) Total Fragrant N % NF 804 75.00 14.77 1291 72.49 12.87 712 62.13 5.87 588 59.33 4.51 3,395 68.04 8.36 Perfumed N % NF 159 14.83 2.92 281 15.78 2.80 189 16.49 1.56 179 18.06 1.37 808 16.19 1.99 Scented N % NF 109 10.17 2.00 209 11.73 2.08 245 21.38 2.02 224 22.60 1.72 787 15.77 1.94 Total N % NF 1072 100 19.70 1781 100 17.75 1146 100 9.45 991 100 7.61 4,990 100 12.28 Therefore, scented, despite being the least common adjective of the set at beginning of the 19th century, increases in frequency at the expense of fragrant, which decreases over time despite still being the default choice to denote the concept SWEET-SMELLING. Fragrant and scented thus seem to have undergone a process of convergence in terms of frequency, particularly at the turn of the century. However, it is not clear whether this gradual confluence is limited to frequency alone or if it is also the case that fragrant and scented have become semantically more similar over time, with scented overtaking some of the functions of fragrant. This will be the focus of sections 3.1 and 3.2, with the former offering a more quantitative bird’s eye view of the evolution of the collocational preferences of the three near-synonyms and the latter presenting the results of a more qualitative in depth analysis of their most prominent collocates per period. 3.1 Diachronic Semantic Vector Space analysis Figure 1 plots the two dimensional MDS map of the collocational preferences of fragrant, perfumed, and scented per period. The figure is interpreted as follows: each point in the graph represents one near-synonym per period, and the closer two points are located in the graph, the more similar the collocational preferences of those adjectives are. The vertical axis in Figure 1 (Dimension 2) arranges the points in the graph according to adjective, with perfumed at the top, fragrant at the bottom, and scented occupying the middle ground but overall being much closer to perfumed (except in P4). The horizontal
9 axis (Dimension 1), on the other hand, arranges the points on a temporal continuum from left to right, with P1 (1810--1859) and P2 (1860--1909) exhibiting negative values and P3 (1910--1959) and P4 (1960--2009) exhibiting positive ones. Independent-samples ttests on the MDS coordinates confirm the visual interpretation of Figure 1. Fragrant differs significantly from both perfumed and scented with respect to their positions on the vertical axis (t = -6.8, df = 4.49, p < 0.01 and t = -4.98, df = 5.12, p < 0.01, respectively), but the difference between perfumed and scented is only marginally significant (t = 2.27, df = 5.72, p = 0.06). However, the differences between the near-synonyms along the horizontal axis are not significant. In fact, on this axis the points are arranged according to period: P1 diverges significantly from P3 and P4 (t = 4.27, df = 2.78, p < 0.05 and t = 5.81, df = 3.17, p < 0.01, respectively), P2 diverges significantly from P3 and P4 (t = 4.28, df = 3.89, p < 0.05 and t = 6.4, df = 3.98, p < 0.01, respectively), but the differences between P1 and P2, on the one hand, and P3 and P4, on the other, are not significant. The differences among the periods on the vertical axis do not reach statistical significance. Therefore, whereas the adjectives seem to be fairly similar in P1 and P2, on the one hand, and in P3 and P4, on the other, concerning their collocational profiles, there is a considerable change from P1/P2 to P3/P4, that is, at the turn of the 19th to the 20th century. Figure 1: Two dimensional MDS map of the collocational preferences of fragrant, perfumed, and scented in four time periods (stress = 0.21) 5 By making use of MDS we have been able to identify differences in collocational behavior between, on the one hand, fragrant and the other two adjectives and, on the other, all three near-synonyms in the 19th century and the 20th and 21th centuries. However important these differences may be, this analysis does still not clarify the nature of such differences. To this end, the 2,682 types of noun collocates were classified into semantic categories by employing the UCREL Semantic Analysis System, or USAS for short (Archer, Wilson, and Rayson, 2002). USAS contains twenty-one major sematic classes that are then further divided into more specific subclasses. 6 This tool allows researchers to automatically analyze strings of words according to their semantics and thus check the domains to which particular words belong. Then, a PMI score for each semantic category and sub-category was calculated for each near-synonym in each period. This was done so as to determine whether correlations existed between the points’ MDS coordinates and
16 Figure 4: Collocational network of fragrant, perfumed, and scented in P2 (1860--1909) Figure 5 displays the collocational network for P3 (1910--1959). With 39 collocates, fragrant is still the most productive adjective of the set but it now seems to start losing some ground as compared with previous periods, particularly at the expense of scented (11 collocates in this period). Perfumed, on the other hand, seems to remain rather stable, although its number of collocates decreases somewhat in comparison to P2, with 11 nouns, 5 less than in P2. The types of collocates the near-synonyms occur with point to a certain degree of specialization between them: fragrant still dominates in the natural sense, with 25 collocates clearly depicting natural smells (e.g. blossom, forest, lily, pine, and rose), while perfumed favors the artificial sense (e.g. garment, handkerchief, powder, and soap). As evident from the collocational network, the noun collocates of fragrant and perfumed mainly correspond with the same semantic categories as in P2. On the other hand, scented does not seem to exhibit a clear preference for any of the two senses, occurring both in the natural sense, especially with nouns in the category L3 (4 types; e.g. flower, garden, and tree) and F1 (FOOD; 2 types, e.g. cake), and in the artificial sense with nouns in different categories, especially in O1 (3 types; e.g. powder and oil) but also B4 (2 types; soap and handkerchief).
17 Figure 5: Collocational network of fragrant, perfumed, and scented in P3 (1910--1959) Finally, in P4 (1960--2009) fragrant, with 33 collocates, loses even more ground at the expense of scented, which now has 13 collocates, while the number of collocates of perfumed also decreases slightly in contrast with P3, from 11 to 9 (cf. Figure 6). In fact, many of the collocates that scented shares with the other two adjectives are now more strongly associated with it, that is, their PMI score with scented is now higher than with either fragrant or perfumed. This is the case of bath, hair, oil, perfume, smell, smoke, warm, and water. Entities that emit a natural smell still collocate more commonly with fragrant (21 natural collocates, e.g. bloom, flesh, jasmine, and shrub), while perfumed is the preferred choice for artificial smells (e.g. bath, handkerchief, oil, and soap). Again, the semantic categories of nouns with which the adjectives collocate remain basically the same: mainly categories F, L3, and W4 in the case of fragrant, and B1, B4, and B5, in the case of perfumed. As in P3, scented seems to occupy a middle ground in this respect, being common with both ‘natural’ (mainly category L3; e.g. geranium and flower) and ‘artificial’ collocates (mainly categories B1, B4, B5, and O1; e.g. hair, perfume, soap, sheet, and oil).
18 Figure 6: Collocational network of fragrant, perfumed, and scented in P4 (1960--2009) The results of the collocational network analysis point to two main conclusions. First, as in the case of the frequency and collocational profiles of the near-synonyms, the most dramatic change takes place from P2 to P3, when the number of significant collocates of fragrant starts to decrease substantially, particularly in favor of scented. It is also in P3 and P4 that scented becomes more semantically neutral than fragrant and perfumed in the sense of not exhibiting a clear preference for either the natural or artificial senses, but being fairly common with both types of nouns. Its more neutral character is probably the reason why scented also occupies an intermediate position in P4 in the MDS in Figure 1. In fact, as we saw in Section 3.1, the value of scented on the vertical axis (Dimension 2) is almost 0 in P4, while perfumed and fragrant are located at opposite ends, the former with a positive value and the latter with a negative one. 4 Conclusions and future research The SVS analysis discussed in Section 3.1 demonstrated that, over time, scented goes from sharing more collocates with perfumed to behaving more similarly to fragrant. Furthermore, the findings of both the SVS and the collocational network analyses showed
19 that the similarities and differences between the synonyms have to do with the sense in which they are more commonly used: whereas fragrant and perfumed seem to progressively have become more specialized towards the natural and artificial senses, respectively, as shown by the nouns they typically cooccur with, scented is more neutral in this respect and becomes even more so at the end of the period. This explains its position in P4 (1960--2009) in the MDS plot (cf. Figure 1). In conclusion, there seems to be a scenario of competition between the near-synonyms, with scented gaining ground at the expense of both perfumed and fragrant over time. However, a process of attraction is also observable, whereby scented becomes more similar to fragrant, as well as simultaneously differentiating from perfumed by virtue of increasingly collocating with ‘natural’ collocates. The semantic specialization of fragrant and perfumed, discernible in the collocational networks, suggests that these two adjectives are also undergoing a process of differentiation, thus progressively moving towards opposite ends of the natural-artificial sense continuum. In this contribution, we have opted for examining the collocational profiles of the three near-synonyms by averaging over all their exemplars, a so called type based SVS analysis, thus providing a description of the general semantics of the adjectives. However, the present dataset could be modified so as to be submitted to so called token based SVS modelling (e.g. Heylen, Speelman, and Geeraerts, 2012), an approach which represents the meaning of each individual occurrence of the target word(s). Type based SVS models such as the one employed here are commonly used to investigate meaning relations between words (e.g. synonymy, or antonymy), as was the original goal of the present study. Token based SVS analysis, on the other hand, is employed to explore the existence of different senses (i.e. polysemy) within the same target word/s. Creating a token based SVS model of the individual occurrences of fragrant, perfumed, and scented would thus be useful to further test whether the distances between the adjectives, as represented in the MDS plot of Figure 1, indeed correspond to them occupying different positions on the natural artificial sense continuum. Unfortunately, this is an issue that must be left for future research. Acknowledgements For financial support, I am grateful to the following institutions: the Regional Government of Galicia (grant ED481A-2016/1687) and the European Regional Development Fund, the Spanish Ministry of Science, Innovation and Universities (grant FFI2017-86884-P). Thanks are also due to Iván Tamaredo for feedback on an earlier version of this paper. Lastly, I would like to express my sincere gratitude to two anonymous reviewers for their fruitful comments and thorough revision, as well as to the editors of the journal for their time and consideration. References Archer, D., Wilson, A., and Rayson, P. (2002) Introduction to the USAS category system, Benedict Project Report: 1–37. http://ucrel.lancs.ac.uk/usas/ (accessed 19 April 2020).
20 Baker, P. (2017) American and British English. Divided by a common language? Cambridge: Cambridge University Press. Doi 10.1017/9781316105313 Bolinger, D. (1977) Meaning and form. London: Longman. Brezina, V., McEnery, T., and Wattam, S. (2015) Collocations in context: A new perspective on collocation networks. International Journal of Corpus Linguistics 20: 139--173. Doi 10.1075/ijcl.20.2.01bre Cambridge Dictionary. (2019). https://dictionary.cambridge.org/ [last accessed 23 November 2019]. Collins online Unabridged English Dictionary. (2012--) https://www.collinsdictionary.com/ [last accessed 23 November 2019]. Church, K. W., and Hanks P. (1990) Word association norms, mutual information, and lexicography. Computational Linguistics 16: 76--83. Doi 10.3115/981623.981633 Church, K. W., Gale, W., Hanks, P., and Hindle, D. (1991) Using statistics in lexical analysis, in Zernik U. (ed.) Lexical acquisition: Exploiting on-line resources to build a lexicon 115--164. Hillsdale: Lawrence Erlbaum. Church, K. W., Gale, W., Hindle, D., and Rosamund M. (1994) Lexical Substitutability, in Levin, B., and Zampolli, A. (eds.) Computational approaches to the lexicon 153-- 177. Oxford and New York: Oxford University Press. Croft, W. (2000) Explaining language change. An evolutionary approach. Essex: Pearson Education. Cruse, A. D. (2000) Meaning in language: An introduction to semantics and pragmatics. Oxford: Oxford University Press. Csardi G., and Nepusz T. (2006) The igraph software package for complex network research. InterJournal Complex Systems 1695. http://igraph.org Davies, M. (2010–) The Corpus of Historical American English (COHA): 400 million words, 1810--2009. https://corpus.byu.edu/coha/ [last accessed: 8 November 2019] De Smet. H., D’hoedt, F., Fonteyn, L., and Van Goetham, K. (2018) The changing functions of competing forms: Attraction and differentiation. Cognitive Linguistics 29: 197--234. Doi 10.1515/cog-2016-0025 Desagulier, G. (2014) Visualizing distances in a set of near-synonyms: Rather, quite, fairly, and pretty, in Glynn, D., and Robinson, J. A. (eds.) Corpus methods for semantics: Quantitative studies in polysemy and synonymy 145--178. Amsterdam and Philadelphia: John Benjamins. Doi 10.1075/hcp.43.06des Divjak, D. (2010) Structuring the lexicon: A clustered model for near-synonymy. Berlin and New York: Mouton de Gruyter. Doi 10.1515/9783110220599 Divjak, D., and Gries, S. Th. (2006) Ways of trying in Russian: Clustering behavioral profiles. Corpus Linguistics and Linguistic Theory 2: 23--60. Doi 10.1515/CLLT.2006.002
21 Divjak, D., and Gries, S. Th. (2008) Clusters in the mind? Converging evidence from near synonymy in Russian. The Mental Lexicon 3: 188--213. Doi 10.1075/ml.3.2.03div Firth, J. R. (1957) Papers in linguistics, 1934-1951. London and New York: Oxford University Press. Geeraerts, D. (1986) On necessary and sufficient conditions. Journal of Semantics 5: 275- -291. Doi 10.1093/jos/5.4.275 Gries, S. Th. (2001) A corpus-linguistic analysis of English -ic vs -ical adjectives. ICAME Journal 25: 65--108. Gries. S. Th. (2003) Testing the sub-test: An analysis of -ic and -ical adjectives. International Journal of Corpus Linguistics 8: 31--61. Doi 10.1075/ijcl.8.1.02gri Gries, S. Th. (2013) 50-something years of work on collocations: What is or should be next… . International Journal of Corpus Linguistics 18 (1): 137--165. Doi 10.1075/ijcl.18.1.09gri Heylen, K., Speelman, D., and Geeraerts, D. (2012) Looking at word meaning. An interactive visualization of semantic vector spaces for Dutch synsets, in Butt, M., Carpendale, S., Penn, G., Prokić, J., and Cysouw, M. (eds.) Proceedings of the EACL- -2012 joint workshop of LINGVIS & UNCLH: Visualization of language patterns and uncovering language history from multilingual resources 16--24. Stroudsburg: Association for Computational Linguistics. Hilpert, M., and Correia Saavedra, D. (2017) Using token-based semantic vector spaces for corpus-linguistic analyses: From practical applications to tests of theoretical claims. Corpus Linguistics and Linguistic Theory, Ahead of Print: 1--32. Doi 10.1515/cllt-2017-0009 Jackson, H. (1988) Words and their meanings. London: Longman. Jansegers, M., and Gries, S. Th. (2017) Towards a dynamic Behavioral Profile: a diachronic study of polysemous sentir in Spanish. Corpus Linguistics and Linguistic Theory, Ahead of Print: 1--43. Doi 10.1515/cllt-2016-0080 Jones, M. A. (1996) Historia de Estados Unidos 1607–1992. Translated by Carmen Martínez Gimeno. Madrid: Cátedra. Kjellmer, G. (2003) Synonymy and corpus work: On almost and nearly. ICAME Journal 27: 19--27. Levshina, N. (2015) How to do linguistics with R. Amsterdam and Philadelphia: John Benjamins. Doi 10.1075/z.195 Lexico (2019). https://www.lexico.com/en [last accessed 23 November 2019] Longman Dictionary of Contemporary English. (2015--). https://www.ldoceonline.com/ [last accessed 23 November 2019].
22 Liu, D. (2010) Is it a chief, main, major, primary, or principal concern? A corpus-based behavioral profile study of the near-synonyms. International Journal of Corpus Linguistics 15: 56--87. Doi 10.1075/ijcl.15.1.03liu Liu, D. (2013) Salience and construal in the use of synonymy: A study of two sets of near-synonymous nouns. Cognitive Linguistics 24: 67--113. Doi 10.1515/cog-20130003 Liu, D., and Espino, M. (2012) Actually, genuinely, really, and truly. A corpus-based behavioral profile study of the near-synonymous adverbs. International Journal of Corpus Linguistics 17: 198--228. Doi 10.1075/ijcl.17.2.03liu MacMillan Dictionary. (2009--). https://www.macmillandictionary.com/ [last accessed 23 November 2019]. Merriam Webster Dictionary and Thesaurus. (2019) https://www.merriam-webster.com/ [last accessed 23November 2019]. Murphy, M. L. (2003). Semantic Relations and the Lexicon. Antonymy, Synonymy, and Other Paradigms. Cambridge: Cambridge University Press. Doi 10.1017/CBO9780511486494.002 Newbury House Dictionary of American English. (2019) http://nhd.heinle.com/home.aspx [last accessed 23 November 2019]. Oxford English Dictionary. 3rd edition (2012--). http://www.oed.com/ [last accessed 23 November 2019]. Peirsman, Y., Heylen, K. and Geeraerts D. (2008) Size matters. Tight and loose context definitions in English word space models”, in Baroni, M., Evart, S., and Lessi A. (eds.) Proceedings of the ESSLLI workshop on distributional lexical semantics: Bridging the gap between semantic theory and computational linguistics 34--41. Hamburg: ESSLII. Pettersson-Traba, D., (2018) A diachronic perspective on near-synonymy: The concept of SWEET-SMELLING in American English. Corpus Linguistics and Linguistic Theory, Ahead of Print: 1--31. Doi 10.1515/cllt-2018-0025 Primahadi-Wijaya-R., G., and Rajeg, I M. (2014) Visualising diachronic change in the collocational profiles of lexical near-synonyms, in Sudipa, I N., and PrimahadiWijaya-R., G. (eds.) Cahaya Bahasa: A Festschrift in honour of Prof. I Gusti Made Sutjaja 247--258. Denpasar: Swasta Nulus. R Core Team. (2017) R: A Language and Environment for Statistical Computing (version 3.4.3). Vienna: R Foundation for Statistical Computing. Sahlgren, M. (2006) The word-space model. Using distributional analysis to represent syntagmatic and paradigmatic relations between words in high-dimensional vector spaces. Ph.D. thesis. Stockholm: Stockholm University. Samuels, M. L. (1972) Linguistic evolution with special reference to English. London and New York: Cambridge University Press. Doi 10.1017/CBO9781139086707
23 Sinclair, J. (1966) Beginning the study of Lexis. in Bazell, C.E., Catford, J.C., Halliday, M. A. K., and Robins, R.H. (eds.) In memory of J. R. Firth 410--430. Harlow: Longman. Stubbs, M. (2001) Words and phrases: Corpus studies of lexical semantics. Oxford: Blackwell. Taylor, J. R. (2003) Near Synonyms as co-extensive categories: ‘High’ and ‘Tall’ Revisited. Language Sciences 25: 263--284. Doi 10.1016/S0388-0001(02)00018-9 1 In this paper, the terms ‘synonym’/‘synonymy’ and ‘near-synonym’/‘near-synonymy’ are used interchangeably given that most scholars agree on the fact that absolute synonyms are uncommon in language. 2 The division of the corpus into four fifty year periods was undertaken due to the relatively low frequency of the near-synonyms, particularly perfumed and scented, which would have made the task of observing changes in their collocational preferences much more complicated if shorter time periods, such as decades or years, had been considered instead. 3 I would like to thank Professor Stefanie Wulff for suggesting the use of collocational networks for the analysis of my data as it has indeed proved to be a useful resource. 4 Brezina et al. (2015) and Baker (2017) make use of the software GraphColl to build collocational networks. However, in this paper, the networks were created by means of the igraph package (Csardi & Nepusz 2006) in R (R Core Team 2017). The reasons for this decision were twofold: (i) loading the corpus into the program required a lot of computation time due to its large size and (ii) R allows a greater flexibility in the customization of the visual parameters of the collocational networks. 5 A better fit is achieved by a three dimensional MDS solution (stress = 0.16). However, for the purposes of the present paper, only a two dimensional solution will be discussed as it allows a much clearer visualization of similarities in collocational behavior between the near-synonyms. 6 For more detailed information about the different classes and sub-classes distinguished, see http://ucrel.lancs.ac.uk/usas/.