scieee AI-readable full text Open interactive document viewer

Hidden communication aspects in the exponent of Zipf's law

Ferrer Cancho, Ramon

Abstract

This article focuses on communication systems following Zipf’s law, in a study of the rel-ationship between the properties of those communication systems and the exponent of the law. The properties of communication systems are described using quantitative measures of semantic vagueness and the cost of word use. The precision and the economy of a communication system is reduced to a function of the exponent of Zipf’s law and the size of the communication system. Taking the exponent of the frequency spectrum, it is demonstrated that semantic precision grows with the exponent, where-as the cost of word use reaches a global minimum between 1.5 and 2, if the size of the communication system remains constant. The exponent of Zipf’s law is shown to be a key aspect for knowing about the number of stimuli handled by a communication system, and determining which of two systems is less vague or less expensive. The ideal exponent of Zipf’s law, it is therefore argued, should be very slightly above 2.

Full text

Glottometrics 11, 2005, 98-119 Hidden communication aspects in the exponent of Zipf’s law Ramon Ferrer i Cancho1 Abstract. This article focuses on communication systems following Zipf’s law, in a study of the relationship between the properties of those communication systems and the exponent of the law. The properties of communication systems are described using quantitative measures of semantic vagueness and the cost of word use. The precision and the economy of a communication system is reduced to a function of the exponent of Zipf’s law and the size of the communication system. Taking the exponent of the frequency spectrum, it is demonstrated that semantic precision grows with the exponent, whereas the cost of word use reaches a global minimum between 1.5 and 2, if the size of the communication system remains constant. The exponent of Zipf’s law is shown to be a key aspect for knowing about the number of stimuli handled by a communication system, and determining which of two systems is less vague or less expensive. The ideal exponent of Zipf’s law, it is therefore argued, should be very slightly above 2. Keywords: Zipf´s law, frequency spectrum, exponent, precision, economy INTRODUCTION Word frequencies in human language arrange themselves according to what is known as Zipf's law. If P(f) is the proportion of words whose frequency is f in a given sample (e.g. a text), we say that a sample follows Zipf's law (Zipf, 1932, 1935, 1949) if β − ffP ~)(, (1) where β is the exponent of the law. We assume that β > 1. The previous equation appears as a straight line when P(f) is plotted on a logarithmic scale. Although different functions have been proposed for modelling P(f) (Chitashvili & Baayen, 1993; Tuldava 1996; Naranan & Basubrahmanyan, 1998), the basic trend described in simplified form by Eq. 1 appears to hold without exceptions in word frequencies. This article uses the functional form in Eq. 1 because its simplicity is extremely helpful for the analytical calculations discussed here. Typically, β ≈ 2 is found (Zipf, 1932, 1935, 1949) but significant deviations from that value have been reported in single author samples: • β > 2 in fragmented discourse schizophrenia. This type of speech is characterized by multiple topics and the absence of a consistent subject. The lexicon of such a text may be varied and chaotic (Piotrowski et al. 1995, Piotrowska et al., to appear). β ∈ [2.11,2.42] is found. Schizophrenic patients of this kind tend to be in the acute phase of the disease. 1 Address for correspondence: Ramon Ferrer i Cancho, Dip. di Fisica, Università ‘La Sapienza’, Piazzale A. Moro 5, ROMA 00185, ITALY. E-mail: ra[email protected]ma1.it Hidden communication aspects 99 •Values suspiciously above the ideal β = 2 have been found in nouns from single author samples. More precisely, β ∈ [2.15,2,32] (Balasubrahmanyan & Naranan, 1996). •1 < β < 2 in advanced forms of schizophrenia (Whitehorn & Zipf, 1943; Zipf, 1949; Piotrowski et al., 1995; Piotrowska et al., to appear). Texts are filled mainly with words and word combinations related to the patient’s obsessional topic. The variety of lexical units employed here is restricted and repetitions are many. β = 1.66 is reported in (Piotrowski et al. 1995; Piotrowska et al., to appear). • β = 1.6 in very young children (Brillouin, 1960; Piotrowski et al., 1995). Older children conform to the typical β ≈ 2 (Zipf, 1942). •Exponents larger than β ≈ 2 can be obtained as a result of deficient sampling from a text with the typical β ≈ 2 (Piotrowski et al., 1995; Piotrowska et al., to appear). Therefore, the exponents that are of interest here seem to be constrained to a very narrow domain, i.e. β ∈ [1.66,2.42] (Ferrer i Cancho, 2005b). Whether Zipf’s law can distinguish between acute and chronic schizophrenic patients is a matter of current research. The main message concerning schizophrenia here is that the disease shows exponents on both sides of the interval of variation in humans and that the value of the exponent may be related to the stage of the disease. Significant variations of β have also been found in multi-author samples (Piotrowski et al., 1994; Ferrer i Cancho, 2005d, Ferrer i Cancho & Solé, 2001; Montemurro, 2001; Montemurro & Zanette, 2002), particular word classes (Balasubrahmanyan & Naranan, 1996) and both (Ferrer i Cancho, 2005a). The focus of the present paper is communicative aspects of single individuals. Significant deviations in multi-author texts will not be considered. The aim of the present paper is to show the connection between the exponents and various types of quantitative measures suggesting that the variation of the exponent may be due to tuning the vagueness and the cost of word use. Most of the measures of vagueness and the measure of cost of word use that are employed here are defined using Shannon’s information theory (Ash, 1990). Support for the hypothesis of the strong association between Zipf’s law and communication, comes from recent models where Zipf’s law and/or the value of the exponent can be explained as the outcome of minimizing or constraining various standard information theory measures (Ferrer i Cancho, 2005a, 2005d; Ferrer i Cancho & Solé, 2003). THE MODEL We assume a general communication system mapping words to stimuli. We have a set of n words S = {s1, ..., si, ..., sn} and a set of m stimuli R = {r1, ..., rj, ..., rm}. We assume that words connect to stimuli to build their meaning. Word-stimuli associations are defined by a binary matrix A = {aij} where aij = 1 if si and rj are linked and aij = 0 otherwise. Let us consider in greater detail what is meant here by “stimuli”. Various experiments have shown that words are associated with the activation of different brain areas (Pulvermüller, 2003). Generally speaking, nouns tend to activate visual areas. Verbs tend to activate motor areas if the corresponding action can be performed by the individual and visual areas otherwise. The activated areas are associated to different types of stimuli experienced with the word. Let us take one of the definitions of the Webster's Revised Unabridged Dictionary (1913)2 for the word write: “to inscribe on any material by a suitable instrument”. In our view, the verb write is associated to the motor stimuli of the action of writing and the visual (tactile, olfactory,...) stimuli of the instruments used for writing. The construction of a complex meaning would 2 www.dict.org. Ramon Ferrer i Cancho 100 involve a structure combining diverse stimuli. From that point of view, a word in S does not refer to stimuli in R, but it is merely associated to them. We do not claim that words in S refer to stimuli in R via A although they may. We do not use the term reference because it is stronger than association. In our example, write can only refer to the motor stimuli of the action of writing. write cannot refer to the instrument used for writing, although it is associated with it. The action and the instrument are both stimuli involved in the construction of the complex meaning of the verb write. Defining “word meaning” is an open problem in various fields ranging from cognitive science to philosophy. In our view, complex meaning would emerge from the interaction between different stimuli. Referential associations may be a subset or a higher order structure of the associations defined by A. It makes sense to assume that the more stimuli a word is associated with, the higher the probability of using that word. It is important to note that when we say that a word has no meaning we usually mean that it has no referential power. Nonetheless, if a word lacks referential power it does not imply that it has no associations with stimuli. Our framework is not inconsistent with the existence of words with no apparent meaning, such as prepositions, conjunctions or articles. Real words with no apparent meaning are the words with the highest frequencies. The five most frequent word in the British National Corpus3, a large collection of text samples, are the, of, and, to and a. The framework here predicts that the most frequent words would have the largest number of connections with stimuli in R. Since those connections are merely associative (and not always referential) there is no inconsistency here. Furthermore, that high number of associations may underlie those words’ lack of referential power or “meaning”. The uncertainty associated with the interpretation of highly connected words is so large (Ferrer i Cancho, 2005c,e) that reference cannot be effectively attributed. Words with no meaning may have two different origins: words that have no links, and words having too many links. It makes sense to suppose that words with no meaning may have an excess of connections rather than a lack thereof, although those connections could be very weak given the high frequency of the words involved (Ferrer i Cancho & Reina, 2002). A first approach to the semantic vagueness of the set of words could be the average number of links per word, that is, <k>. The number of links of a vertex (e.g. a word) is called “degree” in standard graph theory (Bollobás, 1998), so <k> is the mean signal degree. The idea behind the relationship between <k> and vagueness is very simple: the more links a word has, the higher the number of possible interpretations in the context where it appears. The higher the value of <k>, the lower the precision of the communication system. Hereafter we assume that ‘precision’ and ‘vagueness’ have opposite meaning. H(R| S), that is, the average uncertainty (or entropy) associated with the interpretation of every stimulus once the corresponding word is known, is a more precise measure, from the information theory point of view. That measure is defined as ∑ = =n i ii sRHspSRH 1 )|()()|(, (2) where H(R|si) is the uncertainty (or entropy) associated with the interpretation of si , and p(si) is the probability of using si. H(R|S) is the average uncertainty associated with the interpretation of the words in S. The higher the value of H(R|S), the lower the precision of the communication system. Since H(R|S) is mathematically a hard function to manipulate, a simpler version has been considered (Ferrer i Cancho, 2005a): 3 www.natcorp.ox.ac.uk Hidden communication aspects 101 ∑ = =n i i sRH n SRG 1 )|( 1 )|(. (3) G(R|S) is the amount of uncertainty per word associated with the interpretation of the words in S. G(R|S) and H(R|S) have similar properties. The upper and lower bounds are the same, i.e. .log)|(),|(0 mSRHSRG ≤≤ G(R|S) has the virtue of allowing Zipf’s law (Eq. 1) to be derived using the maximum entropy principle (Ferrer i Cancho, 2005a). A possible approach to the cost of word use is H(S), the entropy of the set of words (Ferrer i Cancho, 2005a,d; Ferrer i Cancho & Solé, 2003). This is defined as ∑ = −= n i ii spspSH 1 )(log)()(. (4) Support for H(S) as a measure of the cost of word use comes from two different sources. Firstly, it is known in psycholinguistics that the availability of a word in various linguistic tasks is correlated with the frequency of that word. The availability of a word obeys the socalled word frequency effect, i.e. the more frequent the word, the higher its availability (Akmajian et al., 1995; Carroll, 1994). The best availability is achieved when a word has probability one, which means that the rest of the words have probability 0. In that case, H(S) = 0. The worst case is when all words are equally likely, that is when p(si) = 1/n for each word. In that case, H(S) = log n. That is, H(S) is a good measure of cost of word use. Second, the use of H(S) as a measure of cost is justified by models leading to Zipf’s law when the information transfer is maximized while H(S) is minimized (Ferrer i Cancho, 2005d; Ferrer i Cancho & Solé, 2003). Those models explain Zipf’s law as the outcome of maximizing the communicative efficiency, but saving as much cost as possible. Interestingly, if those models replace H(S) with the effective vocabulary size (i.e. the proportion of words with at least one link) as a measure of cost of word use, Zipf’s law is not reproduced. Vocabulary size is an important ingredient for the cost of a communication system (Köhler, 1986, 1987) but it does not seem to be essential for Zipf’s law. We may assume that the p(si), the probability of occurrence of word si, is proportional to ki, the number of connections of si, that is M k sp i i=)(, (5) where ∑ = =m j iji ak 1 (6) and ∑ = =n i i kM 1 (7) (as in Ferrer i Cancho, 2005a,b,d). Eq. 5 contains the basic assumption that words are used according to the number of semantic associations they have. Eq. 5 states that a word is used with a probability proportional to the number of stimuli it is associated with. Eq. 5 is chosen Ramon Ferrer i Cancho 102 for simplicity and its predictive power: it can explain the interval of variation of β in human language (Ferrer i Cancho, 2005b). We may also assume that P(k), the proportion of words with k links obeys .~)( β− kkP (8) Zipf’s law (Eq. 1) is recovered from Eqs. 5 and 8 (Ferrer i Cancho, 2005a,b). We assume a fixed P(k) or P(f), given the surprising tendency of human language to arrange according to Zipf’s law even in atypical cases. Although there is variation in β for (human) words, the basic trend described by Eq. 1 has essentially no exceptions, as far as we know. From Eq. 5 and ∑ = =m j jii rspsp 1 ),()((9) it follows that the probability that si and rj are associated by the communication system is M a rsp ij ji =),(. (10) We may write Eq. 5 as >< === kn k M k kksp ii )|(, (11) where <...> is the expectation operator over P = {P(1),...,P(k),...P(m)} and P(k) is the proportion of words having k connections. Assuming Eq. 8, the uncertainty (or entropy) associated to the interpretation of si becomes H(R|si) = log k if ki = k (Ferrer i Cancho, 2005a). Thus, H(R|S) in Eq. 2 becomes (Ferrer i Cancho, 2005b) >< >< =k kk SRH log )|((12) and G(R|S) becomes (Ferrer i Cancho, 2005a) (|) logGRS k=< >. (13) Figs. 1-3 show that <k>, G(R|S) and H(R|S) are decreasing functions of β for different values of m. The three functions grow with m for a given value of β . Hidden communication aspects 103 Fig 1. <k>, the mean word degree, versus β , the exponent of the frequency spectrum of Zipf’s law. Series from the bottom to the top are for m = 10, m = 102, m = 103, m = 104 and m = 105 (solid lines). The approximated expected curve for m → ∞ is also shown (dashed line). Fig 2. G(R|S), the uncertainty per word associated with the interpretation of every word, versus β , the exponent of the frequency spectrum of Zipf’s law. Series from the bottom to the top are for m = 10, m = 102, m = 103, m = 104 and m = 105 (solid lines). The approximated expected curve for m → ∞ is also shown (dashed line). Natural logarithms were used. Ramon Ferrer i Cancho 104 Fig 3. )|( SRH , the average uncertainty associated with the interpretation of every word, versus β , the exponent of the frequency spectrum of Zipf’s law. Series from the bottom to the top are for m = 10, m = 102, m = 103, m = 104 and m = 105 (solid lines). The approximated expected curve for m → ∞ is also shown (dashed line). Natural logarithms were used. Here we will define vagueness as the opposite of precision. <k>, G(R|S) and H(R|S) are inverse measures of precision and direct measures of vagueness. As for cost of word use, substituting Eq. 11 into Eq. 4 we get, >< >< −><= k kk knSH log )log()(, (14) where M = n<k> is the total amount of links. Knowing Eq. 12, Eq. 14 can be written as )|()log()( SRHknSH −><= . (15) Fig. 4 shows H(S) for n = 103 and different values of m. H(S) decreases as m grows for a fixed value of β whereas the vagueness measures behave inversely. H(S) has a minimum at β = β *, a critical value of β , such that 1 < β * < 2 for the values of m that we used here. Notice that although the exact value of H(S) depends on n, β * depends only on m (recall Eq. 15). Fig. 5 shows β * versus m. Hidden communication aspects 105 Fig. 4. H(S) versus β where H(S) is the entropy of the set of words S and β is the exponent of the frequency spectrum of Zipf’s law. Series from top to the bottom are for m = 10, m = 102, m = 103, m = 104 and m = 105 (solid lines). n = 103 is used in all cases, although the point where the minimum H(S) is reached is independent of n. The approximated expected curve for m → ∞ is also shown (dashed line). Natural logarithms were used. Fig. 5. β *, the value of β minimizing H(S), versus m. β is the exponent of the frequency spectrum representation of Zipf’s law, H(S) is the entropy of the set of words S and m is the number of stimuli. A. β * versus m calculated without integrals using Eq. 15 (solid line) and with integrals using Table 1 (dashed line) B. β * versus m calculated with integrals till very large values of m. Ramon Ferrer i Cancho 106 RESULTS We can obtain formulae for the measures of vagueness and cost using approximation with integrals (see Appendix A). Results are summarized in Table 1. The measures of vagueness are functions of β and m whereas H(S) is a function of β , m and n. When m → ∞, we can obtain simple mathematical expressions in particular domains of β (Table 2). Table 1 Summary of the relationship between Zipf’s law and communication measures Function Information theory Approximation β ≠ 1 and β ≠ 2 )1)(2( )1)(1( 1 2 −− −− − − β β β β m m β = 1 m m log 1−. >< k - β = 2 m m 1 1 log − β ≠ 1       − +         − − − − − ββ β β 1 1 1 1 log 1 11 1mm m >< klog G(R|S) β = 1 2 logm β ≠ 1 and β ≠ 2       − +         − − − − − ββ β β 2 1 2 1 log 1 12 2mm m β = 1 1 1)1(log − +− m mm >< >< k kk log H(R|S) β = 2 2 logm <k> is the mean word degree, G(R|S) is the uncertainty per word associated with the interpretation of every word and H(R|S) is the average uncertainty associated with the interpretation of every word. <…> is the expectation operator over k. m is the number of stimuli. β is the exponent of the power spectrum of Zipf’s law. <k> when β > 2 and m → ∞ is shown as dashed line in Fig. 1. G(R|S) when β > 1 and m → ∞ is shown as dashed line in Fig. 2. H(R|S) when β > 2 and m → ∞ is shown as a dashed line in Fig. 3. H(S) when β > 2 and m → ∞ is shown as a dashed line in Fig. 4. When β > 1 and m → ∞, we have G(R|S) = 1/( β - 1). Zipf’s law can be alternatively defined as P(i) ~ i-α, where P(i) is the frequency of the i-th most frequent word in a piece of text and α = 1/( β - 1) (Chitashvili & Baayen, 1993). Notice that G(R|S) = 1/( β - 1) tells us that α = G(R|S). The value of β where β = α (and thus β = G(R|S)) can be calculated solving β = 1/(β - 1), which has two solutions Hidden communication aspects 113 When γ ≠ 1, we have γ γ γ γ − − =≈ − − ∫1 1 ),( 1 1 0 m dkkmF m(26) and       − +         − − − =≈ −− ∫ γγγ γγγ 1 1 1 1 log 1 1 log),( 1 1 1mmdkkkmF m . (27) When γ = 1, we have mmF log),( 0≈ γ (28) and 2 log ),( 2 1 m mF ≈ γ .(29) Table 4 Summary of definitions of different functions and their relationships Function Information theory Definition ),( mFx γ - ∑ = − m k xkk 1 log γ c - 1/F0( β ,m) >< k - F0( β -1,m)/F0( β ,m) >< klog G(R|S) F1( β ,m)/F0( β ,m) >< >< k kk log H(R|S) F1( β -1,m)/F0( β -1,m) Table 4 summarizes the relationship between the auxiliary function Fx( γ ,m) and the functions of vagueness. Thus, we may write, 1 0 1 (1,) m k kck cF m ββ − = <>= = − ∑, (30) where c is the normalization constant of Eq. 8, defined as ),( 11 0 1 mF k cm k β β == ∑ = − . (31) Eqs. 30 and 31 together give Ramon Ferrer i Cancho 114 0 0 (1,) (,) Fm kFm β β − <>= . (32) If β ≠ 1, substituting Eq. 26 on Eq. 31 gives (Cohen & Havlin, 2002) 1 1 1− − ≈− β β m c. (33) If β = 1, substituting Eq. 28 on Eq. 31 gives m clog 1 ≈.(34) When β ≠ 1 and β ≠ 2, substituting F0( β - 1,m) with Eq. 26 with γ = β - 1 and F0( β ,m) by Eq. 26 with γ = β into Eq. 32 we obtain 2 1 (1 )( 1) (2 )( 1) m km β β β β − − −− <>≈ −− . (35) When β = 1, substituting F0( β - 1,m) with Eq. 26 (with γ = β - 1 = 0) and F0( β ,m) by Eq. 28 (since γ = β = 1) into Eq. 32 we obtain 1 log m km − <>≈ . (36) When β = 2, substituting F0( β - 1,m) with Eq. 28 (since γ = β - 1 = 1) and F0( β ,m) by Eq. 26 (with γ = β = 2) into Eq. 32 we obtain log 1 1 m k m <>≈ − . (37) When β > 2 and m → ∞ , Eq. 35 becomes 2 1 − − >≈< β β k. (38) The previous equation is shown as a dashed line in Fig. 1. G(R|S) can be written as 1 10 (,) (|) log log (,) mFm GRS k c k k Fm ββ β − =< >= = ∑. (39) Hidden communication aspects 115 When β ≠ 1, substituting F1( β ,m) by Eq. 27 and F0( β ,m) by Eq. 26 (both with γ = β ) into Eq. 39 we get       − +         − − − ≈− − ββ β β 1 1 1 1 log 1 1 )|( 1 1mm m SRG . (40) When β = 1, substituting F1( β ,m) by Eq. 29 and F0( β ,m) by Eq. 28 (since γ = β = 1 in both cases) into Eq. 39 we get 2 log )|( m SRG ≈.(41) When β > 1 and m → ∞ , Eq. 40 becomes 1 1 )|( − ≈ β SRG . (42) (as in Ferrer i Cancho, 2005a). The previous equation is shown as a dashed line in Fig. 2. As for H(R|S), the numerator in Eq. 12 can be expressed as 1 1 1 log log ( 1, ) m kk ck kcF m ββ − <>= =− ∑.(43) Substituting Eqs. 30 and 43 into Eq. 12 we obtain ),1( ),1( )|( 0 1 mF mF SRH − − = β β . (44) If β ≠ 1 and β ≠ 2, substituting F1( β - 1,m) with Eq. 27 and F0( β - 1,m) with Eq. 26 (with γ = β - 1 in both cases) into Eq. 44 we get       − +         − − − ≈− − ββ β β 2 1 2 1 log 1 1 )|( 2 2mm m SRH . (45) If β = 1, substituting F1( β -1,m) with Eq. 27 and F0( β -1,m) with Eq. 26 (both with γ = β - 1 = 0) into Eq. 44 we get 1 1)1(log )|( − +− ≈m mm SRH . (46) If β = 2, substituting F1( β - 1,m) with Eq. 29 and F0( β - 1,m) with Eq. 28 (since γ = β - 1 = 1 in both cases) into Eq. 44 we get 2 log )|( m SRH ≈.(47) Ramon Ferrer i Cancho 116 Eq. 45 with β > 2 and m → ∞ becomes 2 1 )|( − ≈ β SRH . (48) The previous equation is shown as a dashed line in Fig. 3. When β > 2 and m → ∞ , substituting Eqs. 38 and 48 into Eq. 15 we obtain 2 1 2 )1( log)( − −         − − = ββ β n SH . (49) The previous equation is shown as a dashed line in Fig. 4. Since Eq. 49 is an approximation, and <k> and H(R|S) diverge for β = 2, it is convenient to keep β >> 2. APPENDIX B Here we give an outline of proof for Lemma 1 and 2. LEMMA 1. Suppose we have two communication systems A and B, with exponents β A (or α A) and β B (or α B) with β A, β B > 1, and the number of stimuli is mA and mB, respectively. If β A < β B (or α A > α B) and mA ≥ mB then B is a strictly more precise communication system than A. Proof: The proof is based on H(R|S), the reference measure for word vagueness. Assuming β A, β B > 1 we warrant that P(k) is a probability distribution even when m → ∞. In general, there are only four situations: 1) mA and mB are finite. It is easy to see from the approximate equations in Table 1 Appendix A (recall also Fig. 3) that H(R|S) is a monotonically decreasing function of β (when β > 0) when mA and mB are finite. Given a particular β , the larger the value of m, the larger the value of the measure. 2) mA is finite and mB is not. That contradicts mA ≥ mB. 3) mA is infinite and mB is not. That contradicts the notion that A is a communication system if β A ≤ 2. β A > 2 must be satisfied and thus we can proceed as in 1). 4) mA and mB are infinite. That contradicts the notion that A and B are communication systems if β A ≤ 2 and/or β B ≤ 2. β A, β B > 2 must be satisfied and thus we can proceed as in 1). LEMMA 2. If we have two communication systems A and B, with exponents β A (or α A) and β B (or α B) with β A, β B > 0, their number of stimuli is mA and mB, and their lexicon size is nA and nB, respectively. We assume mA, mB, nA and nB are finite. If β * ≤ β A < β B (or α * > α A > α B with α * = 1/( β * - 1)) and mB ≥ mA and nA ≥ nB then it follows that A is a more economical communication system than B. Proof: α * = 1/( β * - 1) comes from the equivalence between, α , the exponent of the frequency versus rank representation and β , the exponent of the frequency spectrum (Chitashvili & Baayen, 1993). It is easy to show from the approximate equations in Table 1 (recall Figs. 4-5) that H(S), the measure of cost, is a monotonically increasing function of β when β > β * and Hidden communication aspects 117 that given a particular β , the larger the value of m, the lower the cost, and that, the larger the value of n, the larger the cost. REFERENCES Allen, H. A., Liddle, P.F. & Frith, C. D. (1993). Negative features, retrieval processes and verbal fluency in schizophrenia. British Journal of Psychiatry 163, 769-775. Akmajian, A., Harnish, R. M., Demers, R. A. & Farmer, A. K. (1995). Linguistics. An introduction to language and communication. Cambridge, MA: MIT Press. Ash, R. B. (1990). Information Theory. New York: Dover Publications, Inc. Balasubrahmanyan, V. K. & Naranan, N. (1996). Quantitative linguistics and complex systems studies. Journal of Quantitative Linguistics 3, 177-228. Bollobás, B. (1998). Modern graph theory. New York: Springer. Brillouin, L. (1960). Science and theory of information (Russian translation). Moscow: Gosudarstvennoe Izdatel´stvo Fiz.-Mat. Literatury. Carroll, D. W. (1994). Psychology of language. Chapter 9, Conversational interaction. pp. 242-248. Pacific Grove, California: Brooks/Cole Publishing Company. Chitashvili, R. J. & Baayen, R. H. (1993). Word frequency distributions. In: L. Hřebíček, G. Altmann (eds.), Quantitative text analysis: 54-135. Trier: Wissenschaftlicher Verlag Trier. Cohen, R. & Havlin, S. (2002). Scale-free networks are ultrasmall. Physical Review Letters 90, 057801. Cormen, T. H., Leisserson, C. E. & Rivest, R. L. (1990). Introduction to algorithms. Cambridge: MIT Press. Damasio, A. (1999). The Scientific American Book of the Brain. New York: Scientific American. See http://hypertextbook.com/facts/2002/AniciaNdabahaliye2.shtml. Devitt, M. & Sterelny, K. (1999). Language and reality: an introduction to the philosophy of language. Cambridge, MA: MIT Press. Dunlap, R. A. (1997). The golden ratio and Fibonacci series. Singapore: World Scientific Publishing. Elvevåg, B., Weinstock, M., Akil, M., Kleinman, J.E. & Goldberg, T. E. (2001) A comparison of verbal fluency tasks in schizophrenic patients and normal controls. Schizophrenia Research 51, 119-126. Elvevåg, B., Fisher, J. E., Gurd, J.M. & Goldberg, T. E. (2002). Semantic clustering in verbal fluency: schizophrenic patients versus control participants. Psychological Medicine 32, 909-917. Elvevåg, B., Storms G., Heit, E. & Goldberg, T. (2005). Category content and structure in schizophrenia: an evaluation using the instantiation principle. Neuropsychology 19, 371380. Ferrer i Cancho, R. (2005a). Decoding least effort and scaling in signal frequency distributions. Physica A 345, 275-284. Ferrer i Cancho, R. (2005b). The variation of Zipf’s law in human language. European Physical Journal B 44, 249-257. Ferrer i Cancho, R. (2005c). The consequences of Zipf’s law for syntax and symbolic reference. Proceedings of the Royal Society of London B 272, 561-565. Ferrer i Cancho, R. (2005d). Core and peripheral lexicon from word length optimization. Submitted to the Journal of Quantitative Linguistics. Ferrer i Cancho, R. (2005e). Zipf’s law from a communicative phase transition. Submitted to European Physical Journal B. Ramon Ferrer i Cancho 118 Ferrer i Cancho, R. and Reina, F. (2002). Quantifying the semantic contribution of particles. Journal of Quantitative Linguistics, 9, 35-47. Ferrer i Cancho, R. & Solé, R. V. (2001). Two regimes in the frequency of words and the origins of complex lexicons: Zipf's law revisited. J. Quantitative Linguistics, 8, 165-173. First appeared as Santa Fe Institute Working Paper 00-12-068. Ferrer i Cancho & Solé, R. V. (2003). Least effort and the origins of scaling in human language. Proceedings of the National Academy of Science USA 100, 788-791. Ghyka, C. M. (1927). Esthétique des proportions dans la nature et dans les arts. Paris: Gallimard. Gribbin, J. (1986). In search of the big bang: Quantum physics and cosmology. New York: Bantam Books. See http://www.sunspot.noao.edu/sunspot/pr/answerbook/universe.html. Goldberg, T. E., Dodge, M., Aloia, M., Egan, M. F. & Weinberger, D. R. (2000). Effects of neuroleptic medications on speech disorganization in schizophrenia: biasing associative networks towards meaning. Psychological Medicine 30, 1123-1130. Johnson, C., Davis. H. & Macken, M. (1999). Symbols and structure in language-acquisition. In: Lock, A. & Peters, C.R. (eds.), Handbook of Human Symbolic Evolution: 686-746. Oxford: Blackwell. Kripke, S. A. (1980). Referring to artifacts. Phylosophical Review 89, 109-114. Köhler, R. (1986). Zur linguistischen Synergetik: Struktur und Dynamik der Lexik. Bochum: Brockmeyer. Köhler, R. (1987). System theoretical linguistics. Theoretical Linguistics 14, 241-257. Macbeth, D. (1995). Names, natural kind terms and rigid designation. Philosophical Studies 79, 259-281. McCowan, B., Hanser, S. F. & Doyle, L. R. (1999). Quantitative tools for comparing animal communication systems: information theory applied to bottlenose dolphin whistle repertoires. Animal Behavior 57, 409-419. McCowan, B., Doyle, L. R. & Hanser, S. F. (2002). Using information theory to asses the diversity, complexity and development of communicative repertoires. Journal of Comparative Psychology 116, 166-172. Montemurro, M. (2001). Beyond the Zipf-Mandelbrot law in quantitative linguistics. Physica A 300, 567-578. Montemurro, M. & Zanette, D. (2002). Frequency-rank distribution in large text samples: phenomenology and models. Glottometrics 4, 87-98. Naranan, S. & Balasubrahmanyan, V. K. (1998). Models for power law relations in linguistics and information science. Journal of Quantitative Linguistics 5, 35-61. Newman, M. E. J. (2005). Power laws, Pareto distributions and Zipf's law. Contemporary Physics, in press. cont-mat/0412004. Piotrowska, X., Pashkovska, W., & R. Piotrowski, R. (to appear). Pathological text and its statistical parameters. Piotrowski, R.G., Pashkovskii, V.E., Piotrowski, V.R. (1995). Psychiatric linguistics and automatic text processing. In: Automatic Documentation and Mathematical Linguistics, 28(5), 28-35. [First published in Naučno-Techničeskaja Informacija, Serija 2, Vol. 28, No. 11. pp. 21-25, 1994]. Pulvermüller, F. (2003). The neuroscience of language. On brain circuits of words and serial order. Cambridge: Cambridge University Press. Tuldava, J. (1996). The frequency spectrum of text and vocabulary. Journal of Quantitative Linguistics 3, 38-50. Whitehorn, J.C. & Zipf, G.K. (1943). Schizophrenic language. Archive of neurology and psychiatry 49, 831-851. Hidden communication aspects 119 Zipf, G. K. (1932). Selected studies of the principle of relative frequency in language. Cambridge, MA: Harvard University Press. Zipf, G. K. (1935). The psycho-biology of language: an introduction to dynamic philology. Boston, MA: Houghton-Mifflin. Zipf, G. K. (1942). Children's speech. Science 96, 344-345. Zipf, G. K. (1949). Human behavior and the principle of least effort. Reading: AddisonWesley.