Porlex, a lexical database in european portuguese
Full text
PORLEX DATABASE IN EUROPEAN PORTUGUESE 1 Referência: Gomes, I., & Castro, S. L. (2003). Porlex, a lexical database in European Portuguese. Psychologica, 32, 91-108. Porlex, a lexical database in European Portuguese Inês Gomes1 and São Luís Castro2 1Professor Auxiliar, Universidade Fernando Pessoa, Porto, Portugal 2Professor Associado, Faculdade de Psicologia e Ciências da Educação, Universidade do Porto, Porto, Portugal Address for correspondence: São Luís Castro FPCE-Universidade do Porto Rua do Campo Alegre, 1021 P 4169 - 004 Porto Portugal Phone: +351 22 607 9756; Fax: +351 22 607 9725; E-mail: [email protected] Running head: PORLEX DATABASE IN EUROPEAN PORTUGUESE
PORLEX DATABASE IN EUROPEAN PORTUGUESE 2 Abstract This paper presents a tool for research in the psychology of language, a computerized lexical database in European Portuguese. Porlex was built on the basis of a middle sized adult lexicon, and provides orthographic, phonological, phonetic, part-of-speech, and neighborhood information for about 30 000 words (uninflected content words and inflected function words). Frequency was included whenever possible. After highlighting the role of lexical databases for experimental research on language, we give an overview of the sources and contents of Porlex 1.0, with a special emphasis on the 44 different types of information it provides. A brief characterization of the corpus is also included. Key words: Database, Lexicon, European Portuguese, Psychology of Language
PORLEX DATABASE IN EUROPEAN PORTUGUESE 3 Lexical Databases as Research Tools Contemporary research about language can hardly be done without an elaborate and sometimes laborious process of selection of stimuli. Words, syllables or phonemes that are presented to the participants have to be carefully chosen not only on the basis of the variables under scrutiny, but also depending on a number of other characteristics that may affect the perceptual and cognitive processes involved in the preparation of the response. Words that differ on semantic category may have to be matched on frequency, imageability, or syllabic structure; syllables that differ on structure may have to be chosen according to their relative frequency in the language or potential position in the word; phonemes that differ on word position may have to be matched on orthographic consistency. In order to accomplish this, an impressive amount of knowledge is required, namely knowledge on aspects of language that are, or are presumed to be, cognitively relevant. A further source of complication is that language as such is a highly abstract entity. In practice, what the researcher deals with are specific languages, such as English, Portuguese, French, etc. Thus, the need of a reliable source of information that is both language specific and cognitively founded has become apparent with the advances of experimental and neuropsychological language research. Computerized lexical databases are a response to this need. Over the last years, several databases have been developed. One of the first was the MRC Psycholinguistic Database (Coltheart, 1981), that gathered psycholinguistic measures on 150 000 English words. Other languages followed; for example, French, with Brulex (Content, Mousty, & Radeau, 1990) and Lexique (New, Pallier, Ferrand, & Matos, 2001), Dutch and German with Celex (Baayen, Piepenbrock, & Gulikers, 1995), and Spanish (Piñeiro & Manzano, 2000). Most were developed from adult vocabularies, but others are based on children's lexica (e.g., Piñeiro & Manzano, ib., Lambert & Chesnet, 2001).
PORLEX DATABASE IN EUROPEAN PORTUGUESE 4 Typically, these databases provide orthographic, phonological, grammatical and frequency information on either lemmas (uninflected wordforms) or inflected wordforms, but they vary widely on format and size. For example, Brulex has ca. 30 000 lexical entries, most of them lemmas (verbs in the infinitive; content words in the singular form), and some inflected words (articles and pronouns; non-homophonic plural and feminine forms, e.g., journal/journaux; petit/petite); 29 informations are given for each entry, ranging from phonological transcription to neighborhood count. Lexique is composed of three related files, one based on written inflected wordforms (ca 130 000 entries), another on lemmas (ca 55 000 entries), and the third compiling frequency measures of letters, bigrams, trigrams, phonemes and syllables of the inflected wordforms. Celex is a trilingual database, that comprises three separate lexica in Dutch, English and German: lemmas, wordforms, and the corpus lexicon that combines both. The number of entries depends on type of lexicon and language, from ca 52,000 German lemmas to 380,000 Dutch wordforms, with English in-between – ca 53,000 lemmas and 160,000 wordforms. Approximately 950 different types of linguistic and psycholinguistic informations are provided for each entry (Burnage, 1990; Piepenbrock, 2001). In a recent review, Nascimento, Rodrigues and Gonçalves (1996) listed 26 corpora in European Portuguese (cf. also http://www.clul.ul.pt/). About half are lexical databases that contain written words; some include morpho-syntactic annotations, others phonetic transcriptions. However, none provides the range of informations that are required for cognitively-oriented research on language (cf. Appendix A). This situation lead us to develop Porlex. Porlex: Sources, Overview, Corpus and Variables Porlex is a computerized lexical database in European Portuguese designed as a tool for research in the psychology of language. It was built on the basis of a middle sized adult
PORLEX DATABASE IN EUROPEAN PORTUGUESE 5 lexicon, and provides orthographic, phonological, part-of-speech, and neighborhood information for each entry. Its current version, Porlex 1.0, contains 29,238 different words and 44 types of information. Porlex 1.0 is in Excel format (Microsoft Corporation, 1998). It is available for non-commercial purposes upon request to the authors. Sources for Porlex Lexical entries, grammatical and morphological classification, syllabication, phonetic transcription and frequency information were collected from various sources. There are listed on a separate section of the References, and will be briefly reviewed here. Porlex words come from the Dicionário Universal Fundamental (Texto Editora, 1998), that was selected because of its size. In the process of compiling the remaining source informations, it became clear that the final selection of the lexical entries had to be fine tuned. For that purpose, we used Porto Editora (Costa & Melo, 1997) and Cândido de Figueiredo dictionaries (1996), as well as the grammars of Cunha and Cintra (1987), Mateus, Brito, Duarte and Faria (1989), and Vilela (1995). The grammars were specifically used to enter inflected pronouns and contractions, as well as prepositions and conjunctions. Morphological classification and grammatical class were also compiled from these sources. Syllabication of the orthographic wordforms was marked according to Michaelis dictionary (Melhoramentos, 1998). Phonetic transcriptions were based on the pocket Langenscheidts Portuguese dictionary (Irmen & Kollert, 1995) and on Vilela (1991). Frequency information comes from Nascimento, Marques and Cruz (1987) and Nascimento, Rivenc and Cruz (1987), the only source of frequency information available to us during the period while Porlex was being compiled. The "Dicionário da Língua Portuguesa Contemporânea" (Academia das Ciências de Lisboa, 2001), that includes phonetic transcription, and CORLEX, a more recent source of frequency information (Nascimento, Casteleiro, Marques, Barreto, & Amaro, n.d.) were available shortly afterwards1.
PORLEX DATABASE IN EUROPEAN PORTUGUESE 6 Porlex entries were inserted automatically whenever possible. However, because we were unable to find source information in compatible electronic format, about a third of these was typed in manually. For example, at the time Porlex was started there were no Portuguese middle-sized dictionaries that included phonetic transcription of the words, and these had to be entered individually. A comprehensive survey of Porlex contents follows. Overview From the source informations, we extracted automatically almost all of the remaining Porlex contents (cf. Table 1). Some computations, like the number of letters in a word, were performed on a single entry. Others were structural, in that they explored the relations between words; these were computed on multiple entries. ____________________ Insert Table 1 about here ____________________ The orthographic representation is a good starting point for a lexical database, because in all likelihood it is the most robust way to represent a word. Being based on widely accepted conventions that are hardly ever updated, the written form of the word rarely changes. Accordingly, Porlex is based on the orthographic wordforms, that dictate the alphabetical order by which words are organized. Several informations are specifically related to this wordform: whether it contains diacritics, how it is divided into syllables, and how many letters and syllables it contains (Variables 3, 10, 11 and 16, respectively), among others. The representation of the spoken word, however, is more prone to variability. In Porlex, we present two transcriptions for each word, the phonetic wordform that differentiates between allophones (e.g., different symbols for the first and last segment of lua and sal, respectively), and the phonemic wordform where that level of phonetic detail is absent (see
PORLEX DATABASE IN EUROPEAN PORTUGUESE 7 next section for a fuller explanation). The symbols used in the transcriptions are presented in Appendix B. Due to the lack of a well established source for the transcription of spoken words in Standard European Portuguese, and also because it is a useful information as such, Porlex includes a variable that signals the words for which more than one broad phonetic transcription is acceptable (cf. infra Variable 6, Variant Pointer), and the alternative phonetic transcriptions are given as separate entries (for a maximum of three different alternatives, cf. Variables 40 to 42). Porlex gives additional informations about the characteristics of spoken words; for example, how many schwas they include, how they are divided into syllables, where is the stressed syllable; also a pointer for ambisyllabicity, the length in number of phonemes or of syllables, and two different types of phonetic patterns, one based on a gross classification of speech sounds, another more distinctive. For some of these, the difference between phonetic and phonemic wordforms is irrelevant. In such cases, we use the label phonological (cf. Variables 14 and 17). Irrespective of whether their form is written or spoken, words fall into different typeof-speech classes. Porlex provides this information in Variables 7 and 8, where words are classified according to grammatical class, and as open or closed, respectively. The grammatical gender of the word is also given; because nouns and adjectives are uninflected, we added a variable that signals whether gender inflexion is permissible and another that marks plural forms (Variables 19 to 21). Porlex also provides structural measures on the similarities between words, that were extracted by comparing the characteristics of a given lexical entry with the remaining Porlex entries. One set of such measures deals with neighborhood similarities: neighborhood densities, uniqueness points and the listing of the neighbors of a given target were computed for orthographic, as well as for phonetic and for phonemic wordforms (Variables 23 to 31). Neighbors were defined according to Luce (1986; also Charles-Luce & Luce, 1990), as words
PORLEX DATABASE IN EUROPEAN PORTUGUESE 8 that differ on a single phoneme either by substitution or by deletion/addition when relative serial position of the segments is maintained, and uniqueness points according to MarslenWilson and Tyler (1980; also Marslen-Wilson, 1990) as the position of the segment that discriminates a word from its neighbors, counted from the beginning. The second set of structural computations involved the comparison between orthographic and phonetic wordforms in order to extract the number of nonhomophonic homographs, homographic homophones and nonhomographic homophones, that are presented in Variables 32 to 34. For a detailed presentation of the computational procedures, see Gomes (2001). Corpus As a tool for research in the psychology of language, Porlex should incorporate words presumably represented in the mental lexicon. This thought guided the criteria that were adopted to fine-tune the selection of lexical entries into Porlex. For practical reasons we had to start with lemmas rather than wordforms, but in the case of articles and pronouns should feminine or plural forms be left out, like, say, the word 'we'? Based on the well established distinction between content vs. functional words (e.g., Fromkin & Rodman, 1998; Segalowitz & Lane, 2000), we decided to include the inflected forms of functional words, and as many different forms of this type of words as possible. Grammars were chosen as the primary source for the inclusion of all forms of articles, pronouns, prepositions, contractions and conjunctions into Porlex. For the remaining word types, dictionaries were used as the source for Porlex entries. ____________________ Insert Table 2 about here ____________________
PORLEX DATABASE IN EUROPEAN PORTUGUESE 9 The main characteristics of the resulting corpus are summarized in Table 2. Nouns account for 56% of the words; adjectives and verbs are the second largest categories, each with ca 20% of the corpus. Even with the criterion of including all forms (inflectionally and lexically) of articles, pronouns, contractions, prepositions and conjunctions, and only the noninflected forms of nouns, adjectives and verbs, the difference in the number of entries for these categories is huge. Adverbs account for 4% of Porlex. Almost 90% of them are derived from adjectives by addition of the suffix –mente (equivalent to –ly). For this reason (cf. Azuaga, 1996), they were classified as open class words, together with nouns, adjectives and verbs. These account for 98% of the corpus. There is an almost even distribution of nouns into feminine vs. masculine forms, and ca 2% of them are in the plural form. Variables In this section, we present a synopsis of the different variables in Porlex. For each, we show Number, Code Name and Name proper, Values (examples for string type variables; codes for classification variables; and range for numeric variables), and a brief explanatory definition (Concept). 1. # - Number VALUES: 1, 2, . . ., 29 238. CONCEPT: Entry number according to alphabetical order of the orthographic wordforms. Homographs are ordered by grammatical class (see variable 7, Grammatical Class) whenever possible. 2. Orto - Orthographic Wordform VALUES: afável, alpercata, boneca, crítico, lenha, musgo, sala, . . . CONCEPT: Orthographic wordforms in lower case including diacritics. These are the primary entries of Porlex. They consist of uninflected content words and inflected function words of a
PORLEX DATABASE IN EUROPEAN PORTUGUESE 16 26. DFot - Phonetic Density VALUES: 0, 1, . . ., 25. CONCEPT: Number of phonetic neighbors, where neighbor is any other phonetic entry that differs on one segment only, by substitution or by addition/deletion, while preserving relative position of the segments. Homophones were excluded from the computation. This value was computed by an automatic serial check of all phonetic entries, excluding homophones, and was not sensitive to stress. E.g., salA has 14 phonetic neighbors (cf. infra). 27. VFot - Phonetic Neighbors VALUES: alA, balA, falA, galA, malA, palA, sakA, saGA, sajA, sElA, silA, sOlA, talA, valA; …; [empty] if no neighbors. CONCEPT: List of phonetic neighbors of the wordform. 28. PUFot - Phonetic Uniqueness Point VALUES: 1, 2, . . ., 19. CONCEPT: Position of the phone, counted from the start, that uniquely identifies the phonetic wordform from its neighbors. E.g., for salA PU = 3 because the preceding phones are shared with saKA (note that due to the velarization of the final /l/, the phonetic transcription of the word sal also differs in the third phone, sa9). 29. DFom - Phonemic Density VALUES: 0, 1, . . ., 28. CONCEPT: Number of phonemic neighbors, where neighbor is any other phonemic entry that differs on one segment only, by substitution or by addition/deletion, while preserving relative position of the segments. This value was computed by an automatic serial check of all phonemic entries, excluding homophones, and was not sensitive to stress. E.g., salA has 16 phonemic neighbors (cf. infra). 30. VFom - Phonemic Neighbors
PORLEX DATABASE IN EUROPEAN PORTUGUESE 17 VALUES: alA, balA, falA, gala, malA, palA, sakA, saga, sajA, sal, salsA, sElA, silA, sOlA, talA, valA; …; [empty] if no neighbors. CONCEPT: List of phonemic neighbors of a given wordform. 31. PUFom - Phonemic Uniqueness Point VALUES: 1, 2, . . ., 18. CONCEPT: Position of the phoneme, counted from the start, that uniquely identifies the phonemic wordform from its neighbors. E.g., for salA PU = 4 because the preceding phones are shared with sal. 32. HGnF - Nonhomophonic Homographs VALUES: 1, 2, 3; [empty] if no nonhomophonic homographs. CONCEPT: For a given orthographic wordform, number of identical entries that differ on phonetic/phonemic wordform. This is one of the two sole instances of repeated orthographic wordforms (for the other, see below). E.g., colher, [ku’LEr], and colher, to [ku’Ler]. Since these cases are scarce, only positive cases are marked. 33. HGHF - Homographic Homophones VALUES: 1, 2, 3, 4; [empty] if no homographic homophones. CONCEPT: Number of entries with identical orthographic and phonetic/phonemic wordforms that differ on grammatical class. This follows from the criterion adopted to establish the database, and it is one of the two sole instances of repeated entries (see above). Since homographic homophones are scarce, only positive cases are marked. 34. HFnG - Nonhomographic Homophones VALUES: 1, 2, 3, 4, [empty] if no nonhomographic homophones. CONCEPT: For a given orthographic wordform, number of entries with identical phonetic/phonemic sequences irrespective of stress position. Strictly speaking, these are
PORLEX DATABASE IN EUROPEAN PORTUGUESE 18 segmental homophones; e.g., túnel, [‘tunE9] and tonel, [tu’nE9]. Since these cases are scarce, only positive cases are marked. 35. PFot1 - Gross Phonetic Pattern VALUES: V’CV.CVC, VC.CVC’CV.CV, . . .; where C = consonant, V = vowel, G = semivowel, and H = homorganic nasal. CONCEPT: Classification of the segments of the phonetic wordform into consonant, vowel or glide types (C, V, G, respectively). Because of the coarticulatory nature of homorganic nasals, these were classified separately (H). Syllable boundaries and stress marks are shown. 36. PFot2 - Detailed Phonetic Pattern VALUES: V’SV.ZVL, VL.PVF’PV.PV, . . .; where P = voiceless stop; B = voiced stop; S = voiceless fricative; Z = voiced fricative; L = lateral approximant; N = nasal; V = oral vowel; M = nasal vowel; G = oral semivowel; W = nasal semivowel; R = trill; F = flap; D = fricatization of /b, d, g/; H = homorganic nasal. CONCEPT: Classification of the segments of the phonetic wordform according to manner (stop, nasal, trill, flap, fricative, approximant; vowel), voicing, and oral vs. nasal quality. Coarticulatory homorganic nasals were included as a separate category. Syllable boundaries and stress marks are shown. 37. InvO - Reverse Orthographic Wordform VALUES: levàfa, atacrepla, acenob, ocitírc, ahnel, ogsum, alas, . . . CONCEPT: Backward sequence of the orthographic wordform (from variable 2, Orthographic Wordform, thus with diactrics). 38. InvF - Reverse Phonetic Wordform VALUES: 9EvafA, Atakr6p9a, AkEnub, ukitirk, ANAl, ugZum, Alas, . . . CONCEPT: Backward sequence of the phonetic wordform. 39. Maius - Uppercase Wordform
PORLEX DATABASE IN EUROPEAN PORTUGUESE 19 VALUES: AFAVEL, ALPERCATA, BONECA, CRITICO, LENHA, MUSGO, SALA, . . . CONCEPT: Orthographic wordform in uppercase without diacritics (but with cedilla). 40. VFot1 - Phonetic Variant 1 41. VFot2 - Phonetic Variant 2 42. VFot3 - Phonetic Variant 3 VALUES: ´leNA, ´lAjNA, ´lENA; …, [empty] if no phonetic variant. CONCEPT: Alternative phonetic transcription(s) of the wordform, [empty] if no phonetic variants or if there are less than 3 phonetic variants. 43. VarLex - Lexical Variant Orthographic Wordform VALUES: alparcata, alpergata, alpargata; bonecra; . . .; [empty] if no lexical variant. CONCEPT: Orthographic wordform of the lexical variant (lower case with diacritics). 44. FotVarL - Lexical Variant Phonetic Wordform VALUES: a9.pAr’ka.tA, a9.p6r’ga.tA; a9.pAr’ga.tA; bu´nEkrA; . . .; [empty] if no lexical variant. CONCEPT: Phonetic transcription of the lexical variant.
PORLEX DATABASE IN EUROPEAN PORTUGUESE 20 Acknowledgments The authors would like to acknowledge Porto Editora for making available the electronic version of the entries from Dicionário da Lingua Portuguesa (Costa, & Melo, 1997). This research was supported in part by a FCT grant to the Center for Psychology at the University of Porto (Language Group). I. Gomes was supported by a FCT grant (PRAXIS XXI/BD/4529/94). Correspondence should be addressed to: São Luís Castro, FPCE-Universidade do Porto, rua Campo Alegre, 1021, P 4169 - 004 Porto, Portugal ([email protected]).
PORLEX DATABASE IN EUROPEAN PORTUGUESE 21 Footnotes 1. Porlex was started in 1998, and the final computations were completed in early 2000. DLPC, the first Portuguese dictionary to provide phonetic transcriptions in European Portuguese for a middle sized vocabulary, appeared in early 2001 (Academia das Ciências de Lisboa, 2001). The entries from the Dicionário da Língua Portuguesa (Costa & Melo, 1997) were made available to us in electronic format, but they had to be checked individually in order to insert source information on phonological characteristics and syllabication, that we were unable to find in a compatible electronic format. CORLEX, a corpus that comprises 26 443 lemmas and 140 315 wordforms, as well frequency information based on 16 210 438 wordtokens, is now available at the Centro de Linguística da Universidade de Lisboa website at http://www.clul.ul.pt/sectores/projecto lmcpc.html.
PORLEX DATABASE IN EUROPEAN PORTUGUESE 22 References A. Sources for Porlex Costa, J. A., & Melo, A. S. (1997). Dicionário da língua portuguesa [Portuguese language dictionary] (7th Ed.). Porto: Porto Editora. Cunha, C., & Cintra, L. F. L. (1987). Nova gramática do português contemporâneo [New grammar of contemporary Portuguese] (4ª Ed.). Lisboa: Edições João Sá da Costa. Figueiredo, C. (1996). Grande dicionário da língua portuguesa [Extended Portuguese language dictionary] (25th Ed.). Venda Nova: Bertrand Editora. Irmen, F., & Kollert, A. M. C. (1995). Langenscheidts Taschenwörterbuch Portugiesisch [Langenscheidt Portuguese pocket dictionary]. Munich: Langenscheidt. Mateus, M. H. M., Brito, A. M., Duarte, I., & Faria, I. H. (1989). Gramática da língua portuguesa [Portuguese language grammar] (4ª Ed.). Lisboa: Editorial Caminho. Melhoramentos (1998). Michaelis: Pequeno dicionário da língua portuguesa [Michaelis: Brief Portuguese dictionary]. São Paulo: Author. Nascimento, M. F. B., Marques, M. L. G., & Cruz, M. L. S. (1987). Português fundamental: Métodos e documentos (Vol. II, Tomo I: Inquérito de frequência) [Basic Portuguese: Methods and documents. Vol. II, Tomo I: Frequency survey]. Lisbon: INIC, Centro de Linguística da Universidade de Lisboa. Nascimento, M. F. B., Rivenc, P., & Cruz, M. L. S. (1987). Português fundamental: Métodos e documentos (Vol. II, Tomo II: Inquérito de disponibilidade) [Basic Portuguese: Methods and documents. Vol. II, Tomo II: Availability survey]. Lisbon: INIC, Centro de Linguística da Universidade de Lisboa. Texto Editora (1998). Dicionário universal fundamental da língua portuguesa [Basic universal Portuguese language dictionary] (1st Ed.). Lisbon: Author.
PORLEX DATABASE IN EUROPEAN PORTUGUESE 23 Vilela, M. (1991). Dicionário do português básico [Elementary Portuguese dictionary] (3rd Ed.). Rio Tinto: Edições Asa. Vilela, M. (1995). Gramática da língua portuguesa [Portuguese language grammar]. Coimbra: Livraria Almedina. fine tune B. Others Academia das Ciências de Lisboa (2001). Dicionário da Língua Portuguesa Contemporânea (DLPC) [Dictionary of contemporary Portuguese language]. Lisboa: Editorial Verbo. Azuaga, L. (1996). Morfologia [Morphology]. In I. H. Faria, E. R. Pedro, I. Duarte, & C. A. M. Gouveia (Eds.), Introdução à linguística geral e portuguesa [An introduction to general and Portuguese linguistics] (pp. 215-244). Lisbon: Editorial Caminho. Baayen, R. H., Piepenbrock, R., & Gulikers, L. (1995). The CELEX lexical database (Release 2) [CD-ROM]. Philadelphia, PA: Linguistic Data Consortium, University of Pennsylvania [Distributor]. Burnage, G. (1990). CELEX – A guide for users. Nijmegen: Centre for Lexical Information, University of Nijmegen. Charles-Luce, J., & Luce, P. A. (1990). Similarity neighbourhoods of words in young children's lexicons. Journal of Child Language, 17, 205-215. Coltheart, M. (1981). The MRC psycholinguistic database. Quarterly Journal of Experimental Psychology, 33A, 497-505. Content, A., Mousty, P., & Radeau, M. (1990). Brulex: Une base de données lexicales informatisée pour le Français écrit et parlé [Brulex: A computerized lexical database for written and spoken French]. L’Année Psychologique, 90, 551-566.
PORLEX DATABASE IN EUROPEAN PORTUGUESE 24 Fromkin, V., & Rodman, R. (1998). An introduction to language (6th Ed.). Fort Worth: Harcourt Brace. Gomes, I. (2001). Ler e escrever em Português Europeu [Reading and writing in European Portuguese]. Unpublished doctoral dissertation, University of Porto, Faculdade de Psicologia e de Ciências da Educação, Portugal. Lambert, É., & Chesnet, D. (2001). Novlex: Une base de données lexicales pour les élèves de primaire [Novlex: A lexical database for elementary school students]. L’Année Psychologique, 101, 277-288. Luce, P. A. (1986). Neighborhoods of words in the mental lexicon (Tech. Rep. No. 6). Bloomington, Indiana: University of Indiana, Speech Research Laboratory, Department of Psychology. Marslen-Wilson, W. D. (1990). Activation, competition, and frequency in lexical acess. In G. T. Altman (Ed.), Cognitive models of speech processing: Psycholinguistic and computational perspectives (pp. 148-172). Cambridge: Bradford Books. Marslen-Wilson, W. D., & Tyler, L. K. (1980). The temporal structure of spoken language understanding. Cognition, 8, 1-71. Microsoft Corporation (1998). Microsoft Office 98 – Macintosh Edition [Computer software]. USA: Author. Nascimento, M. F. B., Casteleiro, J. M., Marques, M. L. G., Barreto, F., & Amaro, R. (n.d.). Léxico multifuncional computorizado do Português Contemporâneo [Multifunctional computacional lexicon of contemporary Portuguese]. Available: http://www.clul.ul.pt/sectores/projecto lmcpc.html/ [2002, Dec. 13]. Nascimento, M. F. B., Rodrigues, M. C., & Gonçalves, J. B. (Eds.). (1996). Actas do XI Encontro Nacional da Associação Portuguesa de Linguística (Vol. I – Corpora)
PORLEX DATABASE IN EUROPEAN PORTUGUESE 25 [Proceedings of the XI National Meeting of the Portuguese Linguistic Association]. Lisbon: Colibri – Artes Gráficas. New, B., Pallier, C., Ferrand, L., & Matos, R. (2001). Une base de données lexicales du Français contemporain sur internet: LexiqueTM [A lexical database for contemporary French on Internet: Lexique]. L’Année Psychologique, 101, 447-462. Piepenbrock, R. (2001, Feb. 18). Celex, The Dutch centre for lexical information [On line]. Available: http://www.kun.nl/celex/ [2002, Nov. 26]. Piñeiro, A., & Manzano, M. (2000). A lexical database for Spanish-speaking children. Behavior Research Methods, Instruments, & Computers, 32, 616-628. Segalowitz, S. J., & Lane, K. C. (2000). Lexical access of function versus content words. Brain and Language, 75, 376-389.