scieee AI-readable full text Open interactive document viewer

Terminografija ir tekstynas

Rūta Marcinkevičienė

Full text

TERMINOLOGY IR PRESENT Ruth MARCINKEVIČIENĖ Vytautas Magnus University TERMINOLOGY IR The linguistics of textbooks, su which appeared with the first textbook more than three decades ago, made a revolution in general lexicography. Today, glossaries are more or less based on textbooks. In ar this respect, terminology has been forgotten in countries that are more advanced in the linguistics of textbooks, let alone in newcomers like us. o ką Here it will be ir tried to show how important, simply indispensable thing is the textbook for the preparation of dictionaries of terms. The era of computer technology and computer textbooks clearly separated the traditional ir terminology terminography from the new one. The approach of terminologists began to differ significantly in ir several things: the very concept of the term, the sources of terms, the criteria for distinguishing a į term from a non-term, the qualification of the terminologist, the standardization of terms and some others. ir With the rapid expansion of the ir interweaving of special sciences, the work of terminologists has greatly increased, because working with traditional methods is hardly possible to keep up with the progress of scientific technology. ir That is why the terms have been looked at differently: more emphasis į is placed on the norm, but on the use. ne Traditional terminology, which emerged at the beginning of this century due to the huge breakthrough in technology, was mainly concerned with naming concepts, organizing the large variety of terms that appeared in a short time, systematizing and standardizing terminology of individual branches of science. ir ir Į from the very beginning of the emergence of terminology as an independent science was viewed differently than other words. The first version of the terminology į (Wūster) (cited by Pearson 1998: 10) argued that su terms should be treated differently than common su language words, because terminology begins with the analysis of concepts. O concepts, of course, exist independently not only of the terms that designate them, but also of different languages. Terms are nothing more than names of concepts, ideally interacting with the concepts themselves in the following relationship: one concept—- one term. The terms were defined mainly by taking into account the place of the terms in the hierarchical system of concepts and their relationships with other concepts. The terminology was based on terms created and/or standardized and approved by experts in their field, which had a clearly defined, established and thus protected meaning. Therefore, it is not surprising that traditional terminologists adhering to prescriptive rules (cf. Wūster, for whom Sein-Norm was far more important than Ist-Norm) were not at all interested in the real state of the language of individual branches of science and the authentic usage of terms. In going from concept to term and thus creating terminology of each individual branch on the onomasological principle, the terminologists of that time did not have and could not have the problem of identifying the term, but only of creating it. In addition, at that time the individual branches of science were still quite autonomous. Terminologists were also more divided into fields and had a better understanding of the terminology of isolated fields. When common subjects are found, their representatives become interested in terms, and new people come into terminology, not necessarily experts in a specific field of science and/or linguists, i.e. programmers, creators of terminology databases and electronic dictionaries, etc. Since the actual use of the language and the number of new terms have long outpaced the available lists of approved terms, terminology of rapidly changing scientific disciplines, which is far behind life due to the traditional method of repositories, does not allow newcomers to rely on older registered terms. They begin to be collected from texts, and the text becomes the beginning of the work of a terminologue (Rogers et al. 1994: 843). In addition, the actual usage of these words is examined, which often differs greatly from the ideal that language norms aim to achieve. It is not clear what is considered a term, when a word is terminologically defined, what use of the word is considered a term, how to distinguish a term from a non-term. Thus, terminology becomes no longer a clearly defined and isolated part of the lexicon. It forms a rather difficult to distinguish group of words, in the center of which are prototypical members, always perceived as terms, followed by those that only under certain circumstances, in a certain context, acquire the status of a term. O in the periphery ir there are words that are not iš similar at first glance, į but should be treated as such. This is the case, for example, with the words of the general dictionary found in popular texts. lo For this other reason, it is very ir important for terminographers to identify terms and choose them to describe authentic texts. iš ir This is why software that facilitates the work with texts becomes ir very su important. The focus on actual use leads terminologists more towards description than norming, the path of prescriptivism and descriptivism divides o terminology into traditional — and modern. į ir As a source of terms, a textbook is useful for terminographers for many reasons. Už jo the authors ir who advocate the creation and use indicate the main ones. iš Perhaps most importantliš y, that one ir to from the textbook itself, three things are immediately obtained that are necessary for thermographers: the terms jų themselves, definitions, i.e. objective information about the concept itself, and examples of use, thus frequency of use, terminology ir etc. ir (Cases 1990, Meyer et et al. 1996, Pearson Of 1998). course, the textbook can be used by ir researchers with a pre-defined list of terms. In this case, they get the iš factual linguistic informatiir on from the text. However, if the terms are derived from a textbook, iš they have one great advantage: they are systematically — interrelated, since the ir individual texts contain sets of terms used by individual authors, in each of which the terms are well thought out and harmoniously iš match each other. For this reason, in order for the resulting — terms to reflect individual systems, the textbook — į must be filled with full texts, 0 ne jų extracts, i.e. samples (Meyer et et al. 1996: 268). The full text is necessary for the factual ir information that would be lost if contented with parts of the texts. ma Today, the texts created by terminographers are still mainly used to collect information about previously held terms. Looks like, whether they are still in use, and if so, how, whether the characteristics of their use have not changed, and with them the meaning, whether the scope of the concepts they refer to has not changed with the emergence of new terms. However, a much more complex text is needed to obtain new terms, perhaps that is why the identification of terms and their automatic or semi-automatic retrieval is one of the most popular topics in electronic terminology today. In this respect, a terminographer differs from a lexicographer. For the latter, the question is not what is a word, but only what words to put in which dictionaries, and for the terminographer, the question of what is a term arises quite often. New terms obtained from the textbook and stored and managed in special databases have another feature — they have very precisely specified and dated sources from which it is possible to judge about the emergence, spread and disappearance of terms. If we imagine the final result of the terminologist’s work as an electronic database (DB) or a paper dictionary prepared on its basis, then the textbook will be important in all stages of this DB preparation. Computers encompass the entire work cycle of a terminographer from computer textbooks through special term search and identification programs to the creation of term databases, term description and its distribution in electronic or traditional book form (Ahmad et al. 1994: 267). The work cycle of a terminographer looks like this. In the first stage, a textbook is created from the electronic text archive according to certain principles, then terms are obtained from it, they are transferred to the DB, described, i.e. in individual DB fields typical examples of their use are taken from the concordance, from the definitions obtained there. Subsequently, based on the linguistic and subject information obtained from the textbook, words related to the descriptive term are indicated: synonyms, antonyms, hyponyms, hyperonyms, cross-references to their columns are given, the term’s dependence on the subject area, its place in the general conceptual system. In addition, grammatical and pragmatic features, conjugation or other restrictions of use, information on the meaning of the term and, if so intended, translation equivalents are provided (for more on the features of the terminographic description, see Sager 1990: 142-163). Thus, the textbook is indispensable from the very beginning, i.e. from acquaintance with the branch of science itself, to the final phase of the dictionary article prepared. Today’s terminologists are very sceptical about the computerization of existing paper dictionaries based on textbooks, i.e. — į ne the creation of databases (Sager 1990: 141). Another reason for preparing special textbooks for terminographers is that authentic language examples are even more important for terminography than for lexicography. They are much more important than introspection, because a terminographer, unlike a lexicographer, is not a specialist, he does not create texts of the language variety he describes. Having encountered a problematic case, he cannot ask: how would I use this word, only: this word means to the su expert in that field, as a scientist, a technologist, an expert would use it. aš o ką jį Thus, because of their poorly adapted sense of language, terminographers are much more dependent on the textbook than lexicographers. All textbook users Atkins et et al. distinguiį shes three a) groups: those who are interested in the b) language of the texts, those who are interested in the content of the texts ir c) those who are interested in texts as texts, i.e. jų structure, ir construction, etc. (quoted by Meyer et et al. 1996: 269). Lexicographers belong to the first group, o terminograph- — ers to the first two, because they are interested in the language ir and the ir subject itself. Having understood the importance of the ir textbook and the possibilities offered by terminology, it is necessary to define what the terminographer’s textbook should be. What text As mentioned, a terminographiš er needs two equally important things: linguistic fact information, so it is important ir that iš jo to be able to get one, the ir other. ir In terms of language, the textbook must be such that it clearly, correctly and, preferably, with definitions, presents the terms. It is desirable that su the terms be as varied as tų possible, ideally all, created and used by specific linguistic communities. ir — ir It is also important that the terms are used in a wide variety of contexts. In ir terms of subject matter, the textbook į should include such texts that highlight the content of the concepts iš referred to by the terms, their interrelationship and dependence. ir It is very important that the terms in those texts are defined. For these reasons, experts suggest that textbooks should include the entire į range of possible texts from the simplest to the most complex, presented in a communicative and pragmatic way. Pearson distinguishes four groups of special language texts according to the probability of finding terms in them: those in which the expert’s addressee is a knowledge-equal partner — the same expert, i.e. articles in peer-reviewed scientific journals, monographs, in į these texts the density of terms a) is the highest; b) texts addressed by the expert to a specialist of a lower level than himself, but in the same field, for example, various instructions, textbooks, here there are somewhat fewer terms; (c) popular science texts written for people unfamiliar with the specific field. They deliberately avoid complicated or any terms, preferring to explain, translate in other words, use analogies and comparisons, so the author, if he really wants, can do without terms altogether; d) texts corresponding to teacher-student communication: textbooks, teaching materials. The terms are there, but they are used slightly differently. An important feature of these texts is the explanation of terms (Pearson 1998: 36-40). It is not sufficient to describe the texts of the terminographic textbook by a communicative pragmatic aspect. Other text parameters and proportions are no less important for the quality and representativeness of the textbook. To the question of how many texts need to be collected in order to be able to rely on the data of the textbook composed from them, scholars answer as follows: “Unlike in the case of textbooks of a general nature, quality here is far more important than quantity” (Meyer et al. 1996: 268). As a rule, special texts are smaller than general texts for understandable reasons. However, although they are smaller, they must be composed of full, unabridged texts. In addition, they must cover the entire branch represented with all its branches, pre-selected by the terminographers. Often it is difficult to draw a clear line between individual branches of science, especially if that branch is related to general subjects and the compilers of the textbook are tempted to go deeper, i.e. to include more texts written by experts for experts, which undoubtedly belong to that field. Therefore, it is suggested that terminologists seek advice from experts in the descriptive field on which texts to include, so that a representative textbook is created in all senses, so that systematically important concepts and the terms that designate them are not omitted along with the texts that are not included, and most importantly, so that the priorities of one compiler do not dominate (Ahmad 1993: 60). In the case of a terminographic textbook, the principle sometimes applied by lexicographers to general textbooks: “What I have, I put in” is not appropriate at all. Taking into account the typology of textbooks and the most important types of textbooks in terms of continuity: final continuous substitutive (monitor), terminographic textbook is proposed to be prepared as a second type, or even better, the latter type of texts. Then the textbook will be constantly supplemented with the latest texts, either by leaving the old ones (in the case of a continuous textbook), or by moving them to the archive, if there is a replacement textbook. The continuous textbook would constantly increase without necessarily maintaining the pre-selected principles of its structure, because, changing publishing trends, new texts could shift to one or the other side. The proportions of the new text would remain unchanged and the new texts would only occupy the space originally allocated to texts of the same nature. In any case, it is necessary to constantly update the textbook, because terms age much faster than everyday words. After considering the general issues of the textbook size, nature and subject limits and having taken decisions, it is necessary to look at which specific texts to include in the textbook so that it is representative and balanced. In each specific case, the genre of the text, the time of writing, the author and the linguistic status must be taken into account in order to achieve the greatest diversity. 10 — — ir iš į ir į ir The wide ne range of texts o from serious to ir popular allows to select both scientific, formal terms, among which there are many author’s neologisms, and popular, much more common and established in colloquial language terms. When assessing the age of the texts, preference is given to the newer ones, but older texts may also be valuable, especially if we look at them from an objective point of view. The authorship should also be very diverse, so that the peculiarities of the style of one or another strongly represented author do not overwhelm. Although the stylistic diversity of scientific texts is not very large, if we compare, for example, with fiction, the author’s style is important not only in terms of the use of language units, but also in terms of their selection. Therefore, it is not surprising that different scientists, especially if they belong to different schools of scientific thought, use different 11 terms. O for the standardization of terms, as well as for the search for neologisms, the greatest ir variety of terms in the textbook is very important. Last but not least, the linguistic status of the texts. Preference is given to originals, translations, written by native ne speakers, už ir o ne o po to not yet peer-reviewed edited texts. ir Textbook analysis tools for terminographers During the last decade, a variety of software has been developed to facilitate word search and other text analysis for lexicographers. Although the work of lexicographer terminologists ir is different, they use ir almost the same software options, i.e. a) lists of word frequencies; concordances, b) i.e. the research word in all instances of use in a context of a certain length, usually one line; and connectivity partners statistics c) (Meyer al. et 1996: 274). These tools help to examine the text from a linguistic point of view. In addition to linguistic software tools used in common with lexicographers, special programs suitable only for terminographers are gradually emerging. They differ from the general programs not in linguistic but in subjective nature, since they are intended for the analysis of information contained in the text. In this case, the textbook is treated not as a set of contexts of use of terms, but as a source of special knowledge of a particular field or branch of science. Here are some of their groups: a) programs for analyzing texts in terms of information contained in them — searching for meaningful words, information-rich places in the text indicating the interrelationships and dependencies of concepts, etc. ; (b) programs that select the texts needed by the terminographer, such as definitions of terms; (c) programs for the analysis of term-named concepts, facilitating the management of knowledge acquired from textbooks (knowledge acquisition systems- ). The latter group, although completely new and very promising, is still not sufficiently studied in the specialized literature. Much more has been written about the possibility of semi-automatic collection of the definitions of terms in the text (Pearson 1998: 121-204). 12 Both definitions and compound ir terms require a morphologically annotated text, i.e. one in which all words are marked by parts of jų speech. The initial stage of work must be carried out by the lexicographer himself: he reads texts collects defined terms. Based ir on this information, morphological models of ta definitions are subsequently formed for all combinations of the — sequence of language parts obtained by defining them. Then the computer searches the entire text for such combinations of words. This results in a list of ne terms and definitions, jų but only those defined by Į the candidates. It is ir submitted to the field Zinovyev for review to delete ir unnecessary sentences. In order to reduce the number of unnecessary, randomly coinciding morphological combinations, the computer is given the most frequently occurring lexical units in the definitions. Similarly, there may be — a textbook obtaiš mi ir compound terms, such as the first generation of automatic term identification programs LEXTER (Bourigault and 1995) TERM (The 1994) essence of Lauriston. Be no doubt, there are some problems with this type of text processing. First of all, far from all the ne selected word combinations are terms. Ši problem solved by compiling the ir į computer by typing a list of gų unnecessary words (stop-list), which į terms candidate lists are not drawn. Be therefore, it is often unclear which words belong to the compound term itself, which only o to the — environmejo nt. But those things ne are always ir clear when selecting terms manually. Semi-automatic term selection method In order to illustrate the possibilities offered to the terminographer by the simplest software tools, we V. Labučio Syntax (T. I, V., electronic 1994) version tyréme ir text su Miko Scott's Linguist Plugin Pack WordSmith Tools, version 2.0 (Oxford, 1997). The purpose of the statistical analysis was to obtain a few words ar jų lists of compounds iš from which candidates can be subsequently selected for the terms themselves. į The ir statistical indicators of the text under analysis are as follows: 37 319 words word ir forms, excluding repetitive ones, remain 13 CAN DO 17 PRINCIPAL DEMUO 17 SPIELINKSNINEM CONSTRUCTIONS 17 WITH A KILMNIK 17 AS IN 16 LANGUAGES SYSTEMS 16 The Interrogator's Dementia 16 SUDETINI SAKINI 16 BE TO 15 AND UNDER ARTICLE 15 ZODZIU JUNGINIO 15 As can be seen from iš the examples given, longer word combinations resemble more sentences ar than phrases, — shorter syntactic word combinations. Unequal is ir jų the degree of accumulation, shown by the iš number of uses, for example, some two-word combinations are used 118 times even in such a short text. In addition, some word chains completely coincide with compound terms, e.g.: second-degree subjunctive sentence, two-way interface relationship, language system, and from others these terms have to be chosen, e.g. :explains another predicate, word form as if applied to, word with main accent, etc. Although they are shortened, they are informative in their own way, because the predicative words (in the examples they are thinned out), belonging not to the term itself, but to its environment, provide important information about the conjugation and the peculiarities of use. The latter list, like the previous ones, is not determined, so the same keyword term is repeated in another grammatical form. Among other things, it is repeated in chains of words of different lengths —from the longest to the shortest: forms of at least two independent words, forms of at least two independent words, forms of independent words, forms of words. However, frequent repetitions do not hinder, but help to notice the most important, most commonly used compound terms. Even if the semi-automatic software tools mentioned above do not select all the terms in the research text, they still allow the most important ones to be identified, and the fact that the lists compiled by the computer are read and managed by a terminographer guarantees the quality of the final result. 20 Conclusion Modern terminology, aiming to collect, describe and ir standardize the terms of various branches of science, cannot do without be computer textbook ir su other related stages of the preparation of the computerized term bank: special software of a general nature, lexical ar databases, etc. The textbook must be prepared in such a way as to cover the widest possible variety of texts, thus representing the language of the descriptive science. ir Using even general ir software for lexicographers, it is possible to automatically obtain lists of candidate terms for word-word combinationir s, from which known new — terms are į selected. iš ir Received 21 June 1999 Literature 1. Ahmad 1994 — Ahmad, K., Davies, A., Fulford, H. and Rogers, M. What is a term? The semi-automatic extraction of terms from text. In: M. Snell-Hornby, F. Pochacker, Kaindl K. (eds.). Translation Studies: An Interdisciplinary. Amsterdam — Philadelphia, PA 1994, P. 267-278. 2. Ahmad 1993 — Ahmad, K. Pragmatics of Specialist Terms: The Acquisition and Representation of Terminology. In: Petra Steffens (ed.). Machine Translation and the Lexicon. Thirds International EAMT Workshop. Heidelberg, Germany, April 26-28. 1993. Proceedings. Berlin — Heidelberg, 1995, P. 51-76. 3. Bourigault 1996 —-Bourigault, D., Gonzalez-Mullier, I., Gros, LEXTER, Natural C. a Language Processing Tool for Terminology Extraction. In: Euralex 96 Proceedings II, Goteborg, 1996, P. 771 —780. 4. Lauriston 1994 — Lauriston, A. Automatic Recognition of Complex terms: Problems and the The TERMINO solution. In: Terminology. Vol 1(1), 1994, P 147-170. 5. Meyer 1996 — Meyer, I. and Mackintosh, The K. Corpus from a Terminographer’s Viewpoint. In: International Journal of Corpus Linguistics. Vol. 1(2), 1996, P. 257-285. 6. Pearson 1998 — Pearson, J. Terms in Context. Amsterdam — Philadelphia, PA 1998. 7. Rogers 1994 — Rogers, M., Ahmad, Computerised K. Terminilogy for Translators: the role text. of In: M. Brekke, O. Andersen, T. Dahl, J. Myking. Applications and Implications of Current LSP Research. Proceedings of the 9" European Symposium on LSP, Bergen, Aug. 2-6, 1993. Vol. II. 1998-1999 Bergen, 1994. P 840-851. 21 8. Sager 1990 -Sager, J.C. A Practical Course in Terminology Processing. Amsterdam—Philadelphia, 1990. TERMINOGRAPHY AND CORPUS Summary The paper deals with role of a corpus in terminology in general and terminography in particular. It argues for the necessity to use not only corpus but also general purpose and specific tools for the extraction of terminology and compilation of terminological databases. Peculiarities of the terminographer’s corpus such as its contents, size and specificity are also discussed. In addition, a wide range of possible general purpose and specific tools for term extraction are presented here. Some of them, based on statistical analyses of a text, were applied for one Lithuanian text on syntax in order to show how different tools can produce different semi-automatic lists of terms. 22