This is a machine-generated translation of the original document and may contain errors, omissions, or changes in meaning. Do not rely on this translation for quotations, citations, or authoritative references. Please consult and cite the original document.
Preparing document pages…
Page 1
DAIVA ŠVEIKAUSKIENĖ
Institute of the Lithuanian Language
ARŪNAS RIBIKAUSKAS
Vilnius Gediminas Technical University
VYTAUTAS ŠVEIKAUSKAS
Institute of the Lithuanian Language
Research areas: grammar, computational linguistics.
A MULTILINGUAL WEBSITE ON THE GRAMMAR OF THE LITHUANIAN LANGUAGE
A multilingual online information system for Lithuanian grammar
ANNOTATION
A year ago, work began at the Institute of the Lithuanian Language on a project whose aim is to make the presentation of detailed data on the grammatical properties of Lithuanian words freely accessible to users on the Internet. In March 2017, a trial version was prepared, comprising only words with the root „bėg“ (cf. the verb „bėgti/laufen“). The database consists of around 25,000 records. The project seeks to avoid the shortcomings of existing work in the field of computational linguistics. These relate to the range of words, the variety of grammatical data, and the clarity and comprehensibility of how the data are presented, among other things. Grammatical information about a word entered by a user is presented in three areas: word structure, morphological data, and morphemic data. This is the first website in Lithuania to provide users with information about types of morphemes. A declension table can be obtained for every word contained in the database. Usage examples are also provided for some words. The website is available in seven languages: Lithuanian, English, German, French, Italian, Russian, and Japanese, so that foreigners interested in the Lithuanian language can use it.
144
Acta Linguistica Lituanica LXXVII
Page 2
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
KEYWORDS: grammar, morphology, morpheme, information system, computational linguistics.
ANNOTATION
The project was started in the Institute of Lithuanian language in 2016. It aims at presenting detailed data about the grammatical features of Lithuanian words. The goal is to make those data freely accessible to any user on the Internet. The preliminary version covering the words with the root “bėg/run” was published in March 2017. The database contains around 25,000 entries. We are trying to avoid drawbacks that are specific to the words done in the field of computational linguistics. It includes word coverage, grammatical data comprehensiveness, clarity and clearness of presented information etc. Grammatical information about the word in question is presented from three aspects: word structure, morphological data, and morphemic data. It is the first website in Lithuania providing morphemic types. One can get an inflection table for each word the database contains. Usage examples are given for some words. The information on the website is available in seven languages – Lithuanian, English, German, French, Italian, Russian, and Japanese – so foreigners interested in Lithuanian language can use it as well.
KEYWORDS: Grammar, Morphology, Morphemic, Information system, Computational linguistics.
1. INTRODUCTION: COMPUTATIONAL LINGUISTICS
Until the middle of the 20th century, all grammars were printed on paper. After the advent of computers (in 1942, the world's first computer, MARK I, was built at Harvard University (Schwanke 1991: 69)), they began to penetrate every area of life. Languages were no exception: it soon became clear that computers could process not only numbers, but also arbitrary symbols. Consequently, they could also process languages. Several linguists began developing formal descriptions of languages so that computers' capabilities could also be applied to languages (Winograd 1983: 23). In the field of artificial intelligence, language processing became one of the most important areas of application (Charis 1989: 71). From 1968 to 1978, the grammatical analysis of languages was the most important topic for artificial intelligence researchers (Nirenburg 1987: 26). Languages, however, which yield so readily to human beings, proved a hard nut for computers (Jensen, Heidorn,
Straipsniai / Articles
145
Page 3
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBOKAS, VYTAUTAS ŠVEIKAUSKAS
Richardson 1993: 2). Deshalb konnten bislang keine guten Ergebnisse auf dem Gebiet der Sprachverarbeitung mit Hilfe von Computern erreicht werden.
1.1. STATISTICAL METHODS
Statistical methods are of great importance in language processing. The first proposals to use them were put forward shortly after the advent of computers. This is especially true of English. David W. Reed discussed quantitative linguistic analysis (Reed 1949: 236). Warren Weaver was the first to suggest using computers for translation. In 1949, he proposed using statistical methods for this purpose. However, the scientists of the time rejected the idea after a short while, probably because of the insufficient volume of electronically readable texts and the limited processing power of computers at the time. A few decades later, when text corpora had been compiled and the application of statistical methods in speech recognition had produced good results, statistical methods were revisited for translation (Brown at al. 1990: 79). Canada offered particularly favorable conditions for this, because all parliamentary material is stored in two languages (English and French) owing to the country's two official languages. Consequently, a sufficient quantity of parallel texts could be collected very quickly (Al-Onaizan, Curin, year 1999: 1). The structures of the two languages, English and French, are very similar: word order is strict (the subject comes first and the predicate second). Similarities in the structure of the languages made it possible to achieve good results in translation using statistical methods. But this is not the case with Lithuanian. Google translations, which use the statistical method, have not yet produced good translations into Lithuanian. According to data from 2017, the results for both English–Lithuanian and Lithuanian–English translation are worse than those for English–Latvian, Latvian–English, English–Estonian, and Estonian–English (Skadinš 2017: 22).
Statistical methods also penetrated other areas of linguistics. The first studies concerned phonetics (Zipf 1929). In 1944, statistical studies of literary lexicon were conducted, i.e., how many words were used once, how many twice, three times, and so on (Yule 194: 9).
146
Acta Linguistica Lithuanica LXXVI
Page 4
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
Since 1973, the term dialectometry has been used to designate the field of linguistics concerned with the quantitative study of languages and dialects (Köhler, Altmann, Piotrowski 2005: 498).
Statistical methods have also been applied in the field of morphology. Goldsmith describes the algorithm for identifying morphemes in a word as follows: „algorithm that takes as input a list of words and provides as output a segmentation of the words into morphemes“ (Goldsmith 2010: 36).
At the end of the last century, attempts to perform automatic syntactic analysis of sentences using statistical calculations were described in Russian literature. Scientists in Russia proposed that the syntactic structure of sentences be determined not on the basis of linguistic analysis, but by using statistical processing of texts. They claim that every linguistic unit (word, syntactic category of a word, syntactic construction) occurs in a text with a certain frequency. For example, according to data from the English frequency dictionary of valency, 195 different patterns of syntactic relations are characteristic of the verb. Yet the ten most frequently occurring patterns alone account for 86% of all cases used in texts. Statistical regularities are also used for the linear structure of the sentence. Each word class has a frequency with which it can occur at a certain distance from the starting point. A verb in English can occur in the second to fifth position from the beginning of a sentence with a probability of 0.7. And with a probability of 0.95, it occurs in the second to tenth position (Церкас 1979: 124).
Later, data from a text corpus were used for syntactic analysis (Köhler 2012: 31).
However, good results have so far been achieved only in processing the English language. Vidas Daudaravičius, who has worked in the field of computerization of Lithuanian syntax, states: „Naivų manyti, kad metodai, kurie sėkmingai taikomi anglų kalbai, tinka ir kitoms kalboms.“1 (Daudaravičius 2012: 3).
Statistical methods can be used to quickly create fast-performing systems, but on one condition: a certain number of errors must be tolerated (Link 1). Therefore, the Institute of the Lithuanian Language decided to rely as little as possible on statistics, since it strives for the greatest accuracy and reliability of data.
1 “It is naive to think that the methods used successfully for English are equally suitable for other languages.”
Straipsniai / Articles
147
Page 5
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBAKAS, VYTAUTAS ŠVEIKAUSKAS
1.2. THE SITUATION IN LITHUANIA
A great deal of work has already been done in the computerization of Lithuanian grammar. Most of it is understandable only to specialists in computational linguistics. Some work is also intended for the general public. However, it reflects only individual properties and characteristics of words and grammar. For example, a database of morphemics based on a text corpus was created at Vytautas Magnus University in Kaunas. Many word forms are missing there, and a great many abbreviations are used that are not understandable to everyone. Morphemes are separated by hyphens, and no data are provided about the type of morpheme. A database for derivation and morphemics was created at the Institute of Mathematics and Informatics in Vilnius. It specifies the type of morpheme as well as the base words and defining words (Murmulaitytė 2012: 96). Unfortunately, the database is not freely accessible. Therefore, it was decided to create a comprehensive information system intended for the general public and capable of providing users with detailed data.
Two extremes can be observed in the presentation of Lithuanian grammatical information on the Internet: the absence even of common word forms, and the presentation of redundant words that are incomprehensible even to native speakers. The Lithuanian grammar information system seeks to avoid both extremes. All word forms are generated automatically by computer from the word’s lemma, yielding the complete set of inflectional forms. All possible derivations and compounds are entered into the database. Thus, no word forms are missing. Then all words and word forms generated by the computer are checked by hand, and non-existent variants are deleted. In this way, redundant words are eliminated.
2. WORK ON COMPUTATIONAL LINGUISTICS IN LITHUANIA
In Lithuania, two methods of computer processing of language can be distinguished. In Kaunas, work based on a text corpus is carried out at the Centre for Computational Linguistics. In Vilnius, at the university and the Institute of the Lithuanian Language, the focus is more on the properties of the Lithuanian language itself. Both methods have advantages as well as disadvantages. The following section illustrates this with several examples.
148
Acta Linguistica Lithuanica LXXVI
Page 6
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
2.1. TEXT-CORPUS-BASED WORK
A corpus of contemporary Lithuanian was compiled at Vytautas Magnus University in Kaunas. It is used for various kinds of language processing. In the field of grammar, a dictionary of morphemes (Rimkutė, Kazlauskienė, Rąkinis 2011) and, some years later, a database were created. Users are provided with both morphemic and morphological data. This is the first freely accessible resource in which a word is divided into its morphemes. Unfortunately, the type of morpheme is not specified, which sometimes leads to a confusing representation of the word’s morphemic structure. This occurs when words with different structures are represented identically. An example is shown in Figure 1. The words „per-ei-ti“ (to cross) and „per-ė-ti“ (to brood) differ very little in appearance. The morphemic structures of both words look the same in the morpheme database. However, in the first word the morpheme „per“ is the prefix, while in the second word that same first morpheme „per“ is the root.
FIGURE 1. Words with different morphemic structures are represented identically (Rimkutė, Kazlauskienė, Rąkinis 2011: 8, 35)
The absence of even common word forms can be mentioned as a second shortcoming. Participles in Lithuanian have six cases. Table 1 shows the presence of the word forms (singular, masculine) of the participle „bėgantis / running, flowing“ in the morpheme database.
It turns out that even the lemma (nominative singular) of this word is missing. Nor can the other missing forms, e.g. the locative (bėgančiame vandenyje / in flowing water), be considered rare or unusual. Thus, only 2 word forms from 6 cases are present.
Researchers of the second Baltic language, Latvian, also report an insufficient number of word forms in the text corpus (Paikens, Rituma, Pretkalniņa 2013: 272).
Straipsniai / Articles
149
Page 7
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBIKAS, VYTAUTAS ŠVEIKAUSKAS
TABLE 1.
Occurrence in the database (Link 2) of the word forms of the participle bēgantis / running, flowing in the singular, masculine
Case
Word
Database
Nominative
bēgantis
absent
Genitive
bēgančio
present
Dative
bēgančiam
absent
Accusative
bēgantį
present
Instrumental
bēgančiu
missing
Locative
bēgančiame
missing
Total:
6
2
2.2. STARTING POINT – THE LITHUANIAN LANGUAGE
At Vilnius University, the properties of the Lithuanian language are taken as the starting point. The website Morfologija.lt (Link 3) even displays the inflection tables of a word. Unfortunately, there are many blank spaces, sometimes for the entire paradigm. Unknown words that are incomprehensible to a Lithuanian speaker also occur sometimes, e.g. “sutikimoji” (Link 4) and similar forms. A database for word formation and morphemics was created at the Institute of Mathematics and Informatics in Vilnius (Murmulaitytė 2012). It contains very valuable information (the type of morpheme, base word, determining word, etc.), but unfortunately the data are not freely accessible.
3. THE INFORMATION SYSTEM FOR LITHUANIAN GRAMMAR
The Lithuanian language belongs to the group of Europe’s least computerized languages (Vaišnienė, Zabarskaitė 2012: 35). Therefore, it was decided to create an information system containing detailed data on the grammatical properties of Lithuanian words. The need to create more software for a broad user group is often emphasized by scholars working on the Baltic languages. The guidelines of the BSNLP (Balto-Slavic Natural Language Processing) conference in 2015 and 2017 state: “submissions
150
Acta Linguistica Lithuanica LXXVI
Page 8
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
describing systems, that are made available to the wider public would be strongly encouraged / systems made available to the general public are particularly encouraged” (Link 5).
For this reason, it was decided to design the information system for Lithuanian grammar for a broad user group and to present the data as simply and clearly as possible.
Accordingly, responsive web design (RW) was chosen to meet users’ needs as effectively as possible. Consequently, grammatical information about a word is laid out in tiles (English: tiles) (Link 6). This also makes the website accessible by mobile phone.
3.1. METHOD OF CREATING THE DATABASE
The aim of computerizing grammar is to have all grammatical information about every word and all its forms in electronic form. There are two ways to achieve this. The first is to create a formal description of the grammatical rules (which would be a kind of tool) and then have the computer generate the necessary information. The second is to store all the grammatical information about all forms of all words in the computer and retrieve it as needed. Which approach requires more work remains open to discussion. Software for morphological analysis of the Lithuanian language was developed at the Institute of Mathematics and Informatics (Zinkevičius 2000), using the first method. However, 100% accuracy could not be achieved: some information is incorrect, and some Lithuanian words are not recognized.
This approach was abandoned in developing the information system for Lithuanian grammar for the following reason: for a computer to process words itself, it must be told very precisely what operations to perform on which words. Thus, all words must first be divided into groups in such a way that the same regularities apply to all members of a group, i.e. so that the same processing method can be used. To assign a word to one group or another, many features must be identified on the basis of which the words are to be grouped. For example, syntactic analysis of the Lithuanian language uses the semantic feature [time] to determine the syntactic function of the accusative case of nouns. In the sentence “Visą naktį ji
Straipsniai / Articles
151
Page 9
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBOKAS, VYTAUTAS ŠVEIKAUSKAS
“She read this book all night” and “She read the whole book that night.” The two words knygą/Buch and naktį/Nacht do not differ with regard to morphological categories—both are feminine, singular, accusative—and only the semantic feature [time], which the word naktis/Nacht has and the word knyga/Buch does not, allows the computer to correctly treat the word naktį/Nacht as an adverbial of time and the word knygą/Buch as the object in the sentence. Something similar must also be done for morphology. One of the errors in the software already developed for the morphological analysis of the Lithuanian language is that the morpheme -ėj- is considered a suffix with roots with which it cannot combine. Features must be identified for words that allow the computer to determine which roots this suffix -ėj- can combine with. But this is very complicated to do. For example, what feature do the words geroė/Wohlhaben, žinovas/Kenner(Expert) and daržovė/Obst have in common? They all have the suffix -ov-. And what distinguishes them from all the other words to whose roots this suffix cannot be added (one cannot say *kėdovė—a derivation with the suffix -ov- from the word kėdė/Stuhl)? And which feature indicates this?
Thus, when developing the information system for Lithuanian grammar, the second approach was chosen: a database was prepared containing detailed grammatical information about all forms of all words in the Lithuanian language. Here, one does not need to start from grammatical categories (as is the case in all printed grammars), but from the word itself: all the information associated with each word must be made available to the user.
It was decided to take a root as a starting point and develop software that would reflect the entire structure of Lithuanian grammar. The root bėg (from the verb bėgti/to run) was chosen as the first test case. Then all possible derivations were generated by computer, and later all words that do not exist in Lithuanian were manually deleted. In total, the database contains around 25,000 entries.
3.2. FEATURES OF “LIGIS”
The information system for Lithuanian grammar (Lietuvių kalbos gramatikos informacinė sistema – LIGIS) (Link 7) has two essential features. First, it provides the user with more detailed information about a particular word than printed grammars do. And
152
Acta Linguistica Lithuanica LXXVI
Page 10
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
second, it covers more words than dictionaries. Grammars state rules that apply to a particular group of words. But they do not mention all the words to which these rules apply. LIGIS collects all the information associated with the word entered by the user and provides it on a website. Not all words in the Lithuanian language are included in dictionaries; many derivations and compounds are missing. LIGIS covers all possible derivations. An example can be given for comparison. The database for Morpheme has around 75,000 entries. But these correspond to different Lithuanian words. The LIGIS database currently consists of 25,000 entries and covers only the words formed from a single root: beg-. The Morpheme database contains 118 words with this root.
3.3. WORD STRUCTURE
After the word is entered, its structure is explained to the user: the lemma and base word are given, as well as the determining element for derivations and compounds. If the word is grammatically ambiguous, each meaning is assigned a lemma number. The morphological and morphemic data are then displayed under this number. An example of an ambiguous word is given in Appendix A.
Other forms. This is the name of the button (English button) located in the word-structure tile. The website for the morphological analysis of Russian (Link 8) displays all word forms for the entered word every time. In the information system for Lithuanian grammar, it was decided to introduce a button and let users retrieve the other forms via a link, since users do not need all the forms of a word every time.
3.4. MORPHOLOGICAL DATA
Morphological information comprises the designation of the part of speech and grammatical categories: case, gender, and number for nouns; person, tense, and mood for verbs, etc. Ambiguous words are numbered when, for example, a noun has several homonymous case forms.
The database for morphemes (Link 2) also contains morphological data about the word. However, the information is presented entirely in abbreviations
Straipsniai / Articles
153
Page 11
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBOKAS, VYTAUTAS ŠVEIKAUSKAS
in the database. One such example is shown in Figure 2. For non-specialists, so many abbreviations can often be incomprehensible. Information presented in this way can sometimes be unclear, especially to foreigners studying or learning Lithuanian.
FIGURE 2. Data about the word begčio, presented entirely in abbreviations (Link 2)
No abbreviations are used in the information system for Lithuanian grammar. All information is written out in full (Appendix A).
3.5. MORPHEMIC DATA
The main difference from the databases that can already be accessed today1 is that the type of each morpheme is specified. Only in this way can an unambiguous representation of the morphemes of a word be achieved. Otherwise, for example, the second root in a compound word cannot be distinguished from the suffix, as is the case in the database for morphemics.
The first and decisive feature was the representation of morphemes in different colors. A color was assigned to each morpheme type (root, prefix, suffix, ending). Prefixes are written in blue. Red was used for the root. Suffixes are shown in green, and black was chosen for the ending.
In addition, the word divided into morphemes is laid out vertically, and alongside it not only is the designation of the type of each morpheme in the word written, but the properties of the morpheme are also indicated, if it has any—for example, for a suffix: derivational or inflectional; for endings: shortened, pronominalized, and the like. In the case of shortened endings, which are very common in spoken language, the full ending is also displayed alongside. As an example, the words in spoken language can be given: šnekamoje kalboje – full endings, šnekamoj kalboj – shortened endings. If the word has the full ending, no length feature is indicated, because most words have full endings and such information would be superfluous. If
154
Acta Linguistica Lithuanica LXXVI
Page 12
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
if the ending is not pronominalized, information about it is not provided either.
Sound changes are considered a property of the root. They may involve either consonant or vowel changes. Another type of sound change can also be observed in the Lithuanian language: the loss of consonants in the imperative. Infixes, too, are among the sound changes in the root of a word. The origin of a prefix is considered one of its properties: it may be of prepositional origin, derived from a particle, or of international origin.
In cases of grammatical ambiguity of words, their morphemic structures are examined. If they are the same, only one structure is described. If, however, they differ, morphemic structures are provided for each meaning of the word (Appendix A).
3.6. ALL FORMS OF THE WORD
For each word contained in the database, the inflection table can be accessed by clicking the “Other forms” button. All the data from the morphological tile are reproduced there. At the top of the table, categories not included in the inflection table are listed; for example, nouns in the Lithuanian language are not inflected for gender, so gender is indicated at the top of the table, along with the lemma and part of speech. Adjectives and participles have three genders, and one peculiarity of the Lithuanian language is that only two of them (masculine and feminine) have all six cases. The third gender has only one form, so it is listed without a case label (Appendix B). For verbs, all tense forms in all moods are presented (Appendix D).
Conjugation classes and declension classes are therefore unnecessary, so no data about them are provided.
For some words, usage examples are also included in the inflection table. QuickInfo (English tooltip) is used for this purpose. Appendix C gives as an example the word “bėgis” (accusative, singular of the noun bėgis / railway track, run).
3.7. THE DATABASE
In developing the database, the aim was to achieve the highest possible quality and reliability. Therefore, a large part of the work was done by hand
Straipsniai / Articles
155
Page 13
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBAKAS, VYTAUTAS ŠVEIKAUSKAS
… done, because information can be made maximally reliable only with human assistance (Sulger 2013: 53).
Research on Lithuanian prefixes was carried out at the Institute of the Lithuanian Language. It was established that there are about 600 combinations of prefixes in Lithuanian (Šveikauskienė 2015: 195). “Combinations” means various sequences of prefixes that can appear before the root, e.g. at-bėgi, nu-bėgi, ne-be-at-bėgti, ne-nu-bėgti, te-be-bė-ti, etc. It turned out that only 220 of these combinations are used with verb roots; the others, e.g. nuo-, are used with nouns: nuo-monė /opinion. The software was designed to automatically combine the root with all possible combinations of prefixes. The results generated by the computer are then checked by people, and the words actually found in Lithuanian are entered into the database. This makes it possible to avoid two extremes: both the excessive number of forms generated by the computer and an insufficient range of words.
All lemmas and the grammatical information pertaining to them are entered into the database manually. Since the software developed at the Institute of Mathematics and Informatics makes no errors when generating word forms for a given lemma, the computer was entrusted with automatically generating the entries for the remaining word forms.
It is noteworthy that LIGIS also processes words with shortened endings. Such words occur very frequently in spoken language, and none of the databases already available processes these word forms. It should also be emphasized that LIGIS includes derivatives that are absent even from the 20-volume Dictionary of the Lithuanian Language and its electronic version (Link 9), e.g. neužbėgti/no longer run uphill.
The information is stored in the database in a computer-friendly format, i.e. the encoding uses Prague Markup Language (Hana, Štepánek 2012). This form of data is not presented on the Internet, as it would be superfluous for non-specialists in computational linguistics.
The fact that all the information is stored in a database also makes LIGIS a useful tool for linguists who study Lithuanian. One can obtain information of various kinds, for example, search for words with a particular combination of prefixes and suffixes, among other things.
156
Acta Linguistica Lithuanica LXXVI
Page 14
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
3.8. MULTILINGUALISM OF THE WEBSITE
The website is available in seven languages: Lithuanian, English, German, French, Italian, Russian, and Japanese. The need for a multilingual website arose from the unsatisfactory Lithuanian translations provided by Google. Machine translation of Lithuanian is unreliable. It was therefore decided to make the website on Lithuanian grammar multilingual and to use only texts translated by humans. The situation with other languages is evidently not much better. This is borne out by a letter from Lingvist dated June 8, 2017, which asks that the label LINGUIST be translated by native speakers and under no circumstances by machine translation (LinguiList 2017: 28.252).
Our idea in deciding to make the website multilingual was as follows: if a foreigner, for example a Norwegian, is learning or studying Lithuanian and knows some Lithuanian words but not others, they can choose a language they know better than Lithuanian, for example German, and all the Lithuanian words they do not know will become understandable to them. To this end, provision was even made for retaining the word entered by the user when switching languages. Thus, when a user chooses another language, they do not have to enter the word again.
4. PLANS FOR THE FUTURE
At Vilnius University, many bilingual dictionaries are available in digitized form. At the Institute of the Lithuanian Language, the large dictionary in 20 volumes has been digitized, as have the other monolingual dictionaries: synonym, antonym, and phraseological dictionaries, among others. It is planned to connect both types of dictionaries with the information system for Lithuanian grammar and create an equivalent of the German website CANOONET (Link 10).
Since the information system for Lithuanian grammar contains data on morphemes, it is possible to search for words on the basis of the morpheme model. And this is precisely what is planned for the future.
A tab (English: tab) on the LIGIS website is intended for theory and is currently under development. It is to provide in-depth academic knowledge of Lithuanian grammar.
Straipsniai / Articles
157
Page 15
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBAKAUSKAS, VYTAUTAS ŠVEIKAUSKAS
The first stage deals with morphology. At the second stage, syntactic data are also to be entered.
5. Conclusions
1. This approach can also be useful and worthwhile for other highly inflected languages.
2. The information system for Lithuanian grammar can be of great benefit to school pupils, university students, foreigners learning Lithuanian, and linguists researching the Lithuanian language.
3. The collected data can also serve as documentation of the Lithuanian language.
Literaturverzeichnis
Al – Onaizan Yaser, Curin Jan, Jahr Michael, Knight Kevin, Lafferty John, Melamed Dan, Och Franz – Josef, Purdy David, Smith Noah A., Yarowsky David 1999: Statistical Machine Translation. – Final Report JHU Workshop. http://mt-archive.info/JHU-1999-AlOnaizan.pdf (Zugriff am 21.11. 2017).
Brown Peter F., Cocke John, Della Pietra Stephen A., Della Pietra Vincent J., Jelinek Frederick, Lafferty John D., Mercer Robert L., Roossin Paul S. 1990: A Statistical Approach to Machine Translation. – Computational Linguistics Volume 16, Number 2, June, 79–85. http://www.aclweb.org/anthology/J90-2002 (Zugriff am 21.11. 2017).
Charniak Christopher E. 1989: Artificial Intelligence & Turbo C. Homewood: Dow Jones-Irwin.
Daudaravičius Vidas 2012: Teksto skaidymas pastoviųjų junginių segmentais. Daktaro disertacijos santrauka. Kaunas: Vytauto Didžiojo universitetas.
Goldsmith John A. 2010: Segmentation and Morphology. The Handbook of Computational Linguistics and Natural Language Processing. Wiley-Blackwell, 364–393.
Han Jirka, Štěpánek Jan 2012: Prague Markup Language Framework. Procedings of the 6th Linguistic Annotation Workshop. Jeju, Republic of Korea, 12–13 July, 12–21.
Jensen Karen, Heidorn George Emil, Richardson Stephen D. 1993: Natural language processing: The PLNLP Approach. Boston / London: Kluwer Academic Publishers.
158
Acta Linguistica Lithuanica XXVII
Page 16
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
Köhler Reinhard 2012: Quantitative Syntax Analysis. De Gruyter Mouton.
Köhler Reinhard, Altmann Gabriel, Piotrowski Rajmund G. 2005: Quantitative Linguistik. Berlin / New York: Walter de Gruyter.
Linguist List 2017: 28.254, Qs: How do you say the LINGUIST List in your language? 08 Jun 2017. [email protected].
Murmulaitytė Daiva 2012: Lietuvių kalbos morfemikos ir žodžių darybos tyrimų perspektyvos. – Žmogus ir žodis 1(14), 96–102.
Nirenburg Sergei 1987: Machine Translation: Theoretical and Methodological Issues. London: Cambridge University Press.
Paikens Pēteris, Rūma Laura, Pretkalniņa Lauma 2013: Morphological Analysis with limited resources: Latvian example. Proceedings of the 19th Nordic Conference of Computational Linguistics. (NODALIDA 2013); Linköping Electronic Conference Proceedings #85, 267–277.
Reed David W. 1949: A Statistical Approach to Quantitative Linguistic Analysis, WORD, 5:3, 235–247, DOI: 10.1080/00437956.1949.11659355.
Rimkutė Erika, Kazlauskienė Asta, Raškinis Gailius 2011: Dažnis lietuvių kalbos žodyne. Kaunas: VDU.
Schwanke Martina 1991: Maschinelle Übersetzung: Ein Überblick über Theorie und Praxis. Berlin: Springer-Verlag.
Skadiņš Raivis 2017: Neural MT and other Language Technologies at TILDE. http://school.gramaticframework.org/2017/slides/Ravis-Neural_MT_at_Tilde.pdf (zugegriffen 2017.11.21).
Sulger Sebastian, Butt Miriam, King Tracy Holloway, Meurer Paul 2013: ParGramBank: The ParGram Parallel Treebank. – Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, Sofia, Bulgaria, 550–560. http://aclweb.org/anthology/P/P13/P13-1054.pdf (zugegriffen 2017.11.21).
Šveikauskienė Daiva 2015: Morphemic structure of the Lithuanian prefixes. – Language: Meaning and Form 6. – Language System and Language Use. Rīga: Latvijas universitāte, 189–197.
Vaišnienė Daiva, Zabarskaitė Jolanta 2012: The Lithuanian Language in the Digital Age. – G. Rehm, H. Uszkoreit White Paper Series. Heidelberg / New York / Dordrecht / London: Springer.
Winograd Terry 1983: Language as a Cognitive Process 1. Syntax. London: Addison-Wesley Publishing Company.
Straipsniai / Articles
159
Page 17
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBAKAS, VYTAUTAS ŠVEIKAS
Yule George Udny 1944: The Statistical Study of Literary Vocabulary. Cambridge MA, Cambridge university press.
Zinkevičius Vytautas 2000: Lemoklis – morfologinė analizė. – Darbai ir dienos 24, 245–273.
Zipf George Kingsley 1929: Relative frequency as a Determinant of Phonetic Change. – Harvard studies in classical philology 40, 1–95.
Сердюков Пётр Иванович 1979: Оптимизация алгоритма поиска синтаксических связей. – Международный семинар по машинному переводу. Москва: ВЦП, 123–125.
LINKS
Link 1: http://www.digitalgrammars.com/ (Zugriff am 21.11. 2017)
Link 2: http://tekstynas.vdu.lt/page.xhtml?id=morfema-db (Zugriff am 21.11. 2017)
Link 3: http://morfologija.lt/ (Zugriff am 21.11. 2017)
Link 4: http://morfologija.lt/zodzio-formos/sutikimas (Zugriff am 21.11. 2017)
Link 5: http://bsnlp-2017.cs.helsinki.fi/cfp.html (Zugriff am 21.11. 2017)
Link 6: https://de.wikipedia.org/wiki/Microsoft_Windows_8#Oberfl.C3.A4che (Zugriff am 21.11. 2017)
Link 7: http://ligis.lki.lt/ (Zugriff am 21.11. 2017)
Link 8: http://goldlit.ru/component/slog (Zugriff am 21.11. 2017)
Link 9: http://www.lkz.lt/Visas.asp?zodis (Zugriff am 21.11. 2017)
Link 10: http://www.canoo.net/ (Zugriff am 21.11. 2017)
160
Acta Linguistica Lithuanica LXXVI
Page 18
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
ANHANG A. GRAMMATISCHE INFORMATION.
Beispiel des grammatisch mehrdeutigen Wortes nubəɡti / hinlaufen-hingelaufen.
Die grammatische Information wird in der Website in Kacheln ausgelegt.
Diese Darstellungsweise ist günstig für responsives Webdesign – RWD.
Bemerkung: Die Farben sind im Bild entfernt.
Straipsniai / Articles
161
Page 19
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBOKAS, VYTAUTAS ŠVEIKAUSKAS
ANHANG B. FORMENtABELLE FÜR DEKLINATION.
Als Beispiel wird die Flexion des Partizips nubėtas / hingelaufene in allen drei Genera und allen Kasus angeführt.
NUBĖTASPARTIZIP II, PASSIV, PRÄTERITUMMännlichSingularPluralNominativnubėtasnubėtiGenitivnubėtonubėtųDativnubėtamnubėtiemsAkkusativnubėtąnubėtusInstrumentalnubėtunubėtaisLokativnubė tamenubėtuoseWeiblichSingularPluralNominativnubėtanubėtosGenitivnubėtosnubėtųDativnubėtainubė tomsAkkusativnubėtąnubėtasInstrumentalnubėtanubėtomisLokativnubėtojenubėtoseNeutrumnubėta
162
Acta Linguistica Lithuanica LXXVI
Page 20
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
ANHANG C. TOOLTIP FÜR VERWENDUNGSBEISPIELE.
Als Beispiel wird die Schnellinfo für die Akkusativ-Singular-Form bėgį des Substantivs bėgis / Lauf, Getriebe, Eisenbahnschiene angeführt.
Straipsniai / Articles
163
Page 21
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBAKAS, VYTAUTAS ŠVEIKAS
ANANG D. FORMENTABELLE FÜR KONJUGATION.
Als Beispiel wird die Flexion des Verbs bėgti / laufen in allen Tempusformen und Modi angeführt.
164
Acta Linguistica Lithuanica LXXVII
Page 22
Eine mehrsprachige Website zur Grammatik der litauischen Sprache
Daugakalbė lietuvių kalbos gramatikos informacinė sistema internete
SANTRAUKA
Lietuvių kalbos kompiuterizavimo darbuose galima pastebėti du pagrindinius trūkumus: trūksta kartais net ir dažnai vartojamų formų ir pateikiami pertekliniai, lietuvių kalboje neegzistuojantys žodžiai. Kuriant Lietuvių kalbos gramatikos informacinę sistemą (LIGIS) stengiamasi išvengti jų abiejų. Visos formos generuojamos kompiuteriu automatiškai ir vėliau gauti rezultatai peržiūrimi žmogaus, kad būtų atmesti lietuvių nesuprantami žodžiai.
Lietuvių kalbos gramatikos informacinė sistema apima visus darinius. Šiuo metu duomenų bazę sudaro tik šakn inės beg-žodžiai. Jų yra apie 25 000. Palyginimui galima pateikti tokius duomenis: morfemos duomenų bazę sudaro yra apie 75 000 įrašų, bet jie surinkti iš visos lietuvių kalbos leksikos. Su šaknimi beg- ten yra 118 žodžių.
Kuriant Lietuvių kalbos gramatikos informacinę sistemą pagrindinis tikslas buvo pateikti išsamią gramatinę informaciją plačiajai visuomenei. Todėl buvo stengiamasi duomenis atvaizduoti kuo aiškiau ir suprantamiau. Svetainė pritaikyta naudotis ir mobiliuosiuose telefonuose.
Vartotojas, pateikęs žodį, gauna trijų tipų informaciją apie jį: žodžio struktūrą (pradinę formą bei pamatinį žodį dariniams), morfologinę informaciją (kalbos dalis bei jos morfologinės kategorijos – linksnis, giminė, skaičius, laikas, laipsnis ir t. t.) ir morfeminiai duomenys – žodis išskaidytas morfemomis, nurodant morfemos tipą bei požymius, jei morfema jų turi, pavyzdžiui, priesaga – darybinė / kaitybinė, galūnė – įvardžiuotinė, sutrumpėjusi ir kt. Kiekviena morfema žymima vis kita spalva. Reikia pabrėžti, kad Lietuvių kalbos gramatikos informacinė sistema yra pirmasis tinklapis Lietuvoje, kur vartotojui pateikiama informacija apie morfemos tipą. Be to, apdorojami ir žodžiai su sutrumpėjusiomis galūnėmis, kurios labai paplitusios, ypač šnekamojoje kalboje.
Kiekvienam duomenų bazėje esančiam žodžiui vartotojas gali gauti visą kaitybinę lentelę. Kai kuriems žodžiams pateikiama vartojimo pavyzdžių.
Kuriant Lietuvių kalbos gramatikos informacinę sistemą siekiama kuo didesnio tikslumo ir patikimumo. Todėl visos temos (žodžių pradinės formos – vardininkas, bendratis) suvedamos į duomenų bazę rankomis. Kaitybines formas sugeneruoti patikėta kompiuteriui, nes šiame jo darbe nebuvo pastebėta klaidų.
Svetainė pateikiama septyniomis kalbomis: lietuvių, anglų, vokiečių, prancūzų, italų, rusų ir japonų.
Ateityje planuojama Lietuvių kalbos gramatikos informacinę sistemą jungti su skaitmenintais žodynais ir sukurti vokiečių tinklapio CANOONET analogą.
Straipsniai / Articles
165
Page 23
DAIVA ŠVEIKAUSKIENĖ, ARŪNAS RIBAUSKAS, VYTAUTAS ŠVEIKAUSKAS