Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti tekstynai
Full text
BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ Vytauto Didžiojo universitetas LITHUANIAN LANGUAGE MORPHOLOGY AND SYNTACSIS ANNOTATIONS KEY WORDS: textbook, automatic morphological analysis, automatic syntactic analysis, language technology. INTRODUCTION For 25 years, language technologies have been created and developed at the Computer Linguistics Center (CLC) of Vytautas Magnus University (VMU) (Utka et al. 2017). The Lithuanian language 1 resources and tools prepared by VDU KLC are the basis without which it is impossible to do without computerization of the Lithuanian language and in July 2015 VDU KLC and VDU Faculty of Informatics pritaikant jai naujas technologijas. 2015 m. pradžioje Lietuva tapo visateise CLARIN ERIC nare, o together with partners from Vilnius University and Kaunas University of Technology founded the CLARIN-LT consortium and started the CLARIN-LT project. vykdyti projektą „Lietuvos narystė tarptautinėje mokslinių tyrimų infrastruktūroje – Bendroji kalbos 2 Šio straipsnio tikslas – pristatyti du lietuvių kalbos išteklius, anotuotus lietuvių kalbos texts prepared in VDU KLC: morphologically annotated texts MATAS and syntactically annotated tekstyną ALKSNIS. Šie ištekliai viešai prieinami CLARIN-LT saugykloje, paiešką galima atlikti per system ANNIS (see more in chapter 3), as well as searching in automatically morphologically annotated 3 technologies. Grammar analysis and statistical data on language are necessary for computerization of tekstyne galima svetainėje http://corpus.vdu.lt 4 . Anotuoti tekstynai – pagrindiniai ištekliai, atliekantys svarbų vaidmenį plėtojant kalbos language, for the development of computer tools for automatic language analysis. Thus, texts supplemented with grammatical notes are like raw material for further automatic language analysis. Such texts are used to develop other natural language resources and tools in areas such as 1 Online access: http://tekstynas.vdu.lt. 2 Online access: http://clarin-lt.lt/. 3 Access Internet: http://158.129.51.247:8080/annis-gui-3.4.4/. 4 It consists of 208 million words, automatically morphologically annotated. It is not further discussed in this article. The main focus is on the 1.6 million words morphologically annotated textbook MATAS, reviewed by a linguist.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 2 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 automatic language recognition systems, machine learning, automated translation, information mining, etc. Below the article separately presents the mentioned texts and data search in them through sistemą. Tekstynais ir jų duomenimis gali pasinaudoti tyrėjai bei studentai, tiriantys lietuvių kalbos the ANNIS grammar system, lecturers and teachers, creating tasks, and everyone who is interested in the language technologijomis. ANNOTED LITHUANIAN LANGUAGE TEXTS AND THEIR COMPOSITION FEATURES 1. Morfologiškai anotuotas tekstynas MATAS Morphological analysis is the first stage of language processing and often prepares the text as tolesnei sintaksinei ar semantinei analizei, todėl labai svarbu, kad morfologiniai anotatoriai veiktų accurately as possible, because the errors made in the morphological domain interfere with the automatic analysis of the language of other domains. The first morphological annotator of Lithuanian language was created about 25 years ago. The author is Vytautas Zinkevičius (see Zinkevičius 2000). Currently, two morphological annotators are publicly available for the Lithuanian language: the above-mentioned V. Zinkevičius program, usually called Lemuoklis (from the word lema – heading form), and the annotator available on the portal Lithuanian 5 language syntactic and semantic analysis information system (hereinafter semantika.lt). This article was prepared using semantika.lt morphological annotator. code. Semantika.lt annotator developed on 6 Hunspell platform as an open source program; this annotator is updated, supplemented with new pristatomas tekstynas MATAS anotuotas naudojant Lemuoklį, o sintaksiškai anotuotas tekstynas words (now includes 171 000 lems), its rules are adjusted (more Dadurkevičius 2017). In 2017, a study Nors Lemuoklis veikia gana gerai, bet šios programos negalima atnaujinti, nes ji yra uždaro conducted and assessed the quality of both annotator, it was found that semantika.lt annotator works more accurately (Kapočiūtė-Dzikienė et al. 2017). The following textbook MATAS was prepared between 2002 and 2014. The textbook of 1 million words, compiled in 2006, was taken as a basis, it was supplemented with new texts, reorganized. First, MATO texts were automatically annotated morphologically, then they were reviewed and arranged by a linguist. 5 Online access: http://tekstynas.vdu.lt/page.xhtml?id=morphological-annotator; this program not only morphologically annotations the text, but can also synthesize, i.e. generate the necessary forms; it is also used to analyze ancient writings. 6 Prieiga internete: http://semantika.lt/SyntaticAndSemanticAnalysis/Analysis.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti tekstynai | 3 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 MATÁ consists of 1.6 million texts, consisting of documents, fiction, and publications mokslinių tekstų. Kaip ir Dabartinės lietuvių kalbos tekstyne, didžiausią dalį (36 proc.) sudaro (see figure 1). 1 PAV. MATAS is available in several formats: the so-called KLC format, which uses somewhat unusual abbreviations of parts of language and grammatical categories (the text available in the http://clarin-lt.lt/ repository is also annotated in this format), e.g.: <word="Tarp" lemma="tarp" type="prln"> <space> <word="dense" lemma="dense" type="first teig nelygin.l"> <space> <word="suaugusių" lemma="suaugti(-ga,-go)" type="dlv teig nesngr veik.r būt.kart.l neįvardž vyr.gim dgsk K"> <space> <word="tree" lemma="tree" type="dktv vyr.gim dgsk K"> <space> <word="nowhere" lemma="nowhere" type="first teig nelygin.l"> <space> <word="vis" lemma="vis" type="prvks teig nelygin.l"> <space> <word="meet" lemma="meet(-to,-tė)" type="vksm teig sngr just.nuos be.d.l vnsk IIIasm"> <space> <word="heaven" lemma="heaven" type="dktv vyr.gim vnsk K"> <space> <word="sleeper" lemma="sleeper" type="dktv vyr.gim vnsk V"> <sep=". "> <p> MAT is also annotated in the international TEI P5 format. It abbreviates parts of speech and grammatical categories by a single letter or number, for example vatstdv3 means respectively: Scientific literature 24% Periodicals 36% Administrative tekstai 21% Fiction texts | 4 19%
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 verb, personified form, affirmative form, recursive, direct, past tense, singular, third person. The following is the same text excerpt in TEI P5 format: <w lemma="between" ana="#r">Between</w> <pc> </pc> <w lemma="density" ana="#ptn">density</w> <pc> </pc> <w lemma="suaugusti(-ga,-go)" ana="#vdtnvknvdk">adults</w> <pc> </pc> <w lemma="tree" ana="#dbvdk">trees</w> <pc> </pc> <w lemma="where not where" ana="#ptn">where not where</w> <pc> </pc> <w lemma="vis" ana="#ptn">vis</w> <pc> </pc> <w lemma="to meet" ana="#vatstdv3">to meet</w> <pc> </pc> <w lemma="heaven" ana="#dbvvk">heaven</w> <pc> </pc> <w lemma="sleeper" ana="#dbvvv">sleeper</w> <pc>.</pc> Another format, in which morphologically annotated texts are available on the http://corpus.vdu.lt website, is based on the example of MULTEXT-East format. According to this format, each part of 7 the language has a different number of morphological categories (from 2 to 14), they are written by abbreviations, for example, in the notation NcmpnnN stands for noun, c stands for general noun, m stands for masculine, p stands for plural, n stands for noun, n stands for imperfect noun, a dash at the end indicates that this word, for example, universities, is not assigned any semantic notation. 8 In the third chapter we will introduce searching in both MATO and ALKSNIO textbooks using ANNIS system. It uses another format for morphological notes, based on the Leipzig glossary notes (see section 3 for details). Different annotation systems are characteristic not only of Lithuanian language technologies. This problem led to the development of the Pepper conversion tool (Zipser et al. 2010). It shows the stages of creation of various assets and the annotation tools used. As can be seen from the above examples, KLC uses four morphological notation systems. Several different annotation formats have emerged as a result of the long-term development of morphologically annotated texts. At one time, no standards of annotation were applied at all, simply abbreviations of parts of speech and grammatical categories created by the creator of the morphological analyzer V. Zinkevičius were used. They shall be easily understood by consumers. After some time, due to international cooperation, the morphologically annotated textbook was re-annotated using 7 Available online at: http://nl.ijs.si/ME/V4/msd/html/index.html. 8 For more information about these certificates see http://corpus.vdu.lt/lt/morph.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 5 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 international standards. As mentioned, semantika.lt (the main purpose of this system was to create online services) morphological parser uses MULTEXT-East format. The certificates are short, but difficult for humans to understand, they are more suitable for computer analysis. The morphologically and syntactically annotated text ALKSNIS in MULTEXT-East format. The ANNIS textbook search tool (see section 3) made it possible to combine ALKSNIO and MATO morphological annotation formats. The first idea was to use Universal Dependency (UD) annotations, because it is a fairly common annotation system and it is easy to read. However, the UD tags are very long and this proved inconvenient when searching through ANNIS: the tags for each word take up a lot of space, so when analyzing the context of the searched word, you had to scroll the window arrow in one direction or another. Because of this practical problem, it was decided to adapt the Leipzig glossing certificates because they are easy to read and relatively short. Despite the variety of annotation labels described here, it should be emphasized that the ANNIS interface of both textbooks uses only one (adapted Leipzig glossary) system, which should be a great relief for users. Note that MATE does not specify sentence limits, this can be done quickly automatically, but you need to check the results. MATTHEW, as well as the textbook ALKSNIS discussed below, is available in the repository of the http://clarin-lt.lt/ portal. It allows searches using the ANNIS system (see section 3 for details). 9 10 2. Syntactically annotated text ALKSNIS ALKSNIS is a syntactically annotated textbook of Lithuanian language developed in 2016. Its development was supported by the Common Language Resource and Technology Infrastructure Consortium of the European Research Infrastructure (CLARIN ERIC) (MTI-02/2015). The textbook is publicly available in the repository of the http://clarin-lt.lt/ 11 portal. How to use it is explained in the user guide. 12 The textbook is designed as a benchmark for Lithuanian syntactic analysis. This type of texts are characterized by the fact that they are small, of optimal genre composition, and syntactic analysis, although performed automatically, is reviewed and corrected by linguists. Based on the syntax analysis benchmark, it is possible to create a statistically-based parser and expand automatic syntax analysis (Bielinskienė et al. 2016). 9 Online access: https://clarin.vdu.lt/xmlui/handle/20.500.11821/9. 10 Access on the Internet: http://158.129.51.247:8080/annis-gui-3.4.4. 11 Online access: https://clarin.vdu.lt/xmlui/handle/20.500.11821/10. 12 Online access: https://youtu.be/PlE0PWurb4Y.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 6 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 Since ALKSNIS is a newer Lithuanian language resource than MATAS, it is less described, so in this article we will pay more attention to ALKSNIS, and also discuss the issues arising from syntactic annotation. 2.1. Parts of the textbook ALKSNIS consists of 30 599 words, 2355 syntactically annotated sentences. In comparison with texts in the languages of neighboring countries, this textbook is quite similar and, it can be said, is of sufficient size: the Latvian textbook consists of 3800 sentences, 53 thousand words (Predkalnina 13 et al. 2016); The PENN (USA) syntactically annotated text of Estonian language consists of much 14 – 1400 sakinių, 10 600 žodžių (Muischnek ir kt. 2014); lenkų kalbos 15 – 8227 sakinių. Pirmąjį more – about 3 million words. 16 ALKSNIO tekstai apima keturias žanrines dalis (žr. 2 pav.). Bendrosios bei specialiosios periodical and fictional literature parts are acircle, and the share of administrative texts is the specifinė lietuvių kalba: sakiniai dažnai būna ilgi, juose gausu dokumentų pavadinimų, nuorodų į kitus dokumentus ir pan. Tokių sakinių anotavimas neatskleistų tipiškos sintaksinės struktūros. lowest, because it is 2 PAV. Syntactically annotated textbook ALKSNIS structure of text genres. A part Tekstynui sudaryti imti ištisi, nesutrumpinti tekstai. Sakiniai atrinkti iš kuo įvairesnių of the fictional literature consists of sentences from works of Lithuanian authors. Selected sentences were prepared for computer analysis: text encoding was checked, tables and pictures were deleted. 13 Internet access: http://sintakse.korpuss.lv/index.html. 14 Accessed online: https://metashare.ut.ee/repository/browse/estoniantreebank/4eb86e5e463411e2a6e4005056b400242c46832883754ad7bd89d07f69d6a0fc/. 15 Internet access: http://zil.ipipan.waw.pl/Sk%C5%82adnica. 16 Online access: https://catalog.ldc.upenn.edu/ldc99t42.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 7 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 2.2. Textbook annotation and visualization tools Automatic syntax analysis is based on the syntax parser developed in Haskell language during the VDU KLC project “Lithuanian language syntax-semantic analysis system for textbook, Lithuanian Internet and public sector applications”. The parser relies on rules to analyze the morphologically annotated files 17 of the modules and as a result produces sentences broken down into roles) and syntactic functions pagrįstu (angl. rule-based) metodu. Šis įrankis skaito semantika.lt segmentavimo ir morfologinės (Boizou et al. 2014: 69). The web services only generate JSON files, but another KLC tool converts the dėmenis ar sintagmas, generuoja sintaksinius medžius, nurodo jų teminius vaidmenis (angl. thematic parsing results into PML (Prague Markup Language) format. Syntactically annotated visualization and editing of dependency trees in PML format using Prague Charles sentence segmentation and morphological analysis, automatic syntax analysis (rule-based tekstynui yra generuojami priklausomybių medžiai (angl. dependency trees). Šis formatas leidžia universiteto UFAL sukurtą TrED redaktorių 18 (žr. 3 pav.). Visi automatiškai suanotuoti sakiniai patikrinti ir ištaisyti kalbininkų. Taigi sintaksinė analizė apima tris lygmenis: automatinį žodžių ir metodu) ir rankinį sakinių tvarkymą (plačiau žr. 2.4). 2.3. ALKSNIO data structure After automatic syntax analysis, the sentences are displayed in a graphic tree (see Figure 3). Each node in the tree corresponds to a sentence word, punctuation mark or other sentence unit (symbol, digit, etc.). Dependency relationships between words are indicated by "branches" or edges. 17 The syntax parser was developed by L. Boizou and F. Zamblera at the Computer Linguistics Center of VDU. 18 Online access: https://ufal.mff.cuni.cz/tred/.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 8 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 3 PAV. Syntactically annotated sentence dependency tree in TrEd editor format The following order is given for all words (see Figure 3): 1) the specific word form used in the sentence (e.g., gyventojai); 2) heading (vocabulary form) – a lemma (e.g., inhabitant); (3) morphological marks (e.g. Ncmpnn-); 4) a syntax function (e.g. Sub) (see explanations of abbreviations below). 2.4. Analytical components Morphological analysis. Morphological annotation is the first step of automatic analysis and is therefore essential for higher-level syntactic or sentence analysis. Morphological analysis was performed using semantika.lt annotator (Dadurkevičius 2017). The morphological references used in ALKSNYJE are based on the example of the MULTEXT-East format (described in more detail in Chapter 1).
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti tekstynai | 9 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 Sintaksinės analizės lygmuo. Sintaksinis analizatorius apdoroja morfologiškai anotuotus files and returns sentences broken down into domains or syntax, generates dependency grammar. It nurodo jų elementų sintaksines funkcijas. Analizė remiasi priklausomybių gramatikos modeliu is usually applied to typologically similar languages that are characterized by conjugation of word forms and free word order, such as Slavic languages (PDT, SynTagRus), Latvian, etc. Syntactic 19 dependencies are expressed according to a hierarchy. It begins with the main element of the 20 sentence – the top of the tree, to which the other elements of the sentence are connected according to their dependence (i.e. syntactic relationships). Abbreviations of syntactic markers and information about syntactic relationships and dependencies are based on Czech work (Hajič et al. 1999); they were among the first to develop syntactic analysis for synthetic Czech. Currently, 18 basic syntactic markings are used in ALKSNYJE (not counting their variants): Sub – subject, Pred – predicate, Obj – object, etc. (see table 1). Variations of the signs occur when two or double signs are combined into one, for example, when a compound noun or verb type adverb is used with two signs: PredN and PredV respectively; when marking the subordinate (attributive) element of a compound sentence, the double mark Pred_Atr is marked next to the subordinate (attributive) element of the sentence; in order to show the connection (coordination) to the homonymous elements of the sentence, _Co is marked as part of the double marking, i.e. to each of the listed homonymous suffixes Obj_Co is written, etc. By the way, in compound sentences there can be triple affirmations, for example, when it is necessary to affirm homogeneous secondary sentences: Pred_Atr_Co – this is the secondary (attribut) affirmation sentence. TABLE 1. Syntactic functions and their markings in text Syntax function Example Sub Subject Pred Predicate He [Sub] said He went [Pred]; (or auxiliary word) PredV was [Pred] satisfied Must PredN Noun part of Was [Pred] satisfied [PredN] I am Verb part of the predicate [Pred] repay [PredV] predicate Obj Object waiting for a guest [Obj]; need to solve [Obj] Atr Attribute (certificate) Secondary [Atr] school; Member State [Atr] Adj Circumstances Now [Adj] will decide; On the street [Adj] 19 Prague Dependency Treebank, see for more details. https://ufal.mff.cuni.cz/pdt3.0. 20 Internet access: http://www.ruscorpora.ru/en/.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 16 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 9 PAV. If there is a multi-level mixed join, then the join in the construction is the vertex, and from it all the other non-jointed elements follow (commas with their enumerated words are in the same dependency domain). Other punctuation marks are marked Aux (see Figure 10).
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 17 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 10 PAV. Example of annotation of a mixed connection structure (the figure shows more cases of a connecting connection) The top of a connective sentence is a connective conjunction. Connecting words, which consist of adverbs, pronouns or particles (e.g., adverbs therefore), depend on the predicate of the second part of the sentence and perform the function of circumstance (Adj), while the function of the Coord of connection is performed by the comma, as in unconnected sentences or sentences with repetitive connections, which are here considered particles (see Figure 12), as in the case of the connection of the parts of the sentence (see Figures 11, 12).
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 18 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 11 PAV. Example of annotating a connective sentence with a connective word 12 PAV. Annotation of connecting structure for text | 19
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 Sudėtinga anotuoti skaitines išraiškas, pvz., valandas, metus. Jos anotuotos kaip aplinkybės with certificates or as certificates (see Figures 13, 14). 13 PAV. Datos anotavimas 14 PAV. Annotation of dates and other numerical expressions (only a syntactically annotated sentence fragment is shown in the picture) Textbooks | 20
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 It is not always easy to annotate compound consonants. In some cases, they are annoted as they are understood in Lithuanian grammar, i.e. as noun and verb-type tarians (e.g., is [Pred] simple [PredN]; can [Pred] be [PredV]), and in some cases the second accent of the tarian is considered an object (e.g., I intend [Pred] to call [Obj]). It was difficult to annotate constructions with paired connectors because both connectors are hierarchically equivalent. In this case, 15 PAVs were considered to occupy a higher position in the hierarchy. Example of annotation of compound sentences with pairs of pagrindinio sakinio dėmens tarinys, nuo jo priklauso šalutinio sakinio dėmuo (žr. 15 pav.). conjunctions It was decided to treat the conjunctions of the connective connection as particles, Galiausiai svarstyta dėl kai kurių kalbos dalių statuso. Pavyzdžiui, kartojamuosius because their function in the structure of the dependency tree is similar to that of particles (see Figure 12). Syntactically annoting with the hands solved a lot of problems, here only a few of them were named. Many of the annotation problems have been solved by the traditional grammar approach, only by agreeing on a certain way of annotation. In other cases, it was necessary to adapt to syntactic solving anotavimo ypatumų ir anotuojant nuosekliai laikytis susitarimo. Tikimasi iškilusias problemas by further developing automatic syntactic analysis. system. This is one of the challenges of using Kitame poskyryje pristatoma duomenų paieška anotuotuose tekstynuose naudojant ANNIS annotated texts, because the data search in such texts, where there are many marks and even several annotation elements, must be diverse. For this reason, the tool ANNIS developed by researchers of Humboldt University (Germany) was chosen.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 21 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 3. ANNIS search Naudoti ANNIS įrankį nėra labai sudėtinga (plačiau žr. Zeldes 2016). Reikia pasirinkti an annotated textbook (SEE part of the textbook (this textbook is divided into documents, fiction, mokslinius tekstus ir periodiką) arba ALKSNIO tekstyną (jis įvardytas taip: Alksnis_paula_1.0) ir literature, etc. The structure of queries is described 3.1 and 3.2. ANNIS provides the corresponding concordance rows with the relevant information (see Figure 16). The concordance line can be extended from 5 to 25 words (punctuation included) both from the left and from the right; also possible to export results. Both texts use grammar notes drawn up according to the Leipzig glossing notes. There are some additional marks that are not included in the above marks, for example, ~COMP stands for the higher degree. This is the sign of the sign of the sign. Parts of speech are indicated by 21 Universal Dependency Format 16 PAV. Paieškos per ANNIS įrankį rezultatų fragmentas MATO documents part of the search for the lemma command 21 Available online: https://www.eva.mpg.de/lingua/resources/glossing-rules.php. Other annotated texts often use Universal Dependency (UD) marks (see http://universaldependencies.org/; see also Figure 18). As mentioned in the first section, the UD labels are very long (e.g., the accusative, feminine and plural are indicated as case=accusative|gender=feminine|number=plural), they take up too much space in the ANNIS interface and therefore are not easy to read.
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti texts | 22 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 17 PAV. Fragment of the results of the ANNIS tool search in the HUNGER text searching for the lema command (to mark the syntactic links with arrows, press the plus sign next to deps) 3.1. Simple search The search for a single attribute is called simple. The annotation provided (the names of the grammatical categories, the grammatical markings used, etc.) depends on the structure of each textbook, not on the ANNIS tool. Atkreipiame dėmesį, kad formuojant užklausą reikia skirti didžiąsias ir mažąsias raides. 3.1.1. Search for exact forms of words (strings of characters) Paprastoji paieška vykdoma nurodant kategorijos pavadinimą, lygybės ženklą ir kategorijos a value enclosed in quotation marks, e.g.: lemma="language" (this query means that the language of the lemma is being searched); pos="ADJ" (ieškoma būdvardžių); gram=".M.SG.GEN." (searches for specific grammatical categories, in this vienaskaitos vyriškąja gimine kilmininko forma pavartotų žodžių).
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti case texts | 23 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 When searching for a specific form of a word in the text, you must specify the category name tok or simply enter the search form in quotation marks, e.g. : tok="pasakė" arba "pasakė". Both annotated texts can be searched for lems, parts of speech, grammatical categories, konkrečių žodžių formų, o ALKSNYJE dar galima ieškoti sintaksinių pažymų, pvz.: syfun="Atr" (searching for attributes). It is also possible to search for deps in ALKSNY, but they are searched according to a different principle (see 3.2.3). 3.1.2. Search for character structures ANNIS allows you to search using regular expressions; they are marked with special symbols. Regular expression searches are basically written in hyphens as well. kaip tikslių žodžių formų paieškos užklausa, tik reikia angliškas kabutes pakeisti pasviraisiais 3.1.2.1. Bet kokių simbolių paieška The most important special character is the dot, which replaces any single character (letter, digit, etc.), e.g. : lemma=/pl.t.s/ (search for, e.g., wide, width, wide); /p.sak.s/ or tok=/p.sak.s/ (search for, e.g., sakys, pasakos, pasakas, pasakęs, posakis); /201./ arba tok=/201./ (ieškoma, pvz., 2010, 2011, 2012…). Since a dot can replace any character, the dot symbol itself must be searched using kombinaciją pasvirasis kairinis brūkšnys + taškas (\.), pvz.: gram=/\..\.SG\.GEN\./ (search for, e.g.,.M.SG.GEN.,.F.SG.GEN.,...). 3.1.2.2. Search for one of several characters lemma=/pl[ao]t.s/ (e.g., wide, width, but not wide); lemma=/pl[ao]t.s/ (e.g., wide, width, but ieškomų simbolių. Jie rašomi laužtiniuose skliaustuose, pvz.: not wide); /201[0124]/ or tok=/201./ (search for 2010, 2011, 2012, 2014); texts | 24 /t[ei]lp./ arba tok=/t[ei]lp./ (ieškoma, pvz., telpa, tilpo, bet ne talpa);
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 gram=/\.[MFN]\.SG\.GEN\./ (seeking.M.SG.GEN.,.F.SG.GEN.,.N.SG.GEN.). Please note that this method searches by one character from the set of given characters: /dirb[aius]/ allows to search for dirba, dirbi, dirbu, dirbs, but not dirbau or dirbsi. 3.1.3. Iteration operators It is possible to repeat (regular and special) characters, e.g. : lemma=/..važiuoti/ (search for, e.g., go, leave, bypass, come, do not go); /dirb[ao][mt]e/ or tok=/dirb[ao][mt]e/ (to search we work, you work, we worked, dirbote). The number of characters to be searched for is constant, for example, two characters (arba [ao][mt]), taigi pagal užklausą lemma=/..važiuoti/ negausite rezultatų su lema įvažiuoti. in the examples given More options are available with the repeat operators, their values are as follows: ? – the symbol is used only once or not at all; + – a certain symbol is used at least once (i.e. one, two, three, four and more times); * – a certain character is used n times (i.e. it may not be used at all, used one, two or more times) These operators help to find the characters that follow or precede them, e.g. : /šild?o/ (seeking for warm or warm); /Ma+u/ (search for Mau, Maau, Maaau, Maaaau, etc.); /oi*/ (search for o, oi, oii, oiii, etc.). The above operators are compatible with searching for any or one of several characters, e.g.: texts | 25 lemma=/.*važiuoti/ (ieškoma, pvz., važiuoti, išvažiuoti, įvažiuoti, nuvažiuoti, pravažiuoti, nepravažiuoti); /šauk[aiu]*/ arba tok=/šauk[aiu]*/ (ieškoma, pvz., šauk, šauki, šaukia, šaukiu);
AGNĖ BIELINSKIENĖ, LOÏC BOIZOU, ERIKA RIMKUTĖ. Lietuvių kalbos morfologiškai ir sintaksiškai anotuoti BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 /šauk[aiu]+/ or tok=/šauk[aiu]+/ (for example, shout, shout, shout, but not shout); gram=/.*\.F\.. */ (looks for string.F. among any other strings of characters). Note that structures such as [aiu]* or [aiu]+ do not affect the order of the searched characters. That is, return results containing a or i, or u + a; either i or u + a; or i, or u… Thus, the following combinations of characters are recognized: aaa, i, uuuuu, uiua, uaa, etc. 3.1.4. Search for alternatives The vertical dash (|) allows you to search for alternatives, e.g. : lemma="bad" | lemma="good"; /d(au|ū)ž. */ or tok=/d(au|ū)ž. */ (search for, e.g., beat, beat, beating); gram=/.*\.~(COMP|SUP)\.. */ (searches for.~COMP. and.~SUP. among any strings of characters); "only" | "only" or such="only" | such="only" (search only and only). 3.1.5. Negatively formed requests Although this option is more useful when combining features in complex searches, it is also possible to specify what is not to be searched for in simple searches, e.g.: pos!="NOUN" (all words that are not nouns are searched); tok!=/. *[aąeęėiįyouųū]/ (searches for all forms of words that do not end in a vowel). Unless a compound search query is selected (see 3.2), negatively formed simple queries can cause problems, especially if they include many words (as in the first case – nouns), so this query should be used with caution. 3.2. Advanced search Complex search consists of several simple searches. There is also an additional section that describes the relationship between simple searches (indicated by position). Each part of a simple lookup is connected by an ampersand (&).
