Page 1AURELIJA TAMULIONIENĖ
Institute of the Lithuanian Language
Research areas: studies of the usage and variation of contemporary Lithuanian.
THE STATUS OF THE LITHUANIAN LANGUAGE IN THE DIGITAL AGE
About the 24th Jonas Jablonskis Conference
Directions for the Development and Use of Digital Language Resources
Jono Jablonskio (1860–1930) – žymiausio lietuvių kalbos normintojas, reikšmingų lietuvių kalbos mokslui ir praktikai darbų autorius. Pasak Zigmo Zinkevičiaus, „[V]isa Jablonskio veikla – tai ištisa epocha lietuvių bendrinės kalbos istorijoje. Iš mūsų rašomą kalbą į tinkamas vėžes [...], jos raidą pasuko sveika linkmė“ (Zinkevičius 1992: 155).
The scientific conference named after Jonas Jablonskis was first held in 1993, organized by the Language Culture Department of the Institute of the Lithuanian Language and the Department of Lithuanian Language at Vilnius University. Since then, the conferences have become a tradition and have been held every autumn (since 2003, alternately at the Institute of the Lithuanian Language and Vilnius University).
On September 29, 2017, the 24th scientific Jonas Jablonskis Conference Directions for the Development and Use of Digital Language Resources (in English, Digital Language Resources, Directions of Their Development and Possibilities to Harness Them) was held at the Institute of the Lithuanian Language. The conference was organized by the Center for Standard Language Research of the Institute of the Lithuanian Language together with the Institute of Applied Linguistics at Vilnius University. In terms of its subject matter, this conference differed from previous ones; its aim was to discuss issues arising at the intersection of languages and digital technologies, encountered in developing digital resources and creating a common digital market. The conference focused on the semantic level of digital resources (the creation of word networks), as well as providing digital resources with an audio form, creating new opportunities for disseminating digital resources in multilingual communities, and so on.
The conference was attended by speakers from abroad: from the European Parliament and Spain, as well as researchers from various Lithuanian institutions: the Institute of the Lithuanian Language, Vilnius
Straipsniai / Articles
313
Page 2AURELIJA TAMULIONIENĖ
the university, Vytautas Magnus University, the Baltic Institute of Advanced Technology, and UAB “Netcode.” Seventeen presentations were delivered at the conference. The presenters gave their talks and shared their research and experience in four sessions.
The conference began with opening remarks. The conference participants and presenters were welcomed by Member of the European Parliament Algirdas Saudargas and Prof. Dr. Jolanta Elena Zabarskaitė, Director of the Institute of the Lithuanian Language. The opening address was delivered by Dr. Rita Miliūnaitė, Chief Researcher at the Centre for Standard Language Research of the Institute of the Lithuanian Language. It was noted that the links between the Lithuanian language and information technology have strengthened considerably in recent years; we have tangible results both in digitising language resources and in developing tools for language analysis, recognition, and synthesis. Yet, compared with the rapid changes taking place worldwide in this field, Lithuanian still has a great deal of catching up to do. Because Lithuanian is an inflected language, it cannot be directly adapted to information technologies developed for English. Consequently, our current machine-translation and other computer tools remain very imperfect, while research worldwide has already moved further, toward the development of artificial intelligence, the incorporation of objects, and penetration into ever deeper layers of language.
At the plenary session, the first presentation, “Living Language in Artificial Intelligence,” was given by Member of the European Parliament Algirdas Saudargas. The speaker raised a fundamental question that every linguistic community must answer: to what extent should it develop resources and technologies for its own native language, and what can be acquired ready-made on the basis of other languages? The presentation raised concerns that Lithuanian, along with the languages of some other small countries, “has little or no technological support.” Language technologies do not receive adequate attention on the agendas of either Lithuanian or European politicians. The presentation offered insights showing that the development of artificial intelligence is converging toward a hybrid form, in which neural networks correspond to unconscious brain mechanisms, while conscious thought is shaped and modelled by traditional, so-called symbolic artificial intelligence. Every linguistic community must ensure that language technologies corresponding to symbolic artificial intelligence reflect the entire structure of the native language comprehensively and precisely, and that a rich environment of the native language (and native culture) is created for the interaction of neural networks.
The second plenary presentation, “Language equality in the digital age: towards a human language project,” was delivered in English by Rafael Rivera, director of the consultancy Claves. The presentation provided a detailed overview of the European Parliament’s Science and Technology Options Assessment Panel study “Language equality in the digital age: towards a human language project” (the study is available online in English: http://www.europarl.europa.eu/RegData/etudes/STUD/2017/598621/EPRS_STU(2017)598621_EN.pdf).). The presentation reviewed the current state of native-language technologies, quantifying
314
Acta Linguistica Lithuanica LXXVI
Page 3Lietuvių kalbos padėtis skaitmeniniame amžiuje
the economic, social, and linguistic consequences of language in the digital age, and discussed the European Union’s information and communications technology policy. In the digital age, language technologies pose a major challenge for countries with small numbers of speakers. Poorly developed language technologies create divides and barriers. This affects cross-border services, worker mobility, and trade. Such barriers are caused by uncoordinated research and insufficient funding. The speaker emphasized that language technologies do not receive appropriate attention from European politicians. Based on an analysis of the current situation, an initiative is being developed to unite the policies of institutions, research, industry, markets, and public services.
The final presentation of the plenary session, “Written and Spoken Lithuanian in Conventional and Electronic Media,” was delivered by Laimutis Telksnys, Professor at the Institute of Mathematics and Informatics of Vilnius University. The presentation argued that methods and tools—hardware and software—must be developed to ensure the correct use of written and spoken Lithuanian in paper documents and electronic media. The speaker raised an important question: how will a few million Lithuanians fare so as not to disappear in an ocean of more than 7 billion people who write and speak, and of emerging smart machines? The professor also offered an answer: to prevent this, methods and tools—hardware and software—must be developed to ensure the correct use of written and spoken Lithuanian in paper documents and electronic media, and to harmonize the use of written and spoken Lithuanian and other languages in paper documents and electronic media. Transcription tools must be created to represent the characters of non-Lithuanian written languages using Lithuanian written-language characters, and to render the sounds of non-Lithuanian spoken languages using Lithuanian spoken-language characters suitable for the automatic synthesis of the sounds of spoken Lithuanian words.
Three presentations were delivered in the first session, “Speech Recognition and Synthesis.”
In their presentation “Beyond Gutenberg’s Printing Press: Lithuanian Speech in the Latest Technologies,” Audrius Valotka (VU) and Gediminas Navickas (VU) discussed Lithuanian speech in the latest technologies and introduced the project funded by the Structural Funds Services Controlled by Lithuanian Speech – LIEPA (available online: https://www.raštija.lt/liepa),), which, among other things, developed a Lithuanian speech synthesizer and recognizer. During this project, the gates to the digital space were opened to Lithuanian speech; for this reason, the project Development of Services Controlled by Lithuanian Speech – LIEPA 2 is planned. The presenters discussed the results of the LIEPA 2 project, which will be open and freely available to everyone, encouraging the use of Lithuanian speech in information technology products. The most important result of the new project will be the ability to naturally
Apžvalgos / Surveys
315
Page 4AURELIJA TAMULIONIENĖ
communicate with devices (mobile phones, tablets, smartwatches, robots), give commands using the human voice, and understand their responses. LIEPA 2 is intended to create typical services demonstrating new possibilities for using Lithuanian voice-controlled services: a controller for an educational robot, a dialer, a taxi-calling service, a mobile synthesizer for blind people, an internet news reader, and an interlingual communicator.
In their paper “Lithuanian Voice-Controlled Services: Current Situation and Prospects,” Laimutis Telksnys (VU) and Gediminas Navickas (VU) discussed the current situation and prospects for Lithuanian voice-controlled services, as well as the systematic work being carried out to develop these services and expand their use. The work brings together the knowledge of information technology specialists and Lithuanian philologists, as well as engineering capabilities. In the first phase, seven Lithuanian voice-controlled services were created, along with the infrastructure needed to ensure that they function. In the future, new Lithuanian voice-controlled services are planned for mobile electronic environments—smartphones, tablets, smartwatches, and robots.
Pius Kasparaitis (VU) and Gintaras Skeris (VU) discussed the present and prospects of Lithuanian speech synthesis. Their paper, “The Present and Prospects of Lithuanian Speech Synthesis,” presented various text-to-speech services—the LIEPA synthesizer, which is distributed freely together with its source texts, enabling others to find new applications for it. For example, websites already provide audio recordings alongside articles: zinios.lt, unikopedarejas.lt, vilnius.lt, m.delfi.lt, vle.lt. The text-to-speech service RoboBraille can automatically convert any text documents into audio files; announcements in synthetic voices are now heard at the Vilkaviškis bus station, virtual assistants are already operating on the website of the Migration Department, and so on.
Four papers were presented at the second session, “Digital Lithuanian Language Resources.”
The first paper, “E. kalba—an Innovation in the Use of Digital Language Resources,” was given by Elena Jolanta Zabarskaitė (LKI), Deimantė Budriūnaitė (LKI), and Skirmantas Šermukšnis (NetCode). It provided a detailed presentation of the Lithuanian Language Institute’s new project to expand the information infrastructure for Lithuanian language resources (LKIIS) (available online at www.lkiis.lt). The paper focused primarily on the services generated by the infrastructure of the Lithuanian Language Institute’s new project, E. kalba: “Search in the Word Network,” “E. Marketing,” “E. Concepts,” and “E. Advice.” The presenters described the functionality and technological aspects of these services. The project’s functionality will incorporate innovative artificial intelligence technologies for analyzing users’ opinions; the “E. Concepts” service will provide expanded search and enrichment of concepts by integrating additional language resources, such as bilingual dictionaries,
316
Acta Linguistica Lithuanica LXXVI
Page 5Lietuvių kalbos padėtis skaitmeniniame amžiuje
word networks in other languages, and an interactive word-formation assistant will make it possible to quickly and conveniently select appropriate word-formation methods based on the desired semantic category and group or on the means of word formation. The paper also presented the main results of the E. kalba project and its anticipated benefits.
In her paper “Search Options in the Online Lithuanian Neologism Database,” Rita Milunaitė (LKI) explained how the database can meet the diverse needs of language users by revealing a broad range of Lithuanian Neologism Database features (available online: http://naujažodžiai.lki.lt/)). At present, searches can be conducted by parameters such as headword, neologism origin, original form, spelling variants, domain of use, and so on. Database administrators also have access to more advanced searches, which are needed to edit the data from different perspectives. An advanced information retrieval system will be developed as part of the LKIS project. The presenter emphasized that the Lithuanian Neologism Database is a flexible digital resource that is open to expansion. This is primarily dictated by the nature of the data itself—the constantly changing and renewing layer of Lithuanian vocabulary—as well as by new data features becoming available to neologism researchers and changing user needs. The resource is intended to be integrated into LKIS (available online: http://lkis.lki.lt/) and included not only in the system-wide search across the data it contains, but also used to create word networks. The paper illustrated how useful this database can be both to language specialists and to all users interested in neologisms.
Daiva Murmulaitė (LKI) continued the discussion of neologisms in her paper “Prospects for Research into Neologism Formation (the Case of the Lithuanian Neologism Database.” She discussed research into neologism formation and its prospects, including how emerging patterns of neologism formation can be studied using the Lithuanian Neologism Database. The presenter explained that, even during the initial stage of creating this database, preliminary preparations had been made for analyzing neologism formation: dedicated fields were planned for specifying the types of derivatives, their forms of formation, derivational categories (meanings), related words, and more; some data had already been collected. In the special-features field, some neologisms are marked as nonce formations (for example, varškėfobija and others), author-created formations (for example, demagogėja and others), or blends (for example, murmanas and others). The presenter emphasized that a thorough, high-quality analysis of word formation must also take into account certain aspects initially planned for recording: the part of speech of the base word, changes to the derivational base, analogies, the role of translation in creating a neologism, atypical manifestations of word formation, and more. It is important to create a good indexing system—comprehensive, flexible, and convenient to use, expand, and revise.
In his paper “The Structure of the Geoinformation Database of Lithuanian Place Names, Its Links to Other Onomastic Databases, and Directions for Development,” Laimutis Bilkis (LKI)
Apžvalgos / Surveys
317
Page 6AURELIJA TAMULIONIENĖ
presented the Geoinformation Database of Lithuanian Place Names (available online: http://lkis.lk.lt/lietuvos-vietovardziu-geoinformacine-duomenu-baze). We asked the speaker to discuss the structure of this database and its links to other onomastic databases. The speaker presented the search fields available on the database’s public-access page: object type, object status, current administrative-territorial affiliation (municipality, eldership, settlement), interwar administrative-territorial affiliation (county, township, settlement), and river tributaries. Place-name information in the database can be searched using the following attributes: name, gender, number, accentuation, word-formation type, origin according to the language of the base word, and origin according to the lexical group of the base word. The database can be used to find authentic place names recorded during the interwar period within the territory of a settlement (village, church village, manor, town, or isolated homestead) and to see their precise or approximate locations on a map. The presentation emphasized that, when the database was integrated into the LKIS system, three types of links to other onomastic databases were created: to another place name from which the given place name is derived; to a personal name in the Database of Lithuanian Surnames from which the place name is derived; and to historical records of that place name in the Historical Place-Names Database. At present, approximately 25,000 place names have been entered and described (or are being described) in the database.
In the afternoon, the speakers and participants gathered for the third session, “Corpus Linguistics.”
The first presentation, “Morphologically and Syntactically Annotated Corpora of Lithuanian,” was delivered by ERIKA RIMKUTĖ (VDU), AGNĖ BIELINSKIENĖ (VDU), LOÏC BOIZOU (VDU), and ANDRIUS UŽA (VDU). It discussed two annotated Lithuanian-language corpora prepared at the Vytautas Magnus University Centre of Computational Linguistics (corpora available online: http://nl.ijs.si/ME/V4/msd/html/index.html; https://ufal.mff.cuni.cz/tred/).). Annotated corpora are essential resources without which the development of language technologies is impossible. They are generally used to create other natural-language resources and tools in fields such as automatic speech recognition systems, machine translation, and so on. The morphologically annotated corpus MATAS was compiled between 2002 and 2014. It comprises 1.6 million words from texts in various styles. The corpus was prepared by applying statistical models to a one-million-word corpus compiled in 2006. The syntactically annotated corpus ALKSNIS, prepared in 2016 as a gold standard for further research and resources, consists of 2,355 sentences (about 30,000 words) drawn from texts in various styles. The corpus annotation is based on principles of automatic morphological and syntactic annotation, using a syntactic dependency model.
The topic of corpora was continued in the presentation “A Database of Lithuanian Multiword Expressions” by the conference’s large team of presenters: ERIKA RIMKUTĖ (VDU), AGNĖ BIELINSKIENĖ (VDU), LOÏC BOIZOU (VDU), IEVA BUMBULIENĖ (Baltijos pažangių technologijų institutas), and JOLANTA KOVALSKAITĖ
318
Acta Linguistica Lithuanica LXXVI
Page 7Lietuvių kalbos padėtis skaitmeniniame amžiuje
(VDU), Tomas Krilavičius (Baltijos pažangių technologijų institutas), Justina Mandravikaitė (Baltijos pažangių technologijų institutas), and Laura Vilkaitė (Baltijos pažangių technologijų institutas). The presentation at the conference was delivered by Erika Rimkutė (VDU), who spoke mainly about the methodology for studying multiword expressions in contemporary written Lithuanian and the corpus-based Lithuanian collocation dictionary being compiled. The project Automatic Recognition of Multiword Expressions in Lithuanian (PASTOVU) (No. LIP-027/2016) (see http://mwe.lt/),) aims to develop a methodology for studying multiword expressions in contemporary written Lithuanian and to prepare a corpus-based Lithuanian collocation dictionary. A 2014–2016 Delfi.lt corpus, comprising 72 million words, was compiled and used in the project. The collocation dictionary will be prepared on the basis of a database. It will include various information about multiword expressions: grammatical and lexical information, frequency of use, text category, concordance examples, and so on.
Corpus research was also presented by Vilnius University representatives Gintarė Judžentytė (VU) and Vilma Zubaitienė (VU). The presentation “Corpus-Based Research on Academic Phrases: Formal Structure and Semantics” discussed the structure and semantics of academic-language phrases, drawing on the Student Written Work Corpus currently being developed at Vilnius University. This is one of the planned outcomes of the project supported by the State Commission of the Lithuanian Language, Research on Phrases in Student Papers and an Interactive Phrase Inventory. The project aims to identify a list of words important for academic writing, distinguish key collocations and their extensions, as well as recurring word sequences in various sections of academic texts; discuss their links to the rhetorical functions of those sections; and describe the meanings and usage of academic words such as aim, objective, methods, conclusions, results; perform, analyze, determine, rely on, among others.
Four presentations were delivered at the final session, “Automated Analysis of Language Structure and Texts.”
The session opened with Danielius Račys’s (VU) presentation “Machine Translation for Lithuanian,” which covered the recent history and achievements of machine translation, projects carried out by various institutions, and new breakthroughs in the use of neural machine translation. The speaker discussed the development and achievements of machine translation, which are now being effectively applied to Lithuanian as well. In 2005–2007, Vytautas Magnus University carried out the EU Structural Funds-financed project Online Information Translation Tool. The result was a publicly accessible online translation service from English into Lithuanian (available online: http://vertimas.vdu.lt/twsas/).). In 2012–2014, Vilnius University carried out the EU-funded project Development of an English–Lithuanian–English and French–Lithuanian–French Machine Translation System Based on Statistical Methods. The result was a publicly accessible online translation service (available online: https://www.versti.eu/). However,
Apžvalgos / Surveys
319
Page 8AURELIJA TAMULIONIENĖ
Even the best machine translations require a translator’s intervention. So can machines translate really well? The past few years have promised new breakthroughs through the use of neural machine translation. Vilnius University is preparing to begin implementing next-generation neural machine translation for English, Lithuanian, Polish, French, Russian, and German in 2018. It was welcomed that the latest achievements in machine translation have not bypassed Lithuanian either.
In his presentation “Lithuanian Grammar in the Digital Open-Source World,” Virginijus Dudkevičius (VU) spoke about the development of next-generation computational morphology for Lithuanian, morphological analysis and synthesis of words, intelligent search for textual information, and so on. With the advent of computers, classical grammar and lexicon had to “put on a new outfit” and become a fully fledged part of new technologies. To this end, next-generation computational morphology for Lithuanian was developed, with the following goals: to use and create only open-source software; only the data, not their form and interpretation, may have properties specific to Lithuanian; all program code must be universal and suitable for any other language; and to make maximum use of solutions successfully adapted for other languages around the world. The foundation was Dabartinės lietuvių kalbos gramatika (Vilnius, 2006) and corpora (about 1.5 billion words in total). At present, about 99% of the average text appearing online is “recognizable”—that is, on average, the morphological analyzer correctly interprets 99 out of every 100 words. This newly developed method of morphological analysis was successfully applied in Vytautas Magnus University’s Syntactic-Semantic Analysis System (available online: https://semantika.lt/SyntacticAndSemanticAnalysis/Analysis) and in the Register of Legal Acts of the Seimas of the Republic of Lithuania (available online: www.e-tar.lt).
The theme of digital grammar was continued by Daiva Šveikauskienė (LKI) and Vytautas Šveikauskas (LKI). Their presentation, “Digital Grammar of the Lithuanian Language,” discussed the digital grammar of Lithuanian whose development had begun at the Institute of the Lithuanian Language and which can also be used in other areas of computational language processing, including grammatical analysis and direct data translation. They presented a project intended to join the international DIGITAL GRAMMARS system, which currently covers 32 languages. The main aim of creating a digital grammar is to use it in machine translation systems that operate by non-statistical methods. This is highly relevant for Lithuanian. According to 2017 data from Tildė, Google’s translation quality into and from Lithuanian is lower even than for Latvian or Estonian. It is therefore very important for Lithuania to develop an alternative machine translation option. A test example of the digital grammar has now been prepared, covering a few words, but even this shows that translation quality is much better here, especially for sentences that reflect specific features of Lithuanian. As long as English word order is used, statistical methods also provide fairly good translations
320
Acta Linguistica Lithuanica LXXVI
Page 9Lietuvių kalbos padėtis skaitmeniniame amžiuje
but when a sentence has an unconventional word order, the Google translation does not lose the sentence’s meaning. Digital grammar can also be used in other areas of computational language processing, including grammatical analysis and direct data translation. The development of digital grammars is led by Professor Aarne Ranta of the University of Gothenburg, Sweden.
The final conference presentation, “Who Copied from Whom? Automating and Visualizing the Study of the History of Bible Translations,” was delivered by Mindaugas Šinkūnas (LKI). The presentation discussed the automation and visualization of research into the history of Bible translations. The speaker illustrated the results of comparing biblical verses with vivid, detailed diagrams. The abundance of Bible quotations in old Lithuanian writings makes it difficult to study their relationships. A convenient data platform is needed to group biblical verses according to features of interest to researchers. Assessing the similarity of quotations presents the greatest difficulty. A combined analysis of lexicon, syntax, and morphology (and, in some cases, spelling) makes this possible, but this philological research is slow and cannot currently be automated. The comparison was automated using iterative algorithms whose task is to detect identical sequences in two strings of characters and calculate what proportion of the compared strings they cover. Automation of the comparison is hindered by the ambiguous orthography of old writings. An assessment of verses transliterated manually and automatically established that the method of transliteration has no essential effect on the results. Results may be disputed when comparing texts written in different dialects (the leveling of dialectal features has not been tested). The resulting similarity scores are summarized in two-dimensional diagrams. The results of computational analysis comparing biblical verses in Bretkūnas’s Postilė (1591) are consistent with conclusions reached by earlier researchers using other methods.
The conference concluded with discussions. Articles based on the conference presentations will be published in the scholarly journals Bendrinė kalba (available online: www.bendrinekalba.lt) and Lietuvių kalba (www.lietuviukalba.lt).
LITERATŪRA
Zinkevičius Zigmas 1992: Lietuvių kalbos istorija 5. Bendrinės kalbos iškilimas. Vilnius: Mokslas, 144–155.
Įteikta 2017 m. lapkričio 3 d.
AURELIJA TAMULIONIENĖ
Lietuvių kalbos institutas
Petro Vileišio g. 5, LT-10308 Vilnius, Lietuva
[email protected]
Apžvalgos / Surveys
321