Mašininis vertimas lietuvių kalbai
Full text
BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 DANIELIUS ALGIRDAS RALYS University of Vilnius MACHINE TRANSLATION TO LITHUANIAN 1 KEY WORDS: Lithuanian language, history, computational linguistics, machine translation, neural networks, artificial intelligence. INTRODUCTION Today, more and more texts are translated by computers. Such automatic computer translation is usually referred to as machine translation (MT). As a rule, machine translation is carried out in order to understand as quickly as possible in the huge flow of multilingual information, as well as to disseminate its information as widely as possible. The exchange of electronic information between nations and languages is becoming an important part of everyday life, opening up new opportunities for our language, as well as presenting it with new challenges and problems. The term "machine translation" was first used by Warren Weaver in the mid-1940s, when he proposed the use of newly developed computers to translate texts (Weaver 1949: 1). Weaver proposed to look at a text written in another language as a cipher that can be decoded. The cryptographic achievements of British and American specialists in the recently ended World War II aroused optimism. At that time, it was sincerely believed that electronic machines would help overcome the Tower of Babel curse in a few years. Unfortunately, translation has improved slowly, and the quality of machine translation is far from being equal to that of human translation. Seventy years later, after undoubted successes and disappointments, after a lot of work done, the machine translation community is again optimistic. This time, the hopes are placed on artificial neural networks and rapid progress in the development of artificial intelligence. Machine translation has not escaped Lithuanian. Today, easily available online machine translation tools translate from the most popular languages into Lithuanian and vice versa. The quality of the translations of most tools is still quite poor, and the translated texts need editing. Only the translation tool Google Translate, which has begun to use 2 1 The article was prepared based on the presentation “Machine translation for Lithuanian language”, delivered at the 24th International Conference “Digital Language Resources, their Development Directions and Use Opportunities” on September 29, 2017 in Vilnius. 2 Online access https://translate.google.com.
DANIELIUS ALGIRDAS RALYS. Mašininis vertimas lietuvių kalbai |2 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 neural translation technology. Most sentences translated by this tool retain their meaning in the original language. Machine translation is widely used to disseminate information in other languages. Often the mašininio vertimo rezultatai neatsakingai įkeliami į elektroninę erdvę visai neredaguoti. Tokių consequences of these actions are already visible: poor Lithuanian texts with distorted information, which were automatically generated by one or another machine translation program, 3 flashed on the Internet. Gradually, such texts can also penetrate into textbooks. Machine translation has become an everyday, sometimes surprising and sometimes eye-catching reality, which deserves not only the closer gaze of linguists, but also their more significant involvement in the development of machine translation. This article briefly presents the history of machine translation (MT) and some of the ideas that have influenced the development of MT. The results of various translation technologies for Lithuanian are presented. The possibilities of adaptation of MV to the Lithuanian language are analyzed using regulative, statistical and neural methods. For lack of space, attempts to combine several of these methods into one translation system are not described. 1. RESULTS OF MACHINE TRANSLATION A good translation from another language is always a kind of intellectual challenge for a person, and translation has never been an easy task. It is not surprising that there was an attempt to automate the translation in some way. In France, between 1932 and 1935, the engineer Georges Artsrouni invented and constructed a universal mechanical device, which he called the brain mécanique, which, among other functions, had the ability to translate the text collected into several languages (Hutchins 2004a: 12). This electromechanical device rotated a long (up to 40 meters long) memory tape made of flexible cardboard. It was written in a multilingual dictionary, which could include even several tens of thousands of words. The French word entered on the keyboard was mechanically encoded, and according to this code the memory tape was rotated to the right place. After a few seconds, the device showed the translation of the entered word (Daumas 1965: 294–295). The device was shown at the Paris World Exhibition in 1937, received great interest and was awarded the Grand Prix. In fact, it was the first in the world 3 Jonathan Swift, in his Gulliver's Travels, sarcastically describes a machine composing meaningless combinations of words that were constantly diligently written into books. Today, the same crap can be made much faster...
DANIELIUS ALGIRDAS RALYS. Machine translation for Lithuanian language |3 working mechanical BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 dictionary. Artsrouni continued to improve his device, but mechanical problems and the outbreak of World War II interrupted his work. In 1933, Pyotr Smirnov-Troyansky, a Russian engineer, formulated many valuable ideas for machine translation, which were realized only many years later, with the elektromechaninei vertimo mašinai 4 . Nors ši mašina taip ir nebuvo pagaminta, tačiau išradėjas advent of computers. Smirnov-Troyanski spoke Russian and therefore understood the peculiarities of flexional languages. He proposed to divide the translation process into three stages: in the first stage, a person must reduce the text to lemas and annotate them with morphological-syntactic marks (the analysis stage); then, in the second stage, the machine automatically finds the lemas of one or more translation languages and assigns them the same tags (transformation stage); in the third stage, the editor must prepare a fluid translation text according to the translation lemma and their tags (generation stage). Smirnov-Trojanskis also proposed various methods of translating synonyms, homonyms and idioms. Smirnov-Troyansky’s ideas were not appreciated in his native country, he never lived to see the era of computers, but the inventor’s ideas were not forgotten and later became known in the West. 5 The era of practical machine translation began with the advent of computers. In 1947, Weaver wrote to the brilliant linguist and computer engineer Norbert Wiener, asking if the latest advances in cryptography and the newly emerging computers could be applied to language 6 translation. Weaver claimed that any Russian text could be interpreted as an English message encoded in Cyrillic letters. It seemed to him that computers could simply decrypt the message. Unfortunately, Mr Wiener's response was skeptical. In 1949, Weaver sent a memorandum to a wider audience, proposing the use of computers for text translation (Weaver 1949: 7). At that time, the ideas of simple literal machine translation were prevalent in various American and British laboratories. The scientist was convinced that computers could translate much better. In the 7 memorandum, Weaver proposed four closest surroundings of the translated word, thus solving the problems of polysemy. This insight forms the basis of statistical machine translation today. The second proposal was based on the assumption that in the speech vertimo problemų sprendimo ir tyrimo kryptis. Visų pirma, jis pasiūlė statistiškai analizuoti 4 In the patent, the device is described as a "machine for selecting and printing words in translation from one language to another or from one language to another at the same time". 5 Smirnov-Troyansky died in 1950. 6 The first British electronic computer, the Colossus, started operation in early 1944, and the first American electronic digital computer, the ENIAC, was built in late 1945. The first Soviet computer MESM was produced in 1950. 7 The term machine translation, used in this Weaver memorandum, has become quite familiar.
DANIELIUS ALGIRDAS RALYS. Machine translation for the Lithuanian language |4 has logical elements, BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 so using neural network models (McCulloch, Pitts 1943: 115– 133), it should be possible to deduce the translation from the assumptions in the original language. Time has shown the genius of this insight – today the neural machine translation method translates the most qualitatively. translation. Forty Trečią problemų sprendimo kryptį turėtų sudaryti kriptografijos metodų taikymas kalbų years later, this was implemented in statistical machine translation algorithms. To illustrate the fourth proposal, Weaver drew an allegorical picture of people living in high closed towers, symbolizing different languages. People manage to communicate by shouting through the thick walls of the towers. Luckily, the towers have a common foundation and underground, you just have to go down there and then it should be easier to communicate. Weaver argued that the languages of the world have a common deep structure that has yet to be revealed. The latter researcher’s proposal prompted the search and construction of a universal language – machine translation interlingua. 8 After Weaver's memorandum, machine translation quickly gained momentum. In 1954, a machine translation system developed in collaboration between IBM and Georgetown University was publicly demonstrated in New York City. It was translated from Russian into English. Only six grammatical rules were programmed in the system, and about 250 words were used for translation (Hutchins 2004b: 102). The system was focused on organic chemistry, the translations pavyzdžiai iš anksto kruopščiai apgalvoti. Po eksperimento viltasi, kad po kelerių metų mašinos were excellent, and generous funding was allocated to continue the work. The U.S. and the Soviet Union rapidly began developing nuclear weapons in order to gain strategic advantage in the Cold War. The most popular translated languages were Russian and English. 2. CORRECTIVE MACHINE TRANSLATION In the mid-1950s, machine (computer-aided) translation systems appeared, which were called rule-based. They were developed with the view that a language can be described using a set of certain rules (including grammatical ones). In this method, the original language sentence is analyzed to a certain selected level, the sentence structure is determined, which is transformed according to established rules into the equivalent structure of the translation language, according to which the sentence of the translation language is generated. The sentence structures of different languages can differ greatly, have some or other characteristics that must be reflected in the MV 8 It should not be confused with the natural language interlinguas – Esperanto, Ido, Volapük, etc.
DANIELIUS ALGIRDAS RALYS. Machine translation for Lithuanian language in |5 system. Daiva BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 Šveikauskienė (2005: 411–417) studied the problems of representing the syntactic structures of Lithuanian sentences in a manner suitable for machine translation. There are various variations of the regular MV systems, which differ in the depth of analysis (Figure 1). The deeper the analysis of 9 the sentence, the simpler the transformation. Ideally, when analyzing a text to the interlingua depth, no transformation is required – the interlingua image of the text should be the same in all languages. 1 PAV. The Vauquois triangle of regular machine translation The problems of regular machine translation have been solved using both empirical and linguistic methods. In the empirical camp, the most popular method was direct translation. This method of translating between two languages uses only a dictionary and simple programming rules, practically no language analysis or syntactic reorganization (this method corresponds to the lowest level in Figure 1). In the 1950s, a group of researchers led by Erwin Reifler in the United States developed special dictionaries for machine translation, which included local word ordering rules in addition to lexical equivalents. Translations of many phrases and cover terms for solving polysemy problems were also provided. In 1964, the Mark II Russian-English direct translation system, developed for the United States Air Force, was put into operation. It is based on E. Reifler’s dictionary, which includes over 170,000 words (Reifler 1960: 312). Linguistic approaches have been applied more widely at Georgetown University, where a large force of machine translation researchers has been assembled, divided into several groups using different approaches. 9 Such illustrations are often referred to as the Vauquois triangle, a method of illustration proposed by the French pioneer of machine translation Bernard Vauquois (1976: 335).
DANIELIUS ALGIRDAS RALYS. Mašininis vertimas lietuvių kalbai |6 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 groups. The Georgetown Automatic Translation (GAT) system was developed by a group led by Michael Zarechnak. The analysis was carried out on three levels: morphological (including the identification of idioms), syntactic and syntagmic. A modified GAT system was successfully introduced at Euratom (Ispra, Italy) in 1963 and at the U.S. Atomic Energy Commission in 1964. The commercially valuable KANT MV system appeared only in the last decade of the last century (Mitamura et al. 1991: 105–118). The system has been used to translate some technical texts. The language was never created to express a common language. For the decade following the IBM–Georgetown MV presentation, there was a great deal of effort to substantially improve the quality of MV translation, but progress was very small. Founded in 1966 in the USA 10 The ALPAC (Automatic Language Processing Advisory Committee) committee decides that the MV has no prospects in the near future because it is of poor quality and still needs to be edited by a translator, which takes more time than translating without any machines. The ALPAC committee ignored the fact that machine translation helps translate huge amounts of information, while many users are satisfied with poor, barely meaningful translations. After the committee's negative conclusions, funding for MV projects in the United States was practically halted for several decades, and funding decreased in other countries as well. Despite the emerging scepticism about MV, there has been steady progress in this area. In 1968, SYSTRAN (Toma 1977: 569–581) was developed in the United States, and it was constantly improved and later commercialized. The system has been used by the US Department of Defense, NASA, Euratom, the European Commission and many others. Under a commercial agreement in 1975, the European Commission was granted the right to develop the system for its own purposes. The system translated some areas of language into many languages fairly well. Since 2010, SYSTRAN has been using hybrid technologies. Relatively well-functioning MV systems have also appeared in other countries: ARIANE (Grenoble, France), a similar MU system (Kyoto, Japan), as well as SUSY (Saarbrücken, Germany). The European EUROTRA project (1982-1992), costing more than 50,000,000 ECU, ended in failure – hundreds of specialists never created a working MV system. This was the 11 beginning of a serious crisis for the regular MV. For many years there was no progress. 10 In the United States, around $20 million has been spent on the development and improvement of MV systems during this period. 11 1994 m. Europos Bendrijų Komisijos (angl. CEC – Commission of the European Communities) pranešime pateikti EUROTRA finansavimo duomenys – Komisija 1982–1992 m. projektui skyrė 37,5 mln. ECU (CEC 1994: 3.10), the sixteen EUROTRA centres received additional funding of more than 20 million euros from various governments. ECU (ECA 1994: A4.2).
DANIELIUS ALGIRDAS RALYS. Machine translation for Lithuanian language |7 Regular machine BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 translation in Lithuania only appeared in this century. In 2003–2004, the machine translation system Česílko (Hajič et al. 2000: 10–12) was adapted for the Lithuanian language (Homola, Rimkutė 2004: 77–81). During this period, Petras Homola and Erika Rimkutė investigated the possibilities of this system for machine translation between the Czech and Lithuanian languages. In 2005-2007 Vytautas Magnus University carried out the project “Online Information Translation Tool” funded by the European Union Structural Funds. The structure of the projected MV system is described by Vido Daudaravičius (2006: 7–18). A public online translation tool from English to Lithuanian has been created. The translation engine was provided by the Russian company PROMT. Part of the language 12 components of the system were prepared in Lithuania. The system translated only some types of sentences perfectly, good translations were a minority. In 2007, a linguistic assessment of this MV system was carried out (Rimkutė, Kovalevskaitė 2008: 257–264). Audronė Daubarienė and Greta Ziezytė (2013: 55–61) evaluated the compliance of the VDU MV system translation with textuality standards (De Beaugrande, Dressler 1981: 3). In their work, the authors found that the studied machine translation texts of different genres did not meet the standards of coherence and some other textuality due to numerous semantic errors, so such translations cannot be acceptable and informative for the reader. As the decades went by, and as things basically stagnated, it became clear that language is an extremely complex phenomenon, which is very difficult to describe by rules. In order to solve the problems of polysemantic and anaphoric translation, it was proposed to use extralinguistic information about the structure of the world and its logic (Bar-Hillel 1958: 197–207). However, there was simply no such systematic information, and the formal description of the meaning of the text seemed to be a more difficult task than the translation itself. It was obvious that the problems of machine translation required revolutionary solutions. 2. STATISTICAL MACHINE TRANSLATION In 1988, a group of IBM researchers (Brown et al. 1988: 71–76) proposed a translation model for naudoti lygiagrečiuose tekstynuose glūdinčią informaciją, ieškant tikimiausių vertimo variantų. machine translation. The proposed translation model is based on the assumption that a noise channel changes a sentence f in the original language into a sentence e in the translation language. The noise in the channel statistically distorts the original sentence f, so in the channel output we can get various sentences e of different lengths, having one or another meaning in the translation language or having no meaning at all. Because that channel stengiasi išversti teisingai, tai pro triukšmą vis dažniau prasiskverbia teisingi vertimo variantai. 12 Online access http://vertimas.vdu.lt/twsas/.
DANIELIUS ALGIRDAS RALYS. In this model it is assumed that when a sentence f is sent to a channel, BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 the most appropriate translation variant ê of all e is the one with the highest conditional probability p(e|f) of appearing in the channel output, i.e. ê = argument p(e|f). The problem here is that in order to find the best translation variant of ê, it is necessary to review not only all possible sentences in the translation language, but also all kinds of word combinations in it. It is very difficult to calculate such a number of probabilities. Applying the Bayesian theorem, an easier-to-compute equivalent expression is obtained (Koehn et al. 2003: 127): argmaxe p(e|f) = argmaxe p(f|e) p(e) This equation is the basic equation of statistical machine translation (Brown et al. 1993). The conditional probability p(f|e) shows the probability that a sentence f was sent through the channel if a sentence e appeared in the output. The second term of the equation p(e) shows the probability of the sentence e in the translation language. This member describes a language model that is calculated from monolingual texts in the translation language. 13 2 PAV. The Vauquois triangle of statistical machine translation The translation method proposed by IBM scientists immediately competed with regular translation in terms of translation quality. Developers of statistical ML systems were fascinated by the possibility of translating into many languages without having an understanding of either translation or grammar, it was enough to train the systems from existing textbooks. During the classical statistical translation, no 13 The language model reduces the amount of calculations, because it helps to immediately reject sentences not found in the textbooks, helps to choose the smoothest translation, and also indirectly shows why fiction is better translated by a native translator.
DANIELIUS ALGIRDAS RALYS. Machine translation for Lithuanian language |9 analysis of BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 sentences in the original language or synthesis of sentences in the translation language (Figure 2), translated only according to the probabilities determined during the system training. The five models proposed by Peter Brown and his colleagues were not flexible enough, allowing the original word to be translated by one or several words, or not translated at all. In 2003 (P. Koehn et al. 2003: 127–133) a more flexible machine translation model was proposed that could also translate phrases in the original language. During this period, the methods of automatic assessment of machine translation quality were rapidly improved, as it was very important to quickly find out which version of the compatible machine translation system translated best. It was also intended that such an assessment would allow an objective comparison of the quality of translations from different systems. One of the most popular metrics has become the BLEU (bilingual evaluation understudy acronym) metric (Papineni et al. 2002: 311–318), which automatically compares how far the machine translation text is from the human translation. This metric does not take into account the legibility or grammatical correctness of the text, but only counts exactly matching words or phrases. When calculating the metric, matching phrases have a relatively higher statistical weight than single matching words. uždavinio, MV kokybei vertinti dažnai naudojamos ir kitokios, vienokią ar kitokią gerą ypatybę Depending on the metrics used: F-measure, METEOR (Banerjee, Lavie 2005: 62–72), NIST (Doddington 2002: 138–145), TER (Snover et al. 2006: 223–231) and many others. Despite the statistical popularity of the MV, the limitations of this method have become increasingly apparent when translating not all possible į morfologiškai turtingas (fleksines) kalbas. Į fleksinių kalbų lygiagrečius tekstynus dažniausiai forms of phrases or even individual rarer words. If any available formats have not been included in the training of the MV system, then such formats will not be displayed. This shortcoming of the classical MV is caused by data insufficiency, which results in a thinned translation probability matrix when the system is clouded. To overcome this shortcoming, Philipp Cohen and Hieu Hoang proposed factorized statistical machine translation (Koehn, Hoang 2007a: 868–876). It was proposed to add additional linguistic information to the translation system by annoting parallel texts with linguistic analysis tags, which are then used to train the system. In 2007, Cohen and his colleagues released a complete statistical MV package for the general public as open source (Koehn et al. 2007b: 177–180). The application of factorized statistical MV slightly improved the translation quality of flexual languages, usually by 1–2 BLEU percent (Bojar 2007: 235–236; Skadiņš et al. 2010: 128–129). On 25 September 2008, the Google Translate statistical MV system began translating into Lithuanian. In 2012-2014 Vilnius University implemented European Union Structural Funds
DANIELIUS ALGIRDAS RALYS. Machine translation for Lithuanian language |16 D e B ea u g ra BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 n d e R. A., D r e s s le r W. U. 1981: Introduction to Text Linguistics, London: Longman, 270. D o d di n g t on G. 2002: Automatic Evaluation of Machine Translation Quality Using Ngram Co-occurrence Statistics. – Proceedings of the Second International Conference on Human Language Technology Research, San Francisco, CA, USA., Morgan Kaufmann Publishers Inc., 138–145. F o rc a d a M. L, G i n e s t í - Ro s ell M., N o rd f a l k J., O'R e gan J., O r t i z - R o ja s S., P é re z - O r ti z J. A., S á n c h h e z -M a r t í n e z F., R a m í r e z -S á n c h h e z G., T ye r s F. M. 2011: Apertium: a free/open-source platform for rule-based machine translation. – Machine Translation 25 (2), Free/Open-Source Machine Translation (June 2011), 127– 144. H a j ič J., H ri c J., K u b o ň V. 2000: Machine Translation of Very Close Languages. – Proceedings of the 6th Applied Natural Language Processing Conference, Association for Computational Linguistics, 7–12. H o m ol a P., R i m k u tė E. 2004: Artimų kalbų mašininis vertimas. – Kalbų studijos 6, 77–81. H u tc h i ns J. 2004a: Two precursors of machine translation: Artsrouni and Trojanskij. – International Journal of Translation 16(1), 11–31. H u tc h i ns J. 2004b: The Georgetown-IBM experiment demonstrated in January 1954. – Proceedings of the 6th Conference of the Association for Machine Translation in the Americas, AMTA, Washington, DC, 102–114. J o hn s o n M., S c h u st e r M., L e V. Q., K r i k un M., W u Y., C h e n Z., T h or a t N., V i é ga s F. B., W a t te n b e rg M., C o r ra d o G., H u g he s M., D e a n J. 2017: Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation. – Transactions of the Association of Computational Linguistics 5 (1), 339–351. K a lch b re n n e r N., B l un s o m P. 2013: Recurrent Continuous Translation Models. – Proc. of EMNLP, October, Association for Computational Linguistics, Seattle, Washington, USA, 1700–1709. K oe hn P., H o an g H. 2007a: Factored translation models. – Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, ACL, 868–876. K oe hn P., H o a ng H., B i r c h h A., C a ll i so n - B u rc h C., F e de r i co M., B e r t o l di N., Co w a n B., S he n W., M o ra n C., Ze n s R., D y e r C., B o j a r O., C o ns ta n tin A., H e rb s t E. 2007b: Moses: Open source toolkit for statistical machine
DANIELIUS ALGIRDAS RALYS. Mašininis vertimas lietuvių kalbai |17 BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 translation. – 45th Annual Meeting of the Association for Computational Linguistics (ACL), Companion Volume, 177–180. K oe hn P., O c h F. J., M a r c u D. 2003: Statistical phrase-based translation. – Proceedings of HLT-NAACL, 127–133. K os K., B o ja r O. 2009: Evaluation of Machine Translation Metrics for Czech as the Target Language. – Prague Bull. Math. Linguistics 92, 135–148. M c C ul l o c h W., P it t s W. 1943: A logical calculus of the ideas immanent in nervous activity. – Bulletin of Mathematical Biophysics 5, 115–133. M i ko l o v T., S u t sk e ve r I., C he n K., C o r ra d o G. S., D e a n J. 2013: Distributed Representations of Words and Phrases and their Compositionality – Proceedings of the 26th International Conference on Neural Information Processing Systems, Lake Tahoe, Nevada, Vol. 2, 3111–3119. M i t am u r a T., N y be r g E., C a r b o ne l l J. 1991: An efficient interlingua translation system for multi-lingual document production. – Proceedings of Machine Translation Summit III, Washington, DC. Reprinted: – Progress in Machine Translation, ed. S. Nirenburg, IOS Press, (1993), 105–118. P ap i n e ni K., R o u ko s S. W a rd T., Z h u W. J. 2002: BLEU: a method for automatic evaluation of machine translation. – ACL-2002: 40th Annual meeting of the Association for Computational Linguistics, 311–318. R e if l e r E. 1960: The solution of MT linguistic problems through lexicography. – Proceedings of the National Symposium on Machine Translation, (February 2–5, 1960), ed. H.P. Edmunson, London: Prentice-Hall, 1961, pp. 312–316. R i mk u t ė E., K o v al e v s ka i t ė J. 2008: Linguistic Evaluation of the First EnglishLithuanian Machine Translation System. – Proceedings of the Third Baltic Conference on Human Language Technologies (2007), Kaunas, 257–264. S k a di ņ š R., G o b a K., Š i c s V. 2010: Improving SMT for Baltic Languages with Factored Models. – Proceedings of the Fourth International Conference Baltic HLT, Frontiers in Artificial Intelligence and Applications, Vol. 219, IOS Press, 125–132. S n ove r M., D o r r B., Sc h w a r t z R., M i cci u l l a L., M a k h ou l J. 2006: A Study of Translation Edit Rate with Targeted Human Annotation. – Proceedings of the 7th Conference of the Association for Machine Translation in the Americas, Morristown, NJ, USA, 223–231.
DANIELIUS ALGIRDAS RALYS. Machine translation for Lithuanian language |18 networks. BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 S u ts k e ve r I., V i n ya l s O., V. L e Q. V. 2014: Sequence to sequence learning with neural – Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, Canada, Vol. 2, MIT Press Cambridge, MA, USA, 3104– 3112. Š v ei k a u sk i e nė D. 2005: Graph Representation of the Syntactic Structure of the Lithuanian Sentence. – Informatica 16 (3), 407–418. T o ma P. 1977: Systran as a multilingual machine translation system. – Proceedings of the Third European Congress on Information Systems and Networks, Overcoming the language barrier, München, 569–581. V a s iļ je vs A., Sk a d i ņ š R., T i e de m ann J. 2012: LetsMT! : a cloud-based platform for do-it-yourself machine translation – Proceedings of the ACL, System Demonstrations, Jeju Island, Korea, 43–48. V a uqu o i s B. 1976: Automatic translation – A survey of different approaches. – COLING-76, Ottawa. Reprinted: – Readings in machine translation, eds. S. Nirenburg, H. L. Somers, Y. Wilks, MIT Press, (2003), 333-337. W e ave r W. 1949: Translation. – Reprinted: – Machine Translation of Languages, MIT Press, Cambridge, MA. (1955), 15–23. Įteikta 2017 11 13 Adopted on 20 December 2017 MACHINE TRANSLATION FOR LITHUANIAN LANGUAGE S u mm a ry The paper presents a historical overview as well as current state of the art of the machine translation. Translation from another language is always a certain intellectual challenge. In 1949 Warren Weaver suggested using computers to translate texts. The term "machine translation" (MT) appears. Machine translation has been rapidly developing during the first decades in order to gain a strategic advantage in the Cold War. Most popular translated languages were Russian and English. Word-to-word translation and large bilingual dictionaries covering more than 170,000 words prevailed.
DANIELIUS ALGIRDAS RALYS. In the 1950s and 1960s, machine translation systems appeared which could BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 be called rule-based. They are based on the assumption that a language can be described using a set of rules (including grammatical). It was a rather optimistic period - it was expected to create a perfect machine translation in a few years. A computer hardly "understands" grammar. Highly rules. Nobody has done this properly yet. The most advanced systems were the SYSTRAN MT system, launched in the inflected languages require tens of thousands thoroughly hand-tuned and mutually consistent European Commission and the Russian PROMT translation system. Subsequently, the rule-based MT progress slowed down. The EUROTRA project (1982-1992), specialists failed to create a functioning MT system. This marked a serious rule-based MT crisis. Development simply stalled for many forthcoming by some estimates costing more than 50,000,000 ECU, fails – even hundreds of recruited years. The question was raised: if we cannot write so many rules, can it be translated without grammar at all? In 1990 a new breakthrough emerges - the research team at the IBM Thomas J. Watson Research Center formulates the basics of statistical machine translation. The translation process has been considered as a transmission of a certain message over a noisy channel. Decoding then has been performed on the basis of the Bayesian theorem. Translation is based on text corpora, especially on large parallel bilingual text corpora. There was a rapid improvement of statistical MT. The EuroMatrix project, supported by the European Commission, has created a universal open source machine translation software package MOSES, based on industry-level MT systems. Good results have been obtained - it turns out that you can translate without any dictionary or grammar! This method has greatly facilitated the translation of highly inflected languages too. The achievements of machine translation today are effectively applied to the Lithuanian language as well. During 2005-2007 Vytautas Magnus University has carried out an EU-funded project "Internet Information Translator". The result was a public online translation service from English to Lithuanian. (http://vertimas.vdu.lt/twsas/). The rule-based translation engine was provided by the Russian company PROMT, while other linguistic resources were prepared in Lithuania. The overall quality of text translation in BLEU metrics (in percent) is about 10. In practice, this means that only every third sentence can be adequately understood. This translation tool still has considerable potential for improving, for example by expanding phrase dictionary. Since September 25, 2008 Google Translate also supports Lithuanian. According to the results of the tests (2014), the BLEU translation quality was estimated to be around 17.
DANIELIUS ALGIRDAS RALYS. 2012-2014 Vilnius University implemented the EU-funded project BENDRINĖ KALBA 90 (2017) www.bendrinekalba.lt ISSN 2351-7204 "Creation of EnglishLithuanian-English and French-Lithuanian-French machine translation system based on statistical methods". The result was a public online translation service (https://www.versti.eu/). According to the tests carried out in 2014, the BLEU estimates of the translation quality exceeded more than twice the rule-based translation results and were practically equivalent to the Google translation system. When translating the documents of certain domain (such as law) the achieved BLEU score is roughly twice as high as translating general texts and far exceeds Google's results (19-09-2014). However, even the best machine translations often require human intervention and final editing in order to get the perfect translation. So, can machines ultimately translate fine? The last few years promise new breakthroughs using a neural machine translation. Neural networks themselves construct transformation rules. It is likely that in the near future the neural MT systems will translate better than an average translator. In 2018 Vilnius University is preparing to launch the EU-funded project of a new generation neural machine translation of English, Lithuanian, Polish, French, Russian and German. well. Thus, the latest achievements in machine translation apply to the Lithuanian language as KEYWORDS: Lithuanian, machine translation, computational linguistics, history, neural networks, artificial intelligence. DANIELIUS ALGIRDAS RALYS Institute of Applied Sciences, Vilnius University M. K. Čiurlionio g. 29, 03100 Vilnius danielius.raly[email protected]m
