scieee AI-readable full text Open interactive document viewer

Comparative analysis of computer-assisted translation tools

Tarilonte Pérez, Sergio

Abstract

Grado en Estudios Ingleses

Full text

FACULTAD de FILOSOFÍA Y LETRAS DEPARTAMENTO de FILOLOGÍA INGLESA Grado en Estudios Ingleses TRABAJO DE FIN DE GRADO COMPARATIVE ANALYSIS OF COMPUTER-ASSISTED TRANSLATION TOOLS Sergio Tarilonte Pérez Tutora: Belén López Arroyo 2018/2019 2 ABSTRACT Nowadays, information technologies are growing rapidly to the point of supplying the real world in many ways. In the world of translation this is also happening with the constant irruption of new programs and web resources that allow an increasingly perfect automatic translation of texts. However, these tools are still far from perfect and, until they are, the best resource that can be used by a translator is computer assisted translation (CAT) tools. This paper tries to discern what each of the most important CAT tools offers, among which any translator can harbor doubts which is more convenient, and to establish a comparison between them. The results will give the reader a vision of both, which tool is more convenient to use, and of what it is that it provides the user when developing a translation. Key words: Translation Studies, Computer-Assisted Translation, Translation memories, Program Comparison, Machine Translation, Parallel Corpora. En la actualidad, las tecnologías de la información están creciendo rápidamente hasta el punto de llegar a suplir al mundo real en muchos aspectos. En el mundo de la traducción esto también está sucediendo con la irrupción constante de nuevos programas y recursos web que permiten una traducción de textos automática cada vez más perfecta. Sin embargo, estas herramientas aún distan de ser perfectas y, hasta que lo sean, el mejor recurso que puede emplear un traductor son las herramientas de traducción asistida por ordenador (TAO). Este trabajo trata de discernir qué ofrece cada una de las más importantes herramientas TAO, entre las que cualquier traductor puede albergar dudas de cual es más conveniente, y establecer una comparativa entre estas. Los resultados servirán para dar al lector una visión tanto de, qué herramienta es más conveniente usar, como de qué es lo que le aporta al usuario a la hora de elaborar una traducción. Palabras clave: Estudios de Traducción, Traducción Asistida por Ordenador, Memorias de Traducción, Comparación de programas, Traducción Automática, Corpus Paralelos. 3 INDEX 1. Introduction ......................................................................................................................... 4 2. Fundamentals of Machine Translation ............................................................................ 6 2.1. What is a corpus ................................................................................................. 6 2.2. What is Machine Translation? ........................................................................... 9 2.3 Classification of translation systems ................................................................ 10 2.3. What is CAT tool? ........................................................................................... 11 2.4. Representativity of corpora .............................................................................. 12 3. Methodology and Corpus ................................................................................................ 14 3.1. Representativity of corpus ............................................................................... 14 3.2. Methodology .................................................................................................... 17 4. Analysis and Evaluation .................................................................................................. 19 4.1. SDL Trados Studio .......................................................................................... 19 4.2. Wordfast ........................................................................................................... 26 4.3. MemoQ ............................................................................................................ 31 4.4. Déjà vu ............................................................................................................. 36 4.5. OmegaT ........................................................................................................... 42 4.6. Evaluation of CAT Tools ................................................................................. 47 5. Conclusion ......................................................................................................................... 53 6. Bibliography ...................................................................................................................... 54 6.1. Works cited ...................................................................................................... 54 6.2. Resources ......................................................................................................... 56 4 1. INTRODUCTION Thirty years ago, the first commercial translation memory tools appeared in everyday life (Lagoudaki, 2006) (See Section 2 below). Nowadays, these tools are still improving constantly with new elements and techniques, like the capability of using online translators, that make the task of the translators, both professional and nonprofessional, easier. This constant evolution makes this technology an interesting field of research. This paper tries to establish an analysis of the features of some of the main Computer Assisted Translation tools, that use translation memories, showing how those tools work and what they can offer to the target users, aiming to create a critical point of view in the reader of the advantages and disadvantages of these programs beyond its price or popularity. Since there are too many Computer Assisted Translation tools (CAT tools), this paper will focus on some of the most used ones to analyze them properly, since the size of this paper is not large enough to analyze all programs in depth. To check which ones are better in order to elaborate the analysis, this paper uses the following chart (Figure 1) elaborated in 2013 by Jared Tabor in a survey made to over three thousand translators about their CAT tools preferences. Figure 1: Favorite CAT tools 5 This pie chart stands out among the others made by Tabor (2013) because of an important aspect; the survey respondents have used more than one CAT software, so their opinions are not biased by some facts like the price or the habit. The price of these tools is usually expensive, so the users don’t want to expend any extra money on a new one, or once they get used to a software, they won't likely switch to a different one unless they find a very good reason. As it can be seen in Figure 1 the five most preferred tools are Trados, Wordfast, MemoQ, Déjà Vu and OmegaT which will be the ones selected for this paper as they offer at least a trial version to work with. In order to understand this work as clearly as possible, this paper is divided into four main sections. The first section of this paper, Fundamentals of Machine Translation, tries to explain the basic concepts necessary to understand this topic. This section is one of the main ones of this work because it briefly explains the complex or difficult to understand concepts that will be discussed in the paper. Shorter than the previous one, the second section, Methodology, deals with the description of the different processes needed to begin the comparative analysis in the following section along with a brief explanation of how the subsequent analysis will be carried out. After that explanation, the third section, and the most important segment of this paper, Analysis and Evaluation, tries to perform an individual study of the five Computer Aided Translation tools preferred by the translators along with a later joint comparison of those tools and a final interpretation of all the collected data. Finally, the fourth and last section of this paper, Conclusion, summarizes all the work carried out and extracts the final results of the research carried out. 6 2. FUNDAMENTALS OF MACHINE TRANSLATION To understand this paper is important to know some basic notions of translation memories and Computer Assisted Translation tools. 2.1. What is a corpus In order to begin with this brief explanation of the basic concepts necessary to understand this paper, there is a key concept that, as will be seen below, will be repeated throughout the document. This is the concept of corpus (corpora in plural). According to many dictionaries, as Oxford, Cambridge or Collins, a corpus, in the field we want to address, is “a collection of material, spoken or written, in machinereadable form, assembled for the purpose of finding out how language is used”. The different kinds of corpora had been classified in many ways by various authors over time. In figure 2 it can be seen an attempt of classification elaborated by Ornia (1996). In her paper, she studies the lack of agreement in the classification of corpora, trying to elaborate her own classification based on the proposals of different authors, as shown below. Her classification divides different corpora in branches resulting in two different branch tree diagrams. The first diagram comprehends mostly qualities common to all corpora such as if their source is in oral or writing format, if they are or not representative of a language (reference corpus), if they are fixed to a period of time or not (synchronic or diachronic corpora), if the corpora could be actualized with new information (only in the case of diachronic corpora, as they are not fixed to a period of time), if their source are full texts, fragments or a combination of both, if those texts are from published or unpublished sources and if their purpose is general or specialized to a topic in particular. The second diagram continues from the division among general and specialized corpora dividing the corpora regarding if they have one or more languages (monolingual and multilingual corpora). At this point, she suggests a division between parallel and comparable corpora, but, as she exposes, the parallel corpora must be the original text along with their respective translations whereas the comparable corpora could be comprised of original and translated texts in one or more languages that do not fulfill the 7 requirement of the parallel corpora, as stated in the final branch of the second classification of figure 2. Once this division is established, Ornia continues explaining that parallel corpora can be classified between unidirectional, if all translations go from one A language to another (b, c, d) and multidirectional languages, if the translations go indistinctly from one language to another, and she also concludes her classification explaining that parallel corpora can be aligned or unaligned, but in turn leaves room for future improvements in the classification such as the mother languages, the translation methods, the translator level, if they are open…. In order to complete and give a little more depth to this concept and its classification it is advisable to finish this section showing some of the most remarkable examples of corpora, among which we could highlight the ones shown by Berglund Prytz (2012). The first remarkable type of corpora, balanced corpora, “try to represent a particular type of language over a specific span of time. In doing so they seek to be balanced and representative within a particular sampling frame” (McEnery and Hardie, 2012). The main exponent of this type of corpus is the British National Corpus (BNC) that includes thousands of sentences and words in British English extracted and categorized according to its source. The monitor corpora, tries on the other hand, as its name suggests, to monitor the changes in a specific language during different periods of time. Some of the maximum exponents of this type of corpus are the Corpus of Contemporary American English (COCA), which includes fragments of texts that go from 1990 to the present days, or the Bank of English (BoE) that includes different categories of modern texts and recordings mainly extracted from newspapers and media. Another kind of corpora that Berglund Prytz suggests is the diachronic corpora collects comparable texts from various periods of a language. A good example of this type of corpus is the Helsinki Corpus of English Texts that includes a huge amount of Old, Middle and Early Modern English texts. As this example does not go beyond the Early Modern English it would not be correct to label it as a monitor corpus and would be more in the category of non-monitored diachronic corpus exposed by Ornia (1996). 8 Figure 2: Corpora classification (Ornia,1996) 9 Other good example of corpora can be found on the International Corpus of English (ICE) which includes texts of various varieties of English throughout the world such as Irish English, Indian English or African English. As it include texts in different languages, or varieties of a language, of the same genres in the same domains and it uses the same sampling method for all the texts it fits perfectly in the multilingual comparable category exposed by Ornia (1996)and the definition of a comparable corpus exposed by McEnery and Hardie (2014). On the other hand, as an example of parallel corpora could be highlighted the open source parallel corpus (OPUS) or the corpus prepared for the web-based translator Linguee. These kinds of corpora are composed only by original texts and their respective translations. The corpus developed for the study made in this paper belongs mainly to this category. To conclude another type of remarkable corpus that have been used in one of the articles mentioned in this paper, is the TURICOR used by Corpas Pastor and Seghiri (2010) (section 2.4) that fits mainly as a specialized corpus as it focus on a specific kind of text (tourism in this case). As the corpus used in this paper is focused also on a specific kind of text (instruction manuals), apart from multilingual parallel corpus, it could also be defined as a specific corpus. 2.2. What is Machine Translation? Once we have briefly explained the concept of corpus and some of the many examples and different types that can be found, it is important to focus on what is machine translation so that later we can focus on how it is classified in order to see how it has evolved and introduced the corpora in the most recent translation systems. According to SDL, developer company of the first worldwide CAT tool, Trados, Machine translation, also known as automated or instant translation, is “The translation of text by a computer, with no human involvement.”. But, in the words of Chiew Kin Quah (2006), Machine Translation is even more than this. She exposes that machine translation is also an important technology in the scientific, commercial and sociopolitical spheres that has become a bridge between languages with the emergence of the Internet. 16 Figure 8: Representativeness of the Spanish component (1-gram) Figure 9: Representativeness of the English component (2-gram) 17 Figure 10: Representativeness of the Spanish component (2-gram) 3.2. Methodology Once the base translation memory is made and its representativeness checked, the final analysis of CAT tools can be done. This analysis is carried out by reviewing individually four different aspects of the various CAT tools. The methodology of the following sections consists of three main steps: individual analysis of each tool, comparative analysis and a final general conclusion. One of the main aspects that a translator takes into account when deciding on a tool or another is the experience that this application has, the time it has been established in the market and the trust it conveys. It is therefore inevitable that the first aspect to be treated in the analysis of each tool is a brief historical introduction of each of the tools. When the historical bases of the tool are established, the second aspect to study is the interface. At this point, we must take into account not only the external aspect of the application but the functionalities and unique features that the tool brings. For this, the analysis of this aspect will be carried out through the use of the translation memory elaborated previously in order to translate a new document of the same scope, an instruction manual of a Qled Samsung television of 2019, and see how that tool of work plays its role (see section 3.1). 18 Subsequently, once this in-depth analysis has been carried out, it is necessary to dedicate a third section, ease of use, in which it will be commented how the program interacts with the user and the facilities that this second person has to be able to carry out the task of translation. Finally, the analysis of each tool has a final conclusion section in which a summary of the analyzed aspects will be made together with a small opinion of aspects to improve and other aspects that may have been leaved out. Once the five tools of computer-assisted translation have been reviewed, this analysis has a final section in which, supported by a table in which the general aspects of each CAT tool are easily seen, these applications will be analyzed comparatively. 19 4. ANALYSIS AND EVALUATION As stated in the previous sections this one will try to explore the main CAT tools (Trados, Wordfast, MemoQ, Déjà Vu and OmegaT) their functionalities and strong points as well as other aspects that could make unique the program in comparison with their rivals. 4.1. SDL Trados Studio Developed by the German company Trados GmbH (TRAnslation and DOcumentation Software limited liability company) and distributed by the British company SDL plc (Software and Documentation Localization public limited company) SDL Trados Studio is the main Computer Assisted Translation tool in the world. History Trados started as a small company created by Jochen Hummel and Iko Knyphausen in 1984 in Germany who worked for IBM as an LSP (Language Service Provider). But it was not till 1990 when Trados launched their first own application, a terminology database called MultiTerm, marking the beginning of the current translation industry giant we know today. It is important to remark that the tool known nowadays include many of the advancements made in the nineties, as the program T Align made by the computational linguist Matthias Heyn, known as WinAlign after 1997 with the acquisition of part of the company by Windows. These advances are still important today and continue to be revised and updated every year, even more after the purchase of the company by the British SDL in 2005, in order to maintain the trust obtained by almost all professional translators. Interface and functionalities The first thing that comes to mind when opening the program for the first time is its incredible resemblance to the design of the current Microsoft Office suite with a clean and well divided interface in differentiated sections. To carry out the analysis of the document to translate, in the following picture (Figure 11) it can be seen how the new project option contains a long series of menus that 20 reveal options and highlight elements of the program, such as the ability to translate into several languages at a time, or the capacity to import translation memories from other CAT tools and to work with terminology databases, whether they be external or handmade, in order to add a glossary of possible translations of every topic related word. Figure 11: SDL Trados Studio new project screen After the nine steps that the program requires when creating a new project, including the establishment of the above-mentioned translation memory, the progress bar shows clearly on figure 12 the first benefits of working with translation memories. Without having an exceptionally large or complex translation memory or using any terminology database, it can be seen how the computer-assisted translation tool has already translated automatically one fifth of the elements of the television manual to be translated. Focusing more on the statistical data we can see how the translator's task has been reduced in 8,275 words of the 39,726 words included in the document, which means 20.83 percent of the words in the document. a slightly higher percentage can be seen in the number of translated characters, a 20.99 percent of the total, with 41,233 characters translated from the 196,377 characters included in the document. However, the most 21 striking data can be seen in the number of translated segments, of which Trados has translated using the small translation memory that we have offered 1,839 of the 4,947 segments contained in the instruction manual, that means 37.17 percent of the text segments, or in other words, more than a third of the segments contained in the document. All these data taken from the experiential analysis will be evaluated in greater depth at the end of this section of the paper, after having shown the five programs to be analyzed. Figure 12: SDL Trados Studio project overview screen In addition to this automatic translation of identical segments, the advantages of this assisted translation tool continue once the translator begins the translation of the remaining segments. After the translator is ready to translate a specific segment, if the program finds the translation memory a possible translation, or a segment that fits more than a previously marked percentage in the properties of the project (over seventy percent in the case of Figure 13), the system shows the user that possible segment or translation so that the only thing that needs to be done is to modify the translation until it fits perfectly with the translator's version. 22 Figure 13: SDL Trados Studio segment translation screen On the other hand, if the translator decides that the translation of a segment must be exactly the same as the original version, in the case, as in the case of the segment between hyperlink labels of figure 13, with a single right click the user can access a menu with multiple options, as shown in figure 14, such useful options as copying the source segment in the destination segment, looking for possible matches to that segment or adding comments if necessary. Figure 14: SDL Trados Studio segment translation drop-down menu 23 When talking about the interface and features offered by Trados Studio there are other small aspects that can make the difference when the translator chooses a CAT tool or another. Aspects such as good automatic segmentation of the program, which also allows modifying the segment by hand, the possibility of entering terminological databases as a dictionary always accessible to the translator and the ability to easily and comfortably edit the translation memory, such as can be seen in Figure 15, can become very relevant to the potential buyer of the product. Figure 15: SDL Trados Studio translation memory screen Ease of use As mentioned above, the interface is very similar to the interface of the current Microsoft Office suite, which makes it easily accessible even for the least experienced user. In addition, the program includes a test document with a tutorial, shown in figure 16, for those users who do not know how to use the tool. Learning throughout the use of the program is quite intuitive and with just a few sessions any translator who has just started using this program could easily navigate by avoiding certain advanced elements such as some of the programmable tasks included in the CAT tool. In addition to this tutorial included in the program and the smooth learning curve of the program, Trados offers various training courses to which any user can sign up if he considers it necessary to strengthen his knowledge about the program. 24 However, although it is clear that Trados tries to combine all the tools that a translator may need, this can also be counterproductive because the excess of available options can overwhelm the user. Figure 16: SDL Trados Studio demo file Conclusions To finish this small analysis of this computer-assisted translation tool, it is worth highlighting some things that could be improved in the face of future revisions of the program. The first aspect to improve, and the one that has given most problems when preparing the analysis, is that when entering the source and destination languages includes an extensive list of languages with multiple varieties, such as international Spanish, modern Spanish, Spanish from the United States ... and if the user does not select the exact language of the translation memory as a translation language, it gives problems to the point of not reacting. This could easily be solved by allowing the installer language to those with which the translator knows that it will work and allow that if the user wants to add a new one afterwards; this can be added from some internal option of the program. Another facet that draws the user's attention, this one much more understandable, is that the program only works with one translation memory at a time in order to avoid conflicts between them. A possible solution to this problem could be the possibility of 25 establishing a priority order within the memories. To give an example, if the translator works in a medical-scientific text and has a translation memory of medical texts and another one of scientific texts the program should go first to the medical one and then if it does not find any good match look in the memory of scientific texts. On the other hand, it should be noted that Trados includes a store of installable applications as add-ons designed both by the company and by various users who have seen the needs of this tool, as shown in figure 17. However, the installation of these applications, including that of free add-ons, is restricted to users who have paid for the version of the program, which ranges between 99 and 2500 euros, which is a great disadvantage for that translator who wants to use the trial version in order to decide if it convinces him or not. We can conclude that in spite of these slight drawbacks, SDL Trados Studio is a fantastic tool for any translator, regardless of their level, which makes it clear that today it has become a benchmark in computer-assisted translation. Figure 17: SDL Trados Studio appstore 32 Figure 22: MemoQ new project screen But, unlike the rest of CAT tools, the new project created is completely empty. Both the document to be translated and the terminological bases that the translator might want to use, such as translation memories must be created and imported manually, as shown in Figure 23. This means that the tool does not automatically translate the text without the user request, since the program may not be able to discern when they have finished putting memories and glossaries to the translation. 33 Figure 23: MemoQ translation memories screen Once all the previous steps have been completed and the pretranslation has been carried out, the translator can start the translation, the verification or the editing, depending on how the segment is translated. As in the case of other CAT tools, MemoQ provides the user, as can be seen in the right part of figure 24, of the various options that the program has found in the translation memory along with data such as where the program is located the precision failures that do not give this one hundred percent matching. Figure 24: MemoQ segment translation screen On the other hand, this tool includes many very eye-catching exclusive features. 34 A good example of these unique features of MemoQ is the management of errors, where the program marks the user the odd things it finds, such as extra spaces, missing labels, spelling errors… as shown in figure 25. These errors are marked near the matching percentage of the translation with a symbol of a lightning bolt, if these are minor errors, or that of an exclamation, if the program considers them serious errors. Figure 25: MemoQ error windows Another useful feature of this program is the ability to separate and join segments during the automatic pretranslation process in order to find more and better matches in the procedure. This is very useful since, although the most common thing is that the program segments the texts from a dot to another or from one line to the next, sometimes these segments are cut in a different way than in the translation memory, so the program cannot find the most optimal translation in the translation memory. As the last striking element of MemoQ, it should also be noted that this program allows the user to review the translations several times separately. This means that before a translation, the first reviewer can decide that it is correct, but a second reviewer can differ and mark it without altering the opinions of the first one. This is extremely effective when carrying out group translations, since not everyone has the same opinion about a translation. 35 The statistics of the translation of the text by using the translation memory are shown in figure 26. It can be seen that this CAT tool is a bit pickier because of the labels, but it does not produce negligible results. Of 40,315 words MemoQ translates with a hundred percent or more of matching 3,972 words, which is about 9.85 of all the words. Results similar to those of the character count, where 198.169 characters the program sees as perfectly translated 19.658, 9.91 percent of the total. Something higher is however the number of concordant segments where of 5,229 finds 794 segments, 15.18 percent of the total segments. In addition to these results slightly lower than those of its competitors, MemoQ, just as Wordfast provides translation data with lower percentage of agreement. So, the user can see that if the program includes all the elements that exceed fifty percent agreement, which may only need to be postedited later, the results grow significantly. Doing this, MemoQ finds 28,955 translatable words, 71.82 percent of the total words of the manual, 142,339 characters, or as with the number of words a 71.82 percent of the total, and 3,569 segments, which is 68.25 percent of the total. All this without counting the data of repetitions that the statistic shows, since these can be or not translated. Figure 26: MemoQ statistics screen Ease of use When it comes to dealing with ease of use, it can be said that, as in the case of Trados, the design of the interface so similar to that of the globally used Microsoft Office suite makes the use of this tool easier for the user, making the user adapt quickly to the different functionalities of the program. 36 In addition, in order to make the learning curve of the program as smooth as possible, the program has on its website a large number of guides and courses related to the different products of the company, useful both for new users and for most experienced users in the world of CAT tools. Conclusions Although, according to the statistics shown, this program is slightly less efficient than the previous programs, MemoQ contains many elements that can turn it into the best option when it comes to choosing one of these tools, such as the ability to perform work and group corrections or the useful error detection system. However, this program does not get rid of having things that can be improved, since things like the excessive slow pace of the program when it comes to reformatting a notplain text into plain text, or the lack of options to translate a text into several languages at the same time can place it in a position worse than that of its competitors. 4.4. Déjà vu Another of the most used CAT tools in the world, and one of eldest on the market is Déjà Vu developed by the French company Atril. History The origin of this tool goes back to 1993 when this kind of tools had just appeared. At that time its creator, the Spanish Emilio Benito, developed an application with Microsoft Word interface as a product to work out the needs of the company he was working for. Even after the death of its creator in 2004, the CAT tool Déjà Vu continues to progress to this day, constantly including new versions and improvements that allow it to keep up with its competitors. Interface and functionalities As with the rest of the programs, the first aspect to analyze is what first catches the attention of any new user of the application. In the case of Déjà Vu, as it happens with 37 the Trados and MemoQ programs, it contains an interface that follows the designs of the Microsoft Office suite in order to be as user-friendly as possible. When creating a new project, this program takes us, as in the case of Trados, through a series of windows, such as the one in figure 27, where a large amount of information is entered, for example: basic data of the project (name, folder, due date…), objectives, target languages, terminology databases or even automatic translation providers. It should be noted that although, in its free version, the system only allows the user to choose a target language, it can already be seen that in the professional version it is possible to choose more than one language at a time. Figure 27: Déjà Vu new project screen In contrast to the other CAT tools, even though the project creation gives the user the option of entering a translation memory, it must be in its own format and the process of introducing it is independent of the project itself. So, to be able to enter a translation memory first, it must be created as if it were a new project and, once the blank translation memory is created, import the data into this new file (only in .TMX format for the test version) using the option shown in Figure 28. 38 Figure 28: Déjà Vu external data menu After the translation memory has been imported and added to the project, the translator's task can be started by using the pretranslation button located in the top menu of the program. This feature is full of options among which the user can access in order to obtain a more or less translated text with which to work; besides, it includes the possibility of automatically marking as valid those translations that give a hundred percent matching if the user is sure of its reliability. As in the rest of tools of this type, the program shows according to the segment all the information about possible translations, as can be seen in the small box on the right in figure 29, as well as its source and percentage of probability, shown just below the box of possible translations. 39 Figure 29: Déjà Vu segment translation screen On the other hand, one of the latest additions to this program is a function to detect possible errors similar to the one contained in the MemoQ tool. Figure 30 clearly shows how the errors are marked with an exclamation mark or an x depending on how serious the program considers it and how the information referring to the error is shown by a popup text when the mouse passes through the segment with errors. Figure 30: Déjà Vu segment translation screen error pop-up text 40 As in the rest of the computer-assisted translation tools, Déjà Vu has an analysis option in which all the numerical information regarding how the tool works is shown through the information provided to it (translation memories, terminological databases, automatic translators...). On this particular aspect, it should be noted that, of the five tools to analyze, this program has the best analysis tool due to its ability to customize the analysis based on the information that the user wants to obtain, however, this capacity also requires a longer processing time to be able to show all the desired data. In order to obtain empirical results to compare the different CAT tools, figure 31 shows the results related to the analysis of pretranslation through the translation memory developed for this purpose, apart from other features of the analysis tool. Thus we can see that this tool translates with a hundred percent of matching 4,280 words of the 38,790 that it finds (11.03 percent), 25,937 characters of 235,582 (11 percent) and 1,055 segments of the 4,653 in which this program separates the document, a 22.67 percent of the total. These data alone indicate a better pretranslation capacity than the previous tool analyzed. In addition, like the other tools analyzed above, with the exception of Trados, Déjà Vu also shows the results with a matching rate of less than one hundred percent. From these data it can be seen how, apart from repetitions that the program does not show if they have some type of percentage, the program is able to find similarity in 24,606 of the words (63.43 percent of the total), in 146,899 characters (62.35 percent) and 2,197 segments (47.21 percent). This allows the program to provide a pseudo-translation in almost 50 percent of the segments, limiting the translator's task in these segments to a simple revision or post-edition. 41 Figure 31: Déjà Vu analysis screen Ease of use In comparison with the rest of computer-assisted translation tools this program can also be considered as easy to use. This is because it also has two elements that make the learning process easy for the user. On the one hand, as it has already been mentioned above, Déjà Vu has an interface that is tremendously similar to that of any application related to Microsoft Office to facilitate its use. On the other hand, Atril has on its page a series of manuals and tutorials that allow the solution of any doubt that may arise during the use of the program. Conclusions The analysis cannot be concluded without mentioning the small flaws that any user can find when using this program. In the case of this program, these defects are mostly due to consumption issues, since, at least, when performing the analysis of the application, it could be observed the consumption of memory was far superior to the rest of the programs analyzed. This is a big problem for any user who works with old equipment because it may not be able to support so much memory and CPU load. 48 Another minor aspect, which may be relevant for some, is that Déjà Vu and Wordfast are left behind in terms of speed; the first one because of its slowness when performing tasks, and the second one because of its high memory and CPU consumption, as it is commented in the individual analysis. Finally, although a larger study could be spoken of other very influential aspects such as how the terminology bases interact or even the price of the program, it is worth noting the capacity to improve the program of both Trados and OmegaT through the inclusion of add -on downloadable in their respective web pages that give these applications an additional value with respect to their competition. 49 Ease of use Pretranslation Target languages at a time Translation suggestion Error Detection Aids for new users Speed Add-ons SDL Trados Studio Easy Automatically add segments with more than the wanted percentage One or more Automatically adds best option. Shows other options below No In app tutorial, training courses and manuals Fast Yes, on their online shop Wordfast Easy Yes, with a button One or more Sentence suggestions No Wiki Slow Not officially MemoQ Easy Yes, with a button One Sentence suggestions plus autogenerated glossary Yes Guides and courses Fast Not officially Déjà Vu Easy Yes, with a button One, more on pay versions Sentence suggestions plus autogenerated glossary Yes Manuals and tutorials Intermediate due to its heavy memory use Not officially OmegaT Mediocre No, only suggestions for the user One Sentence suggestions No Manuals and video tutorials Fast Yes Table 1: Feature evaluation 50 Empirical evaluation Another crucial aspect when comparing all these tools is how they perform their task. To prove that, there is no better way than submitting the same task to all the software and obtain numerical results to establish the evaluation of this data. In this paper, the text that has been given is, in particular, a manual in PDF format (non-plain text) to be able to get additional data on how the tools behave at the time of importing and segmenting something slightly more complex for the program. Surprisingly, the results of the five programs could not be more different in all aspects, as shown in table 2 (on page 49), leading to some conclusions such as those cited below. Just looking at the data of how the words and the total segments of the documents are imported and detected, it is observed that Déjà Vu obtains almost forty thousand characters more than the others, while at the same time, it is the one that finds fewer words in the document; MemoQ, on the other hand is, in this second aspect, the one with the most words with a result slightly superior to the rest. When it comes to taking all the information and dividing it into segments to translate, OmegaT produces more than two thousand segments less than the rest, being in this case Wordfast the one that segments the text into more parts. After the input data is analyzed, the user observes the data on how these tools perform the task of pretranslation; in this sense, Trados, honoring its position in the global market, is the tool that gets best initial results with a hundred percent matching, despite the fact that, unfortunately, in its statistics, it does not show any data on coincidences less than one hundred percent. On the other hand, it can be observed how Wordfast offers results with a hundred percent matching equally worthy in comparison with the rest of its competitors; it can be seen how the next one of the list does not reach 23 percent of maximum matching when it comes to translating segments, ten percent less than Wordfast, and how these competitors stay between ten and twelve percent when it comes to perfectly translating words and characters, around eight percent less than Wordfast. 51 MemoQ on the other hand gives a higher percentage of total translation suggestions (100% and less) but it is the one with the fewest perfect hits, possibly due to how it manages the labels All this information changes if we add the results with lower matching index, of which the computer-assisted translation tool will give a translation suggestion that is not so perfect but may only need a slight edition. In this case, MemoQ stands out over the rest, giving results of 68 percent of the segments and almost 72 percent of the words and characters, or more if we take into account that none of the programs say whether repeated data elements produce coincidence or no. The most unfavorable result in this aspect would be for Déjà Vu that does not manage to give translation suggestions nor to 50 percent of the segments in which it divides the document. 52 * Repetitions are part of the total of elements but not counted for elements with matching greater as 50% Table 2: Empirical evaluation 53 5. CONCLUSION As a conclusion, it can be said that in this paper it has been observed how these computer-aided translation tools greatly help the work of translators by speeding up and even reducing the translation process. Thanks to these tools, the work that could take months of work for the translator can be solved in a few days. This is why practically all translation companies, if not all, use this type of tools. Throughout this paper multiple tools have been seen and compared and it has been observed that there is none that surpasses all aspects to others, therefore, it is already the task of the translator to decide which is the tool that best suits his taste and needs and which is the one that offers him the most benefits with respect to others. Last but not least, it is important to note that, as its name suggests, these applications assist the translator, at no time they replace him. It is the task of the translator to provide and grow the translation memories that these programs use to develop their work and, even in the hypothetical case in which the tool had a translation memory so good that it would allow it to develop a perfect pretranslation, the task of reviewing and post-editing the document, if there was anything to change or improve, would still be in the translator's hands. 54 6. BIBLIOGRAPHY 6.1. Works cited Berglund Prytz, Y. (2012). Types of Corpora and Some Famous (English) Examples. Retrieved from https://weblearn.ox.ac.uk/access/content/group/3a217dfd-a8cd- 4034-8564-c27a58f89b9b/Handouts/CorpusTypes.pdf Biber, D. (1995). Dimensions of Register Variation: A Cross-Linguistic Comparison. Cambridge: Cambridge University Press. Carl, M., & Way, A. (2003). Recent Advances in Example-based Machine Translation. Dordrecht: Kluwer Academic Publishers. Corpas Pastor, G., & Seghiri Domínguez, M. (2010). Size Matters: A Quantitative Approach to Corpus Representativeness. In R. Rabadán (Ed.), Lengua, Traducción, Recepción. En Honor de Julio César Santoyo/ Language, Translation, Reception. To Honor Julio César Santoyo. León: Universidad. De Haan, P. (1992). The Optimum Corpus Sample Size?, in G. Leitner (Ed.), New Directions in English Language Corpora: Methodology, Results, Software Development. Berlin — New York: Mouton de Gruyter. Faya Ornia, G. (1996). Revisión y Propuesta de Clasificación de Corpus. Babel Revue Internationale De La Traduction / International Journal of Translation, 60(2), 234- 252. doi: 10.1075/babel.60.2.06fay Fernández-Rodríguez, M. (2010). Evolución de la Traducción Asistida por Ordenador. de las Herramientas de Apoyo a las Memorias de Traducción. Sendebar, 21, 201- 230. doi:10.30827/sdb.v21i0.374 Forcada, M. (2017). Making Sense of Neural Machine Translation. Translation Spaces. A Multidisciplinary, Multimedia, And Multilingual Journal of Translation, 6(2), 291-309. doi: 10.1075/ts.6.2.06for Heaps, H. (1978). Information Retrieval. New York: Academic Press. Hutchins, W., & Somers, H. (1992). An Introduction to Machine Translation. London: Academic. 55 Lagoudaki, E. (2006). Wayback Machine. Retrieved from https://web.archive.org/web/20070325114619/http://www3.imperial.ac.uk/portal/ pls/portallive/docs/1/7307707.PDF Leech, G. (1991) "The State of the Art in Corpus Linguistics", in Aijmer K. and Altenberg B. (eds.) English Corpus Linguistics: Studies in Honour of Jan Svartvik. London: Longman. McEnery, T., & Hardie, A. (2014). Corpus Linguistics: Method, Theory and Practice. Cambridge: Cambridge University Press. McEnery, T., & Hardie, A. (2012) “Support Website for Corpus Linguistics: Method, Theory and Practice”. Retrieved from http://corpora.lancs.ac.uk/clmtp . Quah, C. (2006). Translation and Technology. Basingstoke [Angleterre]: Palgrave Macmillan. Seghiri, M. (2014). Too Big or Not Too Big: Establishing the Minimum Size for a Legal Ad Hoc Corpus. HERMES - Journal Of Language And Communication In Business, 27(53), 85. doi: 10.7146/hjlcb.v27i53.20981 Tabor, J. (2013). CAT Tool Use by Translators: What Are They Using?. Retrieved from https://prozcomblog.com/2013/03/28/cat-tool-use-by-translators-what-are-they- using/ Zanettin, F. (2002). Corpora in Translation Practice. Retrieved from https://www.researchgate.net/publication/228806527_Corpora_in_translation_pra ctice Web. 5 Feb 2019 The History of SDL's Translation Software. (2019). Retrieved from https://www.sdltrados.com/about/history.html What Is Machine Translation?. (2019). Retrieved from https://www.sdltrados.com/solutions/machine-translation/ 56 6.2. Resources Atril. (2019). Déjà Vu X3 (Version 9.0.765) [Windows]. Lexitrad. (2019). ReCor (Version 2.0) [Windows]. Universidad de Málaga. MemoQ. (2019). MemoQ (Version 8.7.11) [Windows]. OmegaT. (2019). OmegaT (Version 3.6.0) [Windows]. SDL Trados. (2019). SDL Trados Studio 2019 (Version 15.0.0.20974) [Windows]. Wordfast. (2019). Wordfast Pro 5 (Version 5.6.0) [Windows].