scieee AI-readable full text Open interactive document viewer

Repositorio Institucional de Documentos

Abstract

Nuestra meta principal en esta tesis es proponer una solución para construir una ontología multilingüe, a través de la localización automática de una ontología. La noción de localización viene del área de Desarrollo de Software que hace referencia a la adaptación de un producto de software a un ambiente no nativo. En la Ingeniería Ontológica, la localización de ontologías podría ser considerada como un subtipo de la localización de software en el cual el producto es un modelo compartido de un dominio particular, por ejemplo, una ontología, a ser usada por una cierta aplicación. En concreto, nuestro trabajo introduce una nueva propuesta para el problema de multilingüismo, describiendo los métodos, técnicas y herramientas para la localización de recursos ontológicos y cómo el multilingüismo puede ser representado en las ontologías. No es la meta de este trabajo apoyar una única propuesta para la localización de ontologías, sino más bien mostrar la variedad de métodos y técnicas que pueden ser readaptadas de otras áreas de conocimiento para reducir el costo y esfuerzo que significa enriquecer una ontología con información multilingüe. Estamos convencidos de que no hay un único método para la localización de ontologías. Sin embargo, nos concentramos en soluciones automáticas para la localización de estos recursos. La propuesta presentada en esta tesis provee una cobertura global de la actividad de localización para los profesionales ontológicos. En particular, este trabajo ofrece una explicación formal de nuestro proceso general de localización, definiendo las entradas, salidas, y los principales pasos identificados. Además, en la propuesta consideramos algunas dimensiones para localizar una ontología. Estas dimensiones nos permiten establecer una clasificación de técnicas de traducción basadas en métodos tomados de la disciplina de traducción por máquina. Para facilitar el análisis de estas técnicas de traducción, introducimos una estructura de evaluación que cubre sus aspectos principales. Finalmente, ofrecemos una vista intuitiva de todo el ciclo de vida de la localización de ontologías y esbozamos nuestro acercamiento para la definición de una arquitectura de sistema que soporte esta actividad. El modelo propuesto comprende los componentes del sistema, las propiedades visibles de esos componentes, las relaciones entre ellos, y provee además, una base desde la cual sistemas de localización de ontologías pueden ser desarrollados. Las principales contribuciones de este trabajo se resumen como sigue: - Una caracterización y definición de los problemas de localización de ontologías, basado en problemas encontrados en áreas relacionadas. La caracterización propuesta tiene en cuenta tres problemas diferentes de la localización: traducción, gestión de la información, y representación de la información multilingüe. - Una metodología prescriptiva para soportar la actividad de localización de ontologías, basada en las metodologías de localización usadas en Ingeniería del Software e Ingeniería del Conocimiento, tan general como es posible, tal que ésta pueda cubrir un amplio rango de escenarios. - Una clasificación de las técnicas de localización de ontologías, que puede servir para comparar (analíticamente) diferentes sistemas de localización de ontologías, así como también para diseñar nuevos sistemas, tomando ventaja de las soluciones del estado del arte. - Un método integrado para construir sistemas de localización de ontologías en un entorno distribuido y colaborativo, que tenga en cuenta los métodos y técnicas más apropiadas, dependiendo de: i) el dominio de la ontología a ser localizada, y ii) la cantidad de información lingüística requerida para la ontología final. - Un componente modular para soportar el almacenamiento de la información multilingüe asociada a cada término de la ontología. Nuestra propuesta sigue la tendencia actual en la integración de la información multilingüe en las ontologías que sugiere que el conocimiento de la ontología y la información lingüística (multilingüe) estén separados y sean independientes. - Un modelo basado en flujos de trabajo colaborativos para la representación del proceso normalmente seguido en diferentes organizaciones, para coordinar la actividad de localización en diferentes lenguajes naturales. - Una infraestructura integrada implementada dentro del NeOn Toolkit por medio de un conjunto de plug-ins y extensiones que soporten el proceso colaborativo de localización de ontologías. Espinoza Mejía, Jorge Mauricio; Mena Nieto, Eduardo; Gómez Pérez, Asunción

Full text

2014 56 Jorge Mauricio Espinoza Mejía Ontology Localization Departamento Director/es Informática e Ingeniería de Sistemas Mena Nieto, Eduardo Gómez Pérez, Asunción Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA Departamento Director/es Jorge Mauricio Espinoza Mejía ONTOLOGY LOCALIZATION Director/es Informática e Ingeniería de Sistemas Mena Nieto, Eduardo Gómez Pérez, Asunción Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA Departamento Director/es Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA Ontology Localization Jorge Mauricio Espinoza Mej´ıa PhD Thesis Departamento de Inform´atica e Ingenier´ıa de Sistemas Universidad de Zaragoza Advisors : Dr. Eduardo Mena Dra. Asunci´on G´omez-P´erez March, 2014 Acknowledgements I would like to thank my advisors, Eduardo Mena and Asunci´on G´omezP´erez, for having always believed in my work and for having provided insightful feedback during all stages of the research described in this thesis. Their vision was fundamental in shaping my research and I am very grateful for having had the opportunity to learn from them. It is necessary to emphasize that this thesis is the result of a long period of research that started in 2005. In a first stage (24 months) in the Distributed Information Systems (SID) group at University of Zaragoza, we studied and designed the methods used to discover the set of candidate meanings for a given word (or words) from a pool of ontologies available on the Web. This approach was the core of our proposal to automatically discover the translations of an ontology element. The results obtained in the exploratory phase, were the base of a second stage of research more extensive and complex (29 months) within the Ontology Engineering Group (OEG) at Technical University of Madrid. In this period, we defined the methodology that supports the ontology localization activity and we developed the infrastructure that implements the methods, techniques, and tools for the management of ontology localization in distributed and collaborative environments. The friends at the OEG and the SID group provided a precious collaborative working environment that will be hard to forget. I owe special thanks to the contribution of Jorge Gracia, Elena Montiel-Ponsoda, Raquel Trillo, Boris Billaz´on, Ra´ul Palma, M. Carmen Su´arez-Figueroa, Jose Angel Ramos, Miguel Esteban, Victor Saquicela, Ra´ul Garc´ıa-Castro, Manuel Vilches, and Andr´es Garc´ıa. Nelly Mantilla and Hern´an Cuesta made our life in Zaragoza much easier. Dra. Guadalupe Aguado de Cea was always so readily available to help me with everything I needed. Dr. Oscar Corcho shared with me some of his inspiring ideas. Charito was the best person in the world I could have had to share all the good and bad moments of my PhD. I was very lucky to have found her. Cristina and Sofia my beautiful daughters, thanks to be with us. Last, but not least important, I would like to thank my parents, my siblings, and all my family for their great emotional support over all these years. iii A special mention to Frank, who spent some days of his life in the task of verifying the correct usage of the English in this thesis. I am also indebted to Jordi Bernard for his support with regard to the formalization of this work. Finally, I want to mention explicitly that the work contained in this doctoral thesis has been co-financed by an official scholarship from the University of Zaragoza - Banco Santander Central Hispano (citation 2005), for the European Commission in the context of the project Neon (FP6-027595), and the spanish CICYT project TIN2004-07999-C02-02. iv Abstract Our main goal in this thesis is to propose a solution to build a multilingual ontology through automatic localization of an ontology. The notion of localization comes from the Software Development area where it refers to the adaptation of computer software to non-native environments. In Ontological Engineering, the localization of ontologies could be considered as a subtype of software localization in which the product is a shared model of a particular domain, i.e., an ontology, to be used by a certain application. In particular, our work introduces a novel approach for the multilingualism problem, describing the methods, techniques, and tools for the localization of ontological resources and how multilingualism could be represented in ontologies. It is not the goal of this work to advocate one only approach to ontology localization, but rather to show the variety of methods and techniques that can be re-adapted of other knowledge areas to reduce the cost and effort that means enriching an ontology with multilingual information. We are convinced that there is not one unique approach to ontology localization. We concentrate, however, on automatic solutions for ontology localization. The approach presented in this dissertation provides a comprehensive coverage of ontology localization activity for the ontology practitioners. In particular, it gives a formal account of our general localization process by defining the inputs, outputs, and the main steps identified. Also, we consider various dimensions for localizing an ontology. Such dimensions allow us to establish a classification of different translation techniques based on methods taken from the discipline of machine translation. To facilitate the analysis of these translation techniques we introduce a framework that covers their main aspects. Finally, we give an intuitive view of the whole localization activity and we outline our approach to the definition of a system architecture that supports the ontology localization activity. The proposed model comprises the system components, the externally visible properties of those components, the relationships between them, and provides a base from which localization systems can be developed. The principal contributions of this work are summarized as follows: •A characterization and definition of the ontology localization problems, v based on the problems found in related areas. The characterization proposed takes into account three different problems of the localization: translation, information management, and multilinguality representation problems. •A prescriptive methodology for supporting ontology localization activity, based on existing localization methodologies from the fields of Software Engineering and Knowledge Engineering, as general as possible so that the methodology can cover a broad range of scenarios. •A classification of the ontology localization techniques, which can be used for comparing (analytically) different ontology localization systems as well as for designing new ones, taking advantage of state of the art solutions. •An integrated method for building ontology localization systems in a distributed and collaborative environment, which takes into account the more appropriate methods and techniques depending on: i) the domain of the ontology to be localized, and ii) the amount of linguistic information required for the final ontology. •A modular component to support the storage of the multilingual information associated to each ontology term. This approach follows the current trend in the integration of multilinguality in ontologies which suggests the suitability of keeping ontology and linguistic (multilingual) knowledge separated and independent. •A model based on collaborative workflows for the representation of the process usually followed by different organizations to coordinate the ontology localization activity in different natural languages. •An integrated infrastructure implemented within the NeOn Toolkit by means of a set of plug-ins and extensions that supports the collaborative ontology localization process. vi B.2 User Guide for Translators . . . . . . . . . . . . . . . . . . . . 271 B.2.1 Setting-up Translation Preferences. . . . . . . . . . . . 271 B.2.2 Translating Ontology Labels. . . . . . . . . . . . . . . 272 B.3 User Guide for Reviewers . . . . . . . . . . . . . . . . . . . . 274 B.3.1 Setting-up Revision Preferences. . . . . . . . . . . . . 274 B.3.2 Reviewing Translations. . . . . . . . . . . . . . . . . . 276 xiii xiv List of Figures 1.1 Workflow of a typical localization project. Adapted from [Esselink,2000]............................ 8 1.2 Localization levels for systematic product internationalization. Adapted from [Sturm, 2002]. . . . . . . . . . . . . . . . 10 1.3 Automatic translation approach for the Ontology Localization Activity [Espinoza et al., 2009a] . . . . . . . . . . . . . . 11 1.4 Workflow process used to localize an ontology. . . . . . . . . . 13 2.1 Different levels of analysis in an MT system. . . . . . . . . . . 24 2.2 The global architecture of the EWN database [Vossen, 2004]. 27 2.3 Multilingual ontology deployment in the DOSE platform. . . 28 4.1 The Ontology Localization Activity . . . . . . . . . . . . . . . 47 4.2 Ontology Localization levels. . . . . . . . . . . . . . . . . . . 55 4.3 Ontology Localization scenarios. . . . . . . . . . . . . . . . . 58 4.4 Human translator steps. . . . . . . . . . . . . . . . . . . . . . 59 4.5 Automatic translation approach for the Ontology LocalizationActivity. ........................... 61 5.1 Classification of ontology localization techniques. . . . . . . . 74 5.2 Ontology Example. . . . . . . . . . . . . . . . . . . . . . . . . 77 5.3 Parallel combination of translation algorithms. . . . . . . . . 102 5.4 Sequential composition of translation algorithms. . . . . . . . 103 6.1 The Automated Ontology Localization Life-Cycle Model. . . 111 6.2 Typical Ontology Localization Scenario. . . . . . . . . . . . . 117 6.3 General architecture to support Ontology Localization. . . . . 121 6.4 Workflow process used to localize an ontology. . . . . . . . . . 125 6.5 Synchronization of ontology and linguistic model. . . . . . . . 127 6.6 Workflow to the Translation level. . . . . . . . . . . . . . . . 128 6.7 Detailed Ontology Translator in a Localization System. . . . 130 6.8 Ontology Navigator with a selected ontology element. . . . . 137 6.9 Connecting the Ontology Model with the Linguistic Model (taken from [Montiel-Ponsoda, 2011b]). . . . . . . . . . . . . 138 6.10 Linguistic Information page that support the LIR model. . . 139 xv 6.11 User wizars used by the Workflow Localization Manager. . . . 140 6.12 A perspective of the Ontology Localization Activity. . . . . . 142 6.13 Extract of the sample university ontology. . . . . . . . . . . . 143 6.14 LabelTranslator strategy to localize concept, attributes and relation terms represented by simple labels. . . . . . . . . . . 143 6.15 Some translations of the ontology label “chair” into Spanish. 146 6.16 Context of the ontology label “chair”. . . . . . . . . . . . . . 147 6.17 LabelTranslator strategy to localize concept, attributes and relations represented by compound labels. ...........149 6.18 Algorithm to translate the compound label “AssociateProfessor”intoSpanish..........................152 7.1 NeOn Methodology scenarios for building ontology networks). 156 7.2 Actors involved in the Ontology Localization Activity. . . . . 166 7.3 Ontology Localization Filling Card. . . . . . . . . . . . . . . . 167 7.4 Tasks for Ontology Localization. . . . . . . . . . . . . . . . . 168 8.1 Level of correctness in label translations. . . . . . . . . . . . . 183 8.2 Type of errors found in the translation of ontology labels. . . 184 8.3 LabelTranslator Configuration for Collaborative Ontology Localization..............................193 8.4 Results of SUMI Questionnaire for LabelTranslator. . . . . . 195 8.5 Related items for the term “pest control” extracted from FAOTERM.............................201 8.6 Google definitions of the ontology label “pest control”. . . . . 201 8.7 Uses and “possibly” translated documents of the ontology label “pest control”. . . . . . . . . . . . . . . . . . . . . . . . 204 8.8 Final Ontology using an external module. . . . . . . . . . . . 206 8.9 Extract of the sample economy activity ontology. . . . . . . . 207 8.10 Screenshot of the Ontology Navigator view with the Translate action used by the LabelTranslator plug-in. . . . . . . . . . . 207 8.11 Some translations of the Ontology label “Bars” into Spanish. 208 8.12 Equivalent Translations for the Term “Bars” . . . . . . . . . 209 8.13 Linguistic Information associated to the Ontology Term “Bars”210 xvi List of Tables 4.1 Exact equivalent sample . . . . . . . . . . . . . . . . . . . . . 68 4.2 Near equivalence sample . . . . . . . . . . . . . . . . . . . . . 68 4.3 Partial equivalence sample . . . . . . . . . . . . . . . . . . . . 68 6.1 Some lexical templates to translate a compound label from English into Spanish. . . . . . . . . . . . . . . . . . . . . . . . 151 7.1 Common tasks in software localization methodologies . . . . . 163 7.2 Ontology Localization task and its corresponding tool. . . . . 169 8.1 Ontologies corpus statistics. . . . . . . . . . . . . . . . . . . . 177 8.2 Five point scale for fluency and adequacy measures. . . . . . 179 8.3 Results obtained in the three experiments for Spanish. . . . . 179 8.4 Results obtained in the three experiments for German. . . . . 179 8.5 Multilingual Ontologies used in the evaluation. . . . . . . . . 188 8.6 Translation Techniques Comparison. . . . . . . . . . . . . . . 189 8.7 Translation Resources Comparison. . . . . . . . . . . . . . . . 191 8.8 Linguistic information related to the area of “pest control”. . 202 8.9 Linguistic information related to the area of “control (of a pest)” ...............................203 8.10 Ranked translations of the term “pest control” for French and Italian...............................205 8.11 Semantic fidelity evaluation results. . . . . . . . . . . . . . . . 209 xvii xviii Chapter 1 Introduction In the context of the Semantic Web [Berners-Lee et al., 2001], resources on the net can be enriched by well-defined, machine understandable metadata describing their associated conceptual meaning. Ontologies constitute the foundation upon which to build the whole new generation web, and describe human knowledge by specifying concepts related to many specific areas of interest and by modeling relationships between them. As with the World Wide Web (WWW) [Berners-Lee et al., 1992], the success or failure of the Semantic Web will be determined to a large extent by easy access to, and availability of high-quality and diverse content [Benjamins et al., 2002]. In this respect, an important challenge that needs to be addressed is the multilingualism problem, which until now has not been properly investigated [Tjoa et al., 2005]. This problem already exists in the current Web, and should also be tackled in the Semantic Web. Studies on language distribution over WWW content show that even if English is the predominating language for documents, there exists an important amount of resources written in other languages, according to the following distribution: English 26.8%, Chinese 24.2%, Spanish 7.8%, Japanese 4.7%, Portuguese 3.9%, German 3.6%, Arabic 3.3%, French 3.0%, Russian 3.0%, Korean, 2.0%, other languages 17.8%1. In the case of the Semantic Web the problem is similar: most of the ontologies that have been built so far have English as their basis. Nevertheless, although English is now the de facto language for science and technology, other spoken languages are used and it is important to provide methods and tools both to support the definition of ontologies expressed in languages other than English and also to support interoperability across ontologies written in different languages. However, looking at the statistics of two well-known gateways of the Semantic Web as Watson2and OntoS1Obtained on May 31, 2011 from http://www.internetworldstats.com 2http://watson.kmi.open.ac.uk/WatsonWUI 1 CHAPTER 1. INTRODUCTION elect3, we can observe that the number of multilingual ontologies available on the Web is insignificant compared with the number of monolingual ontologies. The work presented in this thesis proposes an alternative to build a multilingual ontology through automatic4localization of an ontology. The notion of localization comes from Software Development where it refers to the adaptation of computer software to non-native environments. From an Ontology Engineering perspective, localization makes it possible to adapt an ontology to different languages and cultures [Su´arez-Figueroa and G´omezP´erez, 2008]. This definition has been subsequently revisited in [Cimiano et al., 2010] to refer to “the process of adapting a given ontology to the needs of a certain community, which can be characterized by a common language, a common culture or a certain geopolitical environment”. We should note here that the starting point of this work in 2006 was the NeOn Methodology (see [Suarez-Figueroa, 2013]) designed in the framework of the NeOn project5. This methodology identifies nine flexible scenarios that covers commonly occurring situations in the ontology development process. This thesis and the research work presented in [Montiel-Ponsoda, 2011b] began at the same time to support the Scenario 9: Localization of Ontologies. A part of the problem definition was carried out together, then, the thesis of Montiel-Ponsoda derived to the definition of a model to represent multilingual information in ontologies (LIR) [Montiel-Ponsoda et al., 2011], whereas this work focused on the methods, techniques, and tools for supporting the automatic localization of ontologies. Our main contributions are: i) the definition and characterization of the ontology localization problem, ii) the identification and implementation of the methods, techniques and tools for the automatic management of ontology localization in collaborative and distributed environments, and iii) the definition of a methodology to support the ontology localization activity. In this chapter we first describe the motivation from which this work arises. Secondly, we explain briefly the current trends for transforming a monolingual resource to multilingual. Thirdly, we introduce some features of the software localization industry that we believe are valid in the ontology localization context. Fourthly, we introduce the main features of our proposal to perform the localization of an ontology. Finally, we present the structure of this thesis. 3http://olp.dfki.de/ontoselect/ 4We will use the term automatic to refer to an efficient process that allows reducing the human effort of translating a domain-specific ontological resource. We are conscious that the specificity of vocabulary terms in most ontologies precludes fully-automatic translation using general-domain translation resources. 5www.neon-project.org 2 1.1. MOTIVATION 1.1 Motivation Currently, a great effort has been done in the construction of ontologies. Although access to top-quality ontologies (e.g., Galen6, CYC7, or AKT8) is in many cases free and unlimited for users all around the world, most of these ontologies can be said to be essentially monolingual, i.e., documented in one natural language only, and this language is often English as an international lingua franca. However, there is a growing need for multilingual ontology resources that overcome communication barriers arising from cultural-linguistic differences, lack of excellent command of English, need for high precision in communication, etc. In fact, multilingual knowledge is even more prevalent in those countries that have more than one official language [Yang and Li, 2003]. For example, Chinese and English are official languages of Hong Kong; French and English for Canada; and Dutch, French, and German for Belgium. Moreover, the use of ontologies has grown not only in terms of the number of application domains but also in the number of natural languages chosen to build domain specific knowledge bases. Thus, multilingual ontologies are nowadays demanded by institutions worldwide with a huge number of resources available in different languages. Basically, usage of multilingual ontologies traverses many disciplines, and has become an urgent need in certain organizations. For instance, in Agriculture, the Food and Agriculture Organization (FAO) has expressed the need for semantically structuring the information they have in different natural languages. Since all FAO official documents must be made available in Arabic, English, Chinese, French, Russian and Spanish, a large amount of research has been carried out in translating large multilingual agricultural thesauri [Chun and Wenlin, 2002], in mapping methodologies for thesauri [Liang et al., 2005, Liang and Sini, 2006], and in defining requirements to improve the interoperability of these multilingual information resources [Caracciolo et al., 2007]. In Education, the Bologna declaration has introduced an ontology-based framework for qualification recognition [Vas, 2007] across the European Union, in an effort to best match labor markets with employment opportunities. In E-Learning, educational ontologies are used to enhance learning experience [Cui et al., 2004], and to empower system platforms with high adaptivity [Sosnovsky and Gavrilova, 2006]. In the Finance domain, ontologies are used to model knowledge in the stock market domain [Alonso et al., 2005] and portfolio management [Zhang et al., 2002]. In Medicine, ontologies are employed to improve knowledge sharing and knowledge reuse. For example, a notable amount of research has focused on the creation of an ontology of traditional Chinese medicine. 6www.co-ode.org/galen/ 7http://www.opencyc.org/downloads 8http://www.aktors.org/publications/ontology/ 3 CHAPTER 1. INTRODUCTION A further factor that has increased the need for multilingual ontologies is the development of some ontology-based systems that need to interact with information in natural languages. Some examples of these applications are: cross-lingual information retrieval [Guyot et al., 2005], multilingual question answering [Pazienza et al., 2005] or knowledge management [Segev and Gal, 2008]. These examples can serve to highlight the importance of adding multilingualism/multilingual information to ontologies, before trying to solve the numerous pending problems that still exist with the current monolingual approach. It is worth mentioning that at the time of starting this thesis there was no well-defined and broadly accepted definitions of what the ontology localization activity entailed. In fact, from the progress made in this work, other authors introduced new adaptations of the definition of localization of ontologies, putting emphasis on adaptation of ontology to the needs of the target community (see [Cimiano et al., 2010]). In 2010, the EU Multilingual Ontologies for Networked Knowledge project (Monnet) is created to continue the research and implementation of services that allow the automatic localization of ontologies to different languages. Some of the translation services implemented in Monnet are based on the results obtained during the development of this thesis. Furthermore, it is important to emphasize that from July 2011, the Ontology Lexica Community Group9at the World Wide Web Consortium (W3C) is working mainly on the development of models for the representation of lexica (and machine readable dictionaries) relative to ontologies. The development of these models can help to enhance existing language resources by linked data principles, and improve the performance of applications as varied as ontology localization, machine translation, information extraction, multilingual access and presentation and natural language generation [McCrae et al., 2012]. Finally, to our knowledge, no other study has focused on the techniques, methods and tools for this activity. For this reason, in this thesis we present our approach for localizing an ontology into different natural languages. 1.2 From Monolingual to Multilingual Ontologies Over the last decade, research on ontologies was concentrated on methodologies and technologies for supporting the creation and management as well as the population of ontologies. There are some well recognized methodological approaches (e.g., METHONTOLOGY [Fern´andez-L´opez et al., 1999], OnTo-Knowledge [Staab et al., 2001], DILIGENT [Pinto et al., 2004], and NeOn Methodology [Su´arez-Figueroa, 2010]) that provided guidelines to help researchers to develop ontologies. However, most of existing methodologies 9http://www.w3.org/community/ontolex/ 4 1.4. ONTOLOGY LOCALIZATION APPROACH However, we believe that an automatic translation process is possible in the case of ontology elements, due that ontologies consisting of concepts and relationships that are stated clearly and succinctly [Espinoza et al., 2008a]. From our point of view this is the main contribution of Ontology Localization in this thesis. Four main steps are followed in order to discover the most appropriate translations: localization step selection, term context extraction, ontology label translation, and translation revision [Espinoza et al., 2008a,Espinoza et al., 2008b,Espinoza et al., 2009a]. Figure 1.3: Automatic translation approach for the Ontology Localization Activity [Espinoza et al., 2009a] . •Localization Selection. The first step involves the selection of the ontology elements that need be localized to different natural languages. This task is especially important when only limited time or recourses are available. In such a case, it might be interesting to know which part of the ontology can best be translated. •Term Context Extraction. Once the terms to be localized are identified, it is necessary to extract up to a certain depth the context of each ontology term to be localized. That is (depending on the type of term), their synonyms, textual descriptions, hypernyms, hyponyms, properties, domains, roles, associated concepts, etc. The context of an ontology term allows discerning among the different meanings that an ontology label may have. •Ontology Label Translation. The goal of the label translation task is to discover the more appropriate translation of each ontology label. In 11 CHAPTER 1. INTRODUCTION our approach, the task of translation is performed by the combination of different translation methods based on MT techniques. These techniques are combined following different translation strategies inspired from multi-engine machine translation approaches. •Translation revision. A revision process is needed to evaluate the quality of the obtained translations. This task includes measuring the adequacy and fluency of all translations. Notice that these steps cover only the translation task of the ontology localization activity. However, in this work we also describe the life-cycle model by means of representation of the major components of this activity and their interrelationships in a graphical framework that can be easily understood and communicated. 1.4.2 Collaborative Localization Management Despite the high level of automation of our processes, we do not dream of full automation and we acknowledge that the human component is critical. We will always need highly qualified individuals to post-edit the MT output, to provide feedback, to perform a manual check of suggested translation resource entries and, most importantly, to encourage and assist one another. In fact, on most large projects today, localization is a collaborative effort, where the number of users participating in localization ranges from a handful to a couple of dozens. Examples of such collaborative localization processes can be found in international institutions like the FAO, who have been developing and localizing the AGROVOC Thesaurus [AGROVOC, 2005], which is used widely to index agricultural information material all over the world. Thus, in the building of the AGROVOC Thesaurus, the translations were provided by the same thesaurus experts, then, a set of specialists whose main work is the translation of agricultural science literature were responsible for checking and approving the thesaurus translation work. With larger groups of users contributing to ontology localization, we believe that it is necessary to define appropriate workflows, strategies and infrastructure to support the process that coordinates the collaborative ontology localization within an organizational setting. This process can be modeled as a collaborative workflow, describing how project participants reach consensus on ontology label translations, who can perform translations, who can comment on them, when ontology label translations become public and so on. The collaborative workflow that we propose in this thesis is designed to support all aspects of the ontology localization activity. However, the details of this process, as well as the configuration of the collaborative scenario can vary from one organization to another. Thus, in some scenarios the collaborative workflow may be 12 1.4. ONTOLOGY LOCALIZATION APPROACH configurated to omit the use of reviewers or may not perform the automatic localization of the ontological labels. In the following we summarize the main steps in the workflow process to localize an ontology (see Figure 1.4): Figure 1.4: Workflow process used to localize an ontology. •An ontology is passed to the Localization Manager for localization. •The Localization Manager manually selects the ontology labels to be localized and sends the selected labels for translation. •A translator downloads the selected labels to be localized and (s)he performs the translations using an automated localization tool (as proposed in this thesis) or an intensive manual process. •Once translation activities have been accomplished, the translators upload the translated ontology labels and send them for review. •The reviewers download the translated labels and check for possible errors. •Finally, the Localization System updates all linguistic information of each localized label. Our contributions about providing collaborative ontology localization can be summarized in the following points: •Analysis of the main requirements to support a localization of ontologies in collaborative and distributed settings. •Design of a formal model based on graphs for the representation of the activities usually followed by different organizations dedicated to localization. 13 CHAPTER 1. INTRODUCTION •Design and implementation of a centralized repository to store the work of the different ontology stakeholders. •Design of an architecture for managing versions of localized ontologies, controlling ontology access (through some form of check in/check out and file locking), and enabling remote or distributed access. •Integration and implementation of an collaborative workflow to manage the sequence of translation/review/edit tasks, providing the status of tasks and processes, and notifying participants of changes in state, new work, or other information. •Development of an architecture for supporting customizable workflows for collaborative ontology localization. •Development of a set of interfaces that allows users collaboratively to perform the different tasks of the localization process. All these aspects have been considered in the design of the LabelTranslator system, our approach to perform an automated localization in distributed and collaborative environments. 1.4.3 Modular Storage of the Linguistic Information In her doctoral thesis, Montiel-Ponsoda [Montiel-Ponsoda, 2011a] analyzes the state of the art on models or formalisms to represent multilingual information in ontologies. She identifies three main ways of obtaining a multilingual ontology, depending on the layer(s) involved in the localization activity: •Including multilingual labels in the ontology is the most widespread modeling option within the ontological community nowadays, because it is well supported by the most popular ontology development languages: RDF(S) [Brickley and Guha, 2000] and OWL [Bechhofer et al., 2004]. It consists of making use of the labeling functionality of RDF(S) and OWL ontology representation languages. In this case labels can be integrated in the ontology in as many languages as the user wishes. However, this approach does not permit to define any relation among the linguistic annotations themselves (e.g., saying that one is a synonym or translation of the other). This results in a bunch of unrelated data whose motivation is difficult to understand even for a human user. •Combining the ontology with a mapping model assumes the existence of an original ontology and one or several ontologies localized to different natural languages, all of them represented as independent ontologies. The localized (monolingual) ontologies may have been obtained after performing the localization activity on the original ontology. This option enables independent conceptualizations in each language, what 14 1.4. ONTOLOGY LOCALIZATION APPROACH may better capture the specificities of each culture, but the establishment of mappings or alignments among conceptualizations in different languages is by no means trivial, since mismatches arise due to each conceptualization capturing the cultural specificities of each language. •Associating the ontology with an external linguistic model allows that the elements of the ontology have links to linguistic data stored outside the ontology. This type of representation allows the enrichment of domain ontologies with linguistically rich and complex models. Since these are external portable models, they can be associated to any domain ontology and published with them. Since there is just one conceptualization, this model is not as flexible as the previous one, in which cultural specificities were captured at the conceptual layer (despite the limitations imposed by interoperability and mapping discovery). In this thesis, we follow the current trend in the integration of multilinguality in ontologies (third approach above), which suggests the suitability of keeping ontology knowledge and linguistic (multilingual) knowledge separated and independent. Three of the models that follow this trend are: LexInfo [Buitelaar et al., 2009], the Linguistic Information Repository (LIR) [Montiel-Ponsoda et al., 2011], and the Lexicon Model for Ontologies (Lemon) [McCrae et al., 2011b], which is being standardized in the W3C Ontology-Lexicon Community Group11. Our contribution here is to supply the support for a modular approach, in which the conceptualization is kept apart from the multilingual information [Espinoza et al., 2009a]. This representation form allows the inclusion of as much linguistic information as wished, as well as the possibility of establishing links among the linguistic elements within one language or across languages. In this sense, nuances or differences between languages can also be reported and even formalized in the terminological layer, in order to avoid the 100% equivalence correspondence among the different names of ontology elements. Relevant information as, e.g., the provenance of the linguistic elements, can also be included. 1.4.4 Automatic Synchronization Process Whereas the translation process of ontology labels per se implies certain difficulties, the maintenance and updating of translated ontology labels throughout the ontology life cycle also requires special attention. The main difficulty in the management activity is to identify policies for managing changes in the ontology terms and their translated labels. In our case, this situation 11http://www.w3.org/community/ontolex/ 15 CHAPTER 1. INTRODUCTION is even more complicated, because we provide a model where sets of ontology terms and linguistic information associated (in different languages) are separately stored (see previous aspect). In order to keep both models synchronized we first need to find out exactly what has been changed in the ontology model, then find the equivalent places in the linguistic model, and only then start the updating. Thus, in this thesis we provide a comprehensive solution to the problems of managing the conceptual knowledge and the linguistic knowledge by means of synchronization techniques [Espinoza et al., 2009a]. 1.4.5 Prescriptive Methodological Guidelines The complexities of localization projects are very different from the complexities of software development projects. Unlike software development projects, in which well-established and precise practices and methodologies exist, the localization projects do not explain the localization process with the same style and granularity as those methodologies for developing software. To facilitate the prompt assimilation of ontology localization by software developers and ontology practitioners, in this work we propose methodological guidelines in a manner non-oriented to researchers. We also include examples of how to use the guidelines in different cases. From a methodological point of view, our contributions can be summarized as follows: •A characterization of the ontology localization problems. •A study of the different strategies for representing multilingual information in ontologies. •A prescriptive guideline to help users in the development of multilingual ontologies. To the best of our knowledge, the study presented here is the first attempt to offer guidelines for the localization of ontologies. 1.5 Structure of the Thesis This thesis is composed of nine chapters, including this one. Chapter 2 reviews the technological context in which this work has been developed. Particularly, we first review the aspects concerning to machine translation as the key to automatically discover the translations of the elements of an ontology. Then, we analyze the methods for the building of multilingual ontologies, followed by a description of some related works in order to do a comparison between different approaches and ours. 16 1.5. STRUCTURE OF THE THESIS In Chapter 3 we provide a presentation of the objectives and contributions of the thesis. We also describe how this thesis can contribute to this field of research with the set of assumptions, hypotheses and restrictions taken into account. In Chapter 4 we first explain the terminology related to ontology localization activity, providing the meaning that will be used for the distinct terms. Then, we propose a characterization of ontology localization based on the problems found in related areas. The chapter also describes the different scales of localization used in ontologies depending on the type of ontology elements to be localized and the level of adaptation required. Also, we analyze which elements or parts of the ontology are to undergo localization. After the foundations, we give a formal account of our general localization process by defining the input, output, and the four main steps identified. In Chapter 5, we present a classification of different translation techniques based on the way of modeling the context used to disambiguate the candidate translations and the type of resources used to localize an ontology into different natural languages. Later in this chapter we introduce the main characteristics of different translation techniques. To facilitate the analysis of these translation techniques we introduced a framework that covers their main aspects. Then, we present at the strategic level, some natural ways to compose and combine the output of different translation techniques for obtaining ontology translations. Finally, we discuss an alternative for classifying the localization approaches. Chapter 6 describe the life-cycle of the ontology localization activity. Then, it details the main modules to allow such an ontology localization approach in distributed and collaborative environments, but first it introduces some basic requirements for an ontology localization system. Finally, it describes general comments and different technical details related to the LabelTranslator system, our approach to performing an automated localization in distributed and collaborative environments. Chapter 7 describes in detail the general methodology used to guide users in the development of multilingual ontologies. First, we describe the design principles contemplated when defining the methodology and the process followed to define it. Finally, it details the methodology by describing its actors, processes and tasks. Chapter 8 is dedicated to the evaluation of our work according to the initial set of hypotheses. We describe a set of experiments that were carried out with the objective of evaluating the methodological and technological aspects of the localization activity. First, we describe the experiments used to evaluate some aspects related to the translation ranking techniques, where the task is to select the most appropriate translation of ontology labels. Then, we describe the study used to assess the usability of the LabelTranslator system for carrying out the Ontology Localization activity in distributed and collaborative environments. Finally, we describe two case 17 CHAPTER 1. INTRODUCTION studies to measure the understanding and usability of the methodological guidelines. In Chapter 9 we present the conclusions as well as our main contributions; we finish with the future research work. 18 Chapter 2 Technological Context Along this chapter, and before presenting in detail our proposal to localize an ontology to different natural languages, we would like to describe first the technology related to our work. We start with an introduction to ontologies, describing what an ontology is from the perspective of Computer Science, and some related issues about the development of ontologies. Secondly, we present an introduction to Machine Translation (MT). We include the motivation for using MT in the localization of ontological resources and the classification of different MT systems. Finally, we describe the different methods used to build the multilingual ontologies, the main goal of the ontology localization activity. We analyze the strengths and drawbacks of each method to identify open research problems and work assumptions. 2.1 Ontologies The Semantic Web [Berners-Lee et al., 2001] tries to achieve a semantically annotated Web, in which search engines can process the information contained in web resources from a semantic point of view, drastically increasing the quality of the information presented to the user. This approach requires a global consensus in defining the appropriate semantic structures (ontologies) for representing any possible domain of knowledge. In this sense, ontologies can be understood as the scaffolding of the Semantic Web. As ontology is one of the key terms in this thesis, this section will define its basics here. 2.1.1 Ontology basics In [Studer et al., 1998] an ontology is defined as a formal, explicit specification of a shared conceptualization. Conceptualization refers to an abstract model of some phenomenon in the world by having identified the relevant concepts of that phenomenon. Explicit means that the type of concepts 19 CHAPTER 2. TECHNOLOGICAL CONTEXT used, and the constraints of their use, are explicitly defined. Formal refers to the fact that the ontology should be machine-readable. Shared reflects the notion that an ontology captures consensual knowledge, that is, it is not private, but accepted by a group. Other approaches have defined ontologies as explicit specifications of a conceptualization [Gruber, 1995] or as a shared understanding of some domain of interest [Uschold and Grunninger, 1996]. Different knowledge representation formalisms exist for the definition of ontologies. However, they share the following minimal set of components: •Classes: represent concepts, which are taken in a broad sense, that is, they can represent abstract concepts (intentions, beliefs, feelings, etc) or specific concepts (people, computers, tables, etc). Classes in the ontology are usually organized in taxonomies through which inheritance mechanism can be applied. •Relations: represent a type of association between concepts of the domain. Ontologies usually contain binary relations. The first argument is known as the domain of the relation, and the second argument is the range. Binary relations are sometimes used to express concept attributes. Attributes are usually distinguided from relations because their range is a data type, such as string, numeric, etc., while the range of a relation is a concept. •Instances: are used to represent elements or individuals in an ontology. There exist several categorizations of ontologies in function of a particular aspect (such as expressiveness [Lassila and McGuinness, 2001] or subject and type of structure [van Heijst et al., 1997]. An interesting classification was proposed by [Guarino, 1998], who classified types of ontologies according to their level of dependence on a particular task or point of view. •Top-level ontologies: describe very general concepts like space, time, event, which are independent of a particular domain. It seems reasonable to have unified top-level ontologies for large communities of users. Some examples are Sowa’s [Sowa, 1999], Cyc’s [Lenat and Guha, 1989], and SUO [Pease and Niles, 2002] •Domain ontologies: describe the vocabulary related to a generic domain by specializing the concepts introduced in the top-level ontology. There are several representative ontologies in the domains of ecommerce (UNSPSC1, NAICS2, SCTG3, RosettaNet4), medicine (GA1http://www.unspsc.org 2http://www.naics.com 3http://www.bts.gov/programs/cfs/sctg/welcome.htm 4http://www.rosettanet.org 20 2.3. METHODS FOR THE BUILDING OF MULTILINGUAL ONTOLOGIES Princenton WordNet11 [Miller, 1995], developed as a monolingual lexical database for American English. The work initiated in the EWN project is now being continued by the Global Wordnet Association (GWA)12. The aim of EuroWordNet was to develop a multilingual lexicon with wordnets for several European languages (English, Dutch, Spanish and Italian), which could be used “to improve recall of queries via semantically linked variants in any of these languages”. The general approach for EWN was to build the multilingual database taking advantage of existing resources in each language. Participants from each country were responsible for a language specific wordnet using their already available tools and resources built up in previous national and international projects. As in WordNet, information about nouns, verbs, adjectives and adverbs was organized in synsets. A synset is “a set of words with the same part-of-speech that can be interchanged in a certain context” [Vossen, 2004]. Synsets are related to each other by semantic relations, such as hyponymy or meronymy, for example. The wordnets in EuroWordNet are considered “autonomous language specific ontologies”. Then, multilingual wordnets are interconnected through an Inter-Lingual-Index (ILI), a list of unstructured meanings mainly from Princenton WordNet, specifically WordNet1.5, that provide the mappings across the wordnets as illustrated in Figure 2.2. Figure 2.2: The global architecture of the EWN database [Vossen, 2004]. MultiWordNet (MWN) is a multilingual lexical database including information about English and Italian words. The model adopted within MWN, 11http://wordnet.princeton.edu/ 12http://www.globalwordnet.org/gwa/gwa grid.htm 27 CHAPTER 2. TECHNOLOGICAL CONTEXT consists of building language specific wordnets keeping as much as possible of the semantic relations available in the WordNet. This was done by building the new synsets in correspondence with the WordNet synsets, whenever possible, and importing semantic relations from the corresponding English synsets; i.e., if there are two synsets in WordNet and a relation holding between them, the same relation holds between the corresponding synsets in the new language. The MWN model minimizes the discrepancies that can appear when two wordnets are built independently for two different languages, by strictly adhering to the WordNet building criteria and subjective choices. However, MultiWordNet explicitly recognizes the presence of “lexical gaps” in the correspondence between different languages, due to missing direct translations of some words. Another approach is given by [Bonino et al., 2004], in which the authors introduce a simple approach to multilingual semantic elaboration using the Distributed Open Semantic Elaboration platform (DOSE). This approach uses a language independent ontology in which concepts are defined as highlevel entities for which language dependent definitions are specified. Such entities are linked to a set of different definitions, one for each supported language, and a set of words that the authors call synset. Figure 2.3 shows the multilingual ontology deployment used in the DOSE platform. Figure 2.3: Multilingual ontology deployment in the DOSE platform. The ontology is physically distinct from definitions and synsets, allowing separate management of concepts and language-specific information, isolating the semantic and the textual layers. This assumption guarantees sufficient expressive power to model conceptual entities typical of each language and, at the same time, reduces redundancy by collapsing all common 28 2.3. METHODS FOR THE BUILDING OF MULTILINGUAL ONTOLOGIES concepts into a single multilingual entity. Synsets and textual definitions are created by human experts through an interactive refinement process. A multilingual team works on concept definitions by comparing ideas and intentions, aided by domain experts with linguistic skills for at least two different languages, and formalizes topics in a mutual learning cycle. Finally, the work presented in [Segev and Gal, 2008] proposes an ontologybased model for building multilingual applications. Their model was based on a global ontology manually designed for a specific domain. Additionally, this model uses local context to specify the ontology. The combination of ontologies and contexts lends itself well to multilingual applications in which a single ontology fails to capture all nuances that stem from language and cultural differences. The procedure used for adapting an existing ontology to the needs of a multilingual environment includes the following four steps, selection,collection,extraction, and adaptation. In the selection step an existing ontology is chosen. In the collection step, sample documents that represent ontology concepts are collected. Contexts are extracted from sample documents in the extraction step. Finally, extracted contexts are associated with ontology concepts. The ontological system works simultaneously in multiple languages, and it is easily expansible and adaptable to other languages. As a final comment, we can say that some of the steps in this approach need intensive human labor. Main advantages and shortcomings An advantage of the works that adopt this approach is that it is easier to ensure language neutrality (i.e. lack of bias towards any one language). However, the costs of producing such ontology, as well as the definition and multilingual equivalence of its terms, have to be established a priori. There are still two critical issues that need to be solved before the building of multilingual from scratch or other similar efforts can be used as a shared conceptual framework for all languages: the scarcity of lexical semantic information (especially from endangered languages), and the lack of a linguistically-motivated shared conceptual core as the basis of multilingual conceptual representation. 2.3.2 Merging of Existing Ontologies This approach for localizing an ontology may be adopted when monolingual ontologies already exist in similar domains and therefore a multilingual ontology could be quickly and robustly constructed from monolingual resources. The aligning and merging of ontologies is actually one of the most active domains of investigation in the Semantic Web community [Euzenat and Shvaiko, 2007, Ehrig, 2007]. However, the issue of mapping ontologies written in different natural languages is still relatively unexplored. 29 CHAPTER 2. TECHNOLOGICAL CONTEXT The works that adopt this localization procedure try to identify the similarities between heterogeneous ontologies and then try to automatically create suitable mappings for transformation. Depending on the way that alignments between ontologies are discovered, these works can be grouped into two categories: cross-lingual ontology alignment approaches and generic approaches that involve machine translation tools and monolingual ontology matching techniques. Note that the ontology localization activity proposed in this thesis can contribute as a plausible solution for the approaches that use MT tools. More details can be consulted in the Section 5.9.4 Cross-lingual Alignment Approaches Dorr et al. [Dorr et al., 2000] and Palmer & Wu [Palmer and Wu, 1995] took a structural approach to this problem. They focused on HowNet verbs and used thematic-role information, which denotes the contexts in which a particular verb may occur. The HowNet thematic-role specifications are mapped to word classes in an existing classification of English verbs called EVCA [Levin., 1993], whose structure is similar to that of the verb classes in HowNet. These mappings are then used to align English EVCA verbs to Chinese HowNet verbs. In [Chen and Fung, 2004] the authors proposed an automatic technique to associate the English FrameNet lexical entries to the appropriate Chinese word senses. Each FrameNet lexical entry is linked to Chinese word senses of a Chinese ontology database called HowNet. First, each FrameNet lexical entry is associated with Chinese word senses whose part-of-speech is the same and Chinese word/phrase is one of the translations. In the second stage of the algorithm, some links are pruned out by analyzing contextual lexical entries from the same semantic frame. In the last stage, some pruned links are recovered if its score is greater than the calculated threshold value. Carpuat et al. [Carpuat et al., 2002] merged thesauri that were written in English and Chinese into one bilingual thesaurus in order to minimize repetitive work while building ontologies containing multilingual resources. A language-independent, corpus based approach was employed to merge WordNet and HowNet by aligning synsets from the former and definitions of the latter. Similar research was conducted in [Malais´e et al., 2007] to match Dutch thesauri to WordNet by using a bilingual dictionary, and concluded a methodology for vocabulary alignment of thesauri written in different natural languages. Monolingual Alignment Approaches with MT Support Asanoma [Asanoma, 2001] aligned the Japanese Goi-Taikei ontology with WordNet by first translating a significant subset of the WordNet synonym sets (synsets) into Japanese, automatically matching these based on (mono30 2.3. METHODS FOR THE BUILDING OF MULTILINGUAL ONTOLOGIES lingual Japanese) lexical overlap, and filling in the gaps for the remaining classes based on their hierarchical positioning relative to the aligned classes. Trojahn et al. propose a multilingual ontology mapping framework in [Trojahn et al., 2008], which consists of smart agents that are responsible for ontology translation and capable of negotiating mapping results. In [Fu et al., 2009a] Fu et al. present the SOCOM Framework which is designed specifically to achieve cross-lingual ontology mapping. The SOCOM framework divides the multilingual mapping task into three phases: an ontology rendering phase, an ontology matching phase and a matching audit phase. The first phase of the SOCOM framework is concerned with the rendition of an ontology labeled in the target natural language, particularly, appropriate translations of its labels. The second phase concerns the generation of matching results in a monolingual environment. The third phase of the framework aids ontology engineers in the process of establishing accurate and confident mapping results. Main Advantages and Shortcomings The principal advantage of this approach is the existence of a great quantity of ontologies that can be used to build a multilingual ontology. However, some problems arise, i.e. the difference in the hierarchical structure of the relevant ontologies, as well as differences in the semantics of the terms. While equivalency is sought between terms, this does not imply that the hierarchical structures themselves must also be equivalent. Therefore, developing a multilingual ontology using this approach involves the risk of getting an unmanageable entity as an outcome, in which great care is required to define relationships between “equivalent ontologies” and to track changes and to coherently update those relations [Bonino et al., 2004]. 2.3.3 Translation of Monolingual Ontologies This localization approach is adopted when an ontology has to be built in a certain target language (e.g., English or Spanish) and there is a monolingual ontology covering the domain of the proposed multilingual ontology. Note that, this approach is the main focus of this thesis. In this section we describe the main features of related works that we consider more relevant. But we first classify them according to the level of automatization and collaboration used to localize an ontology into different natural languages. Basically, the methods for translating an ontology can be grouped in four different categories: by hand, using a community, automatically and semi-automatically. 31 CHAPTER 2. TECHNOLOGICAL CONTEXT Translating an Ontology by Hand One of the benefits of translating an ontology by hand lies in the total control of how the translation is being done, and for small ontologies the amount of work is acceptable. There are however some disadvantages to this method: the translation will be a somewhat subjective one, since the translation is a product of one or few people’s skills, which is affected by their education, experience and temper. Furthermore, the amount of time and effort, and therefore expenses, to complete this task will be very high when translating larger ontologies. On the other hand, in cases of very specialized knowledge, this approach may not be very successful. Using a Community to Translate an Ontology A different approach to localize an ontology into different natural languages is to use a community to translate the ontology. This approach combines the advantages of human translations with speed. If the community is big enough, it would be possible to let them translate a complete ontology. Of course, spammers and people with other bad means as well as inconsistency must be ignored. Therefore a certain threshold could be introduced: a word must have been translated at least a certain number of times, after which the translation is approved automatically. Some works that adopt this approach are: AGROVOC: Caracciolo [Caracciolo et al., 2007], examines some of the issues associated with the development of AGROVOC, a multilingual thesaurus designed to cover the terminology of all subjects of interest to FAO (agriculture, forestry, fisheries, food and related domains such as environment). According to Caracciolo [Caracciolo et al., 2007], translations are provided by native speakers of the target language. Translations are typically made of the English version and sent to FAO for validation and inclusion in the master version. Apparently, terms are assigned a unique number e.g., “Abalone” is assigned to the number five in English. The translation of the word, e.g., in French “ormeau” is also given the number five in the French version. As a result, the result is not alphabetically ordered in each language, but multiple names are attached to a single concept across the languages. The number of terms in the vocabulary varies substantially by language. The differences can be the result of discrepancies in language, but also control over additions to the vocabularies in those languages. Information about the stability of such vocabularies is limited. Information about the impact of that stability on the use or expansion of these vocabularies is also limited. Furthermore, based on these different sizes of the vocabularies, either some vocabularies are under-specified or some are over-specified, or characteris32 2.3. METHODS FOR THE BUILDING OF MULTILINGUAL ONTOLOGIES tics of languages differentiate themselves from other language vocabularies for the same set of concepts. UK Data Archive Thesaurus: In [Balkan et al., 2002] the authors describe an approach for the building of a multilingual thesaurus for the social sciences. The thesaurus was produced by the UK Data Archive (UKDA) as part of the EU-funded LIMBER (Language Independent Metadata Browsing of European Resources) project and was derived from their in-house English monolingual thesaurus, HASSET (Humanities and Social Science Electronic Thesaurus). The multilingual thesaurus is available in four languages English, French, German and Spanish and in various formats, including RDF (Resource Description Framework). Basically, the construction of the thesaurus proceeded in two stages: first the monolingual thesaurus was reduced, and then the translation of the reduced thesaurus was carried out. In the first step reduction of monolingual thesaurus, different policies were adopted for the task of restructuring the UKDA HASSET for use across Europe. The reduction was made with the understanding that the rationale behind ELSST was to produce a common ontology which could be extended via local extensions to cater to the cultural and institutional needs of the individual archives and also allow for inclusion, via mappings, to specialized thesauri in certain subject areas. The translation process was carried out by a team of translators at the UKDA, who met on a regular basis to discuss problems as they arose. They provided feedback to those working on the monolingual thesaurus, so that changes to the monolingual thesaurus, such as the addition of scope notes, could be implemented where necessary. Verification of the translations was carried out by bilingual information experts at the appropriate CESSDA sites. Automatically Translation of an Ontology This approach uses different resources and automatic tools to ensure that information represented in an ontology using one particular natural language achieves the same level of knowledge expressivity if translated into another natural language. The main advantage of this approach is that it does not need human labor to discover the translations of an ontology. However, in the process of achieving this goal, some important challenges have to be addressed. The details on how these challenges are going to be reached and their implementation and evaluation will follow in the next chapters of this thesis. In the following paragraphs we briefly describe the main features of the most relevant works, ordered by similarity with our approach. We wish to point out that to the best of our knowledge works that support the localization of an ontology do not exist. The different approaches that can 33 CHAPTER 2. TECHNOLOGICAL CONTEXT be found in the literature to discover automatical translations of ontology terms offer only a partial solution to the problem: LabelTranslation: LabelTranslation is a strategy and a platform created for supporting the multilingual extension of ontologies existing in just one natural language. This tool was developed in order to support “the supervised translation of ontology labels” [Declerck et al., 2006] and, at the same time, to allow for the semantic annotation of multilingual web documents using the resulting multilingual labels of ontologies. By “supervised translation” is meant that this approach foresees the intervention of the domain expert or translator in case of a lack of results or for validation. Therefore, LabelTranslation offers a semi-automatic strategy. LabelTranslation can be integrated into any ontology engineering platform to enable its users to translate their ontologies inside the application. For the development of LabelTranslation already available multilingual semantic resources and basic natural language processing tools were reused for providing a semi-automatic translation of labels in ontologies. In the current version of the LabelTranslation platform three types of multilingual resources are included: i) EuroWordNet (EWN), a semantic lexical resource, ii) Wikipedia13, the multilingual free encyclopedia on the Web, based on knowledge of the word, and iii) BabelFish14, an on-line translation service used as “fallback position” [Declerck et al., 2006]. The steps for the translation approach are summarized: 1. Upload of an ontology in the LabelTranslator platform 2. Selection of the ontology labels to be translated in one of the target languages (en, es, de) 3. The system accesses the EWN database to find the selected term (or part of a term), and also checks in the WordNet database, only if the source language is English 4. Result(s) (synset and gloss) are displayed, if the matching is successful. Users can then validate the suggestions, modify the translation and save it in the database. A disambiguation problem can as well occur (see Disambiguation problem below) 5. If the matching in EWN is not successful, the system checks in Wikipedia, which also uses a mechanism for relating entries in the various available languages 6. If steps three and five do not provide any results, the system turns to BabelFish 13http://es.wikipedia.org/wiki/Wikipedia 14http://babelfish.altavista.com/ 34 2.3. METHODS FOR THE BUILDING OF MULTILINGUAL ONTOLOGIES 7. If the translation is still not satisfactory, the user can enter a translation, together with partofspeech information and a definition If the same translation session is repeated in the future, the system will return the translation already saved in its memory. Developers of LabelTranslation give priority to the EWN resource because a “high quality in the translation is expected since EWN has been built following semantic considerations and validated by language and/or domain experts” [Declerck et al., 2006]. In the translation step using EWN (step three), sometimes more than just one result (or synset) is returned, which could be the appropriate equivalent translation for the label in the ontology. Then, glosses offered by EWN can be of great help, since the system can use them for disambiguating. Two approaches -or a combination of bothcan be used, and these are the following (Note that LabelTranslator developers suggest the implementation of a hybrid approach combining both strategies): •Rule-based strategy: the terms in the gloss of the target language are also present in the ontology; source and target languages share the same or similar glosses. •Static strategy: based on two gloss-based similarity measure algorithms used in the Perl package WordNet::Similarity. In order to solve the disambiguation problem in Wikipedia (step 5.), the user can go to the Wikipedia encyclopedic articles and manually check that the content, context, etc. of a term match with the ontology content. Ontoling: OntoLing [Pazienza and Stellato, 2006] is a framework for a semi-automatic linguistic enrichment of ontologies. This framework was developed for “supporting manual annotation of ontological data with information from different, heterogeneous linguistic resources” [Pazienza and Stellato, 2006]. The latest version of OntoLing even helps the user with automatic suggestions through the exploitation of different linguistic resources. By exploiting existing bilingual resources, OntoLing helps in the development of multilingual ontologies, “in which different multilingual expressions coexist and share the same ontological knowledge” [Pazienza and Stellato, 2006]. In this sense, if ontologies are already available in one natural language, this tool helps in the process of ontology localization or, as has been defined by its developers, in the “multilingual enrichment process” [Pazienza and Stellato, 2006]. In the current version of OntoLing, two language resources are available for the linguistic or multilingual enrichment, WordNet15, for the linguistic enrichment of ontologies with English labels, and DICT dictionaries16, for 15http://wordnet.princeton.edu/perl/webwn 16http://www.dict.org/links.html 35 CHAPTER 2. TECHNOLOGICAL CONTEXT the linguistic and multilingual enrichment of ontologies. This last resource accesses a compendium of multiple on-line monolingual and bilingual dictionaries, as for example, all bilingual Freedict Dictionaries: English-German, English-Arabic, English-Croatian, English-Hungarian, etc. Since OntoLing has been developed as a plug-in for Prot´eg´e, the user has to upload an ontology in the Prot´eg´e ontology editor in order to use it. Any Prot´eg´e plug-in, exploiting linguistic resources, includes a linguistic watermark package, i.e., a package that contains abstract classes and interfaces for accessing linguistic resources. As already mentioned, the current package contains two implemented linguistic interfaces related to freely available resources, namely: WordNet and DICT dictionaries. Steps and techniques of this localizing tool are summarized in the following: •Open an ontology in the Ontology Panel of the Prot´eg´e editor •Select from the OntoLing menu of available linguistic resources those that will be visualized during the translation task •OntoLing accesses the selected linguistic resources by means of a wrapper called Linguistic Interface. With this Linguistic Interface the user visualizes the linguistic information in the Linguistic Browser Panel embedded in the Prot´eg´e framework. •The ontology can be enriched with: –Additional labels for the selected class, i.e., synonyms –Glosses as descriptions for the selected class –IDs of the selected senses as additional labels for the selected class. This is useful if pointers from ontology concepts to senses from a given linguistic resource are needed. •The user checks the suggestions offered by the linguistic enrichment module and selects the appropriate ones. •Selections are added to the ontology. Regarding the automatic linguistic enrichment of ontologies, this is currently under development. Moreover, this functionality only will be available if the ontology is in OWL (Web Ontology Language)17, and the loaded linguistic resource is a taxonomical lexical resource and/or a linguistic resource with glosses. The enrichment component will exploit the taxonomical structure of the glosses of the linguistic resource to judge which linguistic information can be used to enrich the ontology. 17http://www.w3.org/TR/owl-features/ 36 3.4. HYPOTHESES A3 We assume correctly spelled ontology labels, considering the syntactic rules of the source language. A4 The localization of ontologies can be performed by one ontology engineer or by a team of ontology engineers, translators, and linguists who may be geographically distributed. A5 The collaborative ontology localization within an organization usually follows a well defined process for the coordination of the translation activities. 3.4 Hypotheses Once the assumptions have been identified and presented, the set of hypothesis of our work are described. This set of hypothesis covers the main features of the proposed solutions and they will be validated through this thesis: H1 The characteristics of the ontology labels such as i) the similar lexical formats used to name terms (e.g., concepts as nouns) [McCrae et al., 2011a], ii) the low percentage of spelling errors [Espinoza et al., 2008a], iii) the significantly smaller size than a sentence [McCrae et al., 2011a], make these labels amenable to automatic translation. H2 The use of specific translation methods for localizing simple and compound ontology labels instead of a unified method, could improve the values of precision and recall. H3 The use of more than one resource into the translation process gives a wider range of translation candidates to choose from, and the correct translation is more likely to appear in multiple translation resources than in a single translation resource. H4 An appropriate combination of translation methods leads to better localization results than only using one at a time. H5 Localization methodologies in other areas are general enough to be taken as a starting point to develop an easy to use and understand ontology localization methodology. H6 It is possible to define a unified method to independently localize an ontology to different natural languages of i) the domain of the ontology, and ii) the process used to discover the translations of each ontology element. 43 CHAPTER 3. WORK OBJECTIVES H7 The collaborative process usually followed by organizations for the localization of ontological resources can be modeled by means of collaborative workflows. H8 The implemented infrastructure is usable with regard to the efficiency, effectiveness and users satisfaction. 3.5 Restrictions Finally, the following set of restrictions defines the limits of our contributions and allows the determination of future research objectives. Most of these restrictions are related to the technological aspects of the contribution (R1R3), while R4-R7 are related to the experimentation. R1 The proposed method and technology do not consider the optimization of the localization process of the generated system, neither in terms of the space required during the localization nor in terms of the time needed to complete the localization. R2 An ontology localization systems does not necessarily have to find a translation for each ontology label. The localization of all the labels of an ontology is normally a desirable feature, however in some cases the localization depends on the degree of the shareability of the conceptualizations. R3 The process of localization only considers translations of ontology labels and instances. Translation of annotations labels as rdfs:comments are not supported yet. R4 We do not include support for the argumentation of the selected translations. R5 We are only considering ontologies expressed in OWL as input of the ontology localization activity. R6 The LabelTranslator system proposed for the translation of ontology labels works only with the natural languages: English, Spanish and German. R7 The method for localizing ontologies covers the translation to one target language per time, but does not consider the translation of labels into different natural languages simultaneously. 44 Chapter 4 Ontology Localization Problem Open and dynamic systems, such as the Web and its extension, the Semantic Web, are by nature distributed and heterogeneous. Such characteristics implicate that the ontologies used to describe content and services can be represented using different formats and, more specifically, different natural languages. In this scenario, multilingual ontologies are required. As we described in section 2.3 there are two current trends for the building of multilingual ontologies, however these approaches do not reduce the cost and effort that comes with enriching an ontology with multilingual information. In this chapter we first explain the terminology related to ontology localization activity, providing the meaning that will be used for the distinct terms. Secondly, we describe the problems that characterize the ontology localization activity. Thirdly, we briefly analyze the different scales of localization used in ontologies depending on the type of ontology elements to be localized and the level of adaptation required to make the ontology accessible to speakers of different natural languages. Also, we analyze which elements or parts of the ontology are to undergo localization. Fourthly, we provide the foundations for the thesis, giving a formal account of our general localization process by defining the input, output, and the four main steps identified. 4.1 Definition of Terms In her doctoral work about Multilingualism in Ontologies, Montiel-Ponsoda [Montiel-Ponsoda, 2011a] identifies different definitions about ontology localization (see [Su´arez-Figueroa and G´omez-P´erez, 2008,Cimiano et al., 2010], in the sense of “the adaptation of the ontology and its natural language documentation to the needs of the target users”. She also shows that both Software Localization and Ontology Localization have a very pragmatical 45 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM and economical orientation, since the idea is to reuse software products or ontologies already available instead of developing them from scratch. Based on this premise, she arrived at the conclusion that in “Ontological Engineering, the localization of ontologies could be considered as a subtype of software localization in which the product is a shared model of a particular domain, i.e., an ontology, to be used by a certain application”. Despite the quantity and quality of definitions identified in the research works introduced above, most of the authors do not provide a guide on what really involves the ontology localization activity from a technical point of view. The definition that we propose in this thesis is intended to help understand the process for localizing automatically an ontology. We believe, that ontology localization cannot be fully or correctly understood without being contextualized in reference to a number of interdependent processes. From an Ontology Engineering perspective, these processes can be referred to as a group with the acronym ILT - Internationalization, Localization and Translation. In the following section we define in detail these three processes from the perspective of ontological engineering. 4.1.1 Ontology Localization Definition In this section we introduce the definition of the ontology localization activity exactly as we perceived it in this work. It does not pretend to solve each particular problem nor to strictly cover the complete field. It aims at serving as a guide for this thesis. Thus, rather than attempting to cover the entire spectrum of research in ontology localization, we will concentrate on process, methods and techniques to automatically discover the translations of the elements of an ontology. The automatic process of translating monolingual ontologies into other natural languages is the core of the localization activity and will be explained in detail in the remainder of this chapter. A classification of localization approaches focusing on the translation of the different elements of an ontology will be explained in Chapter 5. Additionally, in Chapter 6 we will give an intuitive view of the whole life cycle of the localization activity. Given one ontology O, ontology localization means that for each entity (concept, attribute, relation, or instance) expressed in a source natural language, we try to find a translation term, which has the same intended meaning, but in a different target natural language(s) L. There are some other parameters that can extend the definition of the localization activity, namely: (i) the localization parameters to accept a translation as suitable, e.g., weights, thresholds; and (ii) external linguistic and semantic resources used by the localization process to obtain the translations, e.g., text corpus, ready-to-use MT systems, or machine readable dictionaries. This can be schematically represented as illustrated in Figure 4.1. The definition is inspired in the work presented by [Euzenat and Shvaiko, 2007] 46 4.1. DEFINITION OF TERMS for ontology matching. The formalization of this process will be introduced in section 4.4.2. Figure 4.1: The Ontology Localization Activity For clarification we provide a short definition of ontology as used in our scenario. So far we have considered ontologies without being precise about their meaning. An ontology can be viewed as a set of assertions that are meant to model some particular domain. Usually, the ontology defines a vocabulary used by a particular application. Definition 1 (Core Ontology) A core ontology is a structure B := (C, ≼C, R, σ,≼R, I), consisting of, - three disjoint sets C,Rand Iwhose elements are called concept identifiers, relation identifiers and instances identifiers (or concepts, relations and instantces, for short). - a partial order ≼Con C called concept hierarchy or taxonomy. - a function σ: R →C×C called signature, where σ(r) =⟨dom(r),ran(r)⟩, where dom(r)and ran(r)are the domain and range of a relation r∈R. - a partial order ≼Ron R, called relation hierarchy . We will denote by OE, the union set, OE =C∪R∪I. An element oe ∈OE is called an ontology element. Relationships between concepts and/or relations as well as constraints can be expressed within a logical language such as first-order logic or Hornlogic. A formal definition of logical language has not been included at this stage of the research, because it is not relevant to the rest of the definitions. Definition 2 (Lexicon) Let L={l1, . . . , ln}be a set of natural languages, and Natli,1≤i≤n, be a set of strings in the language li. A lexicon Lex for a core ontology Bin a set of natural languages Lis a set of functions Lex ={Lexl1, . . . , Lexln} 47 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM Lexli:OE →Natli where OE are the ontology elements of the core ontology B. In this work, an ontology consists of a core ontology, as well as a corresponding lexicon. Definition 3 (Ontology) An Ontology O is therefore defined by the following tuple: O := (B, Lex), consisting of, - the core ontology B. - the lexicon Lex. We will say that O= (B, Lex)is a multilingual ontology if the set Lof natural languages associated to the lexicon Lex has more than one natural language. Once we have defined the concept of Ontology as used in our scenario, our aim is to define other processes related with the ontology localization activity. 4.1.2 Related Terms As we explained before ontology localization cannot be fully understood without being contextualized in reference to two interdependent processes: internationalization and translation. Internationalization. When an ontology is developed, its design is inevitably influenced by the culture and native language of their developers. To adapt an ontology successfully to international regions or markets, the culturally and linguisticallydepend parts of the ontology must be carefully designed, a process referred to as ontology internationalization. This process includes, for example naming conventions of ontology terms and/or hyphenation and morphological rules of the ontology elements. Thus, ontology internationalization can be defined as the process of generalizing an ontology so that it can handle multiple languages and cultural conventions without the need of re-designing it. Internationalization takes place at the design level. There are two key reasons for ontology internationalization: 48 4.1. DEFINITION OF TERMS 1. To ensure that an ontology is properly designed and therefore can be accepted in international markets, and 2. To ensure that an ontology is localizable. In the first case, the labels and descriptions used in the ontology are concise, clear, and they do not contain any jargon or slang. The second reason above mentioned will help to reduce the localization costs by developing the ontology in a way that ensures a smooth localization process. One way to do this is by following a standard for the naming of labels. Some works [Flied et al., 2007, Schober et al., 2007] have proposed naming conventions for ontology terms. These guides are used in specific applications such as ontology verbalization1. We claim that the definition and use of style guidelines should also be extended to ontology engineering. Translation. Translation can be generally defined as the process of “transferring a text from a source language and culture into a target language and culture with a certain purpose” (adapted from [Nord, 1997]). Considering this definition, we agree with [Montiel-Ponsoda, 2011a] in that the translation process may be considered the mother activity that encompasses Ontology Localization. Depending on localization purpose, the process of translation can be categorized in: instrumental and documental [Montiel-Ponsoda, 2011a]. In the first case, the goal of the target ontology can be to have the same function in the target community as the original ontology in the source ontology. The purpose of the translation can also be to “document” the ontology in another language to make it accessible to a community which speaks another language. In both cases and just as it occurs in software localization, in ontology localization the emphasis should be placed on automatic translation tools that allow users to avoid the manual effort of building a multilingual ontology. Some authors believe that this solution is not viable, since machine translation (MT) today suffers from several critical limitations to support the translation of ontology labels [Segev and Gal, 2008]. The general awareness is that automatic translation tools have yet to achieve a level of proficiency comparable to human translation. However, since ontologies consist of concepts, attributes and relations that are stated clearly and succinctly, we hypothesize that ontology components are more readily translatable than full-length text [Espinoza et al., 2008a]. When translating ordinary text, one has to deal with textual phenomena such as anaphora or metaphors, and much care must go into assuring that 1The verbalization makes ontologies accessible to people with no training in formal methods. The goal of the ontology verbalization is to produce natural language from the definition of class or properties. 49 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM one obtains clear and natural-sounding sentences. This is not such a big issue in ontology labels, which tend to have text with single words, compound words, named entities, short phrases, or short sentence fragments. Also, the ontology labels have characteristics, which make these amenable to MT: •Consistency. The lexical formats used for naming ontology terms are very similar [Espinoza et al., 2008a]. Also, the labels used for describing ontology elements commonly use a upper/lower case distinction. It poses some advantages to MT because it allows performing word segmentation. Some works have shown that having a basic word segmenter helps MT performance [Koehn and Knight, 2003,Habash and Sadat, 2006, Chang et al., 2008]. Additionally, we can rely on the initial uppercase letter to identify a phrase initial word. •Accuracy. The spelling accuracy of the labels of an ontology is reported to be approximately 97.0%-99.5% [Espinoza et al., 2008a]. These values are very important because the typographical errors can affect the translation quality. Furthermore, sentence boundaries (used in ontology term comments), which are absolutely crucial for parsing in MT, are usually clear in the ontologies through the use of accurate methods of punctuation. We define the ontology label translation task as finding, for an individual label lin the source language S, the correct translation, either a word or phrase, in the target language T. Clearly, there are cases where lis part of a multi-word term that needs to be translated as a unit. For this case, this approach can be extended by preprocessing the data in Sto find shortphrases, and then executing the entire algorithm treating short-phrases as atomic units. In this thesis, we do not explore the extension of this approach to the translation of sentences (e.g., comments of ontology term). Nevertheless, we focus on the translation of simple and compound labels. The technical details of this process will be explained in section 4.4.2. 4.2 Characterization of Ontology Localization Once the main concepts of the ontology localization activity are defined, we describe in this section the localization problems that must be taken into account when deciding on the method of localizing an ontology. The characterization of the localization problem in ontologies has been analyzed previously in [Montiel-Ponsoda, 2011a]. In this approach, the author provides a classification of different categorization relations (language equivalence problems in our work) that are shared among different cultures, no matter how different the linguistic structures are that express them in 50 4.2. CHARACTERIZATION OF ONTOLOGY LOCALIZATION each language. However, she does not envision problems as: the identification of translation mechanisms that preserve the semantics of the original ontology term, the management of changes in ontology terms and their translated labels, and the representation of the multilingual information. 4.2.1 Language Equivalence Problems Due to the fact that cultures classify the world in a different way, when translating ontologies we may encounter different types of situations: •Existence of an exact equivalence. This is typical of highly specialized technical and engineering fields such as Mechanics, in which there is a direct/complete equivalence among the terms in different languages referring to a certain object or process. In this case there is little place for synonyms or variants. E.g., ‘W¨armekraftmotor’ in German is translated as ‘heat engine’ in English. •Existence of several context-dependant equivalents. When one term in a language can be translated by several equivalents in another language, and the user has to choose the most suitable depending on the context of the ontology and the word connotations. For example, the English term ‘girl’ can be translated into Spanish as ‘ni˜na’, ‘chica’, ‘joven’, o ‘hija’. Each translation reflects different nuances of the concept and it will be necessary to find out which equivalent is needed in the ontology to translate the English term ‘girl’ with the most approximate or suitable equivalent. •Existence of a lexical gap. This is mainly due to mismatches at the conceptualization level, i.e., when a certain culture categorizes reality with a degree of granularity that does not correspond to the granularity degree of the other culture, resulting in a lexical gap in the target language. For example, in French there is a difference between big rivers that flow into the see, which are called ‘fleuves’, and rivers that flow into other rivers, ‘rivi`eres’. In most topography ontologies in English this distinction is not made. 4.2.2 Translation Problems In order for a translation algorithm to be useful in ontology localization, it has to produce reliable translations that exactly correspond to what a human translator would produce. Having a computer assistant that requires the translator to do significant corrections to its suggestions is nearly as good as having no assistant at all. For our purposes, the aim of this process is to suggest terms that are translations of the original concepts in the ontology. This particular task may be performed in at least two ways: 51 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM •The first option involves term translation from source language into target languages(s) followed by a cross-lingual retrieval of the senses of each translated term in the source language. Note that in this case the senses of the translations is in the same language of the term to be localized. To compute the similarity between the senses of the translations and the sense of the term under consideration a similarity measure is necessary. We identified these measures as language dependent because the compared terms need to be defined using the same natural language. •The second option involves term translation from source language into target languages(s) followed by a monolingual retrieval of the senses of each translated term in the target language. In this case, to be able to compare the terms using their contexts, we need to use a similarity measure that will allow cross-lingual comparison of words and their contexts. We identified these measures as language independent. In both cases, this process requires that the full meaning of each term be accurately rendered from a source language into target natural language, with special attention paid to cultural nuances. In the MT literature, some authors have investigated the improvement in the quality of MT approaches, incorporating context-rich approaches from word sense disambiguation (WSD) methods [Carpuat and Wu, 2007, Apidianaki, 2009]. However, to the best of our knowledge, the inclusion of a disambiguation method of ontological terms for improving the ontology localization activity not been tested yet. The formulation of the translation task as a word-sense disambiguation task has multiple advantages. First, if we knew the correct semantic meaning of each word in the source language, we could more accurately determine the appropriate words in the target language. Secondly, the availability of large amounts of resources from which we can infer the senses for disambiguating the words. 4.2.3 Management Problems Whereas the translation process of ontology labels per se implies certain difficulties, the maintenance and updating of translated ontology labels throughout the ontology life cycle also requires special attention. The main difficulty is to identify policies for managing changes in ontology terms and their translated labels. Up to now, none of the works on managing ontology changes [Palma et al., 2008, Tudorache et al., 2008] dealt with changes of ontology elements with multilingual information. Several situations could happen: 52 4.4. ONTOLOGY LOCALIZATION APPROACH 4.4.2 Automatic Localization Approach The localization activity is a complex task that involves different manual tasks (e.g., management, translation, revision, etc.). From these tasks we consider that the translation task definitively requires the most effort. For this reason, one of the main objectives of this work is to integrate an automatic translation process in the localization of ontologies. Identifying the Phases of the Translation Process We propose that the starting point in the design of a general translation process for the ontology localization activity should be guided by the observation of how the translation process is performed by a human expert. Thus, we first consider the nature of the “translation process” itself. Malmkjær [Malmkjær, 2000] points out that the translation process may be used to designate a variety of phenomena, from the cognitive processes activated during translating, both conscious and unconscious, to the more “physical” process which begins when a client contacts a translation bureau and ends when that person declares satisfaction with the product produced as the final result of the initial inquiry. In translation practice, of course, the cognitive aspects are expressed within the physical aspects. Few studies deal specifically with the identification and characterization of the phases or stages for modeling the human translation process (see [Starren and Thelen, 1988, Nord, 2005, Englund, 2005]). For our purposes, we adopt the approach presented in [Starren and Thelen, 1988] which is organized in four steps: 1) meaning discovering, 2) finding receptor(target) language equivalents, 3) checking the meaning of the receptor(target) language item, and 4) formulation of the final translation. Figure 4.4 illustrates the steps above described. Figure 4.4: Human translator steps. In order for a human translator to be able to discern among the different meanings a word may have (homonymys or polysemic words), (s)he needs to analyse the context. Depending on the context in which the word is used, a certain meaning will be selected, whereas the rest of potential meanings of the word will be discarded. This process is performed almost unconsciously in the translator’s mind if (s)he has a good command of the subject and the terminology used in it. For example, if the word to be translated is “bank”, 59 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM and the text is about finances, the translator will undoubtedly assign the meaning “financial institution” to the word “bank”. The next step in the translator’s mind is to look for possible equivalents of the word “bank” with the meaning of “financial institution” in the target language. Assuming the translator’s proficiency on the subject in both, the source and target culture, the translator will look for an equivalent concept in the target culture. If “bank” is going to be translated into Spanish, the translator has to find out if the English word is referring to a “savings bank” or to an “investment bank”, for example, since in the first case, “bank” would be translated into “caja” and in the second into “banco”. Here again the context is essential for the translator to make the right choice. At this stage, it is difficult to separate this action into two steps, finding language equivalents and checking their meaning) because concepts are represented by lexicalizations, and they come together as indivisible items in the translator’s mind. In order to take the final decision on which the most appropriate translation for a certain word is, the translator will have to take into account two additional aspects: 1) which is the concept in the target language that better matches the concept in the source language?, and 2) which is the purpose of the translation? Once the purpose and context have been checked again, the translator is able to select the most appropriate translation for the source word. In the following, we give a formal account of our general translation process by defining the input, output, and the four main steps identified.7 Defining the Automatic Translation Process Figure 4.5 illustrates input, output, and the four main steps of the general automatic localization approach. This detailed stepwise approach of ontology localization is novel and one core contribution. Notice that these steps cover only the translation task in the Ontology Localization Activity. The description of the life-cycle model for localization which involves tasks that extend far beyond the translation process itself will be introduced in Chapter 6. In the following sections, the individual steps will be explained in more detail. Input The input of the process is one or more ontologies, which need to be localized to different natural languages. If more than one ontology is taken, then each ontology is processed individually. Pre-known lexical or multilingual information of the ontology 7It should be noted here that all phases were re-labeled to describe their functionality in the localization activity. 60 4.4. ONTOLOGY LOCALIZATION APPROACH Figure 4.5: Automatic translation approach for the Ontology Localization Activity. terms to be localized may be very useful, giving to the localization algorithm good starting points for discovering the translations of other ontology elements. For example, it may be useful to know that the word “river” is more specific (e.g., as provided by the skos:narrower8property) than the English concept “watercourse” even if the former is not a literal translation of the latter. This lexical term may help to disambiguate the possible candidate translations. Also, if the same concept contains multilingual information indicating that a translation in French of “watercourse” is “cours d’eau” (e.g, as provided by the rdfs:label9property), then, it is possible to use this information as intermediate language of an indirect translation10 in other language. Localization Step Selection Before the localization of ontology elements can be initiated, it is necessary to choose which element actually to consider from the ontology. This step may choose to discover the translations of certain candidate ontology elements and ignore others (e.g., only localize ontology concepts and not ontology relations) Definition 4 (Localization Selection) Given an ontology O, we define 8The Simple Knowledge Organization System (SKOS) is a common data model for sharing and linking knowledge organization systems via the Semantic Web. 9This property allows to include the tag “@lang” to create multilingual labels in RDF. 10Indirect translation is translation into language C based on a translation into language B of a source text in language A [Landers, 2001]. 61 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM a localization selection.SelO, as a subset of the ontology elements OE of O SelO⊆OE To the best of our knowledge, there are no specific methods for selecting the space of candidates to be localized. We consider that the implementation of a specific selection method may depend on various factors, for example: the time and recourses available for performing the localization or the use of the ontology after localization. In [Prins and van den Broek, 2004] the authors propose a method to make a semi-automatic “intelligent” translation of only a part of the ontology. They use as strategy of selection to find out which concepts of the ontology contain the most instances and then translating the concepts one by one. The intuition of this approach is based on the fact that instances are the ontological terms which are used more in their particular cases (multilingual semantic search). To decide which branches are the most important, they use a script that visualizes the amount of concepts and instances in the ontology. This visualization allows the user to identify the parts that have an enormous amount of subclasses and none or only few instances, and then starting with the translation of the selected terms. In our work, we use a similar localization selection strategy, in which the user may choose to localize the complete ontology or only certain ontological elements. Context Extraction In order to translate an ontology element oe ∈OE from the ontology O, one must consider its context11. The context of an ontology term allows discerning among the different meanings that an ontology label (defined in the lexicon Lex of the ontology) may have. Notice that, the inclusion of this phase in our generic translation approach goes in the line of incorporating a word sense disambiguation method to improve the quality of the obtained translations. The clues for discovering the context of an ontology term can be found not only in the surrounding terms, but also in other terms semantically related to the terms under consideration. In other circumstances the clues to extract the context of a term are found in the textual descriptions and also in the practical interpretation of the term. In the rest of this thesis, we will not further distinguish between labels and words, which will be interchangeable. Definition 5 (Ontology Term Context) Let Uoe be a set of ontology elements sucht that oe ∈Uoe. The context of the ontology element oe,ctxoe 11Context is the environment in which a word is used, and context, viz. word usage, provides the only information we have for figuring out the meaning of a new or a polysemous word. 62 4.4. ONTOLOGY LOCALIZATION APPROACH is the set ctxoe =Uoe ∪Lex(Uoe) where Lex(Uoe) = {Lex(u)|u∈Uoe} The selected ontology elements in an ontology term context may vary according to the requirements of the used algorithm of translation. In any case, the context of the ontological term (concepts C, relations R, and instances I) needs to be extracted from intensional and extensional ontology definitions. The main goal of the context of an ontological term is to reduce the “noise” of the word translation, since a single source word can be translated into many words in any target language. For example, the concept term chair of the sample ontology can be translated from English to Spanish as the nouns: silla (seat), and c´atedra (professorship), or as the verb: presidir (take the chair). However, if we use as context the term professor, we can limit the number of obtained translations. We consider two different dimensions for modeling the context ctx of an ontology element oe: •Context interpretation. This dimension is concerned with the way to encode the context used by a particular translation technique. •Context size. This dimension makes reference to the number of elements used to define the context of a term. It is difficult to establish the minimum size that the context should have for determining the meaning of a term. The combination of these two dimensions provides a broad range of possibilities for the translation algorithms. In section 5.1.1 we categorize these dimensions and we show how these can be used to classify different translation techniques. An orthogonal dimension that needs to be considered is the context feature selection. This dimension aims to select the most relevant context features, removing the features least useful and thus improving efficiency or accuracy. There are several techniques to determine which words make up the context of a word: distance-based window, syntactic based-window, relatedness computation between words, etc. [Gamallo, 2007]. Some of these techniques have been applied to a word sense disambiguation domain; see for example [Mihalcea, 2002,Decadt et al., 2004,Gamallo, 2007,Gracia and Mena, 2009]. In our work we use a mechanism for the context selection, which is based on the relatedness computation between words [Espinoza et al., 2008a]. 63 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM Term Translation For a given ontology element oe from ontology O, this step tries to discover the more appropriate translation. The translation computation of an ontology element oe is done by using a wide range of translation functions. Each translation function is composed of the context of the ontology term, the target language(s) in which the ontology should be expressed, and the linguistic and semantic resources to obtain the translations. Definition 6 (Translation similarity function) Let oe be an ontology term, ctxoe be an ontology term context for oe, and lbe a natural language. A function tsl ctxoe :Natl→[0,1] is called translation similarity function. The translation similarity function will use a variety of linguistic and semantic resources to obtain the translations. Definition 7 (Localization ontology) Let O= (B, Lex)be an ontology, SelObe a selection of O,λ∈[0,1] a real number, and lbe a natural language. The localization ontology Ofor the selection SelOto the natural language l with threshold λis a ontology O′= (B′, Lex′) O′=tsλ O that it holds: •B=B′; •Lex′=Lex ∪ {Lexl}; •for all oe ∈SelO,tsl ctxoe (Lexl(oe)) > λ where tsl ctxoe is a given translation similarity function for any oe ∈SelOand given context ctxoe. It is denoted by tloe, the translation of oe, i.e. tloe =Lexl(oe). Different techniques can be used to perform the ontology localization task. A classification of these techniques will be extensively defined and explained in the next chapter. Evaluation In our approach we consider the translation task in a very specific setting of computer-assisted software localization. This setting imposes that the expected translations have a high quality. However, we recognize the need for an adequate procedure to evaluate and to guarantee the quality of the translation. Thus, from the obtained translations, we need to identify their quality. 64 4.4. ONTOLOGY LOCALIZATION APPROACH Definition 8 (translation evaluation) For each obtained translation of an ontology element oe, evaluation is defined as eval: tloe →[0,1] where, tloe is the result of the translation of a ontology element oe into a target natural language. Ideally, without any time or money constraints, translation output could be judged by humans to provide an idea of the system‘s performance. Obviously this is not the case when we need a fast way of evaluating translations. Goutte [Goutte, 2006] reviews a few automatic MT evaluation metrics from two different approaches12: string matching based and information retrieval (IR) based. All metrics presented below rely on a number of reference translations to which the translation output is compared. This does not mean that all words to be translated must have reference translations, only benchmark words. This does however mean that the performance measured automatically on that benchmark may not carry over to a different body of labels, especially in a different domain. String Matching Techniques. These metrics are based on the computation of the minimum edit (Levenshtein) distance . This identifies the minimum number of insertions, deletions and substitution necessary to transform one string into the other. Some metrics that use this approach are: •Word Error Rate (WER) is computed as the sum of insertions, substitutions and deletions, normalised by the length of the reference word. A WER of 0 means the translation is identical to the reference. One problem with WER is that this measure does not guaranteed a value between 0 and 1 and in some settings a wrong translation may yield a WER higher than 1. •WERg [Blatz et al., 2004], normalises the sum of insertions, substitutions and deletions by the length of the Levenshtein alignment path, i.e. insertions, substitutions, deletions and matches. The advantage of this metric is that it is guaranteed to lie between 0 and 1, where 1 is the worst case (no matches). •Position-independent Error Rate does not take into account the ordering of words in the matching operation. In fact it considers the translations and the reference as bag-of-words and computes the differences between them, normalized by the reference length. 12This section is a summary taken from [Goutte, 2006] 65 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM In fact, any string comparison technique may be used to derive similar translation evaluation metrics. One such example relies on the “string kernel”, and allows to take into account various levels of matching depending e.g., on the part-of-speech of the words, or to take into account synonymy relations [Cancedda and Yamada, 2005]. IR-style Techniques These metrics use measures inspired by Information Retrieval. In particular the n-gram precision is the proportion of n-grams from the translation that are also present in the reference. These may be calculated for several values of n and combined in various ways. •BLEU: This metric proposed by [Papineni et al., 2002] is the geometric mean of the n-gram precisions for 4 ≤n≥1, multiplied by an exponentially decaying length penalty. This penalty compensates for short, high precision translations such as “the”. •NIST: This metric was used in the MT evaluation rounds organised by NIST [Doddington, 2002]. NIST computes the arithmetic mean of the n-gram precisions, also with a length penalty. Another significant difference with BLEU is that n-gram precisions are weighted by the n-gram frequencies, to put more emphasis on the less frequent (and more informative) n-grams. •F-measure: The F-measure [Melamed et al., 2003] is the harmonic mean of the precision and recall. It relies on first finding a maximum matching between the translation output and the reference, which favors long consecutive (n-gram) matches. The precision and recall are then computed as the ratio of the total number of matching words in the maximum match over the length of the translation and reference, respectively. •Meteor: The Meteor evaluation system improves upon the F-measure in at least two ways. It uses some linguistic processing to match stemmed words in addition to exact matches, and it puts a lot more weight on the recall in the harmonic mean [Lavie et al., 2004]. BLEU and NIST are the metrics that are currently most widely used, and the ones all other MT evaluation metrics have to be compared with. The F-measure claims to provide higher correlation with human judgements [Melamed et al., 2003], but this is apparently not always the case, especially for smaller segments [Blatz et al., 2004]. Empirical evidence [Lavie et al., 2004] suggests that putting more emphasis on recall further improves the correlation. In fact it shows that recall alone often correlates best with human judgement, at odds with the exclusive use of precision in BLEU and NIST. 66 4.4. ONTOLOGY LOCALIZATION APPROACH Output As we explained before, we consider a multilingual ontology as the output of the ontology localization process. A multilingual ontology express the correspondences between entities belonging to the ontology to be localized and the multilingual terms pertaining to a natural language. A correspondence must consider the two corresponding entities (ontologies entities and multilingual terms) and the relation that is supposed to be held between them. In the following part we first provide the definition of multilingual ontology like it is used in our work. Then, we describe the other important component of the output of the localization, the relation that holds between the source entities with their translations. Definition 9 (multilingual ontology) A multilingual ontology O′= (B′, Lex′)is an association of ontology elements OE with a set of translation terms Tpertaining to a different natural languages L. O′:OE →TL. where, each ontology element oe ∈OE is labeled by a set of translation terms t1, t2, .., tn ∈Tin the language lof the lexicon Lexl. We denote O′(oe) = {t1, t2, ..., tn}. The multilingual ontology defines also the reciprocal relation SL:TL→OE by SL(t) = {oe ∈OE|t∈O′(oe)}. The next important component of a multilingual ontology is the relation that holds between the ontology elements oe and its translations T(see translation problems in the section 4.2). We consider that ontology localization algorithms should primarily use the equivalence relation (=) for expressing synonymy or equivalence relationship. However, according to guidelines for the establishment and development of multilingual thesauri [ISO, 1985] equivalence is divided into: exact equivalence, inexact equivalence, partial equivalence, single-to-multiple equivalence and non-equivalence. In the following part we briefly explain each case. In all examples, the letter X represents the source ontology element that needs to be localized and the letter Y its translation(s): •Exact equivalence (inter-language synonymy): the terms in X and Y are semantic and culturally equivalent. Table 4.1 shows a sample of equivalents terms in different languages. •Inexact or near equivalence (inter-language quasi-synonymy, with a difference in viewpoint): the terms in X and Y express the same general 67 CHAPTER 4. ONTOLOGY LOCALIZATION PROBLEM Table 4.1: Exact equivalent sample German English French Ducth Schienennetz Rail network R´esau ferroviaire Spoorwegnet concept but the meanings of the terms in X and Y are not exactly identical. Often the differences are more cultural than semantic, i.e. there is a difference in connotation or appreciation . In the case of inexact equivalence the terms can be treated as if they were exact equivalents. Table 4.2 shows a sample of inexact equivalence terms in different languages. The terms in Spanish and English are equivalent, however the term in French is only a near equivalence. Table 4.2: Near equivalence sample English Spanish French Historic settlements = Asentamientos hist´oricos ≈Site de peuplement •Partial equivalence (inter-language quasi-synonymy, with a difference in specificity): the term X in one of the languages has a slightly broader or narrower meaning than the preferred term Y in the other language. Table 4.3 shows a sample of partial equivalence terms in different languages. In this case, there are three possible solutions: i) treat the terms as exact equivalents., ii) adopt the terms from each language as loan terms in the other languages, and iii) treat the situation as single-to-many equivalence (see next case). Table 4.3: Partial equivalence sample German English Wissenschaft Science •One-to-many equivalence (too many or not enough terms): to express the meaning of the term X in one of the languages, two or more terms Y are needed in the other language. The issue of one-to-many equivalence can be solved by using “coined terms”. A coined term represents a concept new to the target language, which accepts the concept and constructs a new term in its language to express it. •Non-equivalence: no existing term Y with an equivalent meaning is available in the target language for a term X in the source language. Just like the previous case the solution is the “coined terms”. We believe that the equivalence relationships above described can be represented using relations from the ontology language. For instance, using 68 5.1. CLASSIFICATION OF TRANSLATION TECHNIQUES The overall classification of Figure 5.1 can be read both in descending (focusing on how the translation techniques interpret the context of each ontology term for disambiguating candidate translations) and ascending (focusing on the type of resources used for discovering candidate translations) manner in order to reach the Translation Techniques. 5.1.1 Term Context Interpretation The Term Context Interpretation classification is concerned with the way of modeling the term context used for disambiguate the candidate translations: •The first level is categorized depending on size or depth of the context: without context and with local context. These categories are assumed to be independent of the modalities used for encoding the context of each ontology term to be translated. a. The without context approach uses only the information related to the term itself as context. This option is sometimes disregarded, but it contains important information about the internal structure of the ontology term, e.g., term annotation (see rdsf:comment), or type of term (concept, relation, or instance). b. In the with local context approach the context involves a narrow group of terms centered on the ontology term itself, which fairly well approximates contexts starting from the immediately surrounding of direct relationship terms to the whole ontology. We wish to point out that the division of context into different sizes allows for the showing of their relative influence on translation techniques. One can argue by example that there are no distinct boundaries between local context that uses a small set of terms and a local context that uses many related terms. There are only more or less influential context features, whose general tendency is that their influence diminishes with increasing distance from the ontology term itself. •The second level of this classification decomposes these categories, taking into consideration the way of encoding the context of each ontology term to be translated. There are two different points of view for context pre-processing: linguistic and semantic. a. The linguistic encoding processes the context of an ontology term as linguistic objects. Basically, the linguistic encoding approach uses the information obtained from the lexicon of the ontology in order to generate the term context. 75 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION b. The semantic encoding processes the context as the entities that appear hierarchically organized in an ontological structure. In this approach of encoding, the context is obtained from the entities that are part of both lexicon and core ontology. In other words the semantic encoding makes use of all information of the ontology. •The third level of this classification particularize the categories above mentioned in seven groups of syntactic and semantic context knowledge: term description,term POS tagging,term list association,term description association,term verbalization, and structural context. The first two groups have as context the information of term itself. The rest of the categories use a local context, but with the difference that both term verbalization and structural groups use a semantic encoding; the other groups use a linguistic encoding approach. To illustrate the different ways of modeling the context of an ontological term, this section contains an ontology example of the university domain (see Figure 5.2). Concepts are depicted as rectangular boxes, relations as ellipses, annotation values as hexagons, and instances as rounded boxes. Ontology relations are drawn as solid arrows, whilst the instantiations of concepts and relations are depicted as dotted, arrowed lines. The example contains five concepts person,professor,full professor,associate professor, and faculty; one object relationship belongsTo; two attribute relationships hasFullName and hasName; and three instances Computer,Edu, and Asun. The example is fictitious and any concurrences with the real world are purely by chance. a. Term description. This category is represented by the use of a short description in the natural language of the ontology term under consideration. Usually these descriptions help clarify the meaning of the ontology terms. The rdfs:comment property can be used to define an ontology term description in the natural language (see RDF(S)2for more details). The term description context of the concept professor of our sample ontology can be: ctxprofessor := ( a professor is a member of the faculty ...) b. Term POS tagging. In this case the context is represented by the use of the grammatical category of the term. In order to obtain the grammatical information of a term, the Part-of-Speech (POS) [Church, 1988, DeRose, 1988, Garside, 1987] tagging is a natural option. POS tagging is the process of assigning a partof-speech like noun, verb, pronoun, preposition, adverb, adjective 2www.w3.org/TR/rdf-schema/ 76 5.1. CLASSIFICATION OF TRANSLATION TECHNIQUES Figure 5.2: Ontology Example. or other lexical class marker a word of a text. Most POS taggers3 need at least one short phrase from which it is possible to derive the lexical categories, or parts of speech of each word. We have identified that for the majority of ontology compound labels (e.g., AssociateProfessor) it is not necessary to have an additional processing to determine the part of speech of each token. However, for obtaining the POS of a single term (e.g., Professor) additional information is required, i.e., the relationship with adjacent and related words in a phrase, sentence, or paragraph. One way to solve this problem is to use empirical rules to annotate a simple term. Based on our experience, we propose the following rules: - The concepts, instances and attribute relations are considered nouns. - All the rest of the terms (e.g., object relations) are considered verbs. Another option is to try to generate a natural language sentence from the ontology term. In literature this process is known as 3A Part-Of-Speech Tagger (POS Tagger) is a piece of software that reads text in some language and assigns parts of speech to each word (and other token), such as noun, verb, adjective, etc. Available POS tagger tools can be consulted in the web page of the Stanford Natural Language Processing Group (http://wwwnlp.stanford.edu/links/statnlp.html#Taggers). 77 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION ontology verbalization. Some authors have studied this problem extensively (see [Hewlett et al., 2005, Flied et al., 2007, Schober et al., 2007] for example). This approach can be used for verbalizing the ontology term and then use a part-of-speech tagger to discover the POS of the term. The context of the concept term professor will be: ctxprofessor := ( NN ) Where, NN represents a singular noun. c. Term lists. This category is represented by the use of a bag-ofwords consisting of nwords adjacent to the target ontology term. The list of terms obtained is independent of the semantic relationship between adjacent terms. Thus, for instance the context to depth two of the ontology term professor can be: ctxprofessor := ( person,fullProfessor,associateProfessor,...) d. Term list descriptions. In this category, the descriptions (in the natural language) of surrounding terms in the context are expanded to include descriptions of the terms related to subsumption relations in ontology. The natural language descriptions of each term can be extracted from the rdfs:comment property. For the ontological term professor the context can be: ctxprofessor := ( a professor is a member of the faculty .... ; a person is a human, that has capacities or attributes ....) In the example, the second description belongs to the broader term Person. e. Term verbalization.4To model the context using this approach, it is necessary to transform an ontology term into a natural language sentence. As we commented previously recent works already have studied the way of generating natural language sentences from ontology elements. Intuitively, we can see that this option is an alternative to the approach previously described. Also, the ontology term verbalization has some advantages. In contrast to the term list description approach, where the descriptions not always define the exact meaning of the term, term verbalization reflects exactly the meaning of the ontological term. An example of the term verbalization context for the sample term professor is shown in the following: 4According to the Merriam Webster dictionary, one of the definitions of verbalization is to use words to express or communicate meaning. In this thesis we use this term in the same sense. 78 5.1. CLASSIFICATION OF TRANSLATION TECHNIQUES ctxprofessor := ( a Professor is a Person; a Professor belongs To Faculty; ....) g. Structural term context. The structural context is encoded exactly as the entities appear together in a ontological structure. Also, the structural context uses all logical relations represented in an ontology, such as equivalence, subsumption, disjoint, etc. 5.1.2 Type of Resources Used Following with the explanation of the categories used to classify the different translation techniques introduced in Figure 5.1, in this section we describe the classification of the type of resources used to perform a particular translation technique. The classification is categorized depending on the richness of the internal structure of the resource: linguistic, lexical and terminological and semantic resources. The first type of resources groups together similar words without much distinction in the kind of similarity relation (e.g., linguistic databases, dictionaries, thesauri, Web, etc). The semantic resources on the other hand group together objects denoted by words (or more complex lexical items) according to a principled set of paradigmatic (meta-)relations like synonymy, hyponymy, meronymy, antonymy and syntagmatic (meta-)relations according to the dependency structure (e.g., lexical databases and ontologies). Notice that both types of resources can be used during the search and disambiguation of translation candidates. In the following part we will give a brief overview of these resources (for more details, cf. [Ide and Veronis, 1998, Litkowski, 2005, Agirre and Stevenson, 2006]). Our purpose is not to analyze and compare the existing definitions of these resources, but to justify the convenience of their reuse in the ontology localization activity. Whenever possible, we will be referring to multilingual and online resources.5 Linguistic, lexical and terminological resources The main linguistic, lexical and semantic resources involved in the classification of the translation techniques are the following: a. Corpora. According to [McEnery, 2003] corpora are defined as large collections of general or subject specific documents. Nowadays, corpus primarily means a collection of texts held in electronic form, capable of being analyzed automatically or semi-automatically rather than manually and for different purposes. Elaborate typologies of corpora have been proposed in literature ( [Baker, 1995,Laviosa, 1997]) taking into 5The definitions and justifications of the resources here described, are a short summary of the deliverable found in [Espinoza et al., 2010]. 79 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION consideration aspects such as: the relationship of translations between the different language sections of the corpus, and the number of languages represented in the corpus. According to this, the two main types of corpus are: •Parallel corpora can be defined as corpora that contain source texts and their translations. Parallel corpora can be bilingual or multilingual. They can be uni-directional (e.g., from English into Chinese or from Chinese into English alone), bi-directional (e.g., containing both English source texts with their Chinese translations as well as Chinese source texts with their English translations), or multi-directional (e.g., the same piece of writing with English, French and German versions). These resources can be better exploited by Translation Memory tools, which align translation equivalents. Translation memory is a technology that enables the user to store translated phrases or sentences in a special database for local reuse or shared use over a network [Esselink, 2000]. •Comparable corpora, in contrast, can be defined as corpora containing sets of texts that are collected using the same sampling frame and similar balance and representativeness [McEnery, 2003], e.g., similar features such as the same proportions of the texts of the same genres, in the same domains, in a range of different languages, in the same sampling period. However, the texts of a comparable corpus are not translations of each other. Rather, their comparability lies in their same sampling frame and similar balance. The Web as corpus offers a valuable resource for building and contrasting comparable corpora on the same domain. With the enormous growth of the Information Society, the Web has turned into a reliable test bed of data for natural language processing, not only in terms of data size but also in terms of data type (e.g., multilingual data, link data). b. Glossaries. These resources can be defined as alphabetical lists of terms or words found in or related to a specific topic or text. It may or may not include explanations, and its vocabulary may be monolingual, bilingual or multilingual [Wright and Budin, 1997]. These resources are of interest in the ontology localization activity because they usually contain the specific terminology of a domain. They can be monolingual or multilingual. In the case of monolingual glossaries, the most useful information they provide are definitions of terms, which can be used as contextual information for disambiguation purposes. If they are bilingual, they normally contain lists of translation pairs, which can be used in the translation process and need to be further disambiguated. 80 5.1. CLASSIFICATION OF TRANSLATION TECHNIQUES c. Dictionaries. The dictionaries are, according to [Var´o and Linares, 1997], “books in which lexemes of a language are gathered and explained in the form of headwords or lemmas following an alphabetical order”. A machine-readable dictionary (MRD) is a dictionary in an electronic form that can be loaded in a database and can be queried via application software. It may be a single language explanatory dictionary or a multi-language dictionary to support translations between two or more languages, or a combination of both. MRDs are considered a valuable source of information for use in Natural Language Processing (NLP) because they contain an enormous amount of lexical knowledge. Some examples of MRDs that could be used in ontology localization are WordReference6, Wiktionary7, the Merriam-Webster’s Online Dictionary8or Leo9. d. Encyclopedias. These resources are defined as documents that contain information on all branches of knowledge or treat comprehensively a particular branch of knowledge usually in articles arranged alphabetically often by subject (Glossary of Library Terms). Nowadays, one of the best-known online encyclopedias is Wikipedia. Wikipedia10 defines itself as a free, web-based, collaborative, multilingual encyclopedia project supported by the non-profit Wikimedia Foundation. Others resources of this type are DBpedia11, which is a community effort to extract structured information from Wikipedia and make this information available on the Web or Freebase12, which is a large collaborative knowledge base consisting of metadata composed mainly by its community members. Apart from the detailed information that can be found about a certain article, these resources are interesting for ontology localization because of two major reasons: •They offer a huge source of structured information in different domains. •Most of the articles are multilingual and provide comparable corpora for translational purposes e. Terminological Databases. These resources are databases that contain the specific terminology of one or several domains of knowledge. They are similar to glossaries but usually contain additional data regarding the source from which the data has been obtained, language 6http://www.wordreference.com 7http://es.wiktionary.org 8http://www.merriam-webster.com/ 9http://dict.leo.org/ 10http://en.wikipedia.org/wiki/Wikipedia 11http://dbpedia.org/About 12http://www.freebase.com/ 81 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION usage examples, synonyms and related terms, etc. For the purpose of ontology localization, the multilingual terminology databases are interesting resources. Most international organizations maintain terminology databases to support the writing of technical documentation and its translation, as well as the communication between specialists. An example of a multilingual terminology database is IATE13 , created and maintained by the European Union (EU). f. Lexicons. In a restricted sense, a computational lexicon is considered as a list of words or lexemes hierarchically organized and normally accompanied by meaning and linguistic behaviour information [Hirst, 2003]. One of the best known online English lexicon is WordNet. In addition to this, the EuroWordNet14 lexicon draws on WordNet structure to create wordnets in other languages and link them through a so-called Interlingual index, a list of unstructured meanings that provide the mappings across the wordnets. This kind of resources is very useful because its structure helps in disambiguating the different senses associated to words. In the case of EuroWordNet, it also provides translation candidates. The major drawback is that such resources contain general-purpose lexical entries, although in recent projects drawing on WordNet, wordnets containing the specific terminology of a domain are being developed (see the KYOTO15 project). g. Thesauri. Thesauri are controlled vocabularies of terms in a particular domain with hierarchical, associative, and equivalence relations between terms. Thesauri are mainly used for indexing and retrieving articles in large databases [ISO, 1986]. More specifically in the computer science domain, a thesaurus is defined as “a controlled and dynamic documentary language containing semantically and generically related terms”, which comprehensively covers a specific domain of knowledge. Two well-known multilingual thesauruses are Agrovoc16 and EuroVoc17. Both AGROVOC and EuroVoc have been migrated to semantic web technologies making use of the SKOS (Simple Knowledge Organization System) language. In this way, thesauri are easily queriable in the Web. 13http://iate.europa.eu 14http://www.illc.uva.nl/EuroWordNet/ 15http://xmlgroup.iit.cnr.it/kyoto/ 16The multilingual thesaurus of the Food and Agriculture Organization of the United Nations http://aims.fao.org/website/AGROVOC-Thesaurus/sub 17http://eurovoc.europa.eu/drupal/ 82 5.1. CLASSIFICATION OF TRANSLATION TECHNIQUES Semantic resources The semantic resources that can be used to perform a particular translation technique are the following: a. Taxonomies. These resources comprise an organized list of concepts that are drawn from diverse data sources and organized according to an expert in the domain [Boiko, 2005]. Taxonomies are very common in the biological domain, but we also find many about economical or industrial activities, occupation, etc. See for instance, the Standard Industrial Classification18 (SIC) of the United States, ESCO19 taxonomy for employment in Europe. These classifications can be useful resources to ontology localization because they contain the specific terminology of a certain domain, and with a certain degree of structure. b. SKOS. The Simple Knowledge Organization Systems (SKOS) [Miles et al., 2005] is a W3C recommendation designed for representation of thesauri, classification schemes, taxonomies, subject-heading systems, or any other type of structured controlled vocabulary. Using SKOS, each term in each taxonomy can be represented in a machine readable format containing definitions, labels, and related concepts for the term expressed in SKOS. The SKOS framework allows associating labels and definitions in multiple languages to any concept. This means that we can associate the labels “Dispositivos m´oviles”@es, “appareils mobiles”@fr or “Mobile Ger¨ate to the concept “Mobile device” to include the Spanish, French and German labels. Well-known controlled vocabularies such as EuroVoc have been expressed using an ontology that extends SKOS. All data objects supported by SKOS for handling labels, can be useful to concept identification, disambiguation and translation in ontology localization. c. Linked Data. The term Linked Data refers to a set of best practices for publishing and connecting structured data on the Web [Bizer et al., 2009]. The main objective of Linked Data initiative is connecting data from diverse domains to enable new types of applications. Thanks to the links created between the data, these data can be browsed and queried starting in one data source and navigating along the links to other related data sources. This potentially augments the possibilities of obtaining relevant data. The Linked Data can contain information of many domains of knowledge, being the most represented nowadays: media, geography, publications, e-Government, and live sciences. There are also resources that contain general information, such as DBPedia. 18http://www.sec.gov/info/edgar/siccodes.htm 19http://esco.tenforce.com/esco-browser/ 83 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION These resources in the Linked Data format, and specifically DBPedia, may have an enormous potential for the ontology localization activity. They do not only offer structured information of a certain domain of knowledge, but also numerous links to related information. Those typed links are very useful in the disambiguation process. Although most of the resources on the Linked Data cloud are monolingual in English, in the near future we expect many of them to be in other languages. d. Ontologies. Formally, an ontology consists of terms, their definitions, and axioms relating them [Gruber, 1995]; these resources can be viewed as computational knowledge organization systems for domain specific text and domain specific knowledge. The ontologies have become crucial instruments of knowledge management processes, since they provide a formalized, hence conceptualization of a specific knowledge area that is usually contained in domain specific text corpora. In localization, and particularly in machine translation, ontologies have been used to improve the performance of translation systems [Hovy et al., 2001] by enhancing the knowledge base that supports the linguistic algorithms of source language text analysis and target language text generation. The way of accessing ontologies which are available on the Web is by means of Semantic Web search engines, such as Watson20, or Swoogle21. This allows us to see how a certain concept has been described (by means of properties and relations) in a certain ontology. The ontology localization activity would greatly benefit from the availability of multilingual ontologies to obtain translation candidates. However, multilingual ontologies are still scarce on the Web, and mechanisms should be developed to access and query multilingual ontologies. Once we have presented a brief description of the types of resources that can provide valuable information when localizing an ontology, we now discuss in more detail the main classes of Translation Techniques according to the above classification. 5.2 Basic Translation Techniques In this section, we introduce the main characteristics of the different translation techniques and methods shown in Figure 5.1 (see middle layer). To facilitate the analysis of these techniques we have designed a framework which covers their main aspects: 20http://watson.kmi.open.ac.uk/ 21http://swoogle.umbc.edu/ 84 5.5. CORPUS-BASED TECHNIQUES is that a given input phrase in the source language is compared with the example translations in the given bilingual parallel text to find the closest matching examples that can be used in the translation of that input phrase. One of the main approaches in the EBMT paradigm is to use pattern matching techniques. First, these approaches collect word sequences from each corpus using translation patterns to acquire candidates for bilingual expressions. Second, a search for pairs of words that satisfy the correspondences of the sequences is performed. Therefore, a pre-processing step such as part of speech tagging and syntactic category identification is necessary to apply this method. •Resource pre-processing: Before discovering candidate translations, a bilingual template acquisition from a simple monolingual corpus or parallel corpora has to be completed. In general, this process involves three phases: retrieving local patterns, assigning their syntactic categories with part-of-speech (POS) templates, and making translation patterns. –Retrieving local patterns. In order to retrieve local patterns any method for retrieving word sequences may be used [Kansai et al., 1996,Sato and Saito, 2002]. These methods generate all n-character (or n-word) strings appearing in a text and filters out fragmental strings with the distribution of words adjacent to the strings. This is based on the idea that adjacent words are widely distributed if the string is meaningful, and are localized if the string is a substring of a meaningful string. –Identifying syntactic categories. Since the strings are just word sequences, this task gives them syntactic categories. Thus, this task involves the assignation of part-of-speech tags for each component word discovered in the previous step. A syntactic category can be used to group similar tagged words. For example, the syntactic category NN can be used to group the following sample POS templates, (word) (word) or (word) (preposition) (word). In the example NN represent a noun phrase. –Making translation patterns. The final process is to generate the bilingual translation patterns. In the case of using a monolingual corpus as base to discover the patterns, we need to translate each word (identified in step one) as previous step to identify its syntactic categories. The output of this process is a repository of lexical templates for MT. •Translation candidate extraction: The term POS tagging context could be valuable to filter and prune the extracted candidate translations. 91 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION To retrieve candidate translations, we can collect the n-grams of POSs appearing in a translation pattern (e.g., NN, JN, etc.) from each corpus. As this method simply extracts word sequences according to POS tags, it also collects noisy sequences. However, most meaningless sequences can be eliminated, estimating different types of word similarity correspondences. •Translation selection: After generating a ranked list of translation candidates for each source term, ranking techniques must be used to estimate the coherence of the translated label and decide the best translation. The ranking factor can be estimated using one of the techniques described below: –Ranking through Web. The Web can be considered as an exemplar linguistic resource for decision-making [Grefenstette, 1999,Li et al., 2003]. In this approach, each candidate translation is sent to a Web search engine (e.g., Google) to discover how often the combination of translation alternatives appears. The number of retrieved Web pages in which the translated sequence occurred is used to rank the translation candidates. –Ranking through a test collection. Large-scale test collections could be used to rank the translation alternatives and complete a final translation. We can follow the same steps as the previous technique, replacing the Web by a test collection and a retrieval system to index documents of the test collection. –Ranking through an interactive mode. An interactive mode [Ogden and Davis, 2000] could help solve the problem of identifying final translations. The interactive environment setting should optimize the label translation, select best translation alternatives and facilitate the information access across languages. For instance, the user can access a list of all possible candidates ranked in a form of hierarchy on the basis of word ranks associated to each translation alternative. Statistical-based MT These approaches analyze large collections of texts on a statistical basis and automatically extract the most probable translations in the target language [Peters and Sheridan, 2000]. The recent progress in SMT suggests interesting future development for ontology localization [Stroppa et al., 2007, Gimpel and Smith, 2008]. In particular, phrase-based translation approaches have become the state of the art in SMT, while these approaches have not yet been widely investigated in localization. A recent work [McCrae 92 5.5. CORPUS-BASED TECHNIQUES et al., 2011a] analyzes different translation strategies using statistical machine translation approaches that also utilize the semantic information beyond the label or term describing the concept, that is relations among the concepts in the ontology, as well as the attributes or properties that describe concepts: •Resource pre-processing: Bilingual word/phrase alignment is the first step of most current approaches to SMT. Alignment is a vital issue in the construction and exploitation of parallel corpora. The alignment methodology tries to identify translation equivalence between sentences, words and phrases within sentences. In most literature, alignment methods are either categorized as association or estimation approaches (heuristic and statistical models). Association approaches use string similarity measures, word order heuristics, or co-occurrence measures (e.g., mutual information scores). The major distinction between statistical and heuristic approaches are that statistical approaches are based on well-substantiated probabilistic models while heuristic ones are not. Most current SMT systems use a generative model for word alignment such as the one implemented in the freely available tool GIZA++ [Och and Ney, 2003]. GIZA++ is an implementation of the IBM alignment models [Brown et al., 1993]. These models treat word alignment as a hidden process, and maximize the probability of the observed (e, f) sentence pairs using the Expectation Maximization (EM) algorithm, where e and f are the source and the target sentences. •Translation candidate extraction: To discover candidate translations, SMT-based methods generally use the occurrence frequencies of substrings of the sentence in target-language corpora. The score assigned to each candidate translation depends on both: i) the extent to which the source sentence meaning is also expressed in the candidate translation,and ii) the extent to which the candidate translation is likely to be a valid sentence in the target language regardless of whether or not its meaning bears any relationship to the source sentence. Details of how this score is computed is out of the scope of this thesis, however this information can be consulted in [Hearne and Way, 2011]. •Translation selection: In order to discover the final translations, the approach introduced in [McCrae et al., 2011a] uses word sense disambiguation by comparing the structure of the input ontology to that of an already translated reference ontology. We found this method to be very effective in choosing the best translations. However it is dependent on the existence of a multilingual resource that already has such terms. As such, we view the topic of taxonomy and ontology translation as an interesting sub-problem of machine translation and believe 93 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION there is still much fruitful work to be done to obtain a system that can correctly leverage the semantics present in these data structures in a way that improves translation quality. Translation Memory tools The essential idea behind these techniques is the use of a linguistic database (also called translation memory) in order to reuse previously translated words. These techniques are often used in order to compare segments in the source text with the translated segments in the translation memory. For our purposes, a segment can consist of simple ontology labels, compound labels, or term annotation paragraphs. •Translation candidate extraction:. Linguistic databases provide a number of efficient search options to extract candidate translations: –Fuzzy matching. This is the dominating approach for the retrieval of similar segments from translation memories, because the possibility of exactly repeated segments is small, except in the context of re-translating the labels of a modified resource (in our case an ontology). The method can be based on orthographic similarities, which can be efficiently computed by comparing the number of corresponding substrings (e.g., bior trigrams) of two segments [Willett and Angell, 1983,Rapp, 1997]. Another option to measure the distance between two fuzzy matching content segments is to use the Levenhstein algorithm [Levenshtein, 1965]. The Levenshtein distance between two strings is given by the minimum number of operations needed to transform one string into the other, where an operation is an insertion, deletion, or substitution of a single character. –Syntax trees. This approach requires natural language parsers for both languages to be considered. The parse tree of the segment to be translated is compared to the parse trees of all source language sentences in the linguistic database. If an identical parse tree is found, it is assumed that the parse tree of the correct translation should be identical to the parse tree of the corresponding target language sentence retrieved from the linguistic database [Maruyama, 1992]. The main problem with this approach is that high quality parsers for unrestricted languages are not available for many languages. Also, the disambiguation of semantically ambiguous words is not always possible by only considering the syntax. •Translation selection: Basically, the process of selection is a manual labor, in which the user performs the dominant role and makes the 94 5.5. CORPUS-BASED TECHNIQUES final decisions concerning the chosen translations. A solution to this problem is to describe the linguistic database data with the Translation Memory eXchange format27 (TMX). TMX is an open standard that uses XML for the archiving and mutual exchange of the Translation Memories (TM). In TMX a translation unit28 can contain markup content elements, which can be used to disambiguate the candidate translations. For example, using the term POS tagggig context, we can to select only those annotated translations whose POS match exactly with the POS of the searched term. This assumption is accomplished in the majority of languages. Web-based Particularly for domains where sufficiently large text corpora are not available, or accuracy and coverage of translation dictionaries are rather low, Web-based translation methods are a good alternative. These models propose to mine translations from Web corpora specially for discovering OOV term translations. OOV terms principally consist of short phrases such as named entities (person, location or organization names), book and movie titles, science, medical or military terms and others29. Therefore, we consider that the Web-based methods can be used to discover the translations of instances and very specific domain terms. Anchor-Text Mining Method This approach searches the Web for parallel text and extracts translation pairs among anchor texts pointing together to the same webpage [Lu et al., 2002]. An anchor text is the descriptive part of an out-link of a Web page used to provide a brief description of the linked Web page. For instance, the text part “Apple” is the anchor text in the example below. <a href="http://en.wikipedia.com/wiki/Apple_Computer">Apple</a> This method supposes that for a source term appearing in the anchor text of a Web page, it is likely that its corresponding target translations may appear together in other anchor texts linking to the same page. •Resource preprocessing: As a first step, this method needs to extract the Web pages whose anchor−text sets contain both source and target terms. In order to collect large numbers of pages from the Web and 27http://www.lisa.org/fileadmin/standards/tmx1.4/tmx.htm 28In TMX, an entry consisting of aligned segments of text in two or more languages is called a Translation Unit. 29Some names are single word, which could be regarded as one-word phrases. 95 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION build up a corpus of anchor-text sets, a Web crawler30 needs to be implemented. •Translation candidate extraction: Considering that an anchor text might be a short text, heading, phrase, or URL, the term extraction process needs to extract key terms as translation candidates from the anchor-text corpus. Different methods can be used to extract translation candidates: –PAT-tree-based: The PAT-tree-based keyword extraction method is an efficient statistics-based approach that includes n-gram modeling, and completeness and significance analysis of semantics [Chien, 1997]. The advantage of this method is the ability to extract many significant terms and phrases without the limitations of string length and that of using a dictionary. –Query-set-based: This method takes user queries from real-world search engines as vocabulary sets to segment key terms in anchortext sets. All the query terms in the target language are taken as translation candidates and their similarity to the source query is estimated. –Tagger-based: This method uses a tagger system, to segment the texts into meaningful words and to extract unknown words such as proper nouns and new terms. This method is different from the PAT-tree-based method in that it is more linguistically-based. •Translation selection: The process of selection assumes that a translation candidate has a higher chance of being a translation only if it frequently co-occurred with the source term in the same anchor text sets. To estimate the degree of similarity between a source term and each translation candidate that co-occurs in the same anchor-text sets, any symmetric similarity measure can be used. In literature, different works use a function based on the probabilistic inference model [Wong and Yao, 1995] for these purposes. Search-result mining. These methods are based on the observation that for many source language search-result pages, there are rich snippets of summaries with a mixture of source and target texts. Given an input term in a source language, the search engine searches the translation terms in documents written in other 30According to Wikipedia, a Web crawler is a computer program that browses the Web in a methodical, automated manner or in an orderly fashion. Other terms for Web crawlers are ants, automatic indexers, bots, Web spiders, Web robots, or -especially in the FOAF community-Web scutters. 96 5.5. CORPUS-BASED TECHNIQUES languages. The returned snippets containing the term are collected and translations are extracted from the snippets. Although a quite large amount of term translations can be acquired using a search snippet-based mining scheme, the scheme may fail to extract low frequency term translations. If a term translation pair occurs only a few times on the Web, the translation of the term may not be retrieved by the search engine since the search engine ranks Web pages based on the PageRank algorithm which is irrelevant to the occurrence of its translation. As a result the top-n returned snippets may not contain the translation. •Resource-preprocessing: The collection process of Web pages is performed using a Web search query. Basically, there are two approaches for building the query: i) using a monolingual query for source language pages containing the target language terms [Cheng et al., 2004], or ii) using cross-lingual query expansion [Zhang et al., 2005]. In the last approach to search for pages containing the term to be translated and its translation, the Web search query contains the term and one hint word generated by cross-lingual query expansion. •Translation candidate extraction: The terminology translation mining performs a preprocessing on Web snippet texts by filtering out HTML tags, punctuation marks and non-query source words. Then, it extracts the translation from the processed top-N snippets, and provides confidence scores for each translation candidates. •Translation selection: In order to select the translations these methods rely on different term similarity estimation techniques. Different measures have been proposed in literature for estimating the association between words/phrases based on co-occurrence analysis, including mutual information, the DICE coefficient, and statistical tests, such as the chi-square test and the log-likelihood ratio test. 5.5.2 Advantages Translation corpora are an ideal resource for establishing equivalence between languages since they convey the same semantic content. Also, these techniques can be quickly adapted to new language pairs since the algorithms are almost language independent and most language specific information is automatically derived from parallel corpora. Finally, most of these methods can achieve high translation accuracy. 5.5.3 Disadvantages While this method alleviates the problem of limited scalability found in the previous approaches, it relies on the existence of a parallel corpus in the 97 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION desired domain, which is often an unreasonable requirement. It is not always possible to find corpora of different languages and domains, together with the fact that corpus annotation requires a lot of effort and resources. In case of Web-based translation methods, there are some issues that need to be solved before using the Web information to mine terminology translation: i) how to find more comprehensive results, i.e. mining all possible forms of annotation pairs in the Web, and iii) how to remove the noises formed in the statistics and rank the remaining candidates. 5.6 Semantic-based Techniques The semantic-based techniques take advantage of linked data and ontologies resources that provide a formal description of concepts, terms, and relationships within a given knowledge domain. In this work we have mainly used these techniques to disambiguate the identified candidate translations from the other approaches. 5.6.1 Methods Employed Our intuition in approaching the ontology translation is that the comparison of ontology or taxonomy structures containing source and target labels may help in the disambiguation process of translation candidates [McCrae et al., 2011a]. A prerequisite in this sense is the availability of equivalent (or similar) ontology structures to be compared; we briefly summarize the main steps of our approach, named ontology structure comparison. •Translation candidate extraction: As a first step, this method tries to discover the different senses of a term, using for this purpose semantic descriptions available in different sources of knowledge. The semantic knowledge can be obtained from available online ontologies accessed by means of ontology search engines. These tools crawl the Web to obtain different types of semantic information such as ontologies, instance data, and specific terms i.e., URIs that have been defined as classes and properties. Some ontology search engines require a normalization process before each word is submitted as a query. The normalization process involves rewriting the words in lower-case, removing hyphens, etc. After normalization, the ontology search engines return different ontological terms that match those normalized keywords. The main advantage of using a pool of ontologies instead of just a single one is that many technical or subject-specific senses of a term cannot be found in just one ontology. For each term obtained from the ontology search engines, a sense is built. Each sense is represented by means of the hierarchical graph of hypernyms and hyponyms of synonym terms found in 98 5.7. ANALYSIS OF TRANSLATION TECHNIQUES one or more ontologies. Thus, senses are built with the information retrieved from matching terms in the ontology pool. Notice that the more ontologies or knowledge bases accessed the more chances to find the semantics of a term. As matching terms could be ontology concepts, attributes or instances, three lists of candidate keyword senses are associated with each normalized keyword: concepts, attributes and instances. The result of this process is a list of possible senses for each word. •Translation selection: This step uses the structural context information for ranking the different senses obtained in the previous step according to the similarity with its lexical and semantic context. To estimate the probability of synonymy, in other words the degree in which the words are related, any semantic relatedness measure can be used. These measures consider not only similarity between the words, but any possible semantic relationship between them [Gracia and Mena, 2009]. To avoid the use of a cross-language semantic measure as the source and target senses are expressed in different natural languages, the external ontologies can be limited to those resources that have linguistic information in other languages. 5.6.2 Advantages The main advantage of this approach is the increased availability of online semantic resources. Making the best use of such resources leads to a higher quality translation with lower development costs. 5.6.3 Disadvantages While these techniques allow for the obtaining of more exact translations, the lack of ontologies enriched with linguistic information into different natural languages implicates the use of cross-lingual semantic disambiguation measures. As these measures generally use the Web as multilingual corpus to establish the similarity, the translation process can be very slow. 5.7 Analysis of Translation Techniques Taking into account the advantages and shortcomings of the different techniques introduced in the previous sections, we believe ontology translation techniques do not always clearly fall into one or the other of the four broad categories – online MT-based, knowledge-based, corpus-based and semanticbased; many techniques could combine features of different approaches. For instance, online MT-based techniques may be used to generate candidate 99 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION translations and corpus-based or semantic-based techniques to disambiguate translated ontology terms – these could be thought of as hybrid approaches. In fact, in this work we propose as hypothesis that an appropriate combination of the previous translations techniques leads to better localization results than only using one at a time. Attempts at combining outputs from different systems have proven useful in many areas. For example, people in the speech community pursued the idea of combining off-the-shelf Automatic Speech Recognizers (ASRs) into a super ASR for some time, and found that the idea works (Fiscus [Fiscus, 1997], Schwenk and Gauvain [Schwenk and Gauvain, 2000], Utsuro et al. [Utsuro et al., 2003]). In Information Retrieval (IR), we find some efforts going (under the name of distributed IR or meta-search) to selectively fuse outputs from multiple search engines on the Internet (Callan et al. [Callan et al., 2003]). In Ontology Engineering, some ontology matching systems are using the combining of different matchers to produce a more efficient matching algorithm. In Machine Translation, different multi-engine MT systems have been designed as an attempt to integrate the advantages of different translation systems without accumulating their shortcomings. In the next section we present at the strategic level, some natural ways to compose different translation algorithms to localize an ontology. 5.8 Ontology Localization Strategies In the previous section we described a variety of translation techniques that can be used to localize an ontology to the linguistic level. We also showed that all these approaches have some advantages and disadvantages with regard to discovering the more appropriate translation of an ontology element. With such a wide range of term translation approaches, it would be beneficial to have an effective strategy for combining these models into a localization system that carries many of the advantages of the individual techniques and suffers from few of their disadvantages. From a technical point of view, the different translation models can be seen as the building blocks on which a ontology localization solution is built. In particular, the following aspects of building a working localization system are considered in this section: •organizing the composition of various translation algorithms (section5.8.1). •combining the results of the basic translation algorithms to discover the more appropriate translations for each ontology element (section5.8.2). As the different translation methods focus on the same objective there are several dependencies between them. Nevertheless, certain combinations 100 5.10. SUMMARY OF THE CHAPTER 5.10 Summary of the Chapter Ontology localization has different facets; one of these facets is the translation. To automatize the translation task, a variety of techniques can be used. The classifications discussed in this chapter provide a common conceptual basis to analyze the advantages and shortcomings of each technique with regard to the localization activity. We have provided such classifications based on a way of modeling the context used for the translation on one side and the kind of technology used to localize an ontology into different natural languages on an other. Once the different translation techniques have been identified, we have presented the strategic issues involved in creating localization solutions. In particular, this involves the composition of basic translation techniques and the combinations of their results. We have finished this chapter describing some high level factors that can be used to classify the approaches used to localize an ontology into different natural languages. 107 CHAPTER 5. TRANSLATION TECHNIQUES FOR ONTOLOGY LOCALIZATION 108 Chapter 6 Lyfe-Cycle Model and Architecture In this chapter we discuss two important issues related to ontology localization activity: life-cycle and system architecture. As we discussed in the introduction chapter, a typical localization project involves several tasks that extend far beyond the translation process itself. This is why the first goal of this chapter is to describe the life-cycle model by means of the representation of the major components of this activity and their interrelationships in a graphical framework that can be easily understood and communicated. As second goal of this chapter, we outline our approach to the definition of a system architecture that supports the ontology localization activity. The proposed model comprises the system components, the externally visible properties of those components, the relationships (e.g., the behavior) between them, and provides a base from which localization systems can be developed. First, we give an intuitive view of the whole localization activity, including the translation phase, which was extensively described in the previous chapters. Later in this chapter, we introduce some basic requirements for an ontology localization system. Then, we will propose a system architecture based on the ontology localization life-cycle model, considering also the system requirements identified from different works in related areas. After defining the architecture, we will see the main modules needed to allow such an ontology localization approach in distributed and collaborative environments. Finally, we describe general comments and different technical details related to the LabelTranslator system, our approach to perform an automated localization in distributed and collaborative environments. 109 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE 6.1 Ontology Localization Life-Cycle Localization is a very effort intensive activity and requires a systematic approach covering the entire life-cycle of the localized product [Mudur and Sharma, 2002]. However, based on our investigation of existing academic projects and commercial systems [Esselink, 2000, M¨uuller, 2009, Jevsikova, 2009], we have identified that the current R&D efforts on localization (especially in the software area) suffer from the lack of a comprehensive life cycle model. We consider that the ontology localization is not a once-ina-lifetime activity. It should be viewed as a continuous, iterative activity in which the localization outcomes of the current and past localizations can and should affect the future choice of localization policies and strategies and, thus, the behavior of an automated localization system. A comprehensive localization life cycle model is needed to clearly define the different phases of a localization process and to show: •what information and knowledge should be specified or defined at different phases, and, •how the results of the ontology element translations provide the feedback to other phases of the life cycle. In this section we present an ontology localization model, which identifies the key concepts and elements needed to build an automated ontology localization system. One of the elements in the model is the translation phase, which in many analogous implemented software localization systems is not automated. In the previous chapters of this thesis, we study the key elements of the translation phase, with the dual aim of reducing the localization effort and identifying the steps to produce a general ontology localization model. In fact, the translation phase used to localize an ontology has been the core of our ontology localization life-cycle model. 6.1.1 The Automated Ontology Localization Model The ontology localization life-cycle model is presented in Figure 6.1. This generic model depicts the major issues involved in the automating ontology localization activity. Our approach is inspired on different software lifecycle models [Sheu, 1997,Rajlich and Bennett, 2000,Ruparelia, 2010,Wright, 2011], which are used to illustrate the significant phases or activities of a software project from conception until retirement. Although the order of steps presented in the model is logical, we believe that different ontology localization systems may use a different order, may group two or more steps into a single step or may not implement certain steps at all. The model is also independent of who actually performs the work. For example, if the ontology developer is using a distributed and 110 6.1. ONTOLOGY LOCALIZATION LIFE-CYCLE collaborative team for localizing an ontology, many steps will be performed by the developer and others by the localization team. The value of the model is that it covers the major issues involved in this activity and provides a vocabulary to discuss these issues. In the figure the main phases are represented by a rectangle, whereas the sub-phases are represented by ellipses. The thick line represents the main process flow; the secondary process flow is represented by a solid line. The data access is shown as a dotted line in the figure. Figure 6.1: The Automated Ontology Localization Life-Cycle Model. The proposed ontology localization life-cycle model is concerned mainly with managing the translation and localization of the ontology content into any number of target languages. In the following section we describe the main components involved in the model: 6.1.2 Automated Localization Cycle The ontology localization cycle describes phases of the localization activity and the order in which those phases are executed. Each phase produces deliverables required by the next phase in the life cycle: •Change Detection. This phase monitors the source ontology content and it is responsible for detecting changes and initiating actions. We believe that change monitoring may operate continuously or at regular 111 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE intervals. In addition, this phase starts the Workflow Management process, which is responsible for distributing the work to one or more translators/reviewers in one or more localizations. •Extraction. Each ontology term requires its own extraction method from the ontology. The extraction method is responsible for extracting the ontology labels (representing any ontology element) and its context from the ontology. •Segmentation. Once the label of the ontology element is extracted, it must be segmented into individual short phrases or multiword units1 (MWU) in order to be translated appropriately. •Leveraging. The Leveraging phase tries to translate all source labels using the translations stored in previous ontology localizations. It may use one or more translation memories to store the pre-translated ontology labels. This phase can be performed only when ontologies to be translated have a similar domain to ontologies previously translated. •Work Distribution. Once the ontology localization activity has been initiated, the work must be distributed to one or more translators/reviewers in one or more localizations. This process is carried out by the Workflow Management process. The systems should provide some form of database which stores a list of translators and reviewers along with the language pairs they can handle. We consider that when the ontology to be localized is small and the target languages are known the own ontology editor may execute all tasks. •Translation. In this phase, the translator actually translates the ontology labels received using the localization resources provided by the system or its own tools if the system can interface with them. This is likely the most important step since the main cost of localization is translation and the cost of translation is largely determined by the efficient of an environment provided to the translator. The translator may work online with a browser-based tool or offline on his desktop PC. However, the offline method requires some mechanism to update the realized work. •Review. The translation work is then routed for reviewing (editing and proofing). The work is checked for translation accuracy and for overall term correctness. The system should allow any way of measuring the translation quality. 1A multiword unit (MWU) is a connected collocation: a sequence of neighboring words “whose exact and unambiguous meaning or connotation cannot be derived from the meaning or connotation of its components” [Choueka, 1988]. 112 6.1. ONTOLOGY LOCALIZATION LIFE-CYCLE •Linguistic/Cultural Updating. The goal of this task is to update the ontology with the linguistic information obtained for each ontology term in the target language. The result of this process is a multingual ontology, which expresses the correspondences between entities belonging to the source ontology and the multilingual terms pertaining to a natural language. This phase may require only the adaptation of the ontology to a particular language or an ontology re-engineering process for transforming the conceptual model of an existing and implemented ontology into a new, more correct and more complete conceptual model which is re-implemented. It is at this time that the localization resources (translation memories, glossaries, etc) are updated and that Localization Resource Maintenance is best performed. 6.1.3 Data Structures. All steps shown in the model revolve around two major data structures: Workflow and Localization Resources. The aim of the Workflow repository is to help manage, monitor and control the localization activity, while the Localization Resources repository helps to reduce the cost, increase the quality and increase the consistency of the translation work. They store the basic objects of the ontology localization activity: the participants, and the tools and resources, respectively. These objects require management and maintenance with the appropriate activities: •Workflow Management. This activity refers to the process of defining and maintaining the workflow templates that specify which steps are to be processed by users or by the system, and the conditions under which they are processed. Some ontology localization systems will have wizards with only a few questions to answer, others will require several pages of options to be set, still others will have graphical interfaces that allow for a process to be defined as a flowchart. •Translation Resources Maintenance. The more work that is routed through the ontology localization system, the more translation knowledge is accumulated, promoting more re-use. But as more and more data is accumulated, the system will also accumulate different translations for the same ontology elements. As translation knowledge grows, it becomes less precise and contains more “noise”. Therefore translation resources maintenance is required to avoid the chaotic growth of translation knowledge and ensure that the captured data can be leveraged in a meaningful way. All steps above described are the base of our generic architecture for localizing ontologies and distributed and collaborative environments. The details of our approach will be described in section 6.3. 113 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE 6.2 Key Requirements for an Ontology Localization Infrastructure In this section we describe some desirable requirements for an ontology localization system, motivated by works in related areas and own experiences. To define infrastructure requirements, we took key factors as our starting point. We have been collecting factors from different software localization systems, comparing them with our own observations in the field of ontology localization, and grouping them according to their nature and relationship into three groups: i) collaboration and distribution of the tasks; ii) translation; and iii) extensibility. This grouping has given us a clearer idea of how to convert some of these identified factors into positive influences on ontology localization. This has inspired the definition of the infrastructure requirements. In the following sections we briefly describe the three groups of requirements. 6.2.1 Requirements for Collaborative and Distributed Localization Activity In this section we present the most relevant requirements to support a distributed and collaborative ontology localization based on the analysis of the process (e.g., workflow) typically followed by organizations in the development and localization of ontologies. First, for our analysis, we considered existing processes for collaborative localization used in international institutions. As a case study, we focused on the collaborative localization process followed at FAO for localizing the AGROVOC Concept Server. Secondly, we observed how different software development paradigms and approaches deal with issues like cooperation among distributed team members. Finally, we discuss the main features identified as core requirements to support a collaborative and distributed ontology localization. Collaborative Ontology Localization: AGROVOC Concept Server One of the most important resources for covering the terminology of all subject fields in agriculture domain is the AGROVOC thesaurus (introduced in section 2.3.3), which evolved into a semantic system in order to provide ontology services. This newly reengineered system is called the “AGROVOC Concept Server (ACS)”. The development of the ACS was based on: i) the need of making the development and maintenance of the AGROVOC thesaurus more collaborative and especially more direct for users without the intermediate actions of FAO staff, and ii) the idea to convert AGROVOC into a more complete structure allowing for the representation of more information (such as additional linguistic information, or the ability to have multiple translations for 114 6.2. KEY REQUIREMENTS FOR AN ONTOLOGY LOCALIZATION INFRASTRUCTURE a specific term, etc.). This new infrastructure proposes a system, in which all actors interact collaboratively and concurrently. The collaborative aspect and the number of people that eventually interact via the ACS calls for well defined and well managed workflows to avoid confusion, data inconsistency and assure quality control. Therefore, in the collaborative aspects of the creation and localization of an ontology we need to consider at least: •collaboration over different steps performed by different people, and •collaboration among several participants for every single step Related Software Development Approaches In addition to the case analyzed in the previous section, we have observed how different software development paradigms and approaches deal with issues like cooperation among distributed team members. We have focused our research on techniques from different domains that extensively rely on communication, collaboration and/or coordination techniques. Important influences to our proposal of requirements are: •Collaborative ontology development. The collaborative development of ontologies within an organization usually follows a pre-defined process that specifies who (depending on the user role), when (depending on the ontology state) and how (what actions/operations) an ontology can change. To support this process some authors [Palma et al., 2011, Tudorache et al., 2008] propose the use of workflows to formalize the collaborative ontology development. This same idea could be applied in the collaborative localization of ontologies. •Distributed software development (DSD). We have transposed major characteristics of DSD to the ontology localization context, exploring how they could be extended to localization of ontologies. We have considered reported issues related to global DSD in the wider context [Carmel and Agarwal, 2001,Prikladnicki et al., 2008]. Some issues that arise may apply to collaborative and distributed ontology localization. The following characteristics, already adapted to deal with process improvement have been preserved for our purposes: i) the process management is distributed by the Internet and ii) the process improvement is collaborative and decentralized. •Bug tracking tools. As in software maintenance, it is possible to identify and deal with the weaknesses (translation errors) of a localized version, converting them into improvements to the next version. In this way, one can relate error-handling management to ontology localization. Both approaches follow a similar workflow including submission 115 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE (proposal), evaluation, and approval (or rejection). In the software testing context this error handling is being supported by bug-tracking tools. These tools could be customized to handle ontology translations. •Knowledge Management (KM) practices. The development of software can benefit from many KM practices, and indeed several aspects of KM employed in software development have been studied. There are many tools to support KM practices (e.g., contribution, knowledge dissemination, and collaboration) that can be useful to ontology localization. We are particularly interested in how to promote collaboration and improve participation and as such benefiting from different skills. Building networks and “knowledge communities” powered by accumulated translation knowledge can be a good strategy to facilitate the localization of ontologies. Summary of the Main Features In this section we discuss the main features of both the FAO localization workflow and the distributed software development paradigms. The goal of this discussion is to identify the core requirements to support a collaborative and distributed ontology localization activity. Five main requirements have been identified: •Flexible workflow support. The main common thread for the process that we described in the FAO case is that many steps in these workflows require human actions. Human-centred workflows are different from service workflows that combine software services for automatic execution. For our purposes we consider that a combination of these workflow approaches is a good alternative. We envisage a service workflow that enforces and automatically executes critical ontology localization steps such as ontology submission, change detection, e-mail notification of localization tasks and events, and real-time tracking and reporting of individual localization works. Localization activities such as the selection of ontology labels or reviewing of translations may be controlled by a human-workflow. •User management and provenance of information. With multiple users contributing to the localization of an ontology, it is critical for users to understand where information is coming from. Thus, users must be able to see how localization participants reach consensus on ontology label translations, who can perform translations, who can comment on them, when ontology label translations become public and so on. Any ontology localization system must include these features. •Centralized control. A centralized view on all localization projects should be provided by all ontology localization systems, giving local116 6.4. THE ONTOLOGY MANAGEMENT MODULE to establish links between the linguistic elements within one language or across languages. The NeOnToolkit3is an ontology management tool for engineering contextualized networked ontologies and semantic applications. With NeOnToolkit, we aim to start state of the art ontology localization by developing an ontology localization system. Particularly, we aim at improving the integration of multilinguality in ontologies, using a repository which keeps ontology knowledge and linguistic (multilingual) knowledge separate and independent. The Ontology Repository (OR) is the critical component which supports the association of the ontological model(s) (sources ontologies to be localized) with a multilingual linguistic model. Thus, the ontology repository relies on the combination of two independent modules, the ontological and the linguistic one. In our system, the linguistic information needed to build a multilingual ontology is generated automatically by the Ontology Translator module, which will be explained later on. The rationale underlying OR is not to design a lexicon for different natural languages and then establish links to ontology concepts, but to associate multilingual linguistic knowledge to the conceptual knowledge represented by the ontology. What the Linguistic Repository (LR) does is to associate word senses as defined by Hirst [Hirst, 2003]- in different languages to ontology concepts, although word senses and concepts can not be assumed to overlap. The LR goes along the line of what Pustejovsky [Pustejovsky, 1991] defined as Sense Enumeration Lexicon, in which a unique sense is associated with a word string. It enhances the scalability of the ontology localization approach by avoiding the need for investing time and energy in the development of a multilingual ontology for each target language. The linguistic information stored in the OR is initialized when new ontologies join the Ontology Management module. Also, for each ontology element localized, the OR automatically establishes a link with the ontology term under consideration. To support the translation resources maintenance phase identified in the localization life-cycle model (see section 6.1), we believe that the multilingual ontology, the result of this activity, should be automatically/manually added to the list of resources managed by the Ontology Translation module. The incorporation of this new resource will ensure that the translated data can be leveraged in a meaningful way. 3The NeOnToolkit is the heart of the infraestructure of the NeOn project. NeOn is an large European Research project developing an infrastructure and tool for large-scale semantic applications in distributed organizations. 123 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE 6.5 The Localization Management Module This section describes in detail the goal and functionalities of the Localization Management Module which is the key module for localization managers to monitor, manage and control the localization activity. The Workflow Localization Manager (WLM) is the core component of this module. The main goals of the WLM are: •To manage the timely flow of the localization activity from initiation to delivery, •To detect changes in the ontology model and propagate those changes to the linguistic model, and •To manage the individual localization task performed in the Ontology Translator module. In order to support the first goal the WLM includes a collaborative workflow, which implements the necessary mechanisms to allow the ontology stake-holders to perform the activities of the ontology localization life-cycle. Thus, the collaborative workflow is responsible for the coordination of who (depending on the user’s role) can do what (i.e. what kind of actions) and when (depending on the status of the ontology elements). From a technical point of view, the collaborative workflow is associated with a set of initialization parameters (e.g., user roles, assigned tasks, etc), source and target languages, and a partially ordered set of activities or states. The WLM individually stores the initialization parameters of each ontology. However, the information about user, roles and skills are stored in a shared database, which have two benefits: 1. Improved Project Staffing. The Localization Managers can see all the information related with a participant (e.g., language skills). This saves time and allows for better decisions when staffing a new ontology localization project. 2. Shared Information Across Ontology Projects. The shared information provides a particular benefit to ontology projects that need to localize several ontologies. Maintaining a single user database allows to share users in different ontology projects. Coming back to the description of the workflow, the activities supported are: selecting the ontology elements to be translated, translating the selected elements, reviewing the translations, and updating the ontology with the linguistic information obtained. These activities summarize the localization tasks commonly followed by different organizations (see section 6.2.1 for more details). In the following we summarize the main steps in the workflow process to localize an ontology (see Figure 6.4): 124 6.5. THE LOCALIZATION MANAGEMENT MODULE Figure 6.4: Workflow process used to localize an ontology. •An ontology is passed to the Localization Manager for localization. •The Localization Manager manually selects the ontology labels to be localized and sends the selected labels for translation. •A translator downloads the selected labels to be localized and (s)he performs the translations using an automated localization tool (as proposed in this thesis) or an intensive manual process. •Once translation activities have been accomplished, the translators upload the translated ontology labels and send them for review. •The reviewers download the translated labels and check for possible errors. •Finally, the Localization System updates all linguistic information of each localized label. In the next sections we explain the rest of the associated components. 6.5.1 Synchronization Component The Synchronization component supports the second goal of the workflow localization manager which is detecting the changes in the ontology management module. This component listens the changes in the ontology model and then automatically propagates those changes to the linguistic model using synchronization techniques. Remember that our system follows the current trend in the integration of multilinguality in ontologies, which suggests the suitability of keeping ontology knowledge and linguistic (multilingual) 125 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE knowledge separated and independent [Montiel-Ponsoda et al., 2008,Buitelaar et al., 2009] In order to keep both models, the ontology model (OM) and lexical model (LM), synchronized we first need to find out exactly what has been changed in the ontology model, then find the equivalent places in the linguistic model and only then start the updating. In a previous work [Espinoza et al., 2009a] we introduced the module for managing the conceptual knowledge and the linguistic knowledge by means of synchronization techniques. Hence, we briefly highlight the main features. Addition of new terms in the ontology, or deletion of an existing term can be controlled by some mechanism of change tracking. Change tracking in our approach enables the system to obtain only changes that have been made to the ontology terms, along with the information about those changes. By adopting this feature, our system can accurately identify the minimum set of changes needed to adjust the structure of the linguistic model, a critical first step to ensure that a change is made in the localized ontology. To correctly update the linguistic model, the system needs to identify: 1. all ontology terms in the original ontology whose labels have changed in the updated ontology, 2. any ontology term that has been added to the updated ontology, 3. any ontology term which has been removed from the original ontology, and 4. any ontology term whose position in the updated ontology differs from that in the original ontology. Finding where a translation is required is only part of the problem. We also need to ensure that changes in the ontology structure are accurately propagated to the linguistic information. This requires that elements whose structure need to be updated are clearly flagged in the linguistic model, and that the relevant structural changes are indicated in a form that turns the updating of the translation into a simple process, thus involving minimum work on the part of the linguist user or domain expert. Figure 6.5 illustrates the process used in our system to synchronize the conceptual and linguistic information. In the following we analyze the process in more detail, describing the actions performed by each actor of our scenario. •Ontology expert. (S)he is responsible for editing the changes in the ontology model. All the changes executed in each user session are stored in a repository as a new version. The types of changes that our system can manage are the following: changes of the label content (e.g., ontology label rename) and ontology structure changes (e.g., 126 6.5. THE LOCALIZATION MANAGEMENT MODULE delete or add operations). For each case, the system stores the type of operation executed and its additional information (e.g., the name of the renamed label). This information is used in our system to synchronize the conceptual and linguistic information. •Linguist expert(s). The linguist expert in a specific target language is responsible for performing the localization process. Notice that this process always uses the last version of an ontology. When the linguist needs to update the linguistic model (LM), our system tries to synchronize both models, performing the following actions: (1) obtaining the current version of the LM to be updated, (2) extracting the last version of the changes in the ontology model (OM) from which the last localization was taken (normally the one with the same number as the LM), (3) performing all the actions of the file of changes in the LM, and (4) updating the LM version in the repository. Figure 6.5: Synchronization of ontology and linguistic model. 6.5.2 Localization Component The last goal of the workflow localization manager is supported by the Localization component. This component controls and manages the tasks that the ontology stakeholders are allowed to perform depending on their roles and the status of the ontology elements to be localized. The possible tasks in the collaborative workflow (described previously) apply to different abstraction levels. In our solution we consider two levels: ontology element level and translation level. Although the workflows can be used independently of the underlying ontology model, the specific set of ontology terms depends on the ontology model. In our approach, we are mainly considering the OWL 127 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE ontology model, in which an OWL ontology consists of a set of axioms and facts. Concepts, properties, instances and ontology term comments are the set of ontology elements we are taking into account. The possible states that can be assigned to ontology elements are: •In Use: This is the status assigned to any element when it first passes into the collaborative workflow, or when it was localized and then updated in the Ontology Repository. •New: If the ontology element was added to the ontology after the ontology has been localized, the ontology element is passed to the “New status, and remains there until the element is localized. •Changed: If the original label of the ontology element has changed, then the element is passed to the “Changed” status, and remains there until the element is checked to be localized again. •Unused: If the ontology element has been deleted, then this element is passed to the “Unused” status, and remains there until the ontology is synchronized (see synchronization component in the previous section). The localization component controls also the status of the translations. Figure 6.6 shows the workflow to translation level. States are denoted by rectangles and actions by arrows. The actors in the figure specify the actions that an ontology stakeholder can perform depending on its role. In the following we provide a detailed explanation: Figure 6.6: Workflow to the Translation level. •Not translated: This is the status assigned to any translation when it first passes into the collaborative workflow or when any change has been performed in the element of the ontology under consideration. 128 6.6. THE ONTOLOGY TRANSLATION MODULE •To be Translated: Once the Localization Manager selects the translations with the “Not translated” status, these translations are passed to the “To be Translated” status, and remain there until a “Translator” translates them. •Auto translated: If a “Translator” uses the automatic translation algorithm provided by the system, then the translation is passed to the “Auto-translated” status, and remains there until the own “Translator” sends the translations to the “To Be Reviewed (auto)” status. •Translated: If a “Translator” manually makes a translation, then the translation is passed to the “Translated” status, and remains there until the own “Translator” sends the translations to the “To be Reviewed” status. •To be Reviewed: If a “Reviewer” approves the translations send by the translator, it passes to the “Complete” status. The reviewer knows in advance if translations have been made automatically or manually. For example, the word “automatic” in the message “To Be Reviewed (automatic)” indicates to the reviewer that these translations have been obtained automatically. Additionally, when the translations reach the “Complete” status, they are automatically updated in the Ontology Repository. Note that during the collaborative workflow, actions are performed either implicitly or explicitly. For instance, when a user updates (i.e. modifies) an ontology label, he does not explicitly perform an update action. In this case the action has to be captured from the user interface and recorded when the ontology is saved. In contrast, when Reviewers for example explicitly approve/reject proposed translations, the action is immediately recorded when performed. 6.6 The Ontology Translation Module In this section we explain the ontology translation approach that was already briefly mentioned in Chapter 4. Also, in Chapter 5 we already described the label translation component, showing some natural ways to combine different translation algorithms to localize an ontology. Now we extend the explanation by introducing the additional steps given by the Ontology Translation to discover the more appropriate translations. With the help of the ontology translator module, the translators/reviewers can reduce the effort to manually localize an ontology. For each ontological label, the translation module first uses the Leverage component to try to discover pre-translated ontology labels. If no results are obtained, then, the Translator component performs three pipeline steps to translate a label 129 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE described in a source natural language and obtain the most probable translation of this ontological label in a target natural language. Figure 6.7 shows the components presented in Figure 6.3 in more detail. Figure 6.7: Detailed Ontology Translator in a Localization System. 6.6.1 Leverage Component This component takes as input the ontology label and its context to discover past translations. In its simplest form, this component may rely on a cache to match ontology labels with pre-translated labels from previous ontology lo130 6.6. THE ONTOLOGY TRANSLATION MODULE calizations. The software localization industry recommends the use of translation memories as translation reuse technology [Massion, 2005, Lagoudaki, 2006]. For our purposes, a translation memory (TM) system will transform inventories of past ontology translations into a database by automatically extracting and aligning the source labels with the target language labels. This involves the creation of a simple database of aligned words taken out of context. Since this happens without reference to context, TM technology will require a good deal of manual maintenance by a senior linguist to validate and correct misalignments, especially 1:n and n:1 combinations that are readily apparent to human, but not to automated tools [Kuhns, 2007]. A more serious shortcoming of TM systems is the fact that they have no access to the meaning of the translated text and operate on its surface form. As a result, they fail to match words/sentences that have the same meaning, but a different syntactic structure [Kuhns, 2007]. To overcome these shortcomings, a new generation of TM systems has been proposed, which analyse the segments not only in terms of syntax but also in terms of semantics [Gotti et al., 2005,Pekar and Mitkov, 2007,Mitkov and Corpas, 2008]. Some of these works rely on lexical resources such as WordNet to automatically identify synonyms, and therefore, to make a match between synonymous expressions possible. Due to the rather restricted availability of semantic data in relevant subject areas, the relevance of these approaches for commercial implementations is still rather small [Reinke, 2013]. However, we believe that this type of approaches represents a promising way forward to ensure that translators have a wider range of matches in the ontology localization activity. We envision also that the use of context information in both the input term and the translation memory improves matching algorithms. In our approach, for the time being we only use a cache to avoid translating the same label twice, while we wait until some of the features discussed above can be incorporated into the currently available TM. In the following section we describe the first step of the translator component and briefly the second and third steps. 6.6.2 Translator Component The Translator component relies on three steps: label pre-processing,label translation and label post-processing to discover appropriate translations according the lexical and semantic context of the original ontology label. The three steps identified above are executed if the Leverage component does not return any results. The output of this component is an automatically translated label and manually validated by an expert. 131 CHAPTER 6. LYFE-CYCLE MODEL AND ARCHITECTURE Label Pre-Processing We consider that ontology label pre-processing is essential in an ontology localization system, in order to simplify the core translation processing and make it both quality and time effective. The ontology labels pose different challenges to MT, which can be attributed to two distinct characteristics: •Ontology labels differ linguistically and stylistically from written language: phrases are shorter and in some cases poorly structured, also they can contain ungrammaticality expressions (e.g., Service Transport instead of Transport Service) •The current “standard” for naming the ontology labels is to use a CamelCase4approach. Therefore, we cannot rely on the initial uppercase letter to identify a phrase initial word nor to recognize proper names, since names cannot be identified by an initial capital. These problematic factors are dealt with in a pre-processing pipeline that prepares the input for processing by a core MT technique. Thus, the task of the ontology label pre-processing pipeline is to make the input amenable to a linguistically-principled, domain independent treatment. This task is accomplished in two ways: 1. By normalizing the input, i.e. removing noise, reducing the input to standard typographical conventions, and also restructuring and simplifying it, whenever this can be done in a reliable, meaning-preserving way. 2. By annotating the input with linguistic information, whenever this can be reliably done with a shallow linguistic analysis, to reduce input ambiguity and make a full linguistic analysis more manageable. In the following we describe the functionalities of the different tasks in more detail: Normalization The label normalization groups three components, which clean up and tokenize the input. The text-level normalization phase performs operations at the string level (ontology term comments by example), such as removing extraneous text 4CamelCase (also spelled camel case, camel-case or medial capitals) is the practice of writing compound words or phrases in which the elements are joined without spaces, with each element’s initial letter capitalized within the compound, and the first letter is either upper or lower caseas in “LaBelle”, “BackColor”, or “iPod”. The name comes from the uppercase “bumps” in the middle of the compound word, suggestive of the humps of a camel. The practice is known by many other names. 132 BIBLIOGRAPHY [Guarino, 1998] Guarino, N. (1998). Formal Ontology in Information Systems: Proceedings of the 1st International Conference June 6-8, 1998, Trento, Italy. IOS Press, Amsterdam, The Netherlands, The Netherlands. [Guyot et al., 2005] Guyot, J., Radhouani, S., and Falquet, G. (2005). Ontology-based multilingual information retrieval. In CLEF. [Habash and Sadat, 2006] Habash, N. and Sadat, F. (2006). Arabic preprocessing schemes for statistical machine translation. In NAACL ’06: Proceedings of the Human Language Technology Conference of the NAACL, Companion Volume: Short Papers on XX, pages 49–52, Morristown, NJ, USA. Association for Computational Linguistics. [Hasse et al., 2008] Hasse, P., Lewen, H., Studer, R., and Erdmann, M. (2008). The neon ontology engineering toolkit. In WWW 2008 Developers Track. [Hearne and Way, 2011] Hearne, M. and Way, A. (2011). Statistical machine translation: A guide for linguists and translators. Language and Linguistics Compass, 5(5):205–226. [Hewlett et al., 2005] Hewlett, D., Kalyanpur, A., Kovlovski, V., and Halaschek-Wiener, C. (2005). Effective natural language paraphrasing of ontologies on the semantic web. In End User Semantic Web Interaction Workshop, Galway, Ireland. [Hirst, 2003] Hirst, G. (2003). Ontology and the lexicon. In Handbook on Ontologies in Information Systems, pages 209–230. Springer. [Hovy et al., 2001] Hovy, E., Ide, N., Frederking, R., Mariani, J., and Zampolli, A. (2001). Multilingual Information Management. Pisa, Italy: Giardini Editori e Stampatori and Kluwer Academic Publishers. [Huang and Papineni, 2007] Huang, F. and Papineni, K. (2007). Hierarchical system combination for machine translation. In EMNLP-CoNLL, pages 277–286. [Hudik and Ruopp, 2011] Hudik, T. and Ruopp, A. (2011). The integration of moses into localization industry. In 15th Annual Conference of the EAMT, pages 47–53. [Hutchins, 2007] Hutchins, J. (2007). Machine translation: problems and issues. Presentation. Chelyabinsk, Russia. 18 slides. [Ide and Veronis, 1998] Ide, N. and Veronis, J. (1998). Introduction to the special issue on word sense disambiguation: The state of the art. Computational Linguistics. Special Issue on Word Sense Disambiguation. 235 BIBLIOGRAPHY [IEEE, 2000] IEEE (2000). The Authoritative Dictionary of IEEE Standard Terms. Seventh edition, December. [ISO, 1985] ISO (1985). Documentation – Guidelines for the establishment and development of multilingual thesauri. [ISO, 1986] ISO (1986). Documentation – Guidelines for the establishment and development of monolingual thesauri. [ISO, 1992] ISO (1992). Guidance on usability specification and measures, ISO, CD 9241-11. [Jang et al., 1999] Jang, M.-G., Myaeng, S. H., and Park, S. Y. (1999). Using mutual information to resolve query translation ambiguities and query term weighting. In Proceedings of the 37th annual meeting of the Association for Computational Linguistics on Computational Linguistics, ACL ’99, pages 223–229, Stroudsburg, PA, USA. Association for Computational Linguistics. [Jayaraman and Lavie, 2005] Jayaraman, S. and Lavie, A. (2005). Multiengine machine translation guided by explicit word matching. In Proc. of EAMT, pages 143–152. [Jevsikova, 2009] Jevsikova, T. (2009). Internet Software Localization. PhD thesis, Vilnius University, Vilnius, Lithuania,. [Kansai et al., 1996] Kansai, M. K., Kitamura, M., and Matsumoto, Y. (1996). Automatic extraction of word sequence correspondences in parallel corpora. In Proc. of the 4th Annual Workshop on Very Large Corpora (WVLC-4, pages 79–87. [Kersten et al., 2002] Kersten, G. E., Kersten, M., and Rakowski, W. M. (2002). Software and culture: Beyond the internationalization of the interface. JGIM, 10(4):86–101. [Kilgarriff and Rosenzweig, 2000] Kilgarriff, A. and Rosenzweig, J. (2000). Framework and results for english senseval. Special Issue on SENSEVAL. Computers and the Humanties, pages 15–48. [Kim and Kim, 1997] Kim, S. and Kim, Y. (1997). Sentence segmentation for efficient english syntactic analysis. Journal of Korea Information Science Society. v24 i8. [Kim et al., 2001] Kim, S.-D., Zhang, B.-T., and Kim, Y. T. (2001). Learning-based intrasentence segmentation for efficient translation of long sentences. Machine Translation, 16(3):151–174. 236 BIBLIOGRAPHY [Kim and Oh, 2008] Kim, Y.-S. and Oh, Y.-J. (2008). Intra-sentence segmentation based on support vector machines in english-korean machine translation systems. Expert Syst. Appl., 34(4):2673–2682. [Kirakowski and Corbett, 1993] Kirakowski, J. and Corbett, M. (1993). Sumi: The software usability measurement inventory. British Journal of Educational Technology. [Klein, 2001] Klein, M. (2001). Combining and relating ontologies: An analysis of problems and solutions. [Koehn, 2005] Koehn, P. (2005). Europarl: A Parallel Corpus for Statistical Machine Translation. In Conference Proceedings: the tenth Machine Translation Summit, pages 79–86. AAMT. [Koehn and Knight, 2003] Koehn, P. and Knight, K. (2003). Empirical methods for compound splitting. In EACL ’03: Proceedings of the tenth conference on European chapter of the Association for Computational Linguistics, pages 187–193, Morristown, NJ, USA. Association for Computational Linguistics. [Koehn and Monz, 2006] Koehn, P. and Monz, C. (2006). Manual and automatic evaluation of machine translation between european languages. In Proceedings of the Workshop on Statistical Machine Translation, StatMT ’06, pages 102–121, Stroudsburg, PA, USA. Association for Computational Linguistics. [Kuhns, 2007] Kuhns, R. J. (2007). Advanced leveraging: The new generation of tms. [Lagoudaki, 2006] Lagoudaki, E. (2006). Translation memory systems: Enlightening users perspective. [Landers, 2001] Landers, C. (2001). Literary Translation: A Practical Guide. Topics in translation. Multilingual Matters. [Lassila and McGuinness, 2001] Lassila, O. and McGuinness, D. L. (2001). The role of frame-based representation on the semantic web. In Knowledge Systems Laboratory Report KSL-01-02, Stanford University, 2001; Also appeared as Linkping Electronic Articles in Computer and Information Science, Vol. 6 (2001), No. 005, Linkping University,. [Lavie et al., 2004] Lavie, A., Sagae, K., and Jayaraman, S. (2004). The significance of recall in automatic metrics for mt evaluation. In Proceedings of the 6th Conference of the Association for Machine Translation in the Americas (AMTA-2004. 237 BIBLIOGRAPHY [Laviosa, 1997] Laviosa, S. (1997). How comparable can “comparable corpora” be? In Target 9(2):289-319. [Lenat and Guha, 1989] Lenat, D. B. and Guha, R. V. (1989). Building Large Knowledge-Based Systems; Representation and Inference in the Cyc Project. Addison-Wesley Longman Publishing Co., Inc., Boston, MA, USA. [Lesk, 1986] Lesk, M. (1986). Automatic sense disambiguation using machine readable dictionaries: how to tell a pine cone from an ice cream cone. In Proceedings of the 5th annual international conference on Systems documentation, SIGDOC ’86, pages 24–26, New York, NY, USA. ACM. [Levenshtein, 1965] Levenshtein, V. (1965). Binary codes capable of correcting deletions, insertions, and reversals. doklady akademii nauk sssr, 163(4):845848,. In Russian. English Translation in Soviet Physics Doklady, 10(8) p. 707710,. [Levin., 1993] Levin., B. (1993). English Verb Classes and Alternations: A Preliminary Investigation. University of Chicago Press, Chicago, IL, USA. [Li et al., 2003] Li, H., Cao, Y., and Li, C. (2003). Using bilingual web data to mine and rank translations. IEEE Intelligent Systems, 18:54–59. [Liang and Sini, 2006] Liang, A. and Sini, M. (2006). Mapping agrovoc and the chinese agricultural thesaurus: Definitions, tools, procedures. In New Review of Hypermedia and Multimedia, pp. 51 – 62, 12 (1). [Liang et al., 2005] Liang, A., Sini, M., Chang, C., Li, S., Lu, W., He, C., and Keizer, J. (2005). The mapping schema from chinese agricultural thesaurus to agrovoc. In 6th Agricultural Ontology Service (AOS) Workshop on Ontologies: the more practical issues and experiences. [Litkowski, 2005] Litkowski, K. C. (2005). Computational Lexicons and Dictionaries. In Encyclopedia of Language and Linguistics (2nd ed.). Elsevier Publishers. [Llitjs, 2009] Llitjs, A. F. (2009). Automatic Improvement of Machine Translation Systems. VDM Verlag, Saarbrucken, Germany. [Lu et al., 2002] Lu, W.-H., Chien, L.-F., and Lee, H.-J. (2002). Translation of web queries using anchor text mining. ACM Transactions on Asian Language Information Processing (TALIP), 1(2):159–172. [Malais´e et al., 2007] Malais´e, V., Isaac, A., Gazendam, L., and Brugman, H. (2007). Anchoring dutch cultural heritage thesauri to wordnet: Two case studies. In Proceedings of the Workshop on Language Technology 238 BIBLIOGRAPHY for Cultural Heritage Data (LaTeCH 2007)., pages 57–64, Prague, Czech Republic. Association for Computational Linguistics. [Malmkjær, 2000] Malmkjær, K. (2000). Multidisciplinarity in process research. In S. Tirkkonen-Condit & R. Jaaskelainen (eds.), Tapping and Mapping the Process of Translation: Outlooks on Empirical Research., pages 163–170. [Martin et al., 2008] Martin, H., De Leenheer, P., de Moor, A., and Sure, Y. (2008). Ontology Management, Semantic Web, Semantic Web Services, and Business Applications. Semantic Web and Beyond. Springer-Verlag, Heidelberg. [Maruyama, 1992] Maruyama, H.; Watanabe, H. (1992). Tree cover search algorithm for example-based translation. In Proceedings of the 4th International Conference on Theoretical and Methodological Issues in Machine Translation, Montreal, 173184. [Massion, 2005] Massion, F. (2005). Translation-memory-systeme im vergleich. [Matusov et al., 2006] Matusov, E., Ueffing, N., and Ney, H. (2006). Computing consensus translation from multiple machine translation systems using enhanced hypotheses alignment. In Cambridge University Engineering Department, pages 33–40. [McCrae et al., 2012] McCrae, J., Davis, B., and Gracia, J. (2012). Enriching the web with ontology-lexica. [McCrae et al., 2011a] McCrae, J., Espinoza, M., Montiel-Ponsoda, E., Aguado-de Cea, G., and Cimiano, P. (2011a). Combining statistical and semantic approaches to the translation of ontologies and taxonomies. In Proceedings of the Fifth Workshop on Syntax, Semantics and Structure in Statistical Translation, SSST-5, pages 116–125, Stroudsburg, PA, USA. Association for Computational Linguistics. [McCrae et al., 2011b] McCrae, J., Spohr, D., and Cimiano, P. (2011b). Linking lexical resources and ontologies on the semantic web with lemon. In Proceedings of the 8th extended semantic web conference on The semantic web: research and applications - Volume Part I, ESWC’11, pages 245–259, Berlin, Heidelberg. Springer-Verlag. [McEnery, 2003] McEnery, A. (2003). Corpus linguistics. In R. Mitkov (ed.) Oxford handbook of computational linguistics., pages 448–63. [Melamed et al., 2003] Melamed, I. D., Green, R., and Turian, J. P. (2003). Precision and recall of machine translation. In NAACL ’03: Proceedings 239 BIBLIOGRAPHY of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology, pages 61– 63, Morristown, NJ, USA. Association for Computational Linguistics. [Mihalcea, 2002] Mihalcea, R. F. (2002). Word sense disambiguation with pattern learning and automatic feature selection. Nat. Lang. Eng., 8(4):343–358. [Miles et al., 2005] Miles, A., Matthews, B., Beckett, D., Brickley, D., Wilson, M., and Rogers, N. (2005). Skos: A language to describe simple knowledge structures for the web. In Proceedings of XTech 2005. [Miller, 1995] Miller, G. (1995). WordNet: A Lexical Database for English. Communications of the ACM, 38(11). [Miller and Matthews, 2001] Miller, K. and Matthews, B. (2001). Having the right connections: the limber project. Journal of Digital information, vol. 1 issue 8. [Mitkov and Corpas, 2008] Mitkov, R. and Corpas, G. (2008). Improving third generation translation memory systems through identification of rhetorical predicates. In LangTech Proceedings. [Montiel-Ponsoda, 2011a] Montiel-Ponsoda, E. (2011a). Multilingualism in Ontologies. PhD thesis, Universidad Polit´ecnica de Madrid, Madrid, Espa˜na. [Montiel-Ponsoda, 2011b] Montiel-Ponsoda, E. (2011b). Multilingualism in Ontologies - Building Patterns and Representation Models. LAP Lambert Academic Publishing. [Montiel-Ponsoda et al., 2008] Montiel-Ponsoda, E., Aguado, G., G´omezP´erez, A., and Peters., W. (2008). Modelling multilinguality in ontologies. In Coling 2008: Companion volume - Posters and Demonstrations, Manchester, UK. [Montiel-Ponsoda et al., 2011] Montiel-Ponsoda, E., de Cea, G. A., G´omezP´erez, A., and Peters, W. (2011). Enriching ontologies with multilingual information. Natural Language Engineering, 17(3):283–309. [MSDN, 2012] MSDN (last accessed December 2012). Chapter 1 Understanding Internationalization http://msdn.microsoft.com/enus/library/cc194758.aspx. [Mudur and Sharma, 2002] Mudur, S. P. and Sharma, R. (2002). A reference model for software localisation. 240 BIBLIOGRAPHY [Munt´es et al., 2012] Munt´es, V., Paladini, P., Espa˜na-Bonet, C., and M`arquez, L. (2012). Context-aware machine translation for software localization. In Proceedings of the 16th Annual Conference of the European Association for Machine Translation (EAMT12), pages 77–80, Trento, Italy. [M¨uuller, 2009] M¨uuller, E. (2009). Building quality into the localization process. Multilingual Localization: Getting Started Guide. [Nagao, 1984] Nagao, M. (1984). A framework of a mechanical translation between japanese and english by analogy principle. In Proc. of the international NATO symposium on Artificial and human intelligence, pages 173–180, New York, NY, USA. Elsevier North-Holland, Inc. [Nie, 2010] Nie, J.-Y. (2010). Cross-language information retrieval. Synthesis Lectures on Human Language Technologies, 3(1):1–125. [Nie et al., 2001] Nie, J.-Y., Simard, M., and Foster, G. (2001). Multilingual information retrieval based on parallel texts from the web. In CLEF ’00: Revised Papers from the Workshop of Cross-Language Evaluation Forum on Cross-Language Information Retrieval and Evaluation, pages 188–201, London, UK. Springer-Verlag. [Nomoto, 2004] Nomoto, T. (2004). Multi-engine machine translation with voted language model. In ACL ’04: Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics, page 494, Morristown, NJ, USA. Association for Computational Linguistics. [Nord, 1997] Nord, C. (1997). Translating as a purposeful activity. Functionalist approaches explained. UK: St. Jerome. [Nord, 2005] Nord, C. (2005). Text Analysis in Translation; Theory, Methodology and Didactic Application of a Model for TranslationOriented Text Analysis. GA: Rodopi., Amsterdam - Atlanta, second edition edition. [Noy et al., 2001] Noy, N. F., Sintek, M., Decker, S., Crub´ezy, M., Fergerson, R. W., and Musen, M. A. (2001). Creating semantic web contents with prot´eg´e-2000. IEEE Intelligent Systems, 16(2):60–71. [Oard, 1997] Oard, D. W. (1997). Alternative approaches for cross-language text retrieval. In AAAI Symposium on cross-language text and speech retrieval. American Association for Artificial Intelligence. [Oard and Hackett, 1997] Oard, D. W. and Hackett, P. (1997). Document translation for cross-language text retrieval at the university of maryland. In The Sixth Text REtrieval Conference (TREC-6). National Institutes of Standards and Technology. 241 BIBLIOGRAPHY [Och, 2002] Och, F. J. (2002). Statistical Machine Translation: From SingleWord Models to Alignment Templates. PhD thesis, RWTH Aachen University, Aachen, Germany. [Och and Ney, 2003] Och, F. J. and Ney, H. (2003). A systematic comparison of various statistical alignment models. Computational Linguistics, 29(1):19–51. [OGC, 1996] OGC, editor (1996). The OpenGIS Guide - Introduction to Interoperable Geoprocessing and the OpenGIS Specification. Open GIS Consortium, Inc, Boston. [Ogden and Davis, 2000] Ogden, W. C. and Davis, M. W. (2000). Improving cross-language text retrieval with human interactions. [O‘Sullivan, 2001] O‘Sullivan (2001). A Paradigm for Creating Multilingual Interfaces. PhD thesis, University of Limerick, Irland,. [Palma et al., 2011] Palma, R., Corcho, ´ O., G´omez-P´erez, A., and Haase, P. (2011). A holistic approach to collaborative ontology development based on change management. J. Web Sem., 9(3):299–314. [Palma et al., 2008] Palma, R., Haase, P., Corcho, O., G´omez-P´erez, A., and Ji., Q. (2008). An editorial workflow approach for collaborative ontology development. In 3rd Asian Semantic Web Conference. ASWC 08. Bangkok, Thailand. [Palmer and Wu, 1995] Palmer, M. S. and Wu, Z. (1995). Verb semantics for english-chinese translation. Machine Translation, 10(1-2):59–92. [Papineni et al., 2002] Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. In ACL ’02: Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, pages 311–318, Morristown, NJ, USA. Association for Computational Linguistics. [Park, 2001] Park, S. B. (2001). Computing consensus translation from multiple machine translation systems. In Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop (ASRU-2001, pages 351–354. [Paul et al., 2005] Paul, M., Doi, T., Hwang, Y., Imamura, K., Okuma, H., and Sumita, E. (2005). Nobody is perfect: Atrs hybrid approach to spoken language translation. In Proc. of IWSLT, pages 55–62. [Pazienza et al., 2005] Pazienza, M., Stellato, A., Zanzotto, F., Henriksen, L., and Paggio., P. (2005). Ontology mapping to support ontology based question answering. In Proceedings of the Meaning 05 Workshop. 242 BIBLIOGRAPHY [Pazienza and Stellato, 2006] Pazienza, M. T. and Stellato, A. (2006). Exploiting linguistic resources for building linguistically motivated ontologies in the semantic web. In Second Workshop on Interfacing Ontologies and Lexical Resources for Semantic Web Technologies (OntoLex2006), held jointly with LREC2006, May 24-26, 2006, Genoa, (Italy). [Pease and Niles, 2002] Pease, A. and Niles, I. (2002). Ieee standard upper ontology: a progress report. Knowl. Eng. Rev., 17(1):65–70. [Pedersen et al., 2005] Pedersen, T., Banerjee, S., and Patwardhan, S. (2005). Maximizing Semantic Relatedness to Perform Word Sense Disambiguation. Research Report UMSI 2005/25, University of Minnesota Supercomputing Institute. [Pekar and Mitkov, 2007] Pekar, V. and Mitkov, R. (2007). New generation translation memory: Content-sensitive matching. In Proceedings of the 40th Anniversary Congress of the Swiss Association of Translators, Terminologists and Interpreters, pages 16–30. [Peters and Sheridan, 2000] Peters, C. and Sheridan, P. (2000). Multilingual information access. In ESSIR, pages 51–80. [Pinto et al., 2004] Pinto, S., Staab, S., Sure, Y., and Tempich., C. (2004). Ontoedit empowering swap: a case study in supporting distributed, loosely-controlled and evolving engineering of ontologies (diligent). In Proceedings of the First European Semantic Web Symposium, ESWS 2004, Heraklion, Crete, Greece, pages 16–30. [Prikladnicki et al., 2008] Prikladnicki, R., Damian, D., and Audy, J. (2008). Patterns of evolution in the practice of distributed software development: Quantitative results from a systematic review. In Evaluation and Assessment in Software Engineering (EASE), Bari, Italy,. [Prins and van den Broek, 2004] Prins, J. and van den Broek, P. (2004). Semantic search ilse media towards a dutch semantic web infrastructure. Master’s thesis, Vrije Universiteit Amsterdam, The Netherlands,. [Pustejovsky, 1991] Pustejovsky, J. (1991). The generative lexicon. Computational Linguistics, 17(4):409–441. [Qu et al., 2012] Qu, J., Shimazu, A., and Nguyen, M. (2012). Oov term translation, context information and definition extraction based on oov term type prediction. In Advances in Natural Language Processing, volume 7614 of Lecture Notes in Computer Science, pages 76–87. Springer Berlin Heidelberg. 243 BIBLIOGRAPHY [Quirk, 2004] Quirk, C. (2004). Training a sentence-level machine translation confidence metric. In Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC), pages 825828,Lisbon, Portugal. [Rada et al., 1989] Rada, R., Mili, H., Bicknell, E., and Blettner, M. (1989). Development and application of a metric on semantic nets. IEEE Transactions on Systems, Man and Cybernetics, 19(1):17–30. [Raghavan and Wong, 1986] Raghavan, V. V. and Wong, S. K. M. (1986). A critical analysis of vector space model for information retrieval. J. Am. Soc. Inf. Sci., 37(5):279–287. [Rajlich and Bennett, 2000] Rajlich, V. T. and Bennett, K. H. (2000). A staged model for the software life cycle. Computer, 33(7):66–71. [Rapp, 1997] Rapp, R. (1997). Text-detektor. fehlertolerantes retrieval ganz einfach. In Magazin fr Computertechnik, 4/97, 386392. [Reinke, 2013] Reinke, U. (2013). State of the art in translation memory technology. Translation: Computation, Corpora, Cognition, 3(1). [Resnik, 1999] Resnik, P. (1999). Disambiguating noun groupings with respect to wordnet senses. In Proceedings of the Third Workshop on Very Large Corpora, pages 54-68. Association for Computational Linguistics. [Rosti et al., 2007a] Rosti, A.-V. I., Ayan, N. F., Xiang, B., Matsoukas, S., Schwartz, R. M., and Dorr, B. J. (2007a). Combining outputs from multiple machine translation systems. In HLT-NAACL, pages 228–235. [Rosti et al., 2007b] Rosti, A.-V. I., Matsoukas, S., and Schwartz, R. M. (2007b). Improved word-level system combination for machine translation. In ACL. [Rosti et al., 2008] Rosti, A.-V. I., Zhang, B., Matsoukas, S., and Schwartz, R. (2008). Incremental hypothesis alignment for building confusion networks with application to machine translation system combination. In StatMT ’08: Proceedings of the Third Workshop on Statistical Machine Translation, pages 183–186, Morristown, NJ, USA. Association for Computational Linguistics. [Ruopp, 2010] Ruopp, A. (2010). The moses for localization open source project. In Conference of the AMTA. [Ruparelia, 2010] Ruparelia, N. B. (2010). Software development lifecycle models. SIGSOFT Softw. Eng. Notes, 35(3):8–13. 244 Appendix A Ontology Localization Framework Evaluation Software Usability Measurement Inventory (SUMI) questionnaire to assess the usability of the LabelTranslator system. A.1 Efficiency 1. This software responds too slowly to inputs. •Agree •Undecided •Disagree 2. I would recommend this software to my colleagues. •Agree •Undecided •Disagree 3. The instructions and prompts are helpful. •Agree •Undecided •Disagree 4. The software has at some time stopped unexpectedly. •Agree •Undecided •Disagree 251 APPENDIX A. ONTOLOGY LOCALIZATION FRAMEWORK EVALUATION 5. Learning to operate this software initially is full of problems. •Agree •Undecided •Disagree 6. I sometimes dont know what to do next with this software. •Agree •Undecided •Disagree 7. I enjoy my sessions with this software. •Agree •Undecided •Disagree 8. I find that the help information given by this software is not very useful. •Agree •Undecided •Disagree 9. If this software stops, it is not easy to restart it. •Agree •Undecided •Disagree 10. It takes too long to learn the software commands. •Agree •Undecided •Disagree A.2 Affect 1. I sometimes wonder if Im using the right command. •Agree •Undecided •Disagree 252 A.2. AFFECT 2. Working with this software is satisfying. •Agree •Undecided •Disagree 3. The way that system information is presented is clear and understandable. •Agree •Undecided •Disagree 4. I feel safer if I use only a few familiar commands or operations. •Agree •Undecided •Disagree 5. The software documentation is very informative. •Agree •Undecided •Disagree 6. This software seems to disrupt the way I normally like to arrange my work. •Agree •Undecided •Disagree 7. Working with this software is mentally stimulating. •Agree •Undecided •Disagree 8. There is never enough information on the screen when its needed. •Agree •Undecided •Disagree 9. I feel in command of this software when I am using it. 253 APPENDIX A. ONTOLOGY LOCALIZATION FRAMEWORK EVALUATION •Agree •Undecided •Disagree 10. I prefer to stick to the facilities that I know best. •Agree •Undecided •Disagree A.3 Helpfulness 1. I think this software is inconsistent. •Agree •Undecided •Disagree 2. I would not like to use this software every day. •Agree •Undecided •Disagree 3. I can understand and act on the information provided by this software. •Agree •Undecided •Disagree 4. This software is awkward when I want to do something which is not standard. •Agree •Undecided •Disagree 5. There is too much to read before you can use the software. •Agree •Undecided •Disagree 254 A.4. CONTROL 6. Tasks can be performed in a straightforward manner using this software. •Agree •Undecided •Disagree 7. Using this software is frustrating. •Agree •Undecided •Disagree 8. The software has helped me overcome any problems I have had in using it. •Agree •Undecided •Disagree 9. The speed of this software is fast enough. •Agree •Undecided •Disagree 10. I keep having to go back to look at the guides. •Agree •Undecided •Disagree A.4 Control 1. It is obvious that user needs have been fully taken into consideration. •Agree •Undecided •Disagree 2. There have been times in using this software when I have felt quite tense. •Agree 255 APPENDIX A. ONTOLOGY LOCALIZATION FRAMEWORK EVALUATION •Undecided •Disagree 3. The organisation of the menus or information lists seems quite logical. •Agree •Undecided •Disagree 4. The software allows the user to be economic of keystrokes. •Agree •Undecided •Disagree 5. Learning how to use new functions is difficult. •Agree •Undecided •Disagree 6. There are too many steps required to get something to work. •Agree •Undecided •Disagree 7. I think this software has made me have a headache on occasion. •Agree •Undecided •Disagree 8. Error prevention messages are not adequate. •Agree •Undecided •Disagree 9. It is easy to make the software do exactly what you want. •Agree •Undecided •Disagree 256 A.5. LEARNABILITY 10. I will never learn to use all that is offered in this software. •Agree •Undecided •Disagree A.5 Learnability 1. The software hasnt always done what I was expecting. •Agree •Undecided •Disagree 2. The software has a very attractive presentation. •Agree •Undecided •Disagree 3. Either the amount or quality of the help information varies across the system. •Agree •Undecided •Disagree 4. It is relatively easy to move from one part of a task to another. •Agree •Undecided •Disagree 5. It is easy to forget how to do things with this software. •Agree •Undecided •Disagree 6. This software occasionally behaves in a way which cant be understood. •Agree •Undecided •Disagree 257 APPENDIX A. ONTOLOGY LOCALIZATION FRAMEWORK EVALUATION 7. This software is really very awkward. •Agree •Undecided •Disagree 8. It is easy to see at a glance what the options are at each stage. •Agree •Undecided •Disagree 9. Getting data files in and out of the system is not easy. •Agree •Undecided •Disagree 10. I have to look for assistance most times when I use this software. •Agree •Undecided •Disagree 258 Appendix B Localization User Guides B.1 User Guide for Localization Managers The Localization Manager’s User Guide demonstrates i) how to install the environment for the collaborative ontology localization scenario, ii) how to set preferences of the ontology localization activity, iii) how to import the ontology to be localized, iv) how to set localization parameters, and v) how to select ontology labels to be localized. B.1.1 Installing the Environment. Neon Toolkit installation 1. Open a Web browser and install Neon Toolkit from http://neon-toolkit. org/wiki/download. For Windows, Linux or Mac operating system choose Basic (Installer) version. 2. Fill fields marked with asterisk in the license Web site. 259 APPENDIX B. LOCALIZATION USER GUIDES 3. Wait while Neon Toolkit is downloaded and then save the file into any directory (installation directory). 4. Execute the install file from installation directory and complete the Neon Toolkit setup wizard. LabelTranslator plugins installation 1. Open a Web browser and download LabelTranslator plugins from http://delicias.dia.fi.upm.es/repos/collaborativelabeltranslator/plugins 2. Input login and password to access to plugins repository. 3. Select plugins.zip file and save the file into any directory. 4. Wait while LabelTranslator plugins are downloaded and unpackged the plugins into Neon Toolkit installation directory (see step 3 in Neon Toolkit installation) . 5. Check the installed files into Neon Toolkit plugins directory. Localization server installation 1. Open a Web browser and download Localization Server from http://de licias.dia.fi.upm.es/repos/collaborativelabeltranslator/server 2. Select server.zip file and save the file into any directory. 3. Wait while server.zip file is downloaded and unpackged the file into any directory of the server machine. 260 B.1. USER GUIDE FOR LOCALIZATION MANAGERS 6. On the Ontology Navigator, select the imported ontology and click on save icon to finish the importation. B.1.4 Setting-up Localization Parameters. 1. In the Window/Open Perspective menu, change the perspective to localization. 2. By right click on the imported ontology select the Build Localization Project... option, to configure the parameters of a new ontology localization project. 267 APPENDIX B. LOCALIZATION USER GUIDES 3. Select source and target languages. For our experiment choose English and Spanish as source and target languages respectively. 4. Select as work scenario Freelancer Team. In this scenario there is a team sharing the translation work. 5. Add the actors for executing the localization activity. 268 B.1. USER GUIDE FOR LOCALIZATION MANAGERS 6. Select the users for executing both translation and revision tasks. 7. To finish the configuration, click on Finish button and wait while the ontology project is created. 269 APPENDIX B. LOCALIZATION USER GUIDES B.1.5 Selecting Ontology Labels to be Translated. 1. In the Ontology Navigator, select the classes, object properties, or data properties to be localized. 2. Send the selected ontology labels to translation. 3. The status of the selected labels is changed to For Translation and the labels are disabled. 270 B.2. USER GUIDE FOR TRANSLATORS B.2 User Guide for Translators This user guide demonstrates how to translate automatically ontology labels. The guide is divided into two parts. The first part demonstrates how to login in the system as Translator. The second part contains a complete reference describing the whole process to translate the ontology labels sent by the Localization Manager. B.2.1 Setting-up Translation Preferences. 1. Start Neon Toolkit. 2. Select Ontology Localization Preferences in the Window/Preferences menu. 3. Input the IP address where the Localization Server is executing (see step 8 in the Localization Server installation) and test the connection with the Localization Server. 4. Read the email for obtaining the login and password of the translator account. 271 APPENDIX B. LOCALIZATION USER GUIDES 5. Fill the login and password fields in the Ontology Localization Preferences window and then click on Login button. B.2.2 Translating Ontology Labels. 1. In the Ontology Navigator, select the classes, object properties, or data properties to be translated. 272 B.2. USER GUIDE FOR TRANSLATORS 2. Click on Translator icon to translate automatically the selected labels and wait while the labels are translated. 3. Select the more appropriate translations from the list of translations obtained by the system. 4. Send the ontology labels to Revision status. 273 APPENDIX B. LOCALIZATION USER GUIDES 5. The status is changed to For Review. B.3 User Guide for Reviewers This user guide demonstrates how to review the translations sent by the Translator(s). The guide is divided into two parts. The first part demonstrates how to login in the system as Reviewer. The second part describes the whole process to edit the translations that contain errors. B.3.1 Setting-up Revision Preferences. 1. Start Neon Toolkit. 2. Select the Ontology Localization Preferences in the Window/Preferences menu. 3. Input the IP address where the Localization Server is executing (see step 8 in the Localization Server installation) and test the connection with the Localization Server. 274 B.3. USER GUIDE FOR REVIEWERS 4. Read the email for obtaining the login and password of the reviewer account. 5. Fill the login and password fields in the Ontology Localization Preferences window and then click on Login button. 275 APPENDIX B. LOCALIZATION USER GUIDES B.3.2 Reviewing Translations. 1. In the Localization View, review the ontology labels sent by the Translator and edit the translations that need to be corrected. 2. Send the revised labels to Complete status. 276