Full text
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 511 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: Artificial Intelligence Translation Approaches for Endangered Language Preservation and Revitalization Shahad Alaa Hamza [email protected], [email protected] Abstract As globalization accelerates, endangered languages face increasing vulnerability from dominant world languages. This paper investigates how artificial intelligence (AI) technologies, particularly neural machine translation (NMT), can support the preservation and revitalization of endangered languages. The study examines AI translation technologies including neural machine translation, transfer learning techniques, and multilingual models that facilitate bidirectional translation between endangered and major world languages. It highlights successful applications of AI translation in creating parallel corpora, bilingual dictionaries, and cross-linguistic educational resources. The paper addresses critical challenges inherent to endangered language translation: severe data scarcity, lack of standardized orthography, complex morphological systems, and the imperative of preserving cultural nuance. Through analysis of quality assessment metrics and community-based evaluation approaches, this study emphasizes the essential role of human-AI collaboration in translation workflows. Findings indicate that while AI translation methods offer promising pathways for language preservation, success requires culturally sensitive model development, appropriate quality standards, and most critically, community ownership of both processes and resources to ensure meaningful and sustainable outcomes. 1. Introduction The domain of computational linguistics lies at the crossroads where technology can meet the cultural imperative to preserve speech HWs. Approximately 7,000 languages are spoken around the globe, but linguistic diversity is seriously endangered. Half of all the world's languages will disappear by 2100, and most endangered dialects are spoken by fewer than a thousand people (Haokip, 2022). The linguistic crisis has great influence on the evolution of translation technology, for most traditional machine translation technologies are based on a few successful language pairs and powerful languages, resulting in paralysis of the vast majority of world languages. English’s worldwide distribution has resulted in imbalances of translation. Existing translation pipelines only pull knowledge from endangered languages to top languages and do not support a twoway transference of expressions, which is crucial for the daily use and revitalization of language (Chen & AbdulMageed, 2022). This reverse pattern mirrors a given power relations in which minority languages are the target of study rather than simply being used to communicate in multilingual settings.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 512 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: Neural machine translation as hope and challenge for endangered languages. Despite being able to achieve near-human parity in NMT high-resource language pairs, such systems are predicated on large parallel corpora which is typically millions of aligned sentences (Ranathunga et al., 2021). Many endangered languages do not have these tools, resulting in a problem for translation researchers known as the "low-resource translation problem". The situation calls for a rethinking of translation modalities beyond DIT approaches to novel methods that work well with scarce bilingual resources. Recent advancements in technology can provide solutions via transfer learning, zero-shot translation and multilingual models. transfer learning exploits the knowledge from high resource language pairs and handles low-resource translation (Artetxe et al., 2018; Edunov et al., 2018), zero-shot systems aim at translating between a source-target that does not share any parallel data.The latter has always been considered in a completely unsupervised setting where no parallel data is available for training (Zoph et al., 2016; Kocmi and Bojar, 2018). These methods are especially useful for endangered languages, in which the corpus-based techniques cannot be applied due to the limitation of the available data. Yet language endangerment translation challenges are not limited to data availability alone. There are many endangered languages with complex morphology, difficult grammatical constructions and no direct translation to the dominant language. The absence of standardised orthographies adds further complexity to the development of such systems, leading to translation processing frameworks that operate from oral data and include speech recognition and synthesis facilities. Translation for cultural preservation poses different questions from those of commercial translations. Endangered language translation is unlike commercial or technical translation, where functional equivalence often "does the trick"; it requires cultural nuance, shared traditional knowledge systems and community-specific vocabulary that do not have a ready equivalent in target languages. This imposes a challenge that goes beyond the simplistic path is from lexical and syntactic mapping but rather into cultural, concept preservation. Endangered languages in ILT review ILTs National Flagship Project methodology poses some specific challenges to critically reviewing research on endangered languages. Attention scores) are comparison-based Ones (e.g. statistics such as BLEU scores arise by this means), and large enough evaluation sets of the endangered languages don’t exist. Automatic quality evaluation must account for language variation and dialect diversity, while not being too reliant on reference corpora (Papineni et al., 2002). Good translation to endangered languages, in addition, cannot simply be assessed in terms of linguistic accuracy but also cultural availability and reception within community. Mutual intelligibility is the other side of the coin of translation in endangered language but also poorly attended to. And just as documentation is about translating endangered languages into major ones, we also need the opposite – we need to translate present-day information and educational materials (in those major languages) and everything that each of us knows because we travelled or whatever, which after all (good reality!) adds up to quite a lot! into these vulnerable languages, for community development. It's this kind of two ways like process, that makes the translation system have to be able to tackle effectively so many different kinds of text material all the way from let's say orally transmitted traditions down to modern technical vocabulary.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 513 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: Community ownership and control (Bergman 2012; Brann 1999 [1995]; Thieberger, Persohn & Schnoebelen in press) have become fundamental principles in ethical endangered language work concerning translation technologies. Systems developed in isolation from and without the control of communities are liable to reproduce colonial exploitation, where academic or commercial enterprises profit from indigenous linguistic knowledge with little benefit reaching communities themselves. Successful approaches to other initiatives involve communitysanctioned design processes so that language users retain ownership of their translation technologies and inform development based on what is important to them in terms of both priorities and cultural values. In this paper, we detail a wide range of complex challenges in the unique space that lies at the confluence of AI and endangered language translation and consider both technical needs and cultural imperatives for developing translation systems that are ethical and effective. It can help to inform us how existing technologies might be adapted to endangered language contexts, as well as what the current barriers are and the need for a tradeoff between AI development on the one hand and community autonomy and cultural integrity on the other. 2. Literature Review 2.1 Theoretical Foundations The translation of endangered languages approaches in this paper are informed by the literature from translation studies, computational linguistics and indigenous ancestry. Skopos theory, which puts focus on what a translation is for rather than on strict equivalence becomes especially relevant since endangered language translation has diverse purposes preservation, documentation, and revitalization each with its quality issues. In contrast to commercial translation sector, where the communicative needs of present-day speakers determine translation objectives, endangered language work has an obligation to weigh preservation imperatives against use at time of production. Dynamic equivalence, first used in bible translation, recognises significant cultural differences between source and target languages. Nevertheless, some of the scholarship is critical about cross‐applying Western translation theories to Indigenous contexts and calls for interpretations that are culturally sensitive and align with traditional Native epistemologies and methodologies. In endangered language preservation, the trade-off between adequacy (maintain source language structures and concepts) and fluency (generate understandable target text) must be negotiated with reference to communities' goals rather than externalized standards. 2.2 Neural Machine Translation for Low-Resource Languages Neural machine translation has transformed translation capabilities for well-resourced language pairs but presents significant challenges for endangered languages. State-of-the-art NMT models typically require millions of parallel sentence pairs for training an
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 514 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: insurmountable barrier for most endangered languages, which may have fewer than 1,000 documented sentences in total (Ranathunga et al., 2021). Recent research has explored multiple approaches to address data scarcity. Transfer learning methods to leverage knowledge from high-resource language pairs to improve low-resource translation. Zoph et al. (2016) demonstrated that training a "parent" model on high-resource language pairs, then transferring parameters to a "child" model for low-resource pairs, yielded average improvements of 5.6 BLEU points across four low-resource language pairs. Nguyen and Chiang (2017) extended this work by exploiting source vocabulary overlapping through Byte Pair Encoding (BPE), achieving improvements up to 4.3 BLEU when combining transfer learning with stronger BPE baselines. More recent work has refined these approaches. Gao et al. (2024) proposed a two-step finetuning framework that first adjusts parent model parameters to fit the child language using source data, then transfers adjusted parameters with a distillation loss for efficient optimization, demonstrating significant improvements across five low-resource language pairs. Kocmi and Bojar (2018) showed that even "trivial" transfer learning simply continuing training on a lowresource pair after training on a high-resource pair produces significant improvements, even Another strong technique for exploiting monolingual data is back-translation. Sennrich, Haddow and Birch (2016) showed that training on synthetic parallel data using monolingual target language references with automatic backtranslation leads to large improvements: 2.8-3.7 BLEU for English German and 2.1-3.4 BLEU for Turkish English. Pang et al. (2024) systematically analyzed p(the effectiveness of back-translation, pre-training and multi-task learning in achieving success in lowresource NMT, showing that the combination of these techniques yields consistent improvement across seven translation directions (Figure 1). In the case of endangered languages, multilingual NMT is also advantageous because it facilitates knowledge transfer across typologically similar languages. Yet such models have to be designed with care as regards linguistic diversity in order not to drown the representations of minority languages into that of dominant ones. Figure 1: The Data Scarcity Challenge Parallel Corpus Availability by Language Type.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 515 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: 2.3 Quality Assessment and Evaluation Quality assessment of translation for endangered languages needs specific evaluation frameworks that take into account both peculiar linguistic traits and contextual aspects. Classical automatic metrics (BLEU, METEOR, ROUGE) rely on reference translations and large-scale corpora that can often not be found for endangered languages. Despite its widespread adoption, BLEU suffers from severe limitations insensitivity to grammaticality or well-formedness of reference sentences surface-level string matches are considered equivalent regardless of meaning and low correlation with human judgment for morphologically rich languages (Papineni et al., 2002). The human evaluation of endangered languages has further complexities. Evaluators may not have sufficient numbers of fluent bilinguals to complete evaluations or community members may not be technically trained in translation for systematic quality control. This requires locally adapted evaluation protocols, which should include training for local speakers and consider cultural norms on language use and assessment. Culturally congruent indicators of quality as a research trend is promising. Traditional standards fluency, sufficiency, naturalness will not necessarily reflect the cultural and spiritual aspects Indigenous communities view as important to good translation. Some communities value authenticity above fluent language but others attach cultural protocols to sacred or ceremonial translation that demand specialist processes of assessment. 2.4 Parallel Corpus Development and Resource Creation Building parallel corpora for endangered languages should employ alternative approaches that can address the scarcity of data at a minimum cost and further allow collecting in future sustainable resources. For languages with few professionally translated texts, the construction of corpora often involves elicitation work with bilingual speakers or the collection of natural occurring bilingual data. Crowdsourcing methods have been found to be successful in some endangered languages, but cultural protocols and consent requirements need to be considered. The trick, of course, is to balance the efficient collection of data as a product with ethical community involvement and fair payment for linguistic expertise. The Te Hiku model is an example of best practice in community-led development of such a corpus for Māori. The Kaitiakitanga license, which the organisation created to assert that data is not owned but looked after under principles of guardianship and have any benefits return to sources (Mahelona & Jones, 2022). The Kōrero Māori campaign drew over 2,500 participants providing more than 300 hours of transcribed speech data in just 10 days--all under the guiding constraint of the Kaitiakitanga license to ensure that data serves the benefit of Māori communities.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 516 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: 2.5 Community-Controlled Development and Data Sovereignty There is an increasing focus on community-based translation development that situates in principles of Indigenous self determination and cultural protocols. The development of Māorilanguage technology by Te Hiku Media is an example of community-led efforts, where communities own their language data and cultural values shape technological innovation. On the other hand, the Kaitiakitanga license does not allow use of data that is in contravention of cultural protocols, human rights or community values such as surveillance, discrimination and exploitation for commercial purposes without permission (Jones & Mahelona 2022). Indigenous data sovereignty has become a vital ethical framework in the development of AI. The First Nations Information Governance Centre's OCAP principles/ownership, control, access and possession offers a set of basic rules that are now increasingly accepted as fundamental to addressing digital colonialism. New guidelines from UNESCO emphasize the need for indigenous peoples to own data which is collected, analysed and used by AI development in ways that respect specific forms of indigenous evidence as well as indigenous rights (UNESCO, 2023). The idea of “translation sovereignty” indicates translation priorities, quality needs and usecases for translations should be made by Indigenous communities themselves rather than external researchers or technologists. This departs from a framework where models (and emic concepts) are in control, and the com munities who use them are its passengers: communities become deciders more than subjects, as illustrated in Figure 2. Figure 2: Community-Controlled Development Framework.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 517 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: 2.6 Cultural and Contextual Translation Challenges Culture-bound concepts are difficult to translate for endangered languages, as source languages are likely to contain many concepts and ways of thinking not available in other world cultures. Culturally specific concepts like traditional ecological knowledge, kinship terminology, and language of past practices highlight places where a meaning rather than literal translation or one-size-fits-all rule is called for in order to make output culturally meaningful. Many endangered languages also have particular temporality and aspectual systems that encode time, evidentiality and point of view differently from larger languages. For translation, preserving these grammatical distinctions is crucial and may require special training data or evaluation methods. 2.7 Technological Infrastructure and Deployment Endangered language translation technology needs to be designed with the material and technological (mis)agency of Indigenous communities in mind. For many places, such networks are simply not sufficient: They are without access to a stable internet connection, high-performance computer systems or staff trained in creating and sustaining complex technologies. Translation technologies modified for offline use, low resource consumption and maintainable infrastructure used in community settings. Mobile firstAdLarke said mobile-first design was increasingly important because cell phone technology is often more readily available than desktop computers in Indigenous communities. Mobile translation application development issues include user interface design, being data efficient and support for offline operation to aspects that are often left out of commercial translation applications. Cloud based solutions have benefits but also data sovereignty worries. Although cloud solutions provide access to advanced translation tools without the need for local infrastructure, they can threaten community control of language resources. Other communities may instead wish a locally hosted solution preserving data ownership and processing control in community hands. 2.8 Integration with Language Revitalization Technology for translation is being integrated with more generalized programs of linguistic revitalization well outside the realms of just preservation. Language learning apps, community education courses, and efforts for the intergenerational transmission would benefit from needing pedagogically-effective translation tools, rather than just linguistically-accurate ones. Two-way language-translation systems can encourage both preservation (of endangered languages in the wider languages) and revitalisation (of contemporary information into local ones). Efficient bidirectional systems need to consider different issues of technological and cultural for each translation direction. 2.9 Case Studies in Endangered Language NMT State of the art Recent research indicate developments in translation of endangered languages under various conditions. Chen and Abdul-Mageed (2022) doubled the performance on some of South American Indigenous languages with respect to state-of-the-art via multilingual
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 518 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: transfer learning, this is, while still preserving low-resource constraints than through data augmentation. Alnajjar et al. (2023) built NMT systems for both Moksha and Erzya by ingeniously resolving data sparsity issues through synthetic data creation using statistical machine translation with the aid of transfer learning from related Uralic languages. Hämäläinen & Rueter (2019) used related methodology also for Skolt Sami and North Sami, publishing the discovered cognates via the Online Dictionary of Uralic Languages. Activity on African languages has increased NMT coverage: Ezeani et al. (2020) released the first attention-based high performance English Igbo translation system that BLEU estimates 70% translation accuracy from transfer learning on pretrained models of MarianNMT. These case studies show that while NMT for endangered languages remains a substantial technical challenge, novel techniques that draw upon transfer learning, synthetic data generation, and multilingual modeling can result in translation quality that is at least meaningful even in resource-constrained scenarios. 2.10 Synthesis and Research Gaps So thus far that the literature tells us, endangered language and even headed writing (top-tobottom-roofer scraping) seem to have good hope for survival under such AI-enabled translation—not in the traditional sense of machine translation though. Techniques like transfer learning and back-translation allow for high quality translation with minuscule parallel data; community-controlled frameworks, Te Hiku Media’s Kaitiakitanga license are beacons of hope for ethically responsible technology development. However, there remain many issues to be resolved, such as the lack of data in some lexicons, cultural imperatives for preservation measures; quality control for less-resourced languages without standard references and longterm economic sustainability for non-commercial translation systems and how “ownership” can be ensured by local community members on tools and process. Future research will also have to develop approaches for assessing quality in endangered language contexts, devise funding models that ensure ongoing technology support, and cultivate the technical capacity within Indigenous nations to drive translation technology development. What is even more critical is for the field to realize that endangered langauge translateion is not just technology, BUT a tool of Ind software_needed_ayers new media studiesigenous selfdetermination and epistemologies. Flourishing such collaborations will mean moving away from extractive research models to development based on equitable principles where Indigenous people define the priorities, gate-keep access and receive direct benefit from translation technologies built with their own linguistic knowledge. 3. Methodology This research utilised a scoping review approach along with review of ethno-tourism initiatives from '.gesturing toward' existing endangered language translation projects (Temple and Cavanagh 2012). The review analysed peer-reviewed articles, technical reports and case studies from 2016-2025 with an emphasis on neural machine translation for low-resource and endangered languages as presented in Figure 3.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 519 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: Search pathways aimed at numerous databases, such as ACL Anthology, arXiv, Google Scholar and leading conferences (ACL, EMNLP, WMT). Searches included agreed-upon combinations of the following search terms: neural machine translation, low-resource language, endangered language, indigenous languages, transfer learning, back-translation and data sovereignty/learning and community-controlled AI. Recent articles (2020-2025) were prioritized in the review, as well as fundational papers reporting key methodologies. Great consideration was placed on Indigenous-led products and publications written by or co-written with members of the language communities we served. Case study analysis concentrated on documented endangered language NMT projects, for which public information was available on methods, community engagement practices, quality and outcomes. Special attention was given to projects showing elements of community control and data sovereignty. Figure 3: NMT Translation Pipeline for Endangered Languages. Quality assessment of included literature considered: peer review status, methodological rigor, reproducibility, community involvement in research design and execution, and attention to cultural protocols and ethical considerations beyond standard research ethics requirements. 4. Results 4.1 Technical Approaches and Effectiveness Analysis of recent endangered language NMT literature reveals several effective technical approaches for addressing data scarcity: Transfer Learning: Studies consistently demonstrate that transfer learning from high-resource to low-resource language pairs yields substantial improvements. Zoph et al. (2016) reported average gains of 5.6 BLEU points, while more recent work by Gao et al. (2024) achieved further improvements through refined two-step fine-tuning processes. Transfer learning proves
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 526 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: 5.3 Balancing Preservation and Revitalization Endangered language translation must serve two purposes: it must be part of documentation for preservation, and at the same time uphold revitalization as living languages. These goals sometimes create tension. Properties dedicated to preservation may be focused on archival integrity and exhaustive documentation, even if it means a slower development cycle. Focus of Revitalization The focus of revitalization is on practical utility, development of modern terminologies and quick creation of usable systems. The best solutions combine both attitudes through two-way help-translation techniques that facilitate documentational as well as contemporary use. Yet, building such systems will require to carefully consult the community to establish priorities among what is going to be an uncertain level of acceptable quality, trade-offs in resources between preservation and revivification. 5.4 Cultural Appropriateness and Technical Quality This study underscores noted tensions between generic technical quality measures and cultural relevance. BLEU scores and similar metrics measure some aspects of translation quality, but do not take into account cultural preservation necessities that are central to endangered language work. A translation that is technically ‘better’ according to metrics may be culturally problematic or damaging if it corrupts concepts, disrespects protocols (with respect to content), or neglects the continuing value of dialectal diversity among communities. This tension requires us to create new frameworks for assessing quality that can balance accurate language with cultural integrity. They should include community-based indicators of quality, adherence to cultural protocol, evaluation by qualified members of the community and recognition that standards of quality may vary from one speech community to another for languages in similar states of endangerment. 5.5 Economic Sustainability Economic viability is a vital problem that has not been solved well in literature. The vast majority of endangered language translation projects rely on academic research funds, government support, or volunteer work and all are inherently precarious. Unlike commercial translation for profit-driven markets, endangered language translation serves user communities that generally do not have the economic capacity to support fundamental research on technology development. This phenomenon demands creative investment strategies and infrastructural backing. The opportunities may include government funding that emphasises language preservation as a cultural heritage responsibility, philanthropy that supports community-led projects, and institutional practice that allows technology maintenance in the long term without ongoing research dollars. Without technical success won É even the best innovation is going to left in despair when initial financing runs out.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 527 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: 5.6 Capacity Building and Indigenous Technology Leadership The literature increasingly focuses on the capacity and capability of Indigenous technical expertise rather than research-driven by outsiders. Others such as the First Languages AI Reality project and Lakota AI Code Camp take this training to local communities; aiming to train those within Indigenous communities in AI and software development purposefully for language technology. This transition to Indigenous tech leadership is crucial for long-term sustainability and cultural relevance. Indigenous technologists fill the cultural competence void that exists among external researchers, while providing a counter to conventional technology development paradigms that have under-represented Indigenous influence in the past. But developing Indigenous technical capacity would mean continued investment in education, training and career pathways a longterm commitment beyond the commonly short timescales of research projects. 5.7 Ethical Considerations and Potential Harms There are promising prospects for AI translation and endangered languages, but the possible harms call for great caution: Disinformation Hazard: Low-accuracy systems can produce incorrect translations which spread due to error chains throughout communities, damaging rather than aiding in language transmission. The risk is underscored by documented cases where AI produced language learning materials with false content. False Promises: Exaggerating the power of AI could result in false promises and raise unrealistic expectations, causing communities to direct scarce resources into technologies that end up not being effective instead of focusing on what we know works for revitalization (such as immersion education). Data Utilization: Failing deep considerations of data sovereignty, Indigenous language data could be exploited by corporate entities or researchers similarly to how it has been historically and continue to erode in digital space. Cultural erosion: Technologies that were developed with little awareness of culture can perpetuate stereotypes, smooth over dialectal differences or portray cultural concepts in ways that are harmful to language communities. These risks require thoughtful ethical frameworks, honest communication about what is and isn’t possible, and a commitment to prioritizing community control in development processes. 5.8 Future Directions There are several important issues to be developed further in the field: Technical Innovation: Additional reduction in the required amount of data with methods such as few-shot and zero-shot learning, designing algorithms that are specifically tailored to the properties of rare languages (polysynthetic morphology, complex agreement systems etc), more efficient model architecture for resource-limited deployment.
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 528 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: Quality Assessment: Culturally sensitive evaluation frameworks, Tools for community-based assessment with training materials for local evaluators, field wide standards recognizing that quality assessment must be adapted to endangered language contexts. Infrastructure Development: Developing technological infrastructure that is sustainable and controlled by the community locally hosted systems, mobile first applications and tools that can be maintained by communities without requiring technical expertise. Build: Indigenizing technology education, job paths for Indigenous artificial intelligence professionals, and supporting Indigenous innovation. Policy and Funding: Developing policy frameworks that acknowledge Indigenous data sovereignty, sustaining funding mechanisms for language technology that is communitycontrolled, and putting in place institutional supports to ensure the means for long term sustainability beyond initial research projects. Ethical Frameworks: Continued iteration of ethical guidelines for endangered language technology development, standards for community engagement and consent, as well as accountability mechanisms to ensure the development of AI is in service to the interests of communities. 6. Conclusion This survey has shown that artificial intelligence in general, and neural machine translation in particular, holds much promise for helping to save endangered languages. Recent technological advancements (e.g., transfer learning, back-translation, and multilingual modeling) have dramatically decreased the data requirement of effective translation systems such that NMT is now feasible for languages with extremely minimal parallel data. But motivation and ability aren't enough. Such successes will not come without fundamental changes in how development of language technologies are done. Community control has to be at the center, not the margin. Data sovereignty models such as Te Hiku Media's Kaitiakitanga licence are offering new ways for Indigenous communities to have ownership and control over technologies developed using their linguistic resources. Quality metrics need to move beyond traditional markers and include culturally congruent care and communitydetermined systems for care. Development models need to be built on building the capacity within indigenous communities and not maintaining a reliance in external researcher leadership. The balance between technical innovation and cultural preservation is a tightrope walk. Although there are times when AI translation may be valuable to endangered languages, no amount or quality of machine learning for language technology will serve endangered languages like human transfer and normal use, complementary media that it may be. Translation technology is a single tool, most effective as a part of larger community-led language revitalization efforts. Economic, technical (infrastructural) limitations and the need for capacity building are continuing problems, which require renewed attention and resources. We have to go beyond
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 529 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: short-term research orchard projects and focus on long term institutional support that keeps technology evolving. In the end, endangered language translation technology development must be situated within larger struggles for Indigenous self-determination and cultural survival. Technical success is moot if communities do not own technologies or benefits mainly accrue to outside researchers and institutions. The way forward must be through a model of collaborative partnership in which Indigenous priorities are centred, Indigenous methodologies are respected and AI development for endangered languages serves the broader goals of the community rather than reinforcing patterns of colonial extraction and exploitation. References Alnajjar, K., Hämäläinen, M., Chen, W., & Alnajjar, K. (2023). Neural machine translation for low-resource languages: A survey and case studies. ACL Anthology. Chen, W., & Abdul-Mageed, M. (2022). Improving neural machine translation of indigenous languages with multilingual transfer learning. arXiv preprint arXiv:2205.06993. Ezeani, I., Hepple, M., Onyenwe, I., & Uchechukwu, C. (2020). Igbo-English machine translation: An evaluation benchmark. arXiv preprint arXiv:2004.00648. Gao, Y., Hou, F., & Wang, R. (2024). A novel two-step fine-tuning framework for transfer learning in low-resource neural machine translation. In Findings of the Association for Computational Linguistics: NAACL 2024 (pp. 3214-3224). Hämäläinen, M., & Rueter, J. (2019). Finding Sami cognates with a character-based NMT approach. In Proceedings of the Workshop on Computational Methods for Endangered Languages. Haokip, T. (2022). Artificial intelligence and endangered languages. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4212504 Jones, P., & Mahelona, K. (2022). Kaitiakitanga license and data sovereignty. Te Hiku Media. Retrieved from https://tehiku.nz Kocmi, T., & Bojar, O. (2018). Trivial transfer learning for low-resource neural machine translation. In Proceedings of the Third Conference on Machine Translation (pp. 244-252). Nguyen, T. Q., & Chiang, D. (2017). Transfer learning across low-resource, related languages for neural machine translation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (pp. 296-301). Pang, J., Yang, B., Wong, D. F., Wan, Y., Liu, D., Chao, L. S., & Xie, J. (2024). Rethinking the exploitation of monolingual data for low-resource neural machine translation. Computational Linguistics, 50(1), 25-67. https://doi.org/10.1162/coli_a_00496 Papineni, K., Roukos, S., Ward, T., & Zhu, W. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311-318).
International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 11 November 2025 Received: 2 November 2025 530 Revised: 17 November 2025 Accepted: 27 November 2025 5authors 202 Copyright .17737626ZENODO/10.5281/ORG.DOI://HTTPSDOI: Ranathunga, S., Lee, E. S. A., Prifti Skenduli, M., Shekhar, R., Alam, M., & Kaur, R. (2021). Neural machine translation for low-resource languages: A survey. arXiv preprint arXiv:2106.15115. Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 86-96). UNESCO. (2023). Indigenous people-centered artificial intelligence: Perspectives from Latin America and the Caribbean. UNESCO. Retrieved from https://www.unesco.org/en/articles/new-report-and-guidelines-indigenous-data-sovereigntyartificial-intelligence-developments Zoph, B., Yuret, D., May, J., & Knight, K. (2016). Transfer learning for low-resource neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 1568-1575)