Lightning Proceedings of the 10th International Workshop on Computational Linguistics for Uralic Languages
Abstract
Lightning Proceedings of the 10th International Workshop on Computational Linguistics for Uralic Languages
Full text
Lightning Proceedings of the 10th International Workshop on Computational Linguistics for Uralic Languages Mika Hämäläinen & Khalid Alnajjar (eds.)
Lightning Proceedings of the 10th International Workshop on Computational Linguistics for Uralic Languages Mika Hämäläinen & Khalid Alnajjar (eds.) ISBN 978-952-65595-2-0 Fly for Points Helsinki, Finland 2025 CC BY
Table of Contents UralicMCP: Turning LLMs into Experts in Endangered Languages with MCP Mika Hämäläinen, Jack Rueter 1-7 Diversity of the Saami languages in light of typological datasets Ilya Egorov 8-13 From News Archives to Neural TTS: First steps of the Olonets Karelian TTS project Katri Hiovain-Asikainen, Maria Kuismin 14-18 From Toki Pona to Uralic: A GrammarConstrained Pipeline for Low-Resource Language Generation Artur Roos 19-27
Did Karelian Survive the Year? A Small Data Update Lev Kharlashkin 28-34 Evaluating Finnish Dialect Normalization in GPT Models with and without Reasoning Eiaki V. Morooka, Yuto Omae , Hirotaka Takahashi, Mika Hämäläinen 35-39 Mozilla Data Collective as a Platform for Data on Uralic Languages Niko Partanen, Janine Siewert 40-43 On finding mutual classifiers for nominals and verbs in the enhanced Skolt Sami orthography Jack Rueter 44-63
UralicMCP: Turning LLMs into Experts in Endangered Languages with MCP Mika Hämäläinen1 Jack Rueter2 1Metropolia University of Applied Sciences, Finland 2University of Helsinki, Finland Introduction This paper introduces a new MCP (Model Context Protocol) feature to UralicNLP (Hämäläinen, 2019) Python library, UralicMCP1. The MCP server exposes rulebased tools such as morphological analyzer, inflector, lemmatizer and dictionaries to any LLM tool that supports MCP. With the help of these tools, LLMs have a shot in completing NLP tasks in languages they know virtually nothing about. There has been a lot of research interest in the past in combining traditional rule-based methods with neural networks in the context of endangered languages (Ens et al. 2019; Wiechetek et al. 2021; Alnajjar et al. 2023). 1 https://github.com/mikahama/uralicNLP/wiki/UralicMCP 1
Last year, several researchers studied the use of LLMs for NLP tasks in several endangered Uralic languages with a varying degree of success (Pirinen, 2024; Partanen, 2024; Hämäläinen, 2024a). The limiting factor was the lack of understanding of the endangered languages in question that the LLMs exhibited. As I already stated last year (Hämäläinen, 2024b), I firmly believe that LLMs will be the future way of doing NLP for endangered languages as well. They seem to know a lot about different languages – frankly more than any human ever will – but they just don’t understand languages they have not sufficiently seen in their training data. But what if we gave them tools to interpret languages unknown to them through an MCP server? Setting up MCP Building an MCP compatible tool is a matter of exposing the core features of UralicNLP through an MCP server as tools. We implemented the following tools: analyze_word, morphological_segmentation, inflect_word, lemmatize, dictionary_lookup and list_supported_languages. Each tool has a lengthy prompt-like docstring that explains the LLM how to use the tool. We had to make some changes to UralicNLP to make it easier for LLMs to use the MCP server. Previously, UralicNLP output an error if language models were not present for a given language. Now, instead of an error, UralicNLP will automatically download any missing language models. We also had to make dictionaries faster and more sensical. Previously, UralicNLP dictionaries were quite messy JSON dictionaries, and there was a separate one for each supported language. 2
The best way to combine all these dictionaries into one fast dictionary was to build an FST out of them, as UralicNLP already supports FSTs and this solution does not require any additional dependencies. We combined all the JSON dictionaries into a massive LEXC-file that essentially maps each lemma with a language tag to its translations in another language. Here is an example of one line of the LEXC: myv_васта:eng_husband #; This line also exists in reverse, so that the Erzya (myv) translation can be found with the English word as well. Experiments One of the nicest aspects of MCP is that it is an open standard. Therefore, users are not vendor-locked to a specific LLM by a specific company, but instead UralicMCP can be equally well used through ChatGPT and an open-source model. Here we show some of our early experiments we did using Jan2 desktop application. First, we try one of Jan’s local models, namely Jan-v14B-Q4_K_M. The model is successful at doing small tasks using UralicMCP, but it gets lost in a larger task such as translation of a complete sentence. Figure 1 shows an example of the Jan model successfully using UralicMCP to analyze an Erzya word. 2 https://www.jan.ai/ 3
Figure 1: Jan’s local model using UralicMCP Translation is, however, possible if we use Gemini 1.5 Flash instead of a local model (see Figure 2). Figure 2: Erzya translation task and Gemini 1.5 Flash 4
After several steps of using UralicMCP’s tools, Gemini manages to translate the input Erzya sentence correctly into English (Figure 3). This is remarkable since this task would have been otherwise impossible for the model. Figure 3: Gemini gets the translation right after using UralicMCP several times. Conclusions MCP has a huge potential for endangered languages as well. This is the first time I feel that my early visions of combining rule-based tools with neural networks have actual practical value. 5
Figure 4. MCA plot. The patterns of (dis)similarity visualized in the Figures 2–4 largely correspond to geographical distribution of the Saami languages. Some of the observed similarities and differences can be explained by contact phenomena both within the Saami group and from outside. The proximity of Kildin, Akkala, and Ter Saami likely reflects Russian influence; that between North and Aanaar Saami reflects the influence of North on Aanaar as well as Finnish influence on both; while the greater distance between the western varieties (from Lule westwards) and North Saami and the eastern languages may be attributed to more intensive Scandinavian influence on the former group. In discussing the results, I will briefly compare the observed patterns with the distribution of phonological innovations and the lexicostatistical trees available so far. In the course of further research, a more detailed examination of innovations in historical phonology and morphology, as well as a dedicated lexicostatistical study, will be undertaken. 12
References Donohue, M., Musgrave, S., Whiting, B., & Søren, W. 2011. Typological feature analysis models linguistic geography. Language. 87(2). 369–383. doi:10.1353/lan.2011.0033. Norvik, M., Jing, Y., Dunn, M., Forker, D., Honkola, T., Klumpp, G., Kowalik, R., et al. 2022. Uralic typology in the light of a new comprehensive dataset. Journal of Uralic Linguistics 1(1). 4–42. doi:10.1075/jul.00002.nor. Häkkinen, J., & Piha, M. 2023. Kantasaamesta eteläkantasaameen, osa 2: Äännehistorian todisteita eteläsaamen varhaisesta eriytymisestä. Sananjalka, 65(65). 7–30. doi:10.30673/sja.115746 Sammallahti, P. 1998. The Saami Languages: An Introduction. Kárášjohka: Davvi Girji. Skirgård, H., Haynie, H. J., Blasi, D. E., Hammarström, H., Collins, J., Latarche, J. J., Lesage, J., et al. 2023. Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss. Science Advances 9 (16). eadg6175. doi:10.1126/sciadv.adg6175. 13
From News Archives to Neural TTS: First steps of the Olonets Karelian TTS project Katri Hiovain-Asikainen1,2, Maria Kuismin3 1UiT The Arctic University of Norway 2University of Helsinki 3University of Eastern Finland Introduction This talk presents the first demonstration of the Olonets Karelian Text-to-Speech (TTS) project, a collaborative effort between Divvun (UiT), YLE, Finnish universities (Helsinki and Eastern Finland) and Readit AI aimed at developing a high-quality speech synthesis system for the Karelian language, a Finno-Ugric minority language spoken primarily in Finland and Russia. The project contributes both to ongoing efforts in digital language revitalisation and to the broader field of neural Text-to-speech (TTS) for low-resource languages. We describe our methodology, dataset preparation, model training, and discuss future evaluation of the resulting synthetic voice, highlighting the challenges and opportunities that arise when working with minority languages and limited data. 14
Data and Methods Our speech corpus consists of approximately 11.7 hours of existing Karelian single-speaker news recordings, provided by YLE. These recordings form one of the most substantial publicly available audio resources in Karelian and offer a unique opportunity to explore modern deep-learning-based speech synthesis for the language. Since the corpus was not originally designed for TTS purposes, it has some variation in recording quality and background noise since the recordings were done during several years. Careful pre-processing was therefore required to make the data suitable for training. This included segmenting long-form (~ 5 minutes) news recordings into shorter utterances, performing noise reduction and amplitude normalization and aligning text transcriptions to audio segments using a force-aligner tool WebMAUS [1,2]. We chose FastPitch [3] as our core synthesis architecture, a neural transformer-based model designed for efficient, parallelized speech synthesis with controllable pitch and duration features. FastPitch has proven to be effective in scenarios with limited data availability (see [4,5,6] for the Sámi languages), making it an appropriate choice for our Karelian dataset as well. We trained the model from scratch using the processed YLE corpus and integrated a neural vocoder to generate high-quality waveform outputs. 15
Implications and Future work From a technical perspective, the Karelian TTS project addresses a key research question: to what extent can modern neural architectures like FastPitch deliver intelligible and natural-sounding speech in an endangered language with less than 12 hours of training data? The project also has broader sociolinguistic implications. Karelian is classified as a severely endangered language [7], with limited digital presence and few existing technological resources. The availability of a natural-sounding TTS voice can serve multiple revitalisation and accessibility goals: it enables the creation of digital reading tools, audiobooks, and interactive language learning applications, as well as enhancing the inclusivity of Karelian content in broadcasting, online media and universal accessibility. In addition to the core TTS component, the project lays the groundwork for future research in Karelian speech technologies, including automatic speech recognition (ASR), prosody modelling, and voice cloning for multiple dialects. In future, we could extend the system with additional recordings from other Karelian varieties (e. g. Viena and Tver), increasing both the linguistic diversity and robustness of the model. Moreover, the collaboration between Divvun, YLE, and Readit AI provides a model for cross-sector partnerships that combine academic, media, and industrial expertise in digital language preservation. To our knowledge, this is one of the first-ever neural TTS systems for Karelian, marking a milestone in the language’s digital evolution. The demo we present illustrates that even for small, endangered languages with highly limited audio data, state-of-the-art transformer-based models can produce convincing speech synthesis results. Beyond Karelian, the methods and findings have relevance for other minority and low-resource languages in the Finno-Ugric family and beyond. 16
References [1] Schiel, F. (1999). Automatic phonetic transcription of non-prompted speech. Proceedings of the ICPhS, 607–610. [2] Kisler, T., Reichel, U., & Schiel, F. (2017). Multilingual processing of speech via web services. Computer Speech & Language, 45, 326–347. [3] Łańcucki, A. (2021). FastPitch: Parallel text-to-speech with pitch prediction. ICASSP 2021 – IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6588–6592. IEEE. [4] Hiovain-Asikainen, K., & Moshagen, S. (2022). Building open-source speech technology for low-resource minority languages with Sámi as an example: Tools, methods and experiments. Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages, 169–175. [5] Hiovain-Asikainen, K., & De la Rosa, J. (2023). Developing TTS and ASR for Lule and North Sámi languages. Proceedings of the 2nd Annual Meeting of the ELRA/ISCA SIG on Under-Resourced Languages (SIGUL 2023), 48–52. [6] Hiovain-Asikainen, K., & Suni, A. (2025). Does multilingual and multi-speaker modeling improve low-resource TTS? Experiments on Sámi languages. Proceedings of SSW 2025, 196–201. [7] United Nations Educational, Scientific and Cultural Organisation. (2021). UNESCO Atlas of the World’s 17
Languages in Danger. https://unesdoc.unesco.org/ark:/48223/pf0000187026 18
From Toki Pona to Uralic: A GrammarConstrained Pipeline for Low-Resource Language Generation Artur Roos Metropolia University of Applied Sciences, Finland Introduction Large Language Models (LLMs) have demonstrated impressive multilingual abilities, but their performance often degrades for low-resource languages, particularly for those with rich morphology such as members of the Uralic family. Despite the growing interest in adapting LLMs to under-resourced languages [1][2][3] challenges remain due to the scarcity of high-quality data and the linguistic complexity introduced by morphological richness [4]. To address this gap, we propose a syntaxassisted data generation and training pipeline, prototyped in the controlled setting of Toki Pona. Although Toki Pona is typologically minimalistic — characterized by a very small vocabulary, extremely simple grammar, and few phonemes — its simplicity provides a uniquely tractable environment for testing methodology without the confounding effects of large noisy corpora or massive vocabulary size [5]. Our ultimate goal is to establish a methodology that can be transferred to complex Uralic languages by incrementally re-introducing morphological, compounding, and phonotactic structure in a controlled, interpretable way, thereby bridging the gap between theoretical linguistic supervision and practical LLM training. 19
Motivation Working directly with Uralic languages often demands confronting several interlocking challenges: sparse parallel corpora, highly inflectional morphology, and limited digital resources. For instance, UralicNLP – a common NLP library for Uralic languages – relies on finite-state morphological models because many Uralic languages (e.g., Sámi, Erzya, Komi) have extremely rich inflectional paradigms [6][7]. By contrast, Toki Pona provides an unusually wellcontrolled laboratory environment. It has a very small vocabulary (only around 120-137 root words) and an extremely simple, template-friendly grammar, which makes it ideal for rapid iteration and controlled data generation [8]. Despite its minimalist design, the language exhibits behavior that pushes beyond context-free grammars: for example, some community usage shows a pattern like “X ala X” (i.e., repeating a noun phrase with the negator ALA) in a way that resembles echoic reduplication, creating non-trivial structural constraints. This contrast makes Toki Pona exceptionally well-suited for methodologically isolating the effects of explicit syntactic supervision, without the noise and scale issues that come with large, real-world corpora. Overview of the Pipeline Our pipeline contains four core components, each carefully designed to build synthetic but linguistically faithful data that guides LLM learning under conditions analogous to low-resource Uralic settings. 20
1. Syntax-Assisted Generation via Template Library We begin by constructing a tree-structured grammar and template library that encodes essential linguistic phenomena: closure structures (e.g., embedding or coordinating), valency frames (e.g., intransitive, transitive, ditransitive), modifiers and particle ordering, and allowable reduplication patterns. Using this grammar, we generate Toki Pona sentences that are syntactically valid by construction, avoiding ungrammatical or ill-formed outputs. Each generated candidate is then passed through a structural verification step (e.g., a parser or structural validator) before being admitted into the training pool, ensuring high-quality, well-formed synthetic data. 2. Parallel Corpus Generation via Machine Translation Once we have syntactically verified Toki Pona sentences, we translate them into English with existing machinetranslation (MT) tools. This yields a synthetic, parallel corpus of (Toki Pona | English) sentence pairs. We iterate this generation-translation loop until we reach a target corpus size, ensuring a sufficiently large and diverse synthetic dataset. This technique aligns with recent successful work in generating synthetic parallel corpora for low-resource languages: for example, Scaling Low-Resource MT via Synthetic Data Generation with LLMs demonstrates the utility of LLM-generated synthetic data to improve MT quality [9]. 21
Did Karelian Survive the Year? A Small Data Update Lev Kharlashkin Metropolia University of Applied Sciences, Finland Introduction In 2024 I presented a small, reproducible crawl of Karelian-language news from the OMAMEDIA portal to show that modest, regular collections can yield actionable insights for endangered-language work (Kharlashkin, 2024). Over the past year I extended the same pipeline with one change in scope: alongside the News feed, I added the long-form Articles section. The aim is practical - offer an evidence-based status check of Karelian’s public written presence and indicate where the signal is strongest. This follows broader calls to narrow the digital gap for Finno-Ugric languages (Lindgren & Kinnunen, 2021). 28
Data & Methods I crawled two Karelian-restricted sections on omamedia.ru: News and Articles (OMAMEDIA, 2025a; OMAMEDIA, 2025b). For each item I extracted title, subtitle, publisher, author (if present), date (normalized to YYYY-MM-DD), URL and full text. I produced two CSV datasets and ran consistent descriptive analyses: counts by year and month, distributions by publisher and author, and a coarse length proxy (character count) to contrast short-form news with long-form articles. Cross-set duplicates were checked by URL and a normalized (title, date) key. The pipeline follows basic guidance for ethically collecting cultural content online (Marisova, 2022) and uses lightweight corpus practices common in low-resource NLP (Lewis et al., 2020). Findings: Two Distinct Streams The News and Articles feeds behave as independent pipelines: 0 duplicates by URL or by (title, date). Short-form items drive monthly volume, while long-form items carry richer discourse and consistent bylines - an analytically useful separation. 29
Findings: Volume and Cadence (News) The News dataset contains 344 items spanning 2020-06-08 → 2025-10-24. Activity is decisively 2025-heavy: 280 items in 2025 (≈ 81% of the total), compared with 57 in 2024 - a roughly fivefold increase. The 2025 monthly cadence averages ~28 items/month, peaking in February 2025 (34). This pattern suggests sustained, organized production rather than isolated bursts and offers a concise indicator of digital vitality. Findings: Long-Form Characteristics (Articles) The Articles dataset comprises 119 items over the same window. As expected, texts are substantially longer: mean body length ≈ 4,498 characters (vs. 1,508 for news). This long-form layer offers narrative continuity, denser named-entity contexts and more stable topic development - traits useful for pedagogy, topic modeling and evaluation of downstream tools in low-resource settings. 30
Findings: Publisher Landscape Publisher concentration is a defining feature. News: Oma Mua accounts for 324 items (≈ 94%), with small contributions from Kipinä (9), carelia (7), Karjalan Sanomat (2), Kodima (1) and Oma Media (1). Articles: Oma Mua again dominates (106), followed by Kipinä (13). This concentration simplifies monitoring and collection while underscoring how much the observed Karelian stream on this portal depends on a single editorial hub. Findings: Authorship (Articles) Long-form contributions feature consistent bylines, enabling a clearer view of contemporary Karelian voices. Notable contributors include Irina Zaitseva (10), Aleksandra Lesonen (8), Natalja Sinitskaja (8), Ol’ga Ogneva (8), Nadežda Mičurova (8), Uljana Tikkanen (7), Valentina Mironova (6), Nadežda Vasiljeva (6) and Alina Gapejeva (5). This authorship signal supports stylistic analysis, author-level terminology tracking and targeted community engagement. 31
Interpretation Together, these observations present a concise picture of digital vitality on the portal in 2025. The News stream supplies breadth and tempo - regular, public-facing use that matters for visibility and for routine corpus refreshes. The Articles stream contributes depth - longer, authored texts well suited for evaluation and instruction. Read jointly, the two streams provide complementary affordances: cadence from news, context from long-form, and a practical path to track vitality over time (Väisänen et al., 2019). Practical Next Steps Maintain cadence: Continue the same crawl monthly or quarterly to produce dated deltas; regularity beats sporadic bulk updates. Broaden intake modestly: Add even two additional sources (e.g., municipal pages, community blogs, regional cultural institutions) to diversify genres and reduce single-hub dependence. Add minimal annotation: Sentence segmentation, light LID and basic named-entity tags (people/places/organizations) on a 500–1,000 item slice will materially increase reusability for teaching materials and quick prototypes. 32
Expose a simple interface: A compact viewer (filter by date, publisher, author; keyword search) enables quick discovery and feedback loops. Conclusion A year after the initial pilot, the same lightweight, reproducible pipeline reveals a clear signal: Karelian is being written online with vigor in 2025, particularly in short-form news, while the long-form layer contributes identifiable authors and richer linguistic contexts. The separation between News and Articles matters for analysis and application - from tracking monthly vitality to assembling evaluation sets and classroom materials. Routine crawls, transparent summaries and small, well-chosen annotations provide a durable way to keep pace with a living endangered language. References Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. 33
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. Kharlashkin, L. (2024). Crawling Karelian News: Creating a Dataset to Preserve Cultural Heritage. Lightning Proceedings of the 9th International Workshop on Computational Linguistics for Uralic Languages, 6–13. Lindgren, K., & Kinnunen, T. (2021). Challenges in Finno-Ugric language preservation: The digital gap in Karelian resources. Finno-Ugric Linguistics Studies, 6(2), 45–60. Lewis, J., Schuster, S., & Ho, M. (2020). Leveraging NLP for endangered languages: A systematic study on Karelian.Proceedings of Computational Linguistics, 1124–1130. Marisova, E. (2022). Ethical considerations in the digital collection of cultural heritage. Digital Humanities Quarterly,16(1), 54–68. OMAMEDIA. (2025a). News (Karelian). https://omamedia.ru/news/?language=kar OMAMEDIA. (2025b). Articles (Karelian). https://omamedia.ru/articles/?language=kar Väisänen, R., Saarinen, T., & Huhtinen, A. (2019). Language vitality and revitalization efforts: The case of Karelian.Journal of Endangered Languages, 8(3), 213–230. 34
Evaluating Finnish Dialect Normalization in GPT Models with and without Reasoning Eiaki V. Morooka1, Yuto Omae2 , Hirotaka Takahashi3 and Mika Hämäläinen1 1 Metropolia University of Applied Sciences, Finland 2Nihon University, Japan 3Tokyo City University, Japan Introduction Normalization improves the compatibility of non-standard language—such as dialectal, colloquial, or transcribed spoken Finnish—with tools designed for the standard variety. Building on earlier work showing that normalization by Partanen et al. [1], our study evaluates whether modern GPTbased models can further advance dialect-to-standard Finnish normalization. We first fine-tune a non-reasoning model to establish a baseline, and then fine-tune a reasoning-enabled model to test whether its inference capabilities lead to higher normalization accuracy. 35
Experiment 1 (Non-Reasoning) In our first experiment, we fine-tuned several open models— DeepSeek-R1 1.5B and 7B (Distill Qwen) [2], Gemma 270M and 1B [3], Poro-2 8B (Llama) [4], and Viking 7B and 13B [5]— on 34,822-sample training set from the Institute for the Languages of Finland [6] to produce normalized Finnish text. All models were trained under identical settings using LoRA (r = 16, α = 16) targeting the q_proj, k_proj, v_proj, and o_proj modules, with no hyperparameter tuning, a learning rate of 5e-5, and three epochs. Evaluation on the 7,687-sample test set (Table 1) shows that model size had some effect, but the extent of Finnish pretraining was the primary factor. The Finnish-centric Poro-2-8B outperformed all others, slightly surpassing the normalization quality reported in Dialect Text Normalization to Normative [1] (WER 5.73%). These results demonstrate that GPT-style models can achieve higher accuracy, though at the cost of additional computational resources. Experiment 2 (Reasoning) 36
In the second experiment, we attempted to create a reasoning-capable normalization model by fine-tuning DeepSeek 7B with chain-of-thought style supervision. Using DeepSeek’s official full reasoning model, we generated 3k training examples by providing each dialectal input and its normalized target and prompting the model to explain WHY the normalization was appropriate while “acting” as if it did not know the answer. Each explanation was wrapped in <think>...</think> tags, followed by the correct normalized output, and this sequence formed the fine-tuning sample. The LoRA configuration was identical to Experiment 1 (r = 16, α = 16, same target modules, learning rate, and epoch count). The resulting reasoning-augmented model was evaluated on the same test set as Experiment 1, but using a 100-sample subset to allow close inspection of CoT behavior. The results (Table 2) showed that adding chain-of-thought supervision substantially degraded normalization performance: the CoT-trained DeepSeek 7B reached a WER of 38.85%, compared with 11.17% for the same model in Experiment 1. This drop may stem from the much smaller CoT training set (an order of magnitude less data), the possibility that explicit reasoning traces introduce noise rather than useful structure for this task, or the need for reinforcement-learning–based optimization instead of pure supervised fine-tuning. 37
On finding mutual classifiers for nominals and verbs in the enhanced Skolt Sami orthography Jack Rueter University of Helsinki, Finland Abstract In this article, we begin the discussion and description of enhanced orthographic modeling for Skolt Sami. This is an attempt to gauge how far we have come in the description of this written form of the language, and it is intended to answer the question what the grammar basis is for our description. We briefly discuss five classifiers used in the categorization of Skolt Sami verbal and nominal inflection stems. Applying these classifiers, we examine the stem types exhibiting the most extensive variation in order to establish the grouping of inflections for each stem variant. We then provide two tables and inflection form lists to illustrate how the three stem vowel types align. Next, we discuss how these finding can be plotted 44
on stems with diphthongs. Finally, point out the next phase of comparing nominal and verbal stem types. Introduction The finite-state description of morphologically rich languages requires extensive work with the target language. Ideally, a team can be formed that brings the skills of a language specialist and a person fluent in finite-state description together. In reality, there might just not be enough specialists to go around. Thus, the language technologist is compelled to become more of a specialist in the target language – the language technologist might have to dig into linguistic research alone. A second concern of the language technologist and the language community is that there be enough documentation on the description, so that the description can be maintained even when the original designers are no longer involved. This concern becomes even more relevant when there are no grammar descriptions entirely representing the developments in the finite-state description. In the case of Skolt Sami, which has a growing number of specialists, there are still not enough of them to go around. This may, in part, be due to the decision to describe an enhanced version of the orthography, which entails the inclusion of two extra characters, ‹ẹ› Latin letter E with dot below (U+1EB9), and ‹ˈ› Modifier letter vertical line (U+02C8), which are used in enhanced dictionary paradigms and some teaching materials. To make the normative description of Skolt Sami, these two extra characters are simply filtered out, such that Modifier letter vertical line becomes null, and Latin letter E with dot below becomes ‹e›. There is no consistent grammatical description of this enhanced orthography, but rather there are many instances where it is used, 45
and in the long run one must consult enlightened nativelevel language users. In short, the description of Skolt Sami being constructed in the GiellaLT 1 infrastructure is also a research project of its own. A project whose success in language description can even be seen in its availability through python libraries, such as UralicNLP 2 . Background The finite-state description of Skolt Sami is approached from a synchronic perspective, and there are already a handful of reference descriptions to facilitate this undertaking. In examining features present in the enhanced orthographic representation of the language by Sammallahti & Mosnikoff (1991: 157–202) and subsequent authors Feist (2015) and Satu Mosnikoff et al (2020), we can observe the following five essentials: vowel height variation, suprasegmental palatalization marking, gradation, allegro variation, and a distribution of three stem vowel types. Although these five defining characteristics are present in the three texts mentioned above, they have not been used consistently in the synchronic description of extended regular nominal and verbal paradigms. Therefore, we have set the goal of determining to what extent the stem variation in verbs and nominals could be aligned for a better facilitation of Skolt Sami finite-state description. We see that the creation of tables for consulting in paradigm content validation is very important, as there should be guidelines for determining which stem variant is to be used with which suffixes and inflectional readings. 1 https://giellalt.github.io/ 2 https://github.com/mikahama/uralicNLP 46
The creation of reference tables for regular verbal and nominal inflection involves an understanding of morphophonology slightly beyond that of what is presented explicitly in the most recent Skolt Sami grammars. There must be a working understanding of what variation is synchronically present in the language to ensure the successful construction of an representative analyzer and generator for the enhanced orthography. This understanding includes the integration of vowel height, suprasegmental palatalization, gradation, allegro variation and stem-vowel classifiers into a single set of tables that affords a better comprehension of what morphological readings are to be associated with which characteristics. Thus, we also require a grouping of regular inflection readings that are associated with specific stem variants. To this end, we inspect the most complex stem variation, which is found in single-syllable nominals and verbs with singlesyllable third person singular indicative present forms. Vowel height variation is often presented in the form of a binary high vs low, that is to say, for each individual vowel, there is either a higher or low pair. In monophthongs, the sets are relatively straight forward with a left-to-right correlation high-to-low in ‹u : o› or ‹u : õ›, ‹i : e› or ‹i : ẹ›, ‹o : å›, ‹õ : â› and ‹a : ä›. Here the variation between the pair ‹u : o› vs ‹u : õ› appears to be dialectal, whereas the variation between ‹i : e› vs ‹i : ẹ› is directly associated suprasegmental palatalization, ellaborated below. In diphthongs, however, there is room for ambiguity when it comes to language learning and this ambiguity may present difficulties for speakers of some dialect backgrounds as well. Without palatalization, the diphthong pairs can be presented as follows: ‹uå : uä›, ‹uõ : uâ›, ‹iâ : eä›, ‹iõ : eâ› with left-to-right correlation high-to-low, respectively. In the presence of suprasegmental palatalization, however, ‹uå› and ‹uâ› are always realized as ‹ueʹ›, while ‹uä› is realized as ‹uẹʹ› when it is a mid-length vowel or affected by allegro. A similar 47
alignment can be found for the ‹i› initial diphthongs, too, such that ‹iâ› and ‹eâ› are always rendered as ‹ieʹ› in the presence of suprasegmental palatalization, while ‹eä› is realized as ‹iẹʹ› when it is a mid-length vowel or affected by allegro. As long or short low diphthongs, ‹uä› and ‹eä› do not change. That said, we must note that Satu Mosnikoff et al (2020: 32–33) seem to have an incomplete presentation of diphthong renderings, but this might be due to the fact that they are not always attempting to present enhanced spelling including ‹ẹ› and ‹ˈ›. Suprasegmental palatalization affects both the vowel and consonant centers. In the normative orthography, Modifier letter prime ‹ʹ› (U+02B9) is the indicator of suprasegmental palatalization, but it cooccurs with the letter ‹ǩ› Latin letter K with caron (U+01E9), ‹ǧ› Latin letter G with caron (U+01E7) and ‹j›, which have nonpalatal pairs in ‹k›, ‹g› and ‹ǥ›, respectively. Suprasegmental palatalization, as noted above can affect the rendering of the low vowels ‹å›, ‹â› and ‹ä› as second vowels of a diphthong. While the first two are rendered as neutralized ‹e› in ‹ueʹ› and ‹ieʹ›, ‹ä› only becomes ‹ẹ› in a palatalized diphthong if the vowel center is mid-length or this mid-length diphthong is subjected to allegro variation which renders both the vowel and consonant center as short (see below). Gradation in Skolt Sami has three distinct grades. Gradation is readily observed in the orthography as short vowel aligned with long consonant, mid-length vowel aligned with mid-length consonant, and long vowel aligned with short consonant, i.e., VCC, VVCC and VVC illustrate the simple consonants, and consonant clusters correlate to the first and second, such that VCC is to VYXX what VVCC is to VVYX. Naturally, there are at least two shortcomings to this rule. First, the weak grade of stems with VVCC absolute middle grade observes a complementary dichotomy where one set undergoes quality change while retaining mid-length geminate, and 48
the other set observes a shortening of the geminate to a single consonant. Thus, we have ‹cc : ʒʒ›, ‹čč : jj›, ‹ss : zz›, ‹šš : žž›, ‹kk/ǩǩ : ǥǥ/jj›, on the one hand, and ‹pp : v›, ‹mm : m›, ‹ff : f›, ‹vv : v›, ‹tt : đ›, ‹nn : n›, ‹đđ : đ›, ‹rr : r›, ‹ll : l›, ‹jj : j›, ‹ŋŋ : ŋ›, on the other. Second, there is no way of distinguishing the length of a diphthong preceding a geminate in the literary language. In the enhanced orthography, the Modifier letter vertical line, mentioned above, is inserted between the geminate letters following a diphthong to indicate the consonant is extra long and the diphthong short, as seen in the word for utensil ‹neävˈv› in the nominative singular but ‹neävv› in the genitive singular, where the geminate is short and the diphthong is longer (cf. Feist 2015: 76 ‹siõrr›). The best way to observe distinctions in grades is to investigate stems whose nominative singular or infinitive represents the middle grade in VVCC. Stems of this type will provide us with the most extensive variation, i.e., they will present us with paradigms including all three grades: VCC, VVCC and VVC. If the nominative singular or infinitive stem represents either VCC or VVC, variation will be limited. Allegro variation is where the development of separate descriptions of vowel length and consonant length has paid off. Since separate triggers had been defined for shortening vowels and consonants, it was a simple task to shorten both the vowel center and the consonant center simultaneously. Allegro variation is not obligatory, in fact, where-ever an allegro variant occurs, a largo variant is also possible. The main thing to remember is that allegro variation occurs only in instances where the consonant is mid-length or short and followed by a subsequent foot with an initial consonant cluster, for example in the construction ‹neäˈvstes› (allegro) vs ‹neävvstes› (largo) ‘in his/her utensil’. In the former, allegro form, both the diphthong and the consonant are short, and since the geminate has been reduced to a single consonant Skolt Sami research tradition has moved the placement of the Modifier letter vertical line to where it follows immediately after the diphthong. In 49
the largo form, neither the diphthong nor the consonant are short. Single-syllable nominal and verbal stems can generally be classified according to a three-way distribution of stem vowels, ‹a›, ‹â› and ‹e›. The stem vowel is best observed in the locative singular form of single-syllable nominals or the infinitive form of their verbal counterpart – these are verbs with single-syllable third person singular indicative present forms. The stem vowel, when present, correlates with the quality, height and palatalization of the preceding vowel center (see Satu Mosnikoff et al 2020: 32). Verb stems and inflections The following is a list of regular stem variants associated with the Skolt Sami verb jååʹtted ‹travel›, which can readily be identified as belonging to the set of singlesyllable verbs with ‹e› stems, referred to here as (1E). More specifically, this verb belongs to the subset of verbs with a long vowel in the infinitive and a single consonant in the weak grade (see Sammallahti & Mosnikoff 1991: 197–199). Our choice of this particular subtype derives from the maximal structural variation it provides, i.e., the infinitive takes the vowel and consonant center VVCC, which allows for distinctive VCC (strong) and VVC (weak) counterparts. The choice of an ‹e› stem over ‹â› and ‹a› here is once again attributed to the desire to find the most extensive variation, see Table 1, below. The regular Skolt Sami verb jååʹtted ‘travel’ has eleven distinct stem variants, and a twelfth stem variant has been established on the basis of a distinction found in 50
the ‹â› stem counterpart kaarrâd ‘wrap’ for stem variants (3) and (4), see Table 1, below. The stem variants are given with three classifiers referring to vowel height [high/low], palatalization [yes/no] and vowel length to consonant grade, respectively. Low line ‹_› and undertie ‹ ‿ › will be used in the representation of VCC, VVCC, VVC and VC variation. Vowel center length is indicated by a low line ‹_› for long and an undertie ‹ ‿ › for short. Due to the dichotomy in consonant weak grade, mid-length consonant center is indicated by a low line ‹_›, and weak grade by an undertie ‹ ‿ ›. Next, comes an actual inflectional form of the verb followed by a list of tag sets for analyzed and generated forms, where the first tag set indicates the example form given (1) [high][yes][_ _] jooʹtti : +Act+PrsPrc (2) [low][yes][_ _] jååʹtted : +Inf; +Ind+Prs+Pl1, +Ind+Prs+Pl2; +Imprt+Pl1, +Imprt+Pl2; +Actio, +Actio+Ess (3) [low][no][_ _] jååttam : +Act+PrfPrc, +Ind+Prt+ConNeg, +Der/NomAct (4) [low][no][_ _] jåått : +Ind+Prs+Sg3 (5) [high][yes][ ‿ _] joʹtte : +Ind+Prt+Pl3, +Ind+Prt+Sg1, +Ind+Prt+Sg2, +Ind+Prt+Sg4 (6) [high][no][ ‿ _] jottu : +Imprt+ConNegII, +Pass+PrfPrc (7) [low][yes][ ‿ _] jåʹtte : +Ind+Prs+Pl3 (8) [low][no][ ‿ _] jåttaz : +Imprt+Pl3 (9) [high][yes][_ ‿ ] jooʹđi : +Ind+Prt+Sg3, +Ind+Prt+Pl1, +Ind+Prt+Pl2; +Pot+Sg1, +Pot+Sg2, +Pot+Sg3, +Pot+Sg4, +Pot+Pl1, +Pot+Pl2, +Pot+Pl3, +Pot+ConNeg (10) [low][yes][_ ‿ ] jååʹđet +Ind+Prs+Sg4, +Ind+Prs+ConNeg; +Imprt+Sg2, +Imprt+ConNeg; +VAbess, +Ger+Ess, +Ger+Instr 51
(11) [low][no][_ ‿ ] jååđam +Ind+Prs+Sg1, +Ind+Prs+Sg2; +Cond+Sg1, +Cond+Sg2, +Cond+Sg3, +Cond+Sg4, +Cond+Pl1, +Cond+Pl2, +Cond+Pl3, +Cond+ConNeg; +Imprt+Sg3 (12) [low][yes][ ‿ ‿ ] jåʹđškuẹʹtted +Der/InchL+Allegro+V+Inf The classifiers listed for the Skolt Sami verb jååʹtted ‹travel›, above, are specific to the 1E stem type. Vowel height, indicated by high and low, indicates four high and eight low variants ‹o› and ‹å›, respectively. Suprasegmental palatalization indicated 7 ‹yes› and 5 ‹no› instances. Gradation shows 4 instances of [_ _], 4 [ ‿ _], 3 [_ ‿ ] and 1 [ ‿ ‿ ] (allegro). In reality, there would be one largo variant for each allegro, thus the corrected number for the pattern [_ ‿ ] would be four, with an additional form jååʹđškuẹʹtted, which is can be drawn from stem variant (10) When moving to a table representing the three stem vowel types in ‹e›, ‹â› and ‹a›, certain adjustments had to be made. First, it was noted that the concept <rel.>, indicating relative value as attested in the infinitive, could be applied to both height and suprasegmental palatalization. Suprasegmental palatalization has a concept <if poss.>, which indicates palatalization is present in forms of 1E and 1Â stems but not 1A stems. Finally, where two possible readings are possible, both have been given. Table 1. Comparing the VVCC centers of väällad ‘pour’, kaarrâd ‘wrap’ and ââʹnned ‘use’ 52
Whereas 1E illustrates eleven distinct regular stem variants, 1Â has eight, and 1A only has five. The 1E stem variants include four combinations with high vowels and eight combinations with low. Noun stems and inflections The following is a list of regular stem variants associated with the Skolt Sami noun mââʹnn ‹egg›, which can readily be identified as belonging to the set of single-syllable nouns with ‹e› stems, referred to here as (1E). Once again 1E has been selected, as it provides the most extensive set of stem variants, i.e., it represents the vowel and consonant centers in VVCC and uses distinctive height. 53
As might be expected, there appear to be exceptions to the dichotomy high vs low in diphthongs with the underlying ‹uõ : uâ› and ‹iõ : eâ› pairs. In the 1E verb stem variant (7), we find an unexpected triplet ‹iõ : eâ : eäˈʹ›, where the short diphthong indicates ‹ä› instead of ‹e›. This stem variant is only used in the formation of indicative present third person plural. All three stem types, 1E, 1Â and 1A show this as a low vowel. Even to the extent, that an enhanced form for the verb vueʹlǧǧed ‹set off› is vuẹʹlǧǧe or vuẹʹlǧǧa ‹they are setting off› is given (Sammallahti & Mosnikoff 1991: 169). Is this triplet evidence of something larger than the dichotomy high vs low, and how frequently does this triplet show up? Going ahead In this paper, we have merely scratched the surface, as it were, when searching for a consistent way to classify stem types for nouns and verbs in Skolt Sami. Our aim was to bring basic Skolt Sami concepts to the reader's attention and to outline a point of departure that could feasibly facilitate the construction of both a machineand human readable stem classifier. We illustrated to use of three stem types 1E, 1Â and 1A in alignment with nouns and verbs with the vowel and consonant center structure VVCC [_ _]. Single-syllable verbs and nouns with this structure were demonstrated to have extensive but regular stem variation with challenges in the understanding of diphthongs. We also contemplated the existence of a three-way split in vowel height. 60
Future development of classifiers for inspection of analysis and generation accuracy still require extensive description and documentation. First, we will complete our inspection of single-syllable verbal and nominal stems, we will need to address the types VCC [‿ _], VVC [_ ‿] and VYXX [‿ _]. Second, we will introduce a set 2E, 2Â and 2A for dealing with two-syllable stems such as čeäppat ‹neck›, aalǥât ‹begin!›, on the one hand, and kuåʒʒâlm ‹helm›, viõˈǥâsm ‹grow stronger!›, on the other. Subsequently, further alignments will be sought for explaining parallels between singleand two-syllable stems in the same part of speech, such as [‿ _] : [_ ‿] gradation patterns in toll : tool ‹fire› versus [_ ‿] : [‿ _] in võõnâs : võnnâz ‹boat›. Tags Abe = Abessive, Acc = accusative, Com = comitative, Cond = conditional, ConNeg = connegative, Dimin = diminutive, Ela = elative, Ess = essive, Gen = genitive, Ger = gerund, Ill = illative, Imprt = imperative, Der/InchL = inchoative, Ind = indicative, Instr = instrumental, Loc = locative, Nom = nominative, Par = partitive, Pl = plural, Pl1 = first person plural, Pl2 = second person plural, Pl3 = third person plural, Pot = potential, Prs = presence, Prt = preterit, Px = possessive suffix, PxPl1 = first person plural possessor, PxPl2 = second person plural possessor, PxPl3 = third person plural possessor, PxSg1 = first person singular possessor, PxSg2 = second person singular possessor, PxSg3 = third person singular possessor, Sg = Singular, Sg1 = first person singular, Sg2 = second person singular, Sg3 = third person singular, Sg4 = indefinite person, VAbessive = verbal abessive. 61
Thanks My thanks to the members of the Skolt Sami language and research communities who have provided continuous input into my understanding of the language. Every day brings something new and fascinating. References Alnajjar, K., Hämäläinen, M. & Rueter, J., marrask. 2024, PROCEEDINGS OF THE 9TH INTERNATIONAL WORKSHOP ON COMPUTATIONAL LINGUISTICS FOR URALIC LANGUAGES. Hämäläinen, M., Pirinen, F., Macias, M. & Crespo Avila, M. (toim.). Kerrville: The Association for Computational Linguistics, s. 41–48 8. Fiest, Timothy (2015). A Grammar of Skolt Saami. Suomalais-Ugrilaisen Seuran Toimituksia 273. Helsinki. Giellalt = Sjur Moshagen, Jack Rueter, Tommi Pirinen, Trond Trosterud, Francis M. Tyers (2014). Open-source infrastructures for collaborative work on under-resourced languages. In: Collaboration and Computing for UnderResourced Languages in the Linked Open Data Era. 71–77. https://giellalt.github.io/ Hämäläinen, M., Alnajjar, K., Rueter, J., Lehtinen, M. & Partanen, N., 2021, ELECTRONIC LEXICOGRAPHY IN THE 21ST CENTURY (ELEX 2021). PROCEEDINGS OF THE ELEX 2021 CONFERENCE. Kosem, I., Cukr, M., Jakubíček, M., Kallas, J., Krek, S. & Tiberius, C. (toim.). Brno: Lexical Computing CZ s.r.o., s. 653-664 12 Sivumäärä (Electronic 62
lexicography in the 21st century (eLex 2021). Proceedings of the eLex 2021 conference). Hämäläinen, (2019). UralicNLP: An NLP Library for Uralic Languages. Journal of Open Source Software, 4(37), 1345, https://doi.org/10.21105/joss.01345 Mosnikoff, Satu, Mosnikoff, Jouni, Koponen, Eino, Lehtinen, Miika (2020). Koltansaamen kielioppi = Sääʹmǩiõl ǩiõllvueʹppes. Sääʹmteʹǧǧ. Sammallahti, Pekka & Mosnikoff, Jouni (1991). SuomiKoltansaame sanakirja = Lääʹdd–sääʹm sääʹnnǩeʹrjj. Girjegiisá, Ohcejohka. 63