Full text
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 19 SEMANTIC ANALYSIS IN GLOBAL, TURKIC, AND UZBEK LANGUAGES: RESOURCES, ALGORITHMS, AND DEVELOPMENT PROSPECTS F. Alijonova 1st year Doctoral student, Digital technologies and artificial intelligence development research institute https://doi.org/10.5281/zenodo.18054745 Abstract. This article presents a comprehensive scientific examination of semantic analysis research in global languages, across the Turkic language family, and within the specific context of the Uzbek language. The study evaluates widely used semantic resources — WordNet, SimLex999, WordSim-353 — and their importance in modeling semantic similarity, relatedness, lexicalsemantic relations, and word sense disambiguation (WSD). Semantic tools developed for Turkish, Kazakh and Kyrgyz are systematically compared with the current state of Uzbek computational linguistics. Existing Uzbek resources such as UZWORDNET and SimRelUz are assessed for their strengths and limitations, identifying the lack of WSD corpora, contextual semantic models, and large-scale annotated datasets as the primary research gap. The findings highlight the necessity of creating next-generation semantic algorithms tailored to the agglutinative structure and rich morphology of Uzbek and integrating it into global NLP research. Keywords: semantic analysis; semantic similarity; semantic relatedness; WordNet; WSD; Uzbek language; Turkic languages; NLP. Introduction. Semantic analysis constitutes one of the most theoretically grounded and technically challenging areas of natural language processing (NLP). It involves interpreting the meaning of linguistic units, analyzing semantic relations, resolving lexical ambiguity, and modeling conceptual similarity and relatedness within computational frameworks. The emergence of transformer-based architectures, including BERT, SBERT, GPT, and XLM-R, has revolutionized semantic representation by enabling context-sensitive and fine-grained semantic modeling. In global linguistic research, semantic analysis is strongly supported by extensive resources and benchmark datasets such as WordNet, FrameNet, PropBank, SemCor, SimLex-999, WordSim-353, and MEN. These resources form the empirical foundation for evaluating semantic similarity, relatedness, lexical-semantic relations, and WSD systems. Their presence allows semantic models to be evaluated reliably and consistently across tasks and languages. However, languages with limited computational resources—such as Uzbek—have not yet achieved comparable development. The challenges include insufficient annotated corpora, lack of semantic evaluation datasets, absence of sense-annotated materials for WSD, and limited availability of semantic networks. As a result, Uzbek NLP systems often struggle with polysemy, homonymy, and morphological ambiguity, yielding semantically inaccurate outputs. Uzbek, as an agglutinative language, encodes meaning largely through affixation and derivational morphology. This structure leads to challenges in semantic modeling because morphological variations can dramatically shift semantic interpretation. For instance, lexical items such as ot, bor, and yuz may represent multiple distinct meanings depending on their morphological form or contextual usage. Such characteristics necessitate specialized semantic algorithms that integrate morphological
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 20 disambiguation as a core component. Despite these difficulties, recent advancements—most notably UZWORDNET, SimRelUz, Uzbek morphological analyzers, UzbekBERT, and text classification studies—have created an initial infrastructure for semantic research in Uzbek. These developments provide a critical foundation for designing higher-level semantic systems. This article therefore aims to: 1) analyze global semantic research and leading semantic evaluation resources; 2) review semantic modeling efforts across Turkic languages; 3) critically evaluate existing Uzbek semantic resources; and 4) outline future research directions necessary to achieve international-level semantic modeling for the Uzbek language. Semantic analysis in global languages: approaches, resources, and model development Semantic analysis is a central field within NLP, aiming to uncover meaning, conceptual relations, and lexical ambiguity through formal computational methods. The evolution of semantic modeling has progressed through several key methodological phases—beginning with statistical semantics, expanding into distributional modeling, and culminating in deep contextual representations enabled by transformer architectures. Statistical semantics: the foundational stage of semantic modeling The earliest attempts to computationally represent meaning were grounded in statistical approaches based on word frequency and co-occurrence. Models such as Pointwise Mutual Information (PMI), TF-IDF, and Latent Semantic Analysis (LSA) were designed to capture semantic association using purely distributional signals. Despite pioneering contributions, these models suffer from several limitations: they do not incorporate contextual variation; they cannot resolve homonymy or polysemy; they rely heavily on high-frequency data, performing poorly on rare words. Thus, while important historically, statistical semantics was insufficient for rich, contextdependent meaning representation. Distributional semantics: Word2Vec, GloVe, and fastText The introduction of distributional semantic models marked a conceptual shift. Word2Vec popularized neural embeddings that map words into high-dimensional vector spaces based on the distributional hypothesis—“a word is known by the company it keeps.” Key advantages include: clustering of synonyms; vector direction capturing semantic opposition; emergent semantic groupings; analogical reasoning (e.g., king – man + woman = queen). GloVe further enhanced this by integrating global corpus statistics into word embeddings. However, distributional models share an important weakness: each word receives only one embedding, regardless of how many meanings it carries. This makes them inadequate for: polysemous words (very common in Uzbek), homonym disambiguation, sentence-level semantic tasks. Transformer-based models: deep contextual semantic representation
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 21 Transformers fundamentally redefined semantic modeling. BERT introduced contextsensitive representations, where each token receives a distinct embedding depending on its surrounding context. Advantages include: accurate resolution of word senses (WSD), modeling of subtle semantic relations, improved similarity and relatedness prediction, robust sentence-level embeddings (SBERT GPT extended these capabilities by integrating semantic reasoning within generative architectures. Transformers represent the current state-of-the-art in semantic modeling across global languages. International semantic evaluation resources Global semantic research relies heavily on standardized evaluation datasets, including: WordSim-353 — semantic relatedness benchmark; SimLex-999 — measures true conceptual similarity; MEN — semantic association dataset; RG-65 — one of the earliest semantic similarity corpora. These datasets ensure consistent evaluation of semantic models and allow comparison across languages and methodologies. Semantic analysis in turkic languages: resources, limitations, and perspectives Turkic languages share typological characteristics that significantly impact semantic modeling: agglutinative morphology, rich derivation, flexible syntax, and high degrees of polysemy and homonymy. For these reasons, computational approaches developed for IndoEuropean languages are often poorly suited to Turkic linguistic structure, requiring additional linguistic adaptation. Among Turkic languages, Turkish has the richest semantic infrastructure, while Kazakh, Kyrgyz, and Uzbek remain under-resourced. Turkish the most advanced turkic language in semantic modeling. KeNet is the most comprehensive Turkish semantic network, aligned structurally with Princeton WordNet. It includes thousands of synsets, hierarchical semantic relations, antonymy, synonymy, meronymy, hyponymy, clear sense distinctions. KeNet served as the methodological foundation for constructing UZWORDNET. AnlamVer is one of the largest human-annotated semantic evaluation datasets in the Turkic world. Features: thousands of word pairs, dual scoring (similarity + relatedness), rigorous annotation guidelines, used widely in Turkish semantic model evaluation. SimRelUz was directly inspired by AnlamVer’s methodology. Turkish WSD Systems Several WSD approaches have been explored for Turkish: BERTurk-based models, Modified Lesk algorithms, Hybrid models integrating morphological disambiguation + semantic inference. These advances highlight the feasibility of adapting similar techniques to Uzbek. Semantic modeling in kazakh Kazakh has moderate development compared to Turkish. Available resources: KazNet — a Kazakh WordNet analog KazakhBERT — pretrained contextual model. Missing resources: no semantic similarity dataset;
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 22 no WSD corpus; no large annotated semantic datasets. Thus, Kazakh faces many of the same structural barriers as Uzbek. Semantic research in kyrgyz Kyrgyz is one of the least-resourced Turkic languages: no WordNet equivalent; no WSD dataset; no semantic evaluation corpus; minimal digital linguistic resources. Semantic analysis in the uzbek language: existing resources, limitations, and future directions The Uzbek language is typologically characterized as an agglutinative and morphologically rich language, where meaning is encoded predominantly through affixation and derivational processes. Its flexible syntactic structure, together with widespread polysemy and homonymy, introduces considerable complexity for computational semantic interpretation. Consequently, the development of semantic analysis in Uzbek is heavily dependent on morphological analyzers, lexical-semantic networks, and large annotated corpora. In recent years, several important foundational resources have been created, yet these remain insufficient to support advanced semantic modeling. This section provides a critical assessment of the most significant Uzbek semantic resources — UZWORDNET, SimRelUz, existing morphological analyzers, and UzbekBERT — and outlines the primary challenges and research opportunities for the language. UZWORDNET: the first large-scale lexical–semantic database for uzbek UZWORDNET represents the first systematic attempt to build a WordNet-style lexicalsemantic network for Uzbek. The resource offers structured representations of semantic relations, including synonymy, hypernymy, meronymy, and antonymy. Key characteristics of UZWORDNET: 28,140 synsets, 64,389 senses, 20,683 lexical entries, expert-evaluated accuracy: 75–76%. Scientific significance: it provides the first formal semantic ontology for the Uzbek language. It establishes a basis for semantic search, machine translation, and text classification systems. It enables cross-linguistic comparison with other WordNet-based languages. Limitations: Many synsets are generated through automated translation, leading to semantic inconsistencies. Derivational morphology and morphological variants are not adequately represented. Polysemy distinctions remain incomplete in numerous cases. Some semantic relations do not fully reflect native speakers’ linguistic intuitions. Despite these limitations, UZWORDNET constitutes an essential foundation for the next stage of Uzbek semantic research. SimRelUz: the first semantic similarity and relatedness dataset for uzbek SimRelUz is the first gold-standard evaluation dataset designed to measure semantic similarity and relatedness in the Uzbek language. Key features: 1,418 word pairs; Dual scoring methodology: similarity and relatedness are evaluated separately; Annotators: 11 native Uzbek speakers; Includes rare and out-of-vocabulary (OOV) words.
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 23 SimRelUz provides the only standardized evaluation benchmark for: semantic similarity algorithms, embedding models, transformer-based semantic systems, WSD methods. It is a prerequisite for developing robust semantic models for Uzbek. In Uzbek, morphological structure carries a significant proportion of semantic information. Meaning shifts can occur through simple affix changes, making morphological disambiguation essential for semantic interpretation. Examples of semantic ambiguity: ot → “horse,” “weapon,” “verb (to throw)” yuz → “face,” “hundred,” “surface” bor → “exists,” “to go,” “possession”. Such cases demonstrate the necessity of developing Uzbek-specific WSD (Word sense disambiguation) systems, which are currently absent. Absence of WSD Resources: core limitation in Uzbek NLP Uzbek NLP currently lacks: a sense-annotated corpus, a WSD dataset, transformer-based WSD models, integration of homonym dictionaries into semantic frameworks. Consequences include: inaccurate machine translation outputs, semantically incorrect chatbot responses, contextually irrelevant search engine results, poor performance in semantic text classification. These limitations significantly restrict the development of advanced Uzbek NLP applications. Real-World semantic errors in existing uzbek NLP systems Machine translation error “Oy chiqdi.” → Incorrect: The month went out. → Correct: The moon rose. Chatbot interpretation error “Yuzim achishyapti.” → Incorrect: 100 is burning me. Search engine misinterpretation “Ot haqida kitob.” → Returns results about throwing instead of horses. These errors demonstrate the urgent need for robust WSD and semantic modeling tools adapted to Uzbek linguistic structure. To advance semantic analysis in Uzbek, several key research directions must be prioritized: Expansion and expert revising of UZWORDNET. Creation of a dedicated WSD corpus for Uzbek. Development of SBERT-Uzbek as a sentence-level semantic model. Morpho-semantic disambiguation algorithms integrating Uzbek affixation. Construction of large-scale semantic corpora with human annotation. These steps will form the basis for a sustainable and modern Uzbek NLP ecosystem.
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 24 CONCLUSION This study provides a comprehensive analysis of semantic research in global languages, within the Turkic language family, and in the Uzbek linguistic context. While global NLP benefits from extensive semantic resources and standardized evaluation datasets, Turkic languages — especially Uzbek — lack essential components such as annotated semantic corpora, WSD systems, and contextual semantic models. The creation of UZWORDNET and SimRelUz marks significant progress, yet these resources remain insufficient to support advanced semantic modeling. To position Uzbek NLP on par with global standards, it is imperative to: develop large-scale senseannotated corpora, construct WSD datasets and models, expand and refine UZWORDNET, build contextual semantic models such as SBERT-Uzbek. These steps will enable the formation of a robust semantic infrastructure for Uzbek and open new opportunities for AI-powered linguistic technologies. REFERENCES 1. Miller G.A. WordNet: A lexical database for English // Communications of the ACM. – 1995. – №38(11). – P. 39–41. 2. Fellbaum C. WordNet: An Electronic Lexical Database. — Cambridge: MIT Press, 1998. — 423 p. 3. Agostini A., Usmanov T., Khamdamov U., Abdurakhmonova N. UZWORDNET: A lexicalsemantic database for Uzbek // Proc. Global WordNet Conference. – 2021. – P. 112–120. 4. Salaev U. SimRelUz: Semantic similarity and relatedness dataset for Uzbek // arXiv:2310.18345. – 2023. 5. Mikolov T. et al. Distributed Representations of Words and Phrases // NIPS. – 2013. – P. 3111–3119. 6. Pennington J., Socher R., Manning C. GloVe: Global vectors for word representation // EMNLP. – 2014. – P. 1532–1543. 7. Devlin J. et al. BERT: Pre-training of Deep Bidirectional Transformers // NAACL. – 2019. 8. Reimers N., Gurevych I. Sentence-BERT: Sentence Embeddings // EMNLP. – 2019. – P. 3982–3992. 9. Lesk M. Automatic sense disambiguation // SIGDOC. – 1986. 10. Navigli R. Word Sense Disambiguation: A Survey // ACM Computing Surveys. – 2009. 11. Hill F., Reichart R., Korhonen A. SimLex-999 // Computational Linguistics. – 2015. 12. Finkelstein L. et al. WordSim-353 // ACM Transactions on Information Systems. – 2001. 13. Bruni E. et al. Distributional Semantics // JAIR. – 2012. 14. Luong M.T. et al. Rare Words Problem // ACL. – 2013. 15. Baker C.F., Fillmore C., Cronin B. FrameNet Project // Intl. Journal of Lexicography. – 1998. 16. Tufiș D. et al. BalkaNet // LREC. – 2004. 17. Ercan G., Yıldız O.T. AnlamVer // LREC. – 2018. 18. Schuler K. VerbNet: A Comprehensive Verb Lexicon. – 2005. 19. Bond F., Paik K. Survey of Open Multilingual WordNets. – 2012. 20. Bisazza A., Federico M. Morphological Analysis for Agglutinative Languages. – 2013. 21. Hajiyev E. et al. Kazakh WordNet // Turkic Languages Journal. – 2019. 22. Abduvaliev A. Uzbek Language Morphology // Uzbek Journal of Philology. – 2018. 23. Matlatipov H. Morphological Analyzer for Uzbek. – 2009. 24. Rabbimov K., Kobilov K. Uzbek Text Classification. – 2020.
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 12 DECEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 25 25. Mansurov B. et al. UzbekBERT Pretrained Model. – 2021. 26. Solovyev V. et al. Semantic Network Mining. – 2020. 27. Brown P. et al. Statistical Machine Translation. – 1993. 28. Bahdanau D. et al. Neural Machine Translation. – 2014. 29. Madatov M. et al. Uzbek Stopwords Dataset. – 2021. 30. Kuriyozov J., Matlatipov H. Sentiment Analysis for Uzbek. – 2019.