STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 57 CORPUS-BASED ANALYSIS OF LINGUISTIC MOVES IN DISCOURSE Asrorova Nargiza Isomitdinovna PhD Researcher, Uzbekistan State World Languages University
[email protected] Abstract: This article explores the corpus-based analysis of linguistic moves, a methodology that integrates the quantitative power of corpus linguistics with the functional focus of discourse analysis. By examining large, machinereadable text collections, this approach identifies and validates rhetorical moves— the conventionalized steps used to achieve communicative goals within a genre— through recurrent linguistic patterns. The synergy of computational tools and qualitative interpretation enables a more objective, scalable, and empirically grounded analysis of discourse structure. Key advantages include reduced researcher bias, the ability to uncover subtle functional-form correlations, and facilitation of cross-genre or diachronic comparisons. Despite challenges in automation and corpus design, this hybrid paradigm offers robust insights into the architecture of spoken and written genres, with significant implications for pedagogy and professional communication. Keywords: corpus linguistics, move analysis, discourse structure, genre analysis, rhetorical moves, computational linguistics, text analysis, linguistic patterns The intricate architecture of discourse, how language is structured to achieve communicative goals beyond the sentence level, has long been a central concern in linguistics. Understanding this architecture requires more than just analyzing grammar and vocabulary; it demands an exploration of the functional units that writers and speakers use to construct their messages. These functional units, known
STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 58 as “moves,” are the building blocks of genre, the conventionalized steps or strategies that members of a discourse community employ to achieve a specific purpose. A business proposal, for instance, is built from moves such as “establishing a territory,” “establishing a niche,” and “occupying the niche.” For decades, the identification and analysis of these moves were the domain of qualitative, manual discourse analysis, relying on the close reading and interpretative skill of the researcher. While this approach yields rich insights, it is often limited in scope, susceptible to researcher bias, and difficult to scale. The integration of corpus linguistics methodologies has revolutionized this endeavor, giving rise to a more robust, empirical, and scalable paradigm: the corpus-based analysis of linguistic moves. This synergy allows researchers to ground the functional analysis of discourse in the quantifiable patterns of language use, moving from intuitive description to evidence-based typology. The concept of moves originates from genre analysis, particularly within the English for Specific Purposes (ESP) tradition, where it was used to describe the rhetorical structure of academic and professional texts. A move is a discoursal or rhetorical unit that performs a coherent communicative function within a written or spoken discourse. It is not defined by its length but by its purpose; a move can be a clause, a sentence, or a group of sentences. The primary challenge in move analysis has always been its identification. Traditionally, this was done through manual, iterative reading, where the analyst would segment a text into functional units based on their understanding of the genre’s conventions and the writer’s intent. This process, while invaluable, is inherently subjective. Different analysts might segment the same text differently, and the selection of texts for analysis can be influenced by a researcher’s preconceived notions, potentially overlooking common but less salient patterns or overemphasizing unusual ones. The analysis risks being based on a small, potentially unrepresentative sample, making it difficult to ascertain whether the identified moves are indeed conventional across the entire genre or are idiosyncratic to the selected texts.
STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 59 This is where corpus linguistics provides a powerful corrective. Corpus linguistics, the study of language through large, machine-readable collections of texts, offers a methodology anchored in empirical evidence and quantitative analysis. By applying computational tools to a corpus—a principled collection of texts representing a particular genre or discourse domain—researchers can analyze linguistic patterns across dozens, hundreds, or even thousands of texts simultaneously. The synergy between this data-driven approach and move analysis is profound. Instead of relying solely on top-down, intuitive segmentation, the corpus-based analyst can adopt a bottom-up approach, allowing the frequent and consistent linguistic patterns in the texts themselves to suggest and validate the functional moves. This process often involves the identification of “vocabularybased discourse units” (VBDUs), which are segments of text identified by shifts in lexical cohesion. As Biber, Connor, and Upton (2007) explain, computational algorithms can detect these shifts, effectively providing an initial, data-driven segmentation of texts based on their topical structure, which can then be interpreted functionally as moves. The practical process of a corpus-based move analysis typically follows a recursive cycle. The researcher first compiles a specialized corpus that is representative of the target genre. This corpus is then annotated, sometimes automatically for linguistic features like part-of-speech tags and lemmas, and often manually for a preliminary set of move categories based on existing literature or a pilot study. The power of the corpus approach lies in the subsequent stage, where analytical software like AntConc is used to investigate the linguistic correlates of these proposed moves. The analyst examines frequency lists, keywords, and clusters (n-grams) within each manually identified move category. For example, an analysis of research article introductions might reveal that the move “establishing a niche” is characterized by a high frequency of negative evaluative adjectives like problematic, inadequate, or limited, and phrases like however, a gap in the literature, or little research has been done. These linguistic features serve as
STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 60 empirical fingerprints for the move. Once these key phrases and grammatical patterns are established, they can be used as search queries to find all potential instances of that move across the entire corpus, thereby validating the initial manual coding and ensuring consistency. This iterative process—moving from functional category to linguistic form and back again—creates a much more reliable and explicit coding scheme. One of the most significant advantages of this methodology is its capacity to reveal the intricate connection between communicative function and linguistic form at a scale impossible through manual analysis alone. A qualitative researcher might note that writers often “review previous research,” but a corpus-based analysis can specify the exact lexical and grammatical repertoires that realize this move. Hyland (2009), in his work on academic discourse, has consistently demonstrated how different disciplines and genres employ distinct linguistic strategies to achieve similar rhetorical goals. Through corpus analysis, he can show not just that a move exists, but how it is typically phrased, how frequent it is, and how its linguistic realization varies across sub-fields. This moves the analysis from a descriptive catalog of functions to a precise profile of their linguistic instantiation. It answers not only the “what” and “why” of discourse structure but also the “how,” providing a level of detail that is invaluable for pedagogical applications, such as teaching novice writers the specific, conventionalized language they need to participate effectively in their chosen discourse community. Furthermore, the corpus-based approach brings a level of objectivity and mitigates researcher bias by anchoring the identification of moves in quantitative, observable data. The human brain is not optimized for accurately perceiving frequency distributions across a large set of texts; we tend to notice what is salient or rhetorically striking. A researcher might be drawn to a particularly elegant or unusually direct example of a move, mistakenly assuming it is representative. Corpus tools, however, provide an unbiased audit of the entire dataset. They can reveal that the most frequent way of realizing a “making a claim” move in
STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 61 biochemistry articles is not through bold, innovative statements but through a cautious, heavily hedged phraseology that is statistically dominant but less perceptually salient. As McEnery and Hardie (2012) argue, corpus linguistics provides a solid basis for generalizations about language use, forcing the analyst to account for all the evidence in the corpus, not just the examples that fit a preconceived theory. This empirical grounding enhances the validity and reliability of the move analysis, making its findings more credible and replicable. The scalability of corpus methods also allows for diachronic and comparative studies of move structure that were previously impractical. Once a reliable, linguistically-defined coding scheme for moves has been developed, it can be applied to much larger corpora. A researcher can track how the rhetorical structure of the scientific research article has evolved over a century by analyzing a diachronic corpus, identifying which moves have become more or less prominent and how their linguistic realizations have changed. Similarly, one can compare move structures across cultures or languages. For instance, a study could compile corpora of business letters from American and Japanese companies, using move analysis to identify and quantify cross-cultural differences in rhetorical strategy and politeness. Flowerdew (2023) notes that this ability to uncover patterns across large datasets is a key strength of the corpus-based approach, allowing for generalizations that are far more robust than those based on a handful of texts. This moves the field from a deep understanding of individual texts to a broad understanding of genre as a conventionalized social practice. However, the corpus-based analysis of moves is not without its challenges and limitations. The most significant of these is that computational tools, while excellent at identifying formal linguistic patterns, cannot, on their own, determine communicative function. A computer can easily find all instances of the phrase “more research is needed,” but it cannot definitively label this as part of a “identifying a gap” move without human interpretation. The phrase could, in a different context, be part of a conclusion summarizing limitations. Therefore, the
STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 62 corpus-based approach does not eliminate the need for the qualitative, interpretative skill of the discourse analyst; it repositions it. The analyst’s role shifts from being the sole identifier of moves to being the interpreter of quantitatively-derived linguistic patterns, using their understanding of context and communicative purpose to assign functional labels. This hybrid methodology embodies a powerful form of triangulation, where quantitative trends inform and are interpreted through qualitative insight. Another challenge lies in the initial design and representativeness of the corpus. The findings of any corpus-based study are only as valid as the corpus itself. If the corpus does not accurately represent the genre it purports to model, any move analysis derived from it will be skewed. Biber (1993) emphasizes that corpus design is paramount, requiring careful consideration of the sampling frame to ensure representativeness. Furthermore, the process, while scalable in its automated components, still requires significant manual labor for the development and validation of the coding scheme. This places a practical limit on the size of corpora that can be analyzed in depth within the constraints of a single research project. In conclusion, the integration of corpus linguistics with move analysis represents a significant methodological advancement in discourse studies. By marrying the functional, rhetorical focus of genre analysis with the empirical, datadriven techniques of corpus linguistics, this approach provides a more objective, scalable, and linguistically-grounded framework for understanding the architecture of discourse. It allows researchers to move beyond anecdotal evidence and intuitive segmentation, instead building typologies of moves that are validated by the consistent and frequent patterns of language use within a discourse community. This methodology reveals the crucial link between communicative purpose and linguistic form, offering insights that are not only theoretically valuable for understanding how language constructs social and professional realities but also immensely practical for applications in language teaching, professional
STUDIES IN ECONOMICS AND EDUCATION IN THE MODERN WORLD Vol. 4 No. 2 (2025) 63 communication, and translation. While it requires a careful balance of computational power and human interpretation, the corpus-based analysis of linguistic moves has firmly established itself as an indispensable tool for deconstructing and understanding the complex, conventionalized, and powerful patterns of human discourse. References Biber, D. (1993). Representativeness in corpus design. Literary and Linguistic Computing, 8(4), 243–257. Biber, D., Connor, U., & Upton, T. A. (2007). Discourse on the move: Using corpus analysis to describe discourse structure. John Benjamins Publishing. Flowerdew, L. (2023). Corpus-based discourse analysis. In A. O'Keeffe & M. J. McCarthy (Eds.), The Routledge handbook of corpus linguistics (2nd ed., pp. 275–289). Routledge. Hyland, K. (2009). Academic discourse: English in a global context. Continuum. McEnery, T., & Hardie, A. (2012). Corpus linguistics: Method, theory and practice. Cambridge University Press.