ValiText - Application Example Parliamentary Speeches
Abstract
This application example evaluates parliamentary speeches delivered in the lower chamber of the Russian parliament, the State Duma, between 1999 and 2024, which were web-scraped from the institution’s official website.
Full text
Indicators and Metrics of Data Quality: Evaluating the Language Model Classification of Russian Parliamentary Speeches David Avagian February 15, 2025
1. Introduction In this work, I present an evaluation of my dataset. Parliamentary speeches delivered in the lower chamber of the Russian parliament, the State Duma, between 1999 and 2024 were web-scraped from the institution’s official website. The dataset contains information about the name and party affiliation of the speaker, as well as the legislative stage of the motion under consideration. The speeches were classified as neutral toward, supportive of, or critical of the government. The classification is performed using a BERT-like zero-shot classifier (Laurer et al. 2023). I will specifically focus on the variables obtained after the classification and evaluate its quality. Since I am using data derived from supervised classification, I will apply the ValiTex framework (Birkenmaier 2024). ValiTex requires three types of validation evidence (ibid.). The first is substantive evidence, which focuses on the theoretical underpinnings of the measure. I proceed by completing the checklist related to this dimension. 2. Substantive Evidence 2.1. Documentation of the conceptual background I did not conduct a systematic literature review concerning the concept of government criticism by the parties in authoritarian parliaments. However, I use a case study that observed this phenomenon in 2018 after the announcement of the unpopular pension reform as the motivation for further investigation (Bondarenko 2023). I conceptualize government criticism as 2
blame attribution. I found the existing evidence that authoritarian regimes pragmatically allow criticism of some institutions to shift the blame away from themselves (Klimovich 2024) 2.2. Justification of the operationalization I operationalise criticism of the government as critical statements against the action of the executive in the public speeches delivered in the State Duma. Even though I don’t have enough evidence that citizens are perceptive to those speeches, I treat parliaments as the institutionalised platform for communication with a wide range of political actors. I don’t have a specific template of manifistation of the criticism in the textual data. I leave it for the model to decide which traits of the speech represent critical statements. Since I don’t fine tune the model on my own dataset, this might reduce the quality of the classification and alienate the scores from the construct they are meant to measure. 2.3. Justification of data collection decisions I collected all parliamentary speeches delivered under President Putin’s rule. Since the party landscape changed drastically during his first term with the emergence of the dominant, regime-sanctioned party United Russia, and since the regime transformed from a defective democracy to a personalized dictatorship (Golosov 2023), I believe this time frame is sufficient to infer the behavior of loyal opposition parties in Russia. After collecting the data, I computed summary statistics at different levels. 3
I found that many politicians were repeatedly recorded as speakers on the same parliamentary day. This seemed suspicious, so I investigated further and discovered an error in data collection: a speech by a member of parliament was entered as two separate speeches if the speaker was interrupted. I ensured that such interrupted speeches were merged into a single entry. 2.4. Justification of method choice This part is missing from my analysis. I do not justify the use of zero-shot classifiers over alternative methods, such as APIs of large language models like ChatGPT provided by OpenAI, since I plan to use both in the future. The decision to use a BERT-like classifier was pragmatic: it is easier to implement than traditional text analysis approaches, which require extensive preprocessing. At the same time, previous research has shown it to be sufficient for classifying text into different categories (Laurer et al. 2023), making it a natural first step before employing more advanced models. The specific model 1was chosen mainly due to it’s multilingual support. 2.5. Justification of the level of analysis I focus on the classification of individual speeches. One challenge associated with this is that some speeches exceed the 512-token limit of the model, resulting in truncation. This is a significant limitation that reduces the quality of classification. To address this, I plan to first summarize each speech to under 512 tokens and then apply the classifier to the summarized versions. 1The model is available at https://huggingface.co/MoritzLaurer/deberta-v3-large-zeroshot-v2.0 4
I expect this approach to improve classification quality. 2.6. Justification of preprocessing decisions Each speech begins with the speaker’s name, which I remove in one of the columns. I also perform stopword removal and lemmatisation in other columns, but I have not used those columns for analysis yet. 3. Structural Evidence I do not collect substantive evidence to assess the quality of the classification. I tested the classification results by sorting the observations from most to least critical and visually inspecting the text in relation to its score. In my analysis, I compare the scores between loyal opposition and dominant regime parties. I find that even loyal opposition parties are more critical of the government than the regime-sanctioned party. However, I use this result for inference rather than for assessing the quality of the classification. More work is needed to improve the reliability of the data. 4. External Evidence I do not collect external evidence for the acquired classification. Comparing the measures with a human-annotated test set would require additional resources. However, recent research suggests that Large Language Models outperform human coders in annotation tasks (T¨ ornberg 2024). One potential avenue for improving this dataset would be to first apply zero-shot 5
classification and then use a Large Language Model API to assess the quality of the measure. I also plan to select a random subset of the data for independent human evaluation. 5. Robustness Checks I do not conduct a comprehensive robustness check. However, I compared English and Russian language inputs for the model and found that it is robust to the language of the instruction. Later, I plan to rerun the analysis with different specifications of the task to see how much the classification is dependent on the initial input. 6. Conclusion I applied the ValiTex framework (Birkenmaier 2024) to evaluate my dataset of Russian parliamentary speeches and their classification as containing criticism of the government. Although I acquired some substantive evidence to establish the link between the concept and the measure, the data remains of very low quality. More work is needed to collect structural and external evidence before publishing this dataset. 6
References Birkenmaier, L. (2024). Valitext. Bondarenko, O. (2023). Between loyalty and opposition: The communist party of russia and the growing intra-party cleavage. Communist and Post-Communist Studies 56(4), 143–165. Golosov, G. V. (2023). The place of russia’s political regime (2003–2023) on a conceptual map of the world’s autocracies. Social Science Information 62(3), 390–408. Klimovich, S. (2024). Federalism at war: Putin’s blame game, regional governors, and the invasion of ukraine. Post-Soviet Affairs, 1–16. Laurer, M., W. van Atteveldt, A. Casas, and K. Welbers (2023). Building efficient universal classifiers with natural language inference. T¨ ornberg, P. (2024). Large language models outperform expert coders and supervised classifiers at annotating political social media messages. Social Science Computer Review, 08944393241286471. 7