scieee AI-readable full text Open interactive document viewer

The Numbers of Fear: Pandemic Governmentality and Media Discourse in Italy (2020)

Vardanega, Agnese; Vardanega, Claudia

Abstract

This report presents the methodology and results of a quantitative content analysis of Italian media discourse during the first COVID-19 lockdown (February-May 2020). The study investigates the hypothesis that the pandemic narrative was ‘dataist’—marked by a separation between numerical reporting and scientific context—and carried an anxiety-inducing emotional profile consistent with post-truth dynamics. Based on a corpus of 2,144 news headlines, we employed LDA to identify themes and sentiment analysis to assess their emotional tone. Our findings reveal a clear thematic divide between a “numbers” cluster (themes Numeri, Bollettini) and a “disciplinary” cluster (Scienza, Esperti, Misure). The “numbers” cluster, in particular, is associated with a negative, high-arousal emotional profile. These results support the initial hypothesis that the “datafication” of the pandemic narrative (data without context) functioned as a driver of emotional activation. --- This document accompanies the presentation "I numeri della paura: governamentalità pandemica fra opacità e trasparenza" ("The Numbers of Fear: Pandemic Governmentality between Opacity and Transparency"), presented at the conference "The Great Fear: Epidemics in the Italian Peninsula from the 17th Century to Today," University of Teramo, September 30 - October 1, 2025.

Full text

The Numbers of Fear: Methodological Note and Results Agnese Vardanega Claudia Vardanega October 1, 2025 Abstract This report presents the methodology and results of a quantitative content analysis of Italian media discourse during the first COVID-19 lockdown (February-May 2020). The study investigates the hypothesis that the pandemic narrative was ‘dataist’—marked by a separation between numerical reporting and scientific context—and carried an anxiety-inducing emotional profile consistent with post-truth dynamics. Based on a corpus of 2,144 news headlines, we employed LDA to identify themes and sentiment analysis to assess their emotional tone. Our findings reveal a clear thematic divide between a “numbers” cluster (themes Numeri,Bollettini) and a “disciplinary” cluster (Scienza,Esperti,Misure). The “numbers” cluster, in particular, is associated with a negative, high-arousal emotional profile. These results support the initial hypothesis that the “datafication” of the pandemic narrative (data without context) functioned as a driver of emotional activation. This document accompanies the presentation “I numeri della paura: governamentalità pandemica fra opacità e trasparenza” (“The Numbers of Fear: Pandemic Governmentality between Opacity and Transparency”), presented at the conference “The Great Fear: Epidemics in the Italian Peninsula from the 17th Century to Today,” University of Teramo, September 30 - October 1, 2025. Its purpose is to make the data discussed on that occasion available. 1 Introduction The results of a content analysis performed on the headlines of news related to the coronavirus during the first lockdown period are presented. The first part of this study aimed to explore the discourse surrounding the pandemic, as it was constituted in the first phase of the coronavirus spread (February 22 - May 15, 2020; §2), through the identification of themes, actors, and categories. The hypothesis at the center of this report is that the pandemic data were presented in a “dataist” manner (Van Dijck, 2014), and that this conception of data is closely linked with post-truth (Shelton, 2020), understood as a set of discursive practices and shared beliefs (Ferraris, 2017;Lorusso, 2018). The Latent Dirichlet Allocation (§3) and sentiment analysis presented in this document provided data to support the first hypothesis. The narrative of the data (theme Numeri: Numbers) was: • clearly separated from that of Scienza (Science), Esperti (Experts), and Misure (governement Measures) (§3.1); • prevalent, and grew over time, at the expense of other themes (§3.2), in a process of datafication of the narrative; • characterized by an anxiety-inducing emotional profile, while institutional communication (Scienza, Esperti, and Misure) followed the standards of emergency communication (positive, calm, in control; §4). The data became a tool for emotional activation rather than information. The last part of the presentation concerned the qualitative analysis of the counter-narratives presented in the articles (and not just in the headlines), and the use of data to support them. 1 2 The corpus: headlines on Google We considered the headlines of newspapers and news sites that included the word “coronavirus,” selecting them from the top thirty Google search results, in the period between February 22 and May 15, 2020. This period was chosen in consideration of the volume of user searches for the keyword “coronavirus” on Google (Fig. 1; source: Google Trends). To reduce geolocation (and indexing) bias, we chose a number of results that tended to produce duplicates over the days, ensuring a good variety of local sources. The collected headlines were pre-processed and analyzed with the help of the R software (R Core Team, 2025)1. The choice to focus on the headlines of all search results, rather than on major newspapers, is motivated by two considerations. • The public often limits themselves to reading headlines, especially on social media. • Moreover, while in the first weeks of the pandemic (March 2020) the Audiweb rankings were led by Repubblica and Corriere della Sera, followed by TgCom24 and Il Messaggero (Cazzola, 2020a), those on social interactions show the relevance of free news outlets like Fanpage.it (third, preceded by Corriere and followed in fifth place by Repubblica: Cazzola, 2020c). The Comscore ranking for the same period is led by the publisher of Fanpage (Ciaopeople), while Il Corriere della Sera is only in sixth position (Cazzola, 2020b). While the sample cannot be said to be representative of what was published by news sites, it is indicative of the type of news that had the greatest circulation in the considered period: a total of 2,144 headlines2, for 248 online outlets (national, local, television; Table 1). Figure 1: Web searches for the terms ”coronavirus” and ”covid” Dates represented in the graph. February 23: Establishment of the first red zone (peak of searches); March 9: The “escape from the North”; March 11: “Cura Italia” decree (“I stay at home”); April 4: mandatory masks; April 18: postponement of the easing of restrictive measures; May 16: Phase 2 Decree. 1In particular, the Quanteda (Benoit et al., 2025) and topicmodels (Grün & Hornik, 2024; see also Phan et al., 2008) 2 Table 1: Main news outlets included in the sample News outlet N corriere.it 130 repubblica.it 117 lastampa.it 101 ilsole24ore.com 88 ilmessaggero.it 87 ansa.it 86 ilfattoquotidiano.it 82 tgcom24.mediaset.it 69 rainews.it 63 open.online 53 adnkronos.com 52 lanazione.it 41 tg24.sky.it 41 ilrestodelcarlino.it 39 agi.it 37 quotidiano.net 37 ilpost.it 35 genova24.it 33 ilgiorno.it 33 wired.it 33 ivg.it 30 ilcittadinomb.it 29 lanuovasardegna.it 22 ilmattino.it 18 internazionale.it 17 lagazzettadelmezzogiorno.it 17 liberoquotidiano.it 17 laprovinciacr.it 16 lavocedeltrentino.it 16 altalex.com 15 formiche.net 15 iltempo.it 15 ilgiunco.net 14 quifinanza.it 14 sport.sky.it 14 fanpage.it 13 iltirreno.gelocal.it 13 corrieredellosport.it 12 lapressa.it 12 focus.it 11 la7.it 11 bergamonews.it 10 ilcapoluogo.it 10 ilfoglio.it 10 ilsecoloxix.it 10 riviera24.it 10 salute.gov.it 10 ilfattoalimentare.it 9 m.cronachemaceratesi.it 9 medicalfacts.it 9 News outlet N picchionews.it 9 vanityfair.it 9 vita.it 9 ecodibergamo.it 8 corrierecomunicazioni.it 7 forbes.it 7 ilcentro.it 7 luccaindiretta.it 7 varesenews.it 7 corrieredicomo.it 6 infodata.ilsole24ore.com 6 ispionline.it 6 money.it 6 nationalgeographic.it 6 affaritaliani.it 5 askanews.it 5 avvenire.it 5 chietitoday.it 5 corriereadriatico.it 5 corriereromagna.it 5 cremonaoggi.it 5 fondazioneveronesi.it 5 gazzetta.it 5 ildolomiti.it 5 ilgazzettino.it 5 ilmeteo.it 5 ilpescara.it 5 it.businessinsider.com 5 lasiciliaweb.it 5 milanotoday.it 5 motori.virgilio.it 5 nonsprecare.it 5 rete8.it 5 romatoday.it 5 vaticannews.va 5 bolognatoday.it 4 corrieredellumbria.corr.it 4 ilpiccolo.net 4 iodonna.it 4 osservatoriomalattierare.it 4 ravennanotizie.it 4 sulpanaro.net 4 3 3 Topic modeling: LDA The topic modeling analysis was conducted on a matrix defined by 625 “characterizing” terms, i.e., with high relative frequency but present in a limited number of texts (10%). Following the term reduction, the number of analyzable headlines became 2,102. To determine the number of topics to extract and the distribution’s 𝛼parameter (which regulates the number of topics attributed to each document), we tested several models3. Finally, a 7-topic model with 𝛼equal to 0.2 was chosen4. Figure 2: LDA. Composition of the themes: ”numbers” cluster In LDA (D. Blei et al., 2001;D. M. Blei et al., 2003), a topic is defined as a latent dimension that organizes the vocabulary of a corpus. The procedure reconstructs a posterior distribution of terms for each topic (𝛽matrix), and one of topics for each headline (𝛾matrix)5. This is a technique we could call inductive, thus suitable for the type of sample used, for exploratory purposes. Semantic interpretation of the themes This is achieved by using the information in the 𝛽matrix (terms-topics). Figs. 2and 3present the 10 most significant terms for each theme (the measure is not indicative of frequency). For the purpose of the comparison relevant here, five of the seven themes are divided into two main macro-clusters (§3.1): packages. 22,555 after removing duplicates, which become 2,144 after eliminating those outside the date range or not classifiable as “news.” 3With the ldatuning package (for calculating the optimal number of clusters: Nikita, 2020) and topicdoc (for a posteriori quality measures: Friedman, 2022). 4To reduce the topics assignable to each document, compared to the default parameter (0.1), given the brevity of the texts. 5For the application of this technique to very short texts, cf. Albalawi et al. (2020). 4 • a disciplinary cluster (cool colors): Scienza (Science), Esperti (Experts), and Misure (Measures) and • a numbers cluster (warm colors): Numeri (Numbers) and Bollettini (Bulletins).6 The Cronaca (Daily News) theme largely follows the behavior of the latter, while the Esteri (World) theme is quite varied in its composition. Figure 3: LDA. Composition of the themes: ”disciplinary” cluster Evaluation of the results For the evaluation of the themes, a series of information was considered, also relevant for our discourse, and in particular to show that: • the most widespread themes are Numeri and Scienza; • the Numeri and Bollettini themes are not very varied and not very informative, in the sense that they contain rather common terms (which do not stand out from the corpus); • while, on the contrary, the Scienza theme is the most informative and semantically varied. The measures, presented in Table 2, can be read as follows7: •Prominence: number of documents in which a theme is present (regardless of its weight). •Frequency: number of documents uniquely assigned to the theme (based on the highest value). •Size: the number of terms or fractions of terms present in the distribution of each theme. The total is 625, the number of terms in the matrix. 6Theme glossary: Scienza = Science; Esperti = Experts; Misure = (government) Measures; Numeri = Numbers; Bollettini = (Civil Protection) Bulletins; Cronaca = Daily News; Esteri = World News. 7Other available measures, such as the distance in terms of tf-idf and the coherence index, which refer to the distributions of terms in the texts, are not appropriate in this case, given the brevity of the texts themselves (for details on the calculation and use of the measures, cf. Airoldi et al., 2015, p. 238 et seq.). 5 Table 2: Theme characteristics Theme Prominence (no. headlines) Frequency Size (terms) Exclusivity (terms) Distance from corpus Misure 436 214 109,52 9,68 0,60 Esteri 422 201 94,44 9,64 0,58 Cronaca 461 212 89,32 9,69 0,57 Esperti 446 204 104,40 9,47 0,60 Scienza 465 268 113,13 9,58 0,61 Numeri 539 306 67,27 9,66 0,55 Bollettini 416 226 46,91 9,86 0,59 Total 2.102 1.631 625,00 10,00 – –The most semantically varied theme is Scienza; –The least differentiated is Bollettini, preceded by Numeri (after all, numbers are not included in the term-document matrix). •Exclusivity: how many of the most important terms (the top ten, for example) are exclusive to the theme itself. Considering the technique adopted, the brevity of the texts, and the chosen 𝛼 value, we have a strong specificity of the themes (see also Fig. 5). This means that the themes are well-distinct from each other. •Distance (Hellinger) from the corpus: divergence between the distribution of terms in the theme and that in the entire corpus. The greater the distance, the more informative the identified theme is: –the most distant/informative is Scienza. –the least distant is Numeri. Figure 4: LDA. Dendrogram of the themes 3.1 The separation of “numbers” and science While the two most frequent and transversal themes, Numeri and Scienza represent completely distinct “narratives” of the pandemic. 6 The dendrogram in Fig. 4highlights that the numbers macro-cluster (Numeri and Bollettini) separates clearly and immediately from the other themes, particularly from Scienza and Esperti. This also emerges in more detail in the graph in Fig. 5, which represents the connections between themes and terms (i.e., the 𝛽matrix). The presence of connections (shared terms) between themes can be interpreted as indicative of semantic proximity. Figure 5: LDA. Theme-Term Graph (beta) The table in Fig. 6provides further confirmation of this separation. The values represented are the means of the scores from the 𝛾(gamma) matrix, which is the probability that a headline is associated with a theme (the threshold is 0.14, i.e., 1 / 7 themes): the strong presence of numbers (3 or more occurrences within the headline) characterizes the Numeri and Bollettini clusters, while Scienza and Esperti are characterized by their absence8. Figure 6: Mean gamma values of headlines with numbers 8The numbers variable represents the count of numbers in the headlines, in words or digits. These tokens were identified via POS tagging and then manually corrected, to exclude dates and numbers contained in the expressions such as “covid 19”, “sars-cov 2”, etc. Numbers in digits did not contribute to the definition of the themes. 7 3.2 Numbers prevail over time The average daily distribution of themes shows how they develop over time (Fig. 79): the Numeri theme increasingly accounts for a larger share of the discourse, at the expense, in particular, of Scienza,Esperti, and Misure. Figure 7: Themes over time (presence; 15-day moving average) 4 Sentiment analysis The sentiment analysis was carried out with the affective dictionary ELIta (Di Palma, 2024b,2024a), which provides scores for 6,905 Italian lexical forms, on two classifications of emotions: the VAD model (§4.1) and Plutchik’s wheel of emotions (§4.2). The aggregated scores of the headlines by theme confirm the distance between the two macro-clusters, which have opposite emotional profiles. This result is all the more significant as: • it relies on two different types of classification of the affective tone of the texts (not simply positive/negative); • a theme is attributed by the minimum threshold of the 𝛾matrix value, and this tends to flatten the results (each document can belong to multiple themes). 4.1 The VAD model The VAD model (Russell, 1980;Russell & Mehrabian, 1977), a cornerstone of the dimensional analysis of emotions, is articulated on three dimensions: 9Moving average of the daily 𝛾value; the total for each day is one; the significance threshold is 0.14. 8 Table 3: Mean VAD scores by theme Theme Valence Arousal Dominance Bollettini 0,150 0,534 -0,091 Cronaca 0,052 0,549 -0,190 Esperti 0,223 0,448 -0,012 Esteri 0,105 0,534 -0,105 Misure 0,243 0,453 -0,023 Numeri -0,105 0,560 -0,325 Scienza 0,184 0,504 -0,045 – 0,122 0,512 -0,113 Scale: -4, +4; gamma min.: 0.14 •Valence. Affective polarity (unpleasant/pleasant). •Arousal. Physiological and psychological intensity of an emotion. A more arousing emotion is more intense, regardless of its valence. •Dominance (control). The perceived sense of control over a given emotion or situation10. The average scores, as mentioned above, are very low — close to zero (neutrality: Table 3). However, the distance between the two macro-clusters is evident when considering the profiles, defined by the differences from the means (Figs. 8and 9). Figure 8: VAD: numbers cluster 10They closely resemble the semantic oppositions identified as primary by Osgood et al. (1956): good-bad (evaluation, similar to valence), strong/weak (activity, arousal) and active/passive (power, dominance). 9