scieee AI-readable full text Open interactive document viewer

Quantifying the drivers behind collective attention in information ecosystems

Calleja Solanas, Violeta,Pigani, Emanuele,Palazzi, María J.,Solé-Ribalta, Albert,Suweis, Samir,Borge-Holthoefer, Javier,Meloni, Sandro

Abstract

MJP, AS-R and JB-H acknowledge the support of the Spanish MICINN Project PGC2018-096999-A-I00. EP acknowledges a fellowship funded by the Stazione Zoologica Anton Dohrn (SZN) within the SZN-Open University PhD program. SS thanks the support of UNIPD through ReACT Stars 2018 grant and INFN LINCOLN grant. SM and VC-S acknowledge funding from the project PACSS RTI2018-093732-B-C22 of the MCIN/AEI /10.13039/501100011033/ and by EU through FEDER funds (Away to make Europe), and also from the Maria de Maeztu program MDM-2017-0711 of the MCIN/AEI /10.13039/501100011033. All authors acknowledge the support to the TEAMS project of the Cariparo Visiting Program 2018 (Padova, Italy).

Full text

Journal of Physics: Complexity PAPER • OPEN ACCESS Quantifying the drivers behind collective attention in information ecosystems To cite this article: Violeta Calleja-Solanas et al 2021 J. Phys. Complex. 2 045014 View the article online for updates and enhancements. You may also like Competition-induced increase of species abundance in mutualistic networks Seong Eun Maeng, Jae Woo Lee and Deok-Sun Lee - The dynamical behaviour of type-K competitive Kolmogorov systems and its application to three-dimensional type-K competitive Lotka–Volterra systems Xing Liang and Jifa Jiang - Phase transitions in mutualistic communities under invasion Samuel R Bray, Yuhang Fan and Bo Wang - This content was downloaded from IP address 161.111.10.232 on 13/04/2022 at 07:52 J.Phys.Complex. 2(2021) 045014 (15pp) https://doi.org/10.1088/2632-072X/ac35b6 OPEN ACCESS RECEIVED 30 June 2021 REVISED 17 October 2021 ACCEPTED FOR PUBLICATION 2 November 2021 PUBLISHED 26 November 2021 Original content from this work may be used under the terms of the Creative Commons Attribution 4.0 licence. Any further distribution of this work must maintain attribution to the author(s) and the title of the work, journal citation and DOI. PAPER Quantifying the drivers behind collective attention in information ecosystems Violeta Calleja-Solanas1,∗,EmanuelePigani 2,6, María J Palazzi3, Albert Sol´ e-Ribalta3,4, Samir Suweis2,5, Javier Borge-Holthoefer3and Sandro Meloni1,∗ 1IFISC, Institute for Cross-Disciplinary Physics and Complex Systems (CSIC-UIB), 07122, Palma de Mallorca, Spain 2LIPh Lab, Physics and Astronomy Department, University of Padova, via Marzolo 8, 35131, Padova, Italy 3Internet Interdisciplinary Institute (IN3), Universitat Oberta de Catalunya, Barcelona, Catalonia, Spain 4URPP Social Networks, University of Zurich, Zurich, Switzerland 5Padova Neuroscience Center, University of Padova, Via Orus 2/B, 35131 Padova, Italy 6Stazione Zoologica Anton Dohrn, Villa Comunale, Naples, 80121, Italy ∗Authors to whom any correspondence should be addressed. E-mail: violeta@ifisc.uib-csic.es and sandro@ifisc.uib-csic.es Keywords: information processing, competition for attention, information ecosystems Abstract Understanding human interactions in online communications is of paramount importance for our society. Alarming phenomena such as the spreading of fake news or the formation of echo-chambers can emerge in unhealthy communication environments and, ultimately, undermine the democratic discourse. In this context, unveiling the individual drivers that give rise to collective attention can help to conserve the health of our information ecosystems. Here, following a recently proposed analogy between natural and information ecosystems, we explore how competition for attention in online social networks and the strategies adopted by the users to maximize their visibility shape our communication dynamics. Specifically, by analyzing large-scale datasets from the micro-blogging platform Twitter and performing numerical modeling of the system dynamics, we are able to measure the amount of competition for attention experienced by users and how it changes when exogenous events captivate collective attention. The work relies on topic modeling to extract users’ interests and memes context from the data and a framework based on ecological niche theory to quantify the strength of negative (competitive) and positive (mutualistic) interactions for both users and memes. Interestingly, our findings show two different behaviors. While memes undergo a sharp increase in competition during exceptional events that can lead to their extinction, users perceive a decrease in effective competition due to a stronger effect of mutualistic interaction, explaining the focus of collective attention around specific topics. Finally, to confirm our results we reproduce the observed shifts with a data-driven model of species dynamics. 1. Introduction Online social networks (OSNs) have not only transformed the way in which we communicate, but also how we access and process information. In fact, since their appearance, they have played a double role: as spaces to build social interactions and as news media platforms. Consequently, our communication model has switched over from a centralized mass media environment and face to face interactions, to an era when all the actors are, at the same time, information sources and receivers. This duality and the new paradigm it induced, make also OSNs a perfect example of social information processing, leading to emergent phenomena such as viral information spreading [1], fake news [2] and the formation of echo-chambers [3,4]. These changes have also exposed our cognitive limitations. Our brain passed from hearing few broadcast information sources, to be bombed by millions of messages demanding our attention. This has led to a variety of phenomena, such as an acceleration of social dynamics [5], and terms like ‘competition for attention’ entered © 2021 The Author(s). Published by IOP Publishing Ltd J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al our everyday vocabulary. Moreover, OSNs are often designed to captivate our brain and maximize the time we spend on them by providing reinforcing feedbacks [6] and instant gratification [7,8]. These limitations along with the huge amount of information we produce every second, induce competition between ideas/memes for visibility, and users tend to adopt different strategies to increase the chances that their messages would be read. To understand how these low-level drivers shape the way our society processes information, in recent years, different approaches have been proposed. From a data analytic perspective, several works demonstrated the role of competition in information diffusion [9,10], our social interactions [11–14] and the quality of the information we share [15]. All these results also found confirmation from theoretical models, for example for the distribution of memes popularity [16–18] or how echo-chambers emerge [19]. This strong emphasis on the effects of competition and on how users respond to it, also led researchers to draw an analogy between information and natural ecosystems [5,20–22]. Actors (e.g. users and memes) of OSNs are seen as species in ecological communities where they seek to maximize their abundance—visibility in this case—and all are competing for limited resources (e.g. individuals’ attention). Moreover, the communication strategies adopted by the users like, for example, the use of specific hashtags to provide context to their messages can be represented as mutualistic, as they favor both visibility of users and growth of certain memes. One example of how this analogy can lead to insights on the functioning of information ecosystems has been the work of Borge-Holthoefer et al [20], who studied the organization of online discussions around social protests in Spain in 2011 as an ecological network, finding that its structure evolved toward a nested architecture, very close to the typical organization of natural mutualistic assemblages [23,24]. Moreover, building on these results, in a recent work Palazzi et al [21] proposed an ecology-inspired model [25,26]toexplainthe structural flexibility showed by OSNs in response to external shocks such as breaking news. In [5], instead, the authors were able to explain the acceleration in collective attention they found in the digital streams and cultural items by employing a mathematical model based on Lotka–Volterra dynamics—a theoretical framework often used in population dynamics and theoretical ecology [27,28]. Finally, from a more theoretical perspective, models based on the concept of ecological neutrality [29,30], were able to reproduce several emergent patterns found in online communications [22]. Here, following this research line, we exploit the similarities between natural and information ecosystems to understand the main drivers that shape collective attention. Specifically, analyzing different datasets extracted from the online platform Twitter and an ecology-inspired numerical model, we are able to quantify the intensity of the competitive and mutualistic interactions experienced by each user, and how these change when exogenous events focus collective attention. We start by relying on topic modeling to study the evolution of users’ interests over time, along with the diversity of the discussion around the different topics. In this way, we can measure the similarity between users and memes and then, employing a modeling framework based on ecological niche-theory [26], build interaction networks similar to ecological communities with both competitive and mutualistic terms. We then analyze the obtained networks with the aim of understanding how variations in the amount of competition and mutualism experienced by the users can explain the focus of collective attention around one or few dominating topics. These analyses allow us to explain these shifts in attention in terms of a reduction of the effective competition experienced by the users. We then confirm this finding by reproducing the patterns we observe in the data with numerical simulations of an ecological model based on species abundance maximization [21]. Our results not only shed light on how information processing drives users’ behaviors in social media but, in a broader sense, they spotlight the opportunities that an ecological approach can offer to the study of information ecosystems for understanding collective phenomena in techno-social systems The rest of the work is organized as follows. In the next section, we introduce the datasets we will employ in the analysis, the topic modeling technique used to extract information topics from Twitter discussions, the ecologically inspired numerical model and the methods to quantify the strength of competitive and mutualistic interactions from the data. Then, in sections 3.1 and 3.2 we present our results for the time evolution of topics and users’ interests respectively while, in section 3.3, we study the user–meme interaction networks. Moreover, in section 3.4, we confirm our previous results with the numerical simulations of our model. Finally, with section 4we summarize the main findings of this work and draw our conclusions. 2. Methods and data 2.1. Twitter events datasets Since we are interested in quantifying the changes in collective attention during exceptional events, we decided to focus on data from the online social platform Twitter. Twitter, along with being a social network and due to its microblogging nature, is also often considered as a news and information media. Thus, social, political 2 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al Table 1. Summary of the datasets used in the study. Dataset Starting date Ending date # of tweets # of unique hashtags # of unique users Spanish elections 20-04-2019 04-05-2019 4882 546 2952 41 314 Catalan referendum 01-09-2014 13-11-2014 220 364 18116 78 270 Nepal earthquake 09-05-2015 18-05-2015 5641 719 41 589 1706 559 Figure 1. Time evolution of the number of tweets for the three datasets considered in the text: Spanish general elections in 2019 (panel (a)), self-determination referendum in Catalonia in 2014 (panel (b)) and the earthquake that hit Nepal in 2015 (panel (c)). Relevant events are highlighted in gray. or natural events will be likely reflected in its users’ activity, allowing us to record both the change in their interests and in their interactions. 2.2. Information topics For our analyses, we focus on three large datasets describing Twitter activity: the consultations and protests around the self-determination referendum organized by the Catalan government in November 2014 [21]; the general elections held in Spain in 2019 [21] and the response to the earthquake that hit Nepal in 2015 [31]. All the datasets have been collected through the Twitter streaming API and are publicly available (see appendix A for details on the data collection process and their availability). Table 1summarizes the main features of the datasets used. We choose those datasets because they cover both expected and unexpected events [32,33]: from political discussions with a fixed timeline (Spanish elections), to political/social protests like the Catalan referendum with a mixture of organized activities and improvised protests, and an inherently unexpected natural disaster (Nepal’s earthquake). In particular, since our aim is to study how exceptional events shape the communication dynamics in online discussions, we focus only on events that are exogenous to the classical OSNs dynamics and where we can have a detailed view on their evolution—e.g. salient periods can be easily identified from the news and other media. In order to extract information topics from the tweets, for all the datasets, we only considered tweets with at least one hashtag (with the only exception of the Nepal earthquake dataset, where due to the extremely large size we had to restrict our analysis to only hashtags that have been posted at least 100 times). From each tweet, the information considered is: the timestamp of the tweet, the user-ID (previously anonymized) of the user who posted the tweet and all the hashtags present in it. In this way, we were able to extract users’ interests and their evolution throughout the discussion, along with hashtags co-occurrences, fundamental to infer the information topics. Figure 1presents the time evolution of the number of tweets for the datasets considered. In all the cases, it is easy to see peaks in the activity, usually related to external events, like the television debate between candidates or the polling day for the Spanish elections or the referendum day for the Catalan referendum. These high activity periods are usually preceded and followed by calmer periods characterized by a lower and constant activity. Following this separation, we split each dataset into different parts, each corresponding to a ‘peak’ or ‘rest’ period, and study how users’ interests and competition for attention varies between them. In particular, our aim in splitting the datasets is to have periods small enough to clearly distinguish between peaks, containing only a single event, but also to have enough data during the rest intervals. Thus, for each dataset, we identified the largest spike in activity and made its duration the size of the period. For example, for the Catalan referendum, since the largest spike lasts around 10 days, we decided to divide the dataset in intervals of exactly 10 days. We applied a similar reasoning for the Nepal earthquake dataset, in which the activity peak lasts 3 days so, we have 3 periods of 3 days each. However, for the Spanish elections dataset, due to its short duration and 3 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al Figure 2. Example of topics extracted from the Spanish general elections dataset. The word-clouds represent the most abundant hashtags for four selected topics detected with our method. Hashtags’ size in each figure is proportional to the number of times it has been tweeted. the effects of circadian rhythms, we were forced to have slightly uneven intervals to guarantee enough data during the resting periods. Once the datasets and the different periods are defined, we focus on how to extract users’ interests from the raw data. This step is fundamental to measure the similarity between users and, eventually, quantifying the competition for attention that both users and memes experience over the development of the events. To extract users’ interests from their timeline, we rely on a network theory approach based on hashtags co-occurrence, designed specifically to infer information topics in Twitter [34–36]. Hashtags are used in many social media as keywords to indicate the content of a message and they often represent memes, whose meaning is known to the users. They thus provide a concise indication of a tweet’s semantic context. In this way, the appearance of two or more hashtags in the same tweet is often the sign of a semantic association between them—they belong to the same topic—as it happens to words co-occurring frequently in texts [37,38]. Based on this association, we extract information topics as cohesive clusters of hashtags that significantly appear together in different tweets. To do so, for each period considered in the data, we build a weighted co-occurrence network between hashtags, where a link is laid if two hashtags appeared together in the same tweet and its weight represents the number of different tweets where they co-occurred. In order to eliminate spurious links and assure the significance of the semantic associations between the hashtags, we then remove all the links which weight is equal or smaller than 3. Once the networks have been built, information topics can be detected as clusters of densely connected hashtags—i.e. communities in the graph. Following the same procedure proposed in [35,36], we thus employ the OSLOM tool [39] to detect communities in the different networks. We decided to use OSLOM for community detection because it is able to extract overlapping communities where nodes can belong to more than one community at a time. Thus, in our case, it is able to catch the fact that hashtags with different meanings can be part of several topics. Figure 2represents an example of topics extracted from the Spanish elections dataset. In the four topics the names of the distinct national parties are clearly identifiable along with abbreviations of the most important dates. Once the information topics are extracted from the data, we can use them to detect users’ interests from the hashtags they posted, and their membership to the different topics. For each user, we build a feature vector u whose entries uiquantify user’s interest toward topic i. Specifically, uiaccounts for the number of times the user posted a hashtag belonging to topic iin their tweets. Moreover, since hashtags can be part of different topics, we split the weight of each hashtag according to the number of communities it belongs to. For example, a hashtag part of only one topic twill contribute to the pertinent utby adding one. On the contrary, one belonging to 4 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al Figure 3. Sketch of the user–topic and hashtag–topic vectors extraction method and creation of the competitive and mutualistic matrices. (panel (a)) Hashtags are classified into topics using the method proposed in [35]. (panel (b)) Hashtag–topic vectors are built directly with the membership of each hashtag to one or more topics. User–topic vectors can be calculated from the hashtags tweeted by each user. If a hashtag belongs to more than one topic its weight is split evenly between all the topics it participates. (panels (c), (d) and (e)) From the user and hashtag vectors, cosine similarity is employed to calculate topic overlap. (panels (c) and (e)) The hashtag–hashtag and user–user competitive matrices are calculated proportionally to topic similarity between users (hashtags). (panel (d)) Finally, the mutualistic matrix is built as the similarity between users and hashtags. two topics will contribute 0.5 to each of the two corresponding entries of u.Infigures3(a) and (b) we provide a diagram showing how the hashtags posted by the users are classified into topics and then used to build the user vector u. The same procedure can be then applied to hashtags to build an equivalent hashtag vector hfor each of them, accounting for their participation in the different topics. User vectors quantify how individual attention is distributed across the different topics and how external eventscan focus it over a specific theme. However, individual interests are also a proxy of competition for attention. If everyone is interested in a specific topic, hashtags on that topic will experience a stronger competition for spaces in the users’ screens. In the same way, users compete with their peers to get their messages read. On the other side, a user can decide to post about popular topics to have the advantage of reaching a potentially larger audience, leading to a sort of mutualistic interaction. To quantify how much competition users experience in peak and resting periods in our datasets, we measure how similar the user vectors are one to the other. For each couple of user vectors uand vwe calculate the cosine similarity: simcos(u,v)=u·v uv.(1) 5 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al In this way, we have a measure that ranges from 0 (users with totally disjoint interests) to 1 (totally aligned users). Finally, hashtag competition can be estimated in the same way using the hashtag vectors hinstead of u (see section 2.4 for details). 2.3. Dynamical model of users’ attention Defining users and hashtags similarity allows us to quantify the amount of competition and mutualism experienced by each node. However, to get further insights on the driving mechanisms behind these changes in collective attention, we also employ an ecology-inspired visibility optimization model proposed to explain structural changes in the user–hashtag interaction networks [21]overthecourseofanevent. Following the analogies between ecological communities [25,26] and information ecosystems [21,22], the model is based on the assumption that the attention dynamics recorded in the data is the result of an optimization process where users aim to maximize their visibility. Users and hashtags are thus seen as species of a mutualistic assemblage belonging to two different classes—e.g. plants and pollinators in natural ecosystems. Competition takes place between species of the same guild (user–user or hashtag–hashtag), while mutualistic interactions occur between species of different guilds (user–hashtag), with the species dynamics modeled by Lotka–Volterra equations with a Holling-type II functional response [24,25]: dnU i dt=nU i⎛ ⎝ρU i− j βUU ij nU j+kγUH ik θUH ik nH k 1+hkθUH ik nH k ⎞ ⎠, dnH i dt=nH i⎛ ⎝ρH i− j βHH ij nH j+kγHU ik θUH ik nU k 1+hkθHU ik nU k ⎞ ⎠, (2) where nU iand nH istand for the abundance (visibility) of species ipart of the users’ (U) or hashtags’ (H)guild, while ρU iand ρH irepresent the respective growth rates and hthe handling time of the Holling-type II mutualistic functional response. The intensity of the users and hashtags competitive and mutualistic interactions is defined by matrices βUU ,βHH and γUH , respectively. Finally, θis the adjacency matrix of the mutualistic interactions accounting for the hashtags produced by each user in their posts. TheoptimizationprocessatthebasisofthemodelisthesameastheoneproposedinSuweiset al [25], where species (users) rewire their mutualistic connections to randomly selected partners (hashtags) and the new links are kept only if they lead to an increase in abundance (visibility). Otherwise, the original connection is restored. Specifically, at constant time intervals, a random user Uis selected and one of its existing connections to hashtag His rewired to a new hashtag H. The link to be rewired is selected with probability pUH ∝1−k−1 H where kHis the degree of H. After the rewiring, we let the system evolve until the equilibrium is reached. If at the end of this period U’s abundance is greater than the previous one, the link is kept; otherwise, the previous configuration is restored. It is important to notice that in our model only users maximize their abundance (visibility) since they choose the hashtags to post in their tweets. Thus, changes in the hashtag networks are due only to users’ actions. Finally, information topics are modeled as ecological niches [40] associated to each species. For the users, they represent the set of individual interests while, for the hashtags, they define their semantic context. 2.4. Estimating competition and mutualism from data Once the species dynamics are defined, we need to estimate the interaction strength among species (βUU ,βHH and γUH ) from our data. To do so, we follow the approach proposed by Cai et al [26], where competition and mutualistic strengths are proportional to the niche overlap between species of the same (competition) and opposite (mutualism) guild. This way, we can estimate niche overlap Ggg ij between species ibelonging to guild gand species jof gfrom the data as the cosine similarity between the topics vectors of iand j.Then,the elements of the competition matrices βHH ij and βUU ij are simply proportional to the ovelap: βHH ij ∝Ωc·GHH ij and βUU ij ∝Ωc·GUU ij ,withΩca global factor to tune the absolute strength of the interaction. In a similar way, we can define the mutualistic strength as: γUH ik ∝Ωm·GUH ik , with the only difference that, in the model, the mutualistic strength γUH ik is then multiplied by θUH ik toincludethefactthatuserimay (θUH ik =1) or may not (θUH ik =0) interact with hashtag k.Finally,tobeabletocomparethetwostrengthsweimposeΩc=Ω m. 3. Results We start our analyses by dividing our datasets in different periods, depending on their activity (i.e., the overall number of tweets produced). We thus distinguish between what we define as peak periods, when external events produce spikes in Twitter activity, such as the television debate day for the Spanish elections dataset 6 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al Figure 4. (panels (a)–(c)) Time evolution of the number of tweets in the different datasets divided by activity periods. (panels (d)) Distribution of activity, in terms of number of tweets generated by each topic, for the different periods for the Spanish elections dataset. To ease visibility only the largest ten topics are showed. Peak periods are marked by the gray bars. or the referendum day for the Catalan dataset; and resting periods, where the discussion is still active but not driven by major external events. In this way, we can use the resting periods to define a baseline for activity and collective attention and compare them with the peak periods. Figures 4(a)–(c) show the split of the three datasets we are considering. For the Spanish elections dataset (figure 4(a)) we define five different periods with the second and fourth covering large events, while the first and fifth represent our baseline. The third one is a mixture of peak and resting periods, since it is characterized by an increased activity but only smaller events took place. We adopt a similar split also for the Catalan selfdetermination referendum dataset. In this case, we defined seven periods, with the first and seventh as peaks, while the second, fourth and sixth as resting states. The remaining two (third and fifth intervals) are classified as mixed. Finally, in the Nepal earthquake dataset we identify three periods, where the first one represents the resting state, the second the highest activity and the third one a mix of peak and rest. Once defined our baseline and active periods, we apply the topics extraction procedure to each period and study how the hashtags communities and users’ interests evolve through time. To do so, for each period we compute the hashtag co-occurrence network and extract the most significant communities from it, defining our topics and hashtags’ and users’ membership as described in section 2.2. 3.1. Topics evolution From the networks it is possible to study how information topics evolve, and the corresponding discussions re-arrange around the external events. Table 2summarizes the main features of the co-occurrence networks obtained from the data. A result that first catches the eye is the fact that, even if the networks for different periods have different sizes—e.g., the number of nodes/hashtags for the Catalan referendum dataset ranges from 310 to 1602 between the least and most active periods—the number of communities/topics is more or less constant. This contrasts with our assumption that, during relevant events, collective attention focuses only on one or a few ‘important topics’. This discrepancy is solved if we consider not only the number of topics but also the activity around them, in terms of the number of tweets they generate. Figure 4(d) shows how tweets are distributed among the ten largest topics for the different periods of the Spanish elections dataset. Although activity in general is highly heterogeneous, however, during peak periods (marked by gray bars) this tendency 7 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al Table 2. Features of the co-occurrence networks extracted from each considered dataset. Dataset Period # of hashtags # of links # of topics Spa. elec. 1 1344 6866 75 2 1714 11587 68 3 1899 12026 77 4 1736 12373 63 5 1313 5114 60 Cat. ref. 1 703 3170 9 2 310 863 20 3 586 2070 10 4 350 1037 21 5 338 911 16 6 362 1112 13 7 1602 5837 12 Nepal 1 471 1564 29 2 771 4739 27 3 670 2921 33 Figure 5. Alluvial diagram showing the evolution of topics for the Catalan self-determination dataset. Colors represent the different periods and each box represents a topic (a community in the co-occurrence network). Box sizes is proportional to the number of tweets generated by each topic and, for each period, only the four largest topics are show. Flows represent the number of tweets moving from one topic to the other. exacerbates with one or two topics generating almost the 90%of the tweets (we obtain similar results for the other two datasets, not shown). Taken together, the last two results (i.e., the almost constant number of topics over different periods, as shown in table 2, and the higher heterogeneity in attention during peak events presented in figure 4(d)) depict a clear scenario. During normal activity periods, users’ interests are focused on specific topics and hashtags compete for attention inside individual discussions. When an external event captures collective attention, original topics remain active but generate way less activity than before. To visualize how attention flows between the different topics over time, figure 5shows an alluvial diagram of the activity around the topics in the Catalan referendum dataset (we obtain similar results for the other datasets, not shown). In each block (relative to a time period) the size, in terms of number of tweets, and the flows of activity for the four largest topics are shown. As a confirmation of the results in figures 4(c) and (d), during peak periods (first, third and seventh) all flows converge around one single topic. When external events fade out, flows split back to different discussions or, some hashtags stop to be used during calm intervals (e.g. the second period) to regain interest in high activity periods as shown, for example, by the large flows moving from the first period to directly the third. 3.2. Users’ similarity Once studied how exceptional events shape hashtags dynamics, we can now focus on the other side of the coin: users’ behavior. We use the cosine similarity between user–topic vectors to quantify interests overlap between 8 J.Phys.Complex. 2(2021) 045014 (15pp) V Calleja-Solanas et al [36] Mussi Reyero T, Beir´ o M G, Alvarez-Hamelin J I, Hernández L and Kotzinos D 2021 Evolution of the political opinion landscape during electoral periods EPJ Data Sci. 10 31 [37] Turney P D and Pantel P 2010 From frequency to meaning: vector space models of semantics J. Artif. Int. Res. 37 141–88 [38] Martinez-Romo J, Araujo L, Borge-Holthoefer J, Arenas A, Capitán J A and Cuesta J A 2011 Disentangling categorical relationships through a graph of co-occurrences Phys. Rev. E84 046108 [39] Lancichinetti A, Radicchi F, Ramasco J J and Fortunato S 2011 Finding statistically significant communities in networks PLoS One 61–18 [40] Williams R J and Martinez N D 2000 Simple rules yield complex food webs Nature 404 180–3 [41] Crowley P H and Cox J J 2011 Intraguild mutualism Trends Ecol. Evol. 26 627–33 [42] Stanton M L 2003 Interacting guilds: moving beyond the pairwise perspective on mutualisms Am. Nat. 162 S10–23 15