scieee AI-readable full text Open interactive document viewer

New index for the analysis of emotions in social networks (SNAI): Application in the political context [Review]

Sorrel, Miguel A

Abstract

Autonomous University of Madrid

Full text

New index for the analysis of emotions in social networks (SNAI): Application in the political context Comments to the authors The authors introduce an easy-to-understand index to summarize the variables of motivation and emotion in Tweets, so that it can be predictive of the variables of interest (percentage of votes won/lost in elections). The November 11, 2019 national elections held in Spain are presented as a validation case. I enjoyed reading the article and I think it is a promising proposal. The data were publicly available by the authors. Without diminishing the good work conducted by the authors, I believe it would be advisable to address some of the following points. I hope the authors find my comments useful. Major comments 1. More information on the quality of the data should be provided. Specifically: 1.1. It is commented in the introduction that the presence of very active users, spam generated by bots, or individuals specifically dedicated to propaganda may distort the results when analyzing data on politics. How has this fact been handled in this article? I understand that this is one of the reasons for quantifying indirect motivation (i.e., presence of retweets) as 0 or 1. However, the authors probably have access to the number of tweets per user (assuming the “id” variable in the database was this identifier would be the number of unique values in “id”) or the number of users generating the total number of tweets (frequency distribution of “id”). It would be helpful to break down and discuss this information in the subsection Sample of the Method. 1.2. It would also be useful to provide in Table 3 the raw number of tweets per politician. It may be interesting to break down the PANAS test into words to discuss that result later in the discussion. Unless I am mistaken, the distribution of tweets by PANAS word is very non-uniform (9951 tweets for "Proud" but only 3 for "fearful". Is it possible that negative content is being captured worse? (the latter could already be anticipated from Table 3 presented by the authors). 2. Choices regarding index formulation could be better justified. Specifically: 2.1. The article equates "number of words" to "Direct motivation", "presence of retweet" to "Indirect Motivation" and "presence of emojis and exclamations" to "Emotional Intensity". I believe that these equivalences should be better substantiated, either with a theoretical basis (linking, for example, text length with motivation) or with previous studies that have made these operationalizations. 2.2. Likewise, the choice of the SNAI operationalization as DM*(IM + EI) should be justified. For example, why give the same importance to the use of emojis as to the use of exclamations? What other variables (e.g., use of superlatives) could be used but are not included? 2.3. In relation to this, it seems strange not to take into account in some way the number of times RT has been made (having 1 RT is equated to having 1,000). I understand the authors' stance to preserve the index metric manageable. To support their decision, could they describe the number of RTs per tweet? That way the reader could assess the impact of this decision. In general, please identify the decisions made regarding the operationalization of the index (such as those mentioned here) and provide theoretical or practical justifications. 3. The high sample size makes interpreting statistical significance meaningless. For example, in this text, virtually zero correlations are interpreted: “a direct and significant relationship with EI (No. of emojis/exclamations; r = 0.022, p < .0001); in other words, the longest tweets… have the highest number of emojis and exclamations. By contrast, IM has an inverse relationship with EI (r = -0.049, p < .0001); this means that those tweets that are shared have fewer emojis/exclamations”. The same is true for the ANOVA results, where sometimes effects are interpreted with size 0. It would be advisable to completely revise these interpretations. 4. Finally, since the tool (the index) has been generated, I believe that validating it in a single scenario with only 5 values to predict (change in % vote for 5 politicians) may be insufficient. I understand the cost of doing this but the message would be much more powerful if more evidence in favor of the index was shown. It would be useful for the authors to present other validation scenarios (e.g., other regional or other elections). At least, this is something that should be mentioned in the discussion if you choose to leave this work as preliminary validation and wish to include more validation scenarios in a follow-up article. Minor comments 1. The index metric goes from 0 to 280 but in the data it goes from 0 to 58. I think it would be worth discussing this. It seems that the tweets were somewhat short, with about 24 words on average. 2. Table 5 should be translated into English.