scieee AI-readable full text Open interactive document viewer

Identifying and characterizing social media communities: a socio‑semantic network approach to altmetrics

Arroyo Machado, Wenceslao,Torres Salinas, Daniel,Robinson García, Nicolás

Abstract

Funding for open access charge: Universidad de Granada/CBUA. This work has funded by the Spanish Ministry of Science and Innovation grant number PID2019-109127RB-I00/SRA/10.13039/501100011033. Wenceslao Arroyo-Machado has an FPU Grant (FPU18/05835) from the Spanish Ministry of Universities. Daniel Torres-Salinas is supported by the Reincorporation Programme for Young Researchers from the University of Granada. Nicolas Robinson-Garcia is funded by a Ramon y Cajal grant from the Spanish Ministry of Science and Innovation (REF: RYC2019-027886-I).

Full text

Vol.:(0123456789) Scientometrics (2021) 126:9267–9289 https://doi.org/10.1007/s11192-021-04167-8 1 3 Identifying andcharacterizing social media communities: asocio‑semantic network approach toaltmetrics WenceslaoArroyo‑Machado1 · DanielTorres‑Salinas1 · NicolasRobinson‑Garcia1 Received: 26 April 2021 / Accepted: 16 September 2021 / Published online: 12 October 2021 © The Author(s) 2021 Abstract Altmetric indicators allow exploring and profiling individuals who discuss and share scientific literature in social media. But it is still a challenge to identify and characterize communities based on the research topics in which they are interested as social and geographic proximity also influence interactions. This paper proposes a new method which profiles social media users based on their interest on research topics using altmetric data. Social media users are clustered based on the topics related to the research publications they share in social media. This allows removing linkages which respond to social or personal proximity and identifying disconnected users who may have similar research interests. We test this method for users tweeting publications from the fields of Information Science & Library Science, and Microbiology. We conclude by discussing the potential application of this method and how it can assist information professionals, policy managers and academics to understand and identify the main actors discussing research literature in social media. Keywords Network analysis· Socio-semantic networks· Altmetrics· Twitter· Information science and library Science· Microbiology Introduction Research literature is increasingly mentioned, shared and discussed on social media. This represents a substantial challenge as well as an opportunity to anyone trying to study the interactions that take place in the digital environment (Stieglitz etal., 2018). It provides researchers with major opportunities to develop novel methodological solutions by which to inform policy managers, journalists and information professionals on the way by which scientific literature is consumed. In vastly differing fields, many ad hoc solutions exemplify the growing interest in social media. In the field of science communication, for example, research has been conducted into the anti-vaccine movement on Twitter (van Schalkwyk etal., 2020), the dissemination of fake medical news (Waszak etal., 2018), or political * Wenceslao Arroyo-Machado [email protected] 1 EC3 Research Group, Department ofInformation andCommunication Sciences, Faculty ofCommunication andDocumentation, University ofGranada, Granada, Spain 9268 Scientometrics (2021) 126:9267–9289 1 3 communication and the influence of Twitter (Davis etal., 2017). In marketing, a substantial, growing number of social media metrics and analytics have been applied (Misirlis & Vlachopoulou, 2018). In disaster management, information propagated by social media such as Facebook and Twitter has formed the basis for new proposals (Kim & Hastak, 2018); and the digital humanities’ community on Twitter has been identified and analyzed (Grandjean, 2016). In scientometrics, these studies have led to the emerging sub-field of altmetrics (Priem et al., 2010), in which mentions to scientific literature on social media are tracked to explore the social reception of research findings. However, this line of research has not been free of controversy. Initial high expectations of the potential value of tracking aspects of social or broader impact on research (Bornmann etal., 2019; Haustein, 2016) were soon rejected in the face of hard evidence (Robinson-Garcia etal., 2017; Sugimoto etal., 2017). Nonetheless, the relevance of social media in scholarly communication remains unquestioned (Robinson-Garcia etal., 2018; Wouters etal., 2019), leading to a new scenario in which novel metrics are being developed to understand and describe aspects of science communication that transcend traditional academic channels. The rich variety of social platforms (Wikipedia, Mendeley, Twitter, and so on) has given rise to the development of altmetric data aggregators that provide data on a variety of social media sources. These include Altmetric.com, CrossRef Event Data, or Plum Analytics, among others. Despite the evident advantage of offering unique data access points, they do have limitations. Zahedi and Costas (2018) systemically compared altmetric data providers’ coverage, metrics and sources. They found differences in data collection, the identification and merging of different versions of a single publication, and data update periodicity. These can be added to other limitations directly related to the nature of social media and the concept of altmetrics, namely heterogeneity, quality and dependencies (Haustein, 2016). For a variety of reasons, Twitter is the social media platform that has received most attention since the earliest days of altmetric studies. In part, this is because it is the public forum with the second-highest figures for coverage of scientific literature mentions after Mendeley (Robinson-Garcia etal., 2014). Nonetheless, while it is widely used by the general public, it has a relatively low level of acceptance among scientists. Most studies report that around 15% of academics have a Twitter account (Haustein, 2019), although the annual growth rate is constant (Joubert & Costas, 2019). After initially promising results (Eysenbach, 2011), studies report that Twitter mentions to scientific papers poorly reflect citation impact (Haunschild & Bornmann, 2018). Furthermore, the inclusion of automated bots (Haustein etal., 2016) and the un-informative way in which scientific papers are tweeted (Robinson-Garcia etal., 2017) question the extent to which simple counts of tweets mentioning papers can be informative. Many studies have focused on characterizing the Twitter profiles of individuals who tweet scientific literature to better understand who they are (Díaz-Faes etal., 2019; Ke etal., 2017). The present study adds to this growing trend in the literature by proposing a methodological approach through which communities of actors can be identified on the basis of their scientific preferences. Our goal is to develop tools that can inform on targeted groups interested in specific topics which can later be characterized by other methods, as mentioned earlier. To achieve this, we build on previous studies that investigated differences in topics of interest across social media platforms (Arroyo-Machado etal., 2019; Robinson-Garcia etal., 2019). The paper is organized as follows: first, we briefly review the literature and focus on three specific topics, Altmetric studies, studies specifically about Twitter, and studies 9269 Scientometrics (2021) 126:9267–9289 1 3 relating to mapping and visualization techniques. Secondly, we formulate our objectives. We then describe our data retrieval and data processing and present our methodological proposal. We apply this in the field of Information Science & Library Science and in the field of Microbiology. We conclude by discussing our findings. Background Altmetric studies Altmetrics were formally proposed in 2010 with the publication of the Altmetrics Manifesto (Priem etal., 2010), although similar proposals had appeared previously (Neylon & Wu, 2009; Nielsen, 2007; Taraborelli, 2008). The emergence of altmetrics led to a fundamental transformation of the field of scientometrics. This occurred at a time when different metrics, sources and indicators co-occurred, moving the field from an almost universal dependence on certain bibliometric databases to a heterogeneous range of data sources. Although scientometricians acknowledged the technical limitations of altmetrics from the very beginning (Torres-Salinas etal., 2013), an overall optimism led many to consider them an alternative to citation metrics and compared and analyzed their relationship with traditional metrics (Costas etal., 2015; Thelwall, 2018). But, apart from Mendeley (Thelwall, 2018), evidence only suggests the existence of a weak positive correlation. This led to a change in the discourse and altmetrics began to be presented as a complement to citations (Haustein etal., 2015), rather than an alternative. While acknowledging their potential to inform on other indicators of scientific information consumption, there seems to be a consensus that they cannot be interpreted uniformly and that context plays an important role in their interpretation. This has led many to refer to altmetric indicators as metrics that capture an ‘unknown impact’ of scientific outputs (Bornmann etal., 2019; Kassab etal., 2020). Since then, effort has been directed at studying the context in which this unknown impact is produced, identifying new channels of scholarly communication that go beyond the traditional (Holmberg etal., 2019). This shift has led some authors to refer to these new studies as studies on social media metrics (Wouters etal., 2019) and define them as ‘second generation metrics’ (Díaz-Faes etal., 2019). While the previous one transferred the citation model to social media, here the focus is on the activity and interactions that take place on social media. This leads to a new scenario in which the altmetric research is focused on the relational attributes of the social media activity rather than focusing on features (i.e., impact) related to scientific publications. To do so, the methodological framing has also changed, focusing now on techniques which help discover and analyze different kinds of social interactions (Costas etal., 2020) that allow a better understanding of science-society relations. However, these new approaches focus mainly on researchers discovering and topic visualizations in social media. But how can communities of social actors with the same interests be identified? Can communities of social actors who consume scientific literature outside the scientific realm be identified? Numerous examples of these novel approaches to the use of altmetrics can be found in the literature. Table1 summarizes 14 such methodological proposals. Essentially, these fall into three categories of application or approach: identification and characterization of researchers; visualization of topics discussed; and knowledge maps, which center on descriptive analyses and co-citation and co-word network analyses. Also, most of these 9270 Scientometrics (2021) 126:9267–9289 1 3 Table 1 Main altmetric studies and methodological proposals by source of literature Application Methodology Source Scope Zahedi and van Eck (2018) Profiling of Mendeley readers Descriptive analysis and overlay visualizations Mendeley Multidisciplinary Costas etal. (2017) Identify researchers on Twitter Rule-based methods Twitter Multidisciplinary Ke etal. (2017) Profiling of Twitter users Descriptive and network analysis Twitter Multidisciplinary Alperin etal. (2018) Effectiveness of Twitter dissemination on outreach Descriptive and network analysis Twitter Biomedicine Robinson-Garcia etal. (2018) Profiling researchers on Twitter Social network analysis Twitter Multidisciplinary Díaz-Faes etal. (2019) Characterize Twitter communities interacting with science Descriptive analysis and overlay visualizations Twitter Multidisciplinary Haunschild etal. (2019) Topic visualizations based on Twitter discussions Co-word analysis Twitter Climate change Hellsten and Leydesdorff (2019) Topic and actor visualizations based on Twitter discussions Co-occurrence of hashtags and mentions Twitter Biomedicine Haunschild etal. (2020) Topic visualizations based on Twitter discussions Co-word analysis Twitter Library and Information Science Robinson-Garcia etal. (2019) Topic visualizations based on Twitter discussions Social network analysis and overlay visualizations Twitter Microbiology Torres-Salinas etal. (2019) Mapping knowledge relationships in Wikipedia Co-citation analysis Wikipedia Humanities Arroyo-Machado etal. (2020) Mapping knowledge relationships in Wikipedia Co-citation analysis Wikipedia Multidisciplinary Colavizza (2020) Coverage of research topics in Wikipedia Topic modeling and regression analysis Wikipedia COVID-19 Piccardi etal. (2020) Measuring interactions with Wikipedia references Engagement metrics Wikipedia Multidisciplinary 9271 Scientometrics (2021) 126:9267–9289 1 3 studies revolve around the use of Twitter and Wikipedia. Colavizza (2020) estimated how well Wikipedia, as a tool communicating scientific knowledge to the general public, reflects current scientific progress on COVID-19. Similarly, science mapping techniques haven been used to analyze how Wikipedia structured science in comparison with global science maps based on bibliometric databases (Arroyo-Machado et al., 2020); and the humanities (Torres-Salinas etal., 2019). In addition to Wikipedia, other social media sources have also been used to study the dissemination of scientific activity. For instance, Mendeley has been studied to identify its user types’ interests in and their patterns of use of scientific publications (Zahedi & van Eck, 2018). However, in this respect, Twitter is the platform that has most frequently been studied. Twitter Regarding the use of Twitter data, we find a first stream of studies that focus on identifying researchers or users who mention scientific publications and contextualize their activity. Among these we refer to studies like Ke etal. (2017), which identifies scientists from different disciplines; Robinson-Garcia etal. (2018), which proposes the use of mapping techniques to contextualize academics’ engagement in social media; or Díaz-Faes etal. (2019), which characterizes Twitter profiles mentioning scientific publications and identifies four dimensions of social media communication patterns. Secondly, we find studies that focus on using Twitter activity to identify topics of interest. These studies attempt to explain differences between the way scientists communicate research and how research is perceived or characterized by Twitter users. They compare differences between Twitter hashtags and author keywords in tweeted publications (Haunschild etal., 2019, 2020); compare topics of interest by social media platform (Noyons, 2019; Robinson-Garcia etal., 2019); or associate instances of interaction and topic by comparing hashtags co-tweeted by the same profiles (Hellsten & Leydesdorff, 2019). A third line of research is related to the diffusion of scientific publications. These studies aim to determine the social outreach attained by publications disseminated through Twitter (Alperin etal., 2018). Mapping andvisualization techniques One feature common to most of the aforementioned studies is their extensive use of mapping and visualization techniques. Based on network analysis, these techniques seek to construct n-dimensional spatial representations of science (Small, 1999). Most such representations are based on the co-occurrence of given events and are easily interpreted. From a bibliometric point of view, science maps are constructed from three elements: actors, resources and contents (Noyons, 2005), each of which offers a different level of analysis. In recent years, interest in mapping has grown as computational and methodological advances have extended their use. Furthermore, the number of visualization tools has increased considerably (cf. Cobo etal., 2011). Originally, two types of co-occurrence links between similar publications were proposed: co-citation (Small, 1973) and bibliographic coupling (Kessler, 1963). Both were applied at different levels of aggregation (i.e., co-citation networks of authors [White & Griffith, 1981] or bibliographic coupling for journals [Small & Koenig, 1977]). But the number of co-occurrence types has grown to include co-author networks (Glänzel, 9272 Scientometrics (2021) 126:9267–9289 1 3 2001) or co-word maps (Callon etal., 1983), among others. Co-word maps facilitate the exploration of structures across the scientific landscape (Waltman & van Eck, 2012) as an alternative to citation networks (Boyack etal., 2005; Leydesdorff etal., 2013). The emergence of new data sources and indicators, including but not exclusively from altmetrics, has led scientometricians to adapt these mapping techniques to the new metrics. Hence, we find proposals to map scientific literature on the basis of the co-occurrence of publications downloaded by users (Torres-Salinas etal., 2014); to adapt the concepts of co-citation and bibliographic coupling to meet the context of the social media (Costas etal., 2020); and to create thematic landscapes by geographical region (Wouters etal., 2019). These methods can all be used in different contexts. For instance, ArroyoMachado etal. (2020) created different levels of co-citation networks from Wikipedia entry references. Similarly, Haunschild etal. (2019) built thematic landscapes from cotweets to visualize public discussion of specific research topics, while Díaz-Faes etal. (2019) used them to characterize the profiles of Twitter users who participate in scientific discussions on the social network. The co-use of hashtags in tweets mentioning scientific literature has also been proposed (Hellsten & Leydesdorff, 2019), as have follower-following networks of scientists who use Twitter (Robinson-Garcia etal., 2018). Clearly, scientific mapping techniques are being adapted to new environments and gaining complexity. These techniques are based on the social network analysis of actors, relationships and structures (Wasserman & Faust, 1994). They represent any type of entity through nodes and establish relationships between entities that respond to co-occurrences, mentions, or any other type of interaction. Consequently, we can represent science-centered debates on social media at different levels and from different perspectives (Costas etal., 2020). The rationale behind social network analysis is that by combining co-occurring events, actors can be linked in a 2-mode (bipartite) network. Any such network is based on an asymmetrical matrix in which rows and columns are composed of different entities. Recently, Hellsten etal. (2019) suggested that by aggregating bipartite matrices different combinations could produce additional matrices. Figure1a shows a 3-mode network that reflects differing but inter-related entities (actors, objects and concepts). Fig. 1 n-dimensional matrix constructed by combining the 2-mode matrices of objects, concepts and actors. This representation is based on the conceptual framework proposed by Hellsten etal. (2019) 9273 Scientometrics (2021) 126:9267–9289 1 3 Figure1b shows how these matrices are constructed. Furthermore, the sub-matrices that appear in diagonal, show how entities of the same category are related through the combination of interactions between the other entities. Objectives In the present paper we build on our literature review to better refine methods by which communities with common scientific interests can be identified on social media. We test our methodological proposal using Twitter mentions to scientific papers in two research fields: Information Science & Library Science and Microbiology.1 Our main objective is to present a methodological proposal based on social network analysis that allows us to identify cognitive communities by grouping actors who may not necessarily be socially connected but, rather, who are connected through their interests. A proposal that aims to contribute to the new generation of social media metrics (Wouters etal., 2019) as it allows to discover the implicit social and semantic relationships between actors based on the discussion around scientific publications through social media. To this end, we seek to achieve the following objectives: 1. To introduce a novel methodological proposal by which actors in a given network can be grouped on the basis of their cognitive interests thus, to some extent, removing social relationships that could potentially blur the boundaries between communities. 2. To test our methodical approach in a specific case study: Twitter mentions of scientific literature in the field of Information Science & Library Science. 3. To replicate this approach in a different field—Microbiology—to observe potential inconsistencies in the methodology and discuss differences between the two case studies. Our study closely follows recent work in which a genuine effort has been made to conceptually define and then build a framework in which methodological solutions in the field of altmetrics can be expanded. For instance, Costas etal. (2020) recently proposed the concept of heterogeneous coupling in a study in which, from a theoretical perspective, they explored the potential of social network analysis to reveal links between the social media and science communication. Similarly, Hellsten etal. (2019) present their heterogeneous n-mode method which explores different combinations of interaction between actors. Our proposal could fit well into either of these two except for one noteworthy issue. The goal of our paper is to provide a practical application, showcasing a methodological innovation by which communities can be identified on the basis of common interests. The present study builds on previous work which analyzed differences in interests of topic by social media platform (Robinson-Garcia etal., 2019) and by clusters (ArroyoMachado etal., 2019). These earlier studies detected communities of actors who specifically mentioned the same publications and identified the topics that interested them. 1 We selected two categories as distant as possible from each other. Information Science & Library Science and Microbiology belong to very different scientific areas (Social Sciences and Health Sciences) and have significant differences, both, in terms of volume of publications, and communication and collaboration patterns. 9274 Scientometrics (2021) 126:9267–9289 1 3 Data andmethods Software The data needed to reproduce our analyses are available at http:// doi. org/ 10. 5281/ zenodo. 41489 41. We have included supplementary materials at https:// doi. org/ 10. 5281/ zenodo. 43329 21. Network manipulation of co-word maps (semantic maps) was conducted using Gephi 0.9.2 visualization software (Bastian etal., 2009). As we want an easily replicable methodology fully based on social network analysis, the popular Louvain algorithm is used for community detection (Blondel etal., 2008). Social networks and the overlapping social and semantic networks were constructed using the igraph R package (Csárdi, 2020), and the Louvain algorithm was again used to detect social communities. Both social and semantic networks were tested with the Leiden algorithm (Traag etal., 2019) in Gephi and igraph. In both case studies, the results showed no significant improvements with respect to those derived from applying the Louvain algorithm, so we opted for the original version. Visualizations of intersection sets were constructed using UpSet R software (Lex etal., 2014), a visualization technique that defines the characteristics of the entities studied in order to group them. A detailed description of the data processing and the application of the entire process is available in an R Notebook at https:// github. com/ Wence s91/ social_ media_ commu nities. All methods have been automated and gathered under the R package ‘altanalysis’ (https:// github. com/ Wence s91/ altan alysis). Data gathering We downloaded publication data for two research fields: Information Science & Library Science and Microbiology. We used the former as a case study to test our methodological approach. We then replicated the method in the latter field to compare results and analyze discrepancies in different contexts. On 17 July 2019 we retrieved all records indexed in the Web of Science (WoS) InCites database (excluding the Emerging Sources Citation Index) published between 2012 and 2018 in the WoS categories of Information Science & Library Science (84,568 publications); and in Biotechnology and Applied Biochemistry (250,577 publications) and Microbiology (187,013 publications)—these two represent a combined total of 413,910 publications, henceforth referred to as ‘Microbology’. From Altmetric.com’s Altmetric Explorer portal, we extracted all social media mentions of these records by using their DOIs as our search item. Information Science & Library Science has 35,695 publications with DOI (42.21%), and Microbiology has 366,449 (88.53%). Table2 summarizes the processing tasks undertaken prior to data analysis. We obtained the following datasets: • Information Science & Library Science: 14,475 publications were mentioned by at least one altmetric source, giving a total of 167,110 mentions from Altmetric.com. Some 151,505 of these (90.66%) were Twitter mentions of 13,458 (92.97%) publications. • Microbiology: 192,836 publications were mentioned by at least one altmetric source, giving 1,876,599 mentions from Altmetric.com. Some 1,599,315 of these (85.22%) were Twitter mentions of 173,406 (89.92%) publications. 9275 Scientometrics (2021) 126:9267–9289 1 3 Table 2 Summary of data processing of publications mentioned on social media in Information science and library science and microbiology Information science and library science Microbiology Publications % Twitter mentions % Publications % Twitter mentions % 1. Download web of science’s in cites records from 2012 to 2018 84,568 100 – – 413,910 100 – – 2. Recover all Altmetric.com mentions to InCites publications 14,475 17.12 150,806 100 192,836 46.59 1,585,313 100 3. Data cleaning and filter mentions to only made from Twitter 13,446 15.9 150,723 99.94 173,306 41.87 1,579,896 99.66 4. Remove retweets 13,227 15.64 65,933 43.72 171,085 41.33 695,429 43.87 5. Retrieve web of Science author keywords and data cleaning 8452 9.99 35,336 23.43 101,206 24.45 327,449 20.66 9282 Scientometrics (2021) 126:9267–9289 1 3 • Omics and Phylogenic Classification: a community consisting of 26,654 Twitter accounts, disseminating 35,450 publications through 143,604 tweets and sharing 591 keywords. It included publications covering studies of genetic material, bacterial microorganisms and biodiversity. • Immunology and Viral Diseases: a community consisting of 18,695 Twitter accounts, disseminating 23,499 publications through 85,030 tweets and sharing 490 keywords. It is related to viral diseases, their diagnosis, novel treatments, and vaccines. • Bioengineering: a community consisting of 11,523 Twitter accounts, disseminating 17,625 publications through 47,743 tweets and sharing keywords. It includes publications in biotechnology, metabolic engineering, and synthetic biology. • Bacteria: a community consisting of 19,077 Twitter accounts, disseminating 33,805 publications through 111,915 tweets and sharing 660 keywords. It includes publications related to diseases of bacterial origin, epidemiology, and outbreaks of infectious diseases. • Stem Cell Development: a community consisting of 7206 Twitter accounts, disseminating 11,208 publications through 31,081 tweets and sharing 223 keywords. It includes publications on regenerative medicine, gene therapy, and cancer treatment. • Tick transmitted diseases: a community consisting of 1048 Twitter accounts, disseminating 1044 publications through 4477 tweets and sharing 30 keywords. It includes publications relating to tick and flea transmitted diseases. When assigning Twitter users to each of these six topic groups (Fig.8), we found a much more complex and varied picture than in the previous case study. We identified 58 communities of interest. Although Twitter user groups relating to a single topic still stand out (38.84% of all users), most groups show an interest in more than one topic. Some 7909 Twitter users only mentioned keywords relating to omics and phylogenic classifications Fig. 7 Microbiology thematic landscape. This map shows the main component of the network and those terms that co-occur 5 times or more. It shows a total 2309 WoS author keywords 9283 Scientometrics (2021) 126:9267–9289 1 3 (16.44%); 3666 mentioned keywords relating to bacteria (7.62%); 3309 immunology and viral diseases (6.88%); 1920 bioengineering (3.99%); 1297 stem cell development (2.7%); and 104 tick transmitted diseases (0.22%). The presence of ‘mixed’ profiles was much more common than in Information Science & Library Science. For instance, only 29.67% of Twitter users who mentioned keywords related to omics and phylogenic classifications solely discussed this topic. This fell to 19.22% in the case of bacteria, 18% for stem cell development; 17.7% for immunology and viral diseases; 16.66% for bioengineering; and 9.92% for tick transmitted diseases. Discussion In the present study we propose a methodological approach to the identification of social media communities on the basis of common scientific interests. It enables us to link social media users on the basis of the keywords of the publications they mention and then group users by topic. We first applied this to Twitter users who mention publications in the fields of Information Science & Library Science. We then tested its feasibility by replicating the study in the field of Microbiology. Our proposal responds to the need for new efforts in social network analysis (Fu & Lai, 2020), is based on Fig. 8 Intersecting sets with more than 100 actors in Microbiology. A corresponds to all combinations of actors and topics. B shows intersections after introducing a 10% cut-off for the number of times a keyword is mentioned. C shows intersections with a 20% cut-off point 9284 Scientometrics (2021) 126:9267–9289 1 3 recently-published conceptual frameworks, especially the so-called heterogeneous couplings defined by Costas etal. (2020) and n-mode networks proposed by Hellsten etal. (2019), and previous studies in which we looked into differences in topics of interest on social platforms (Robinson-Garcia etal., 2019). This method is in line with the second generation of social media metrics (Díaz-Faes etal., 2019). Twitter mentions are not used here in a quantitative way, not even to filter keywords or actors. The focus of the paper is on social media-objects (Twitter users and tweets) and the papers are treated abstractly as keywords. The resulting socio-semantic network of this proposal has significant differences with respect to other kinds of networks. 2-mode networks can reflect direct and explicit relationships, such as social actors mentioning publications, as well as implicit ones, such as social actors that are connected by co-mention of the same publications. All of them are easily readable, but when an n-mode network is constructed combining 2-mode networks it becomes complex to interpret. Not only do the nodes represent different kinds of entities, but the relationships that exist between them can be of a different nature. This hinders the analysis, especially when network pruning or community detection methods are applied. Our proposal is to overlap instead of adding 2-mode networks. In this way, communities are detected independently, and then joined. While the n-mode network communities are composed of different types of elements, for example social actors and keywords, in ours the social actors have two types of groupings, one based on their social relationships and the other on keywords mentioned by them. The overlap between the two allows determining if their social relations and interests are in line or differ. Our study has not been free from limitations. Firstly, some tweets or accounts in our data sample were subsequently removed from Twitter or blocked. Consequently, they were excluded from our study. Second, to create the semantic maps, we initially extracted terms from publication titles. However, these proved too generic and included many distractors, generating widely varying communities. We resolved this by using WoS author keywords even though this limited the publications included to those present in the WoS database and having associated author keywords. Although actors were correctly assigned to the topic mentioned in most publications and people profiles prevail, bots are also present. In our Microbiology case study, given the complexity of the socio-semantic network, due to the variety of topics and social communities, this was not included. Altmetrics has a number of well-known limitations—for example, the fact that data aggregators only retrieve tweets that include identifiers such as a DOI. The present study represents a step forward in the creation of applied solutions that use altmetrics beyond mere counting. Elsewhere, studies have already identified researchers (Costas etal., 2017; Ke etal., 2017) and communities on Twitter (Robinson-Garcia etal., 2018) or visualized the topics discussed on social media by using WoS author keywords and hashtags (Haunschild etal., 2019, 2020). Indeed, the thematic landscapes in this study seem more granular and more detailed than those generated elsewhere (Robinson-Garcia etal., 2019) due to our use of WoS author keywords instead of title noun phrases. Our study used both methods but integrates them into a single visualization. In this context, Hellsten etal. (2019) and Hellsten and Leydesdorff, (2019) proposed heterogeneous networks and applied these, respectively, to scientific journals and their attributes and Twitter and user mentions and hashtags. These proposals were based on networks produced by aggregating bipartite matrices that combine actors and objects in the same network. Our proposal also combines co-occurrence relationships of actors, publications and author keywords but we do not directly integrate them all into a network. Instead, we take the co-occurring keyword network and the co-tweeted keyword network and overlap these. Thus, the network is only 9285 Scientometrics (2021) 126:9267–9289 1 3 formed of actors linked by social relations and their social communities are delimited through overlapping areas. Concluding remarks Our proposed methodology allows us to identify communities of users in an inclusive way, reflecting a complex reality in which actors may be interested in different aspects of a research field. This is especially evident in the case of Microbiology, where there are many groups consisting of only a few individuals assigned to more than one area. This study furthers our understanding on the use of social media to inform on scientific literature consumption by the general public. By isolating communities of common interest as well as finding those with overlapping interest we can narrow the target audience who is discussing scientific literature in social media. This is potentially useful to assess on the effectiveness of social outreach of scientific research, identify social stakeholders or analyze communication strategies. Further research should consider combining methods such as the one proposed with those strictly focused on characterizing user types (cf. Díaz-Faes etal., 2019). By focusing on concepts (i.e. keywords) rather than objects (i.e. publications), we minimize potential relationships derived from social relations between actors rather than from common research interests (e.g. colleagues from the same institution). This methodology has the potential of being applied in other scenarios from the ones proposed here. Other social media platforms could be considered, as well as other types of contents shared through social media. Some of the many and varied contexts in which it can be applied are political participation and political engagement (Halpern etal., 2017), trolling interactions in the online gaming sphere (Cook etal., 2019), experiences of mental disorders shared in forums (Yoo etal., 2019), or social communities discussing eating disorders (Wang etal., 2017). Moreover, it is possible to use other social objects and links to construct the social network and other kinds of semantic maps, for example Reddit posts as social object, co-mentioned hashtags for social network, and topic modelling for semantic map. In the specific case of altmetrics, a future line of study is the application of this methodology to different social media and the use of other terms to create the semantic maps. This is an initial approach only using Twitter mentions due to their enormous coverage and the extension of altmetrics studies. However, we would hope to study its applicability further by using altmetric sources other than Twitter, to study source-related differences in the type of users who discuss scientific literature. Acknowledgements The authors are grateful to Lydia Robinson-Garcia of the CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences for assessing our description and interpretation of the Microbiology semantic map. Authors’ contributions Conceptualization: NRG; Methodology: WAM; Formal analysis and investigation: WAM, DTS, NRG; Writing—original draft preparation: WAM, NRG; Writing—review and editing: DTS; Funding acquisition: DTS. Funding Funding for open access charge: Universidad de Granada / CBUA. This work has funded by the Spanish Ministry of Science and Innovation grant number PID2019-109127RB-I00/ SRA/10.13039/501100011033. Wenceslao Arroyo-Machado has an FPU Grant (FPU18/05835) from the Spanish Ministry of Universities. Daniel Torres-Salinas is supported by the Reincorporation Programme for Young Researchers from the University of Granada. Nicolas Robinson-Garcia is funded by a Ramón y Cajal grant from the Spanish Ministry of Science and Innovation (REF: RYC2019-027886-I). 9286 Scientometrics (2021) 126:9267–9289 1 3 Availability of data and material (data transparency) All data are available at http:// doi. org/ 10. 5281/ zenodo. 41489 41 Code availability All the code with the data processing is available at https:// github. com/ Wence s91/ social_ media_ commu nities Declarations Conflicts of interest The authors declare that they have no conflict of interest. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creat iveco mmons. org/ licen ses/ by/4. 0/. References Alperin, J. P., Gomez, C. J., & Haustein, S. (2018). Identifying diffusion patterns of research articles on Twitter: A case study of online engagement with open access articles. Public Understanding of Science, 28(1), 2–18. https:// doi. org/ 10. 1177/ 09636 62518 761733 Arroyo-Machado, W., Torres-Salinas, D., Herrera-Viedma, E., & Romero-Frías, E. (2020). Science through Wikipedia: A novel representation of open knowledge through co-citation networks. PLoS ONE, 15(2), e0228713. https:// doi. org/ 10. 1371/ journ al. pone. 02287 13 Arroyo-Machado, W., Torres-Salinas, D., & Robinson-Garcia, N. (2019). Identifying communities of interest in social media: Microbiology as a case study. In G. Catalano, C. Daraio, M. Gregori, H. F. Moed, & G. Ruocco (Eds.), Proceedings of the 17th International Conference on Scientometrics and Informetrics, ISSI 2019 (pp. 1201–1209). http:// issisocie ty. org/ proce edings/ issi_ 2019/ ISSI% 202019% 20-% 20Pro ceedi ngs% 20VOL UME% 20I. pdf Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: An Open Source Software for Exploring and Manipulating Networks. Third International AAAI Conference on Weblogs and Social Media. Third International AAAI Conference on Weblogs and Social Media. https:// www. aaai. org/ ocs/ index. php/ ICWSM/ 09/ paper/ view/ 154 Blondel, V. D., Guillaume, J.-L., Lambiotte, R., & Lefebvre, E. (2008). Fast unfolding of communities in large networks. Journal of Statistical Mechanics: Theory and Experiment, 2008(10), P10008. https:// doi. org/ 10. 1088/ 17425468/ 2008/ 10/ P10008 Bornmann, L., Haunschild, R., & Adams, J. (2019). Do altmetrics assess societal impact in a comparable way to case studies? An empirical test of the convergent validity of altmetrics based on data from the UK research excellence framework (REF). Journal of Informetrics, 13(1), 325–340. https:// doi. org/ 10. 1016/j. joi. 2019. 01. 008 Boyack, K. W., Klavans, R., & Börner, K. (2005). Mapping the backbone of science. Scientometrics, 64(3), 351–374. Callon, M., Courtial, J.-P., Turner, W. A., & Bauin, S. (1983). From translations to problematic networks: An introduction to co-word analysis. Social Science Information, 22(2), 191–235. https:// doi. org/ 10. 1177/ 05390 18830 22002 003 Cobo, M. J., López-Herrera, A. G., Herrera-Viedma, E., & Herrera, F. (2011). Science mapping software tools: Review, analysis, and cooperative study among tools. Journal of the American Society for Information Science and Technology, 62(7), 1382–1402. https:// doi. org/ 10. 1002/ asi. 21525 Colavizza, G. (2020). COVID-19 research in Wikipedia. Quantitative Science Studies, 1, 1–32. https:// doi. org/ 10. 1162/ qss_a_ 00080 Cook, C., Conijn, R., Schaafsma, J., & Antheunis, M. (2019). For whom the gamer trolls: A study of trolling interactions in the online gaming context. Journal of Computer-Mediated Communication, 24(6), 293–318. https:// doi. org/ 10. 1093/ jcmc/ zmz014 9287 Scientometrics (2021) 126:9267–9289 1 3 Costas, R., de Rijcke, S., & Marres, N. (2020). “Heterogeneous couplings”: Operationalizing network perspectives to study science-society interactions through social media metrics. Journal of the Association for Information Science and Technology, 72(5), 595–610. https:// doi. org/ 10. 1002/ asi. 24427 Costas, R., van Honk, J., & Franssen, T. (2017). Scholars on Twitter: Who and how many are they? Costas, R., Zahedi, Z., & Wouters, P. (2015). Do “altmetrics” correlate with citations? Extensive comparison of altmetric indicators with citations from a multidisciplinary perspective. Journal of the Association for Information Science and Technology, 66(10), 2003–2019. https:// doi. org/ 10. 1002/ asi. 23309 Csárdi, G. (2020). igraph: Network Analysis and Visualization. https:// CRAN.Rproje ct. org/ packa ge= igraph Davis, R., Bacha, C. H., & Just, M. R. (2017). Twitter and elections around the world: Campaigning in 140 Characters or Less. Routledge. Díaz-Faes, A. A., Bowman, T. D., & Costas, R. (2019). Towards a second generation of ‘social media metrics’: Characterizing Twitter communities of attention around science. PLoS ONE, 14(5), e0216408. https:// doi. org/ 10. 1371/ journ al. pone. 02164 08 Eysenbach, G. (2011). Can tweets predict citations? Metrics of social impact based on Twitter and correlation with traditional metrics of scientific impact. Journal of Medical Internet Research, 13(4), e123. Fu, J. S., & Lai, C.-H. (2020). Are We Moving Towards Convergence or Divergence? Mapping the Intellectual Structure and Roots of Online Social Network Research 1997–2017. Journal of ComputerMediated Communication, 25(1), 111–128. https:// doi. org/ 10. 1093/ jcmc/ zmz020 Glänzel, W. (2001). National characteristics in international scientific co-authorship relations. Scientometrics, 51(1), 69–115. https:// doi. org/ 10. 1023/A: 10105 12628 145 Grandjean, M. (2016). A social network analysis of Twitter: Mapping the digital humanities community. Cogent Arts & Humanities, 3(1), 1171458. https:// doi. org/ 10. 1080/ 23311 983. 2016. 11714 58 Halpern, D., Valenzuela, S., & Katz, J. E. (2017). We Face, I Tweet: How Different Social Media Influence Political Participation through Collective and Internal Efficacy. Journal of Computer-Mediated Communication, 22(6), 320–336. https:// doi. org/ 10. 1111/ jcc4. 12198 Haunschild, R., & Bornmann, L. (2018). Fieldand time-normalization of data with many zeros: An empirical analysis using citation and Twitter data. Scientometrics, 116(2), 997–1012. https:// doi. org/ 10. 1007/ s111920182771-1 Haunschild, R., Leydesdorff, L., & Bornmann, L. (2020). Library and Information Science Papers Discussed on Twitter: A new Network-based Approach for Measuring Public Attention. Journal of Data and Information Science, 5(3), 5–17. https:// doi. org/ 10. 2478/ jdis20200017 Haunschild, R., Leydesdorff, L., Bornmann, L., Hellsten, I., & Marx, W. (2019). Does the public discuss other topics on climate change than researchers? A comparison of explorative networks based on author keywords and hashtags. Journal of Informetrics, 13(2), 695–707. https:// doi. org/ 10. 1016/j. joi. 2019. 03. 008 Haustein, S. (2016). Grand challenges in altmetrics: Heterogeneity, data quality and dependencies. Scientometrics, 108(1), 413–423. https:// doi. org/ 10. 1007/ s111920161910-9 Haustein, S. (2019). Scholarly Twitter Metrics. In W. Glänzel, H. F. Moed, U. Schmoch, & M. Thelwall (Eds.), Springer Handbook of Science and Technology Indicators (pp. 729–760). Springer International Publishing. Doi: https:// doi. org/ 10. 1007/ 978-303002511-3_ 28 Haustein, S., Bowman, T. D., Holmberg, K., Tsou, A., Sugimoto, C. R., & Larivière, V. (2016). Tweets as impact indicators: Examining the implications of automated “bot” accounts on Twitter. Journal of the Association for Information Science and Technology, 67(1), 232–238. https:// doi. org/ 10. 1002/ asi. 23456 Haustein, S., Costas, R., & Larivière, V. (2015). Characterizing Social Media Metrics of Scholarly Papers: The Effect of Document Properties and Collaboration Patterns. PLoS ONE, 10(3), e0120495. https:// doi. org/ 10. 1371/ journ al. pone. 01204 95 Hellsten, I., & Leydesdorff, L. (2019). Automated analysis of actor–topic networks on twitter: New approaches to the analysis of socio-semantic networks. Journal of the Association for Information Science and Technology, 71(1), 3–15. https:// doi. org/ 10. 1002/ asi. 24207 Hellsten, I., Opthof, T., & Leydesdorff, L. (2019). N-mode network approach for socio-semantic analysis of scientific publications. Poetics, 78, 101427. https:// doi. org/ 10. 1016/j. poetic. 2019. 101427 Holmberg, K., Bowman, S., Bowman, T., Didegah, F., & Kortelainen, T. (2019). What Is Societal Impact and Where Do Altmetrics Fit into the Equation? Journal of Altmetrics, 2(1), 6. Joubert, M., & Costas, R. (2019). Getting to Know Science Tweeters: A Pilot Analysis of South African Twitter Users Tweeting about Research Articles. Journal of Altmetrics, 2(1), 2. Kassab, O., Bornmann, L., & Haunschild, R. (2020). Can altmetrics reflect societal impact considerations?: Exploring the potential of altmetrics in the context of a sustainability science research center. Quantitative Science Studies, 1(2), 792–809. https:// doi. org/ 10. 1162/ qss_a_ 00032 9288 Scientometrics (2021) 126:9267–9289 1 3 Ke, Q., Ahn, Y.-Y., & Sugimoto, C. R. (2017). A systematic identification and analysis of scientists on Twitter. PLoS ONE, 12(4), e0175368. https:// doi. org/ 10. 1371/ journ al. pone. 01753 68 Kessler, M. M. (1963). Bibliographic coupling between scientific papers. American Documentation, 14(1), 10–25. https:// doi. org/ 10. 1002/ asi. 50901 40103 Kim, J., & Hastak, M. (2018). Social network analysis: Characteristics of online social networks after a disaster. International Journal of Information Management, 38(1), 86–96. https:// doi. org/ 10. 1016/j. ijinf omgt. 2017. 08. 003 Lex, A., Gehlenborg, N., Strobelt, H., Vuillemot, R., & Pfister, H. (2014). UpSet: Visualization of Intersecting Sets. IEEE Transactions on Visualization and Computer Graphics, 20(12), 1983–1992. https:// doi. org/ 10. 1109/ TVCG. 2014. 23462 48 Leydesdorff, L., Carley, S., & Rafols, I. (2013). Global maps of science based on the new Web-of-Science categories. Scientometrics, 94(2), 589–593. Misirlis, N., & Vlachopoulou, M. (2018). Social media metrics and analytics in marketing – S3M: A mapping literature review. International Journal of Information Management, 38(1), 270–276. https:// doi. org/ 10. 1016/j. ijinf omgt. 2017. 10. 005 Newman, M. E. J. (2004). Fast algorithm for detecting community structure in networks. Physical Review E, 69(6), 066133. https:// doi. org/ 10. 1103/ PhysR evE. 69. 066133 Neylon, C., & Wu, S. (2009). Article-Level Metrics and the Evolution of Scientific Impact. PLOS Biology, 7(11), e1000242. https:// doi. org/ 10. 1371/ journ al. pbio. 10002 42 Nielsen, F. A. (2007). Scientific citations in Wikipedia. First Monday. https:// doi. org/ 10. 5210/ fm. v12i8. 1997 Noyons, C. M. (2005). Science Maps Within a Science Policy Context. In H. F. Moed, W. Glänzel, & U. Schmoch (Eds.), Handbook of Quantitative Science and Technology Research: The Use of Publication and Patent Statistics in Studies of S&T Systems (pp. 237–255). Springer Netherlands. Doi: https:// doi. org/ 10. 1007/140202755-9_ 11 Noyons, E. (2019). Measuring societal impact is as complex as ABC. Journal of Data and Information Science, 4(3), 6–21. Piccardi, T., Redi, M., Colavizza, G., & West, R. (2020). Quantifying Engagement with Citations on Wikipedia. Proceedings of the Web Conference, 2020, 2365–2376. https:// doi. org/ 10. 1145/ 33664 23. 33803 00 Priem, J., Taraborelli, D., Groth, P., & Neylon, C. (2010). Altmetrics: A manifesto. In Altmetrics. http:// altme trics. org/ manif esto/ Robinson-Garcia, N., Arroyo-Machado, W., & Torres-Salinas, D. (2019). Mapping social media attention in Microbiology: Identifying main topics and actors. FEMS Microbiology Letters, 366(7). Doi: https:// doi. org/ 10. 1093/ femsle/ fnz075 Robinson-Garcia, N., Costas, R., Isett, K., Melkers, J., & Hicks, D. (2017). The unbearable emptiness of tweeting—About journal articles. PloS One, 12(8), e0183551. Robinson-Garcia, N., Torres-Salinas, D., Zahedi, Z., & Costas, R. (2014). New data, new possibilities: Exploring the insides of Altmetric.com. El Profesional de La Information, 23(4), 359–366. https:// doi. org/ 10. 3145/ epi. 2014. jul. 03 Robinson-Garcia, N., van Leeuwen, T. N., & Ràfols, I. (2018). Using altmetrics for contextualised mapping of societal impact: From hits to networks. Science and Public Policy, 45(6), 815–826. https:// doi. org/ 10. 1093/ scipol/ scy024 Small, H. (1973). Co-citation in the scientific literature: A new measure of the relationship between two documents. Journal of the American Society for Information Science, 24(4), 265–269. https:// doi. org/ 10. 1002/ asi. 46302 40406 Small, H. (1999). Visualizing science by citation mapping. Journal of the American Society for Information Science, 50(9), 799–813. https:// doi. org/ 10. 1002/ (SICI) 10974571(1999) 50:9% 3c799:: AIDASI9% 3e3.0. CO;2-G Small, H. G., & Koenig, M. E. D. (1977). Journal clustering using a bibliographic coupling method. Information Processing and Management, 13(5), 277–288. https:// doi. org/ 10. 1016/ 03064573(77) 90017-6 Stieglitz, S., Mirbabaie, M., Ross, B., & Neuberger, C. (2018). Social media analytics—Challenges in topic discovery, data collection, and data preparation. International Journal of Information Management, 39, 156–168. https:// doi. org/ 10. 1016/j. ijinf omgt. 2017. 12. 002 Sugimoto, C. R., Work, S., Larivière, V., & Haustein, S. (2017). Scholarly use of social media and altmetrics: A review of the literature. Journal of the Association for Information Science and Technology, 68(9), 2037–2062. https:// doi. org/ 10. 1002/ asi. 23833 Taraborelli, D. (2008). Soft peer review: Social software and distributed scientific evaluation. Thelwall, M. (2018). Early Mendeley readers correlate with later citation counts. Scientometrics, 115(3), 1231–1240. https:// doi. org/ 10. 1007/ s111920182715-9 9289 Scientometrics (2021) 126:9267–9289 1 3 Torres-Salinas, D., Clavijo, Á. C., & Contreras, E. J. (2013). Altmetrics: New Indicators for Scientific Communication in Web 2.0. Revista Comunicar, 21(41), 53–60. Doi: https:// doi. org/ 10. 3916/ C41201305 Torres-Salinas, D., Jiménez-Contreras, E., & Robinson-Garcia, N. (2014). Tendencias en mapas de la ciencia: Co-uso de información científica como reflejo de los intereses de los investigadores. Torres-Salinas, D., Romero-Frías, E., & Arroyo-Machado, W. (2019). Mapping the backbone of the Humanities through the eyes of Wikipedia. Journal of Informetrics, 13(3), 793–803. https:// doi. org/ 10. 1016/j. joi. 2019. 07. 002 Traag, V. A., Waltman, L., & van Eck, N. J. (2019). From Louvain to Leiden: Guaranteeing well-connected communities. Scientific Reports, 9(1), 5233. https:// doi. org/ 10. 1038/ s4159801941695-z van Schalkwyk, F., Dudek, J., & Costas, R. (2020). Communities of shared interests and cognitive bridges: The case of the anti-vaccination movement on Twitter. Scientometrics. https:// doi. org/ 10. 1007/ s1119202003551-0 Waltman, L., & van Eck, N. J. (2012). A new methodology for constructing a publication-level classification system of science. Journal of the American Society for Information Science and Technology, 63(12), 2378–2392. https:// doi. org/ 10. 1002/ asi. 22748 Wang, T., Brede, M., Ianni, A., & Mentzakis, E. (2017). Detecting and characterizing eating-disorder communities on social media. Proceedings of the Tenth ACM International Conference on Web Search and Data Mining, 1, 91–100. https:// doi. org/ 10. 1145/ 30186 61. 30187 06 Wasserman, S., & Faust, K. (1994). social network analysis: Methods and applications. Cambridge University Press. https:// doi. org/ 10. 1017/ CBO97 80511 815478 Waszak, P. M., Kasprzycka-Waszak, W., & Kubanek, A. (2018). The spread of medical fake news in social media—The pilot quantitative study. Health Policy and Technology, 7(2), 115–118. https:// doi. org/ 10. 1016/j. hlpt. 2018. 03. 002 White, H. D., & Griffith, B. C. (1981). Author cocitation: A literature measure of intellectual structure. Journal of the American Society for Information Science, 32(3), 163–171. https:// doi. org/ 10. 1002/ asi. 46303 20302 Wouters, P., Zahedi, Z., & Costas, R. (2019). Social Media Metrics for New Research Evaluation. In W. Glänzel, M. Henk F, U. Schmoch, & M. Thelwall (Eds.), Springer Handbook of Science and Technology Indicators (pp. 687–713). Springer International Publishing. Yoo, M., Lee, S., & Ha, T. (2019). Semantic network analysis for understanding user experiences of bipolar and depressive disorders on Reddit. Information Processing and Management, 56(4), 1565–1575. https:// doi. org/ 10. 1016/j. ipm. 2018. 10. 001 Zahedi, Z., & Costas, R. (2018). General discussion of data quality challenges in social media metrics: Extensive comparison of four major altmetric data aggregators. PLoS ONE, 13(5), e0197326. https:// doi. org/ 10. 1371/ journ al. pone. 01973 26 Zahedi, Z., & van Eck, N. J. (2018). Exploring Topics of Interest of Mendeley Users. Journal of Altmetrics, 1(1), 5.