scieee AI-readable full text Open interactive document viewer

Replication files for: Discourses about Sustainability and Digitalization in Europe on Twitter Over Time by Mario Angst and Nadine Strauß (GAIA,forthcoming)

Mario Angst

Abstract

Replication files for: Discourses about Sustainability and Digitalization in Europe on Twitter Over Time by Mario Angst and Nadine Strauß in the journal GAIA: Angst, M., & Strauß, N. (2023). Discourses surrounding sustainability and digitalization in Europe on Twitter over time. Gaia, 32(1), 10–20. https://doi.org/10.14512/gaia.32.s1.4 See README.md or README.html or README.pdf for detailed instructions.

Full text

Overview The code and data in this repository makes it possible to replicate the analysis in Discourses about Sustainability and Digitalization in Europe on Twitter Over Time by Mario Angst and Nadine Strauß in the journal GAIA, using R and Python. The replicator should expect the code to run for about 5 minutes to replicate the analysis from preprocessed data. Data Availability and Provenance Statements The analysis relies on tweet data queried via the social media company Twitter’s academic API endpoint. In accordance with Twitter’s developer policy regarding content redistribution (https://developer.twitter.com/ en/developer-terms/policy, as of 26.4.2022, when the data was queried), we cannot share the actual raw data in this repository but only the tweet IDs. The list of tweet IDs used should make it possible to download the exact dataset of tweets we analyzed and classified via the Twitter API. To make it possible to replicate our results even without access to actual tweets: •we provide processed datasets of tweet ids labeled by our classifiers. Additionally, for transparency on the labeling procedure, we provide: • the trained classifier used to label the tweets and all training and test data used to create and evaluate the classifier. • the set of patterns used to identify topics in tweets in a rule-based classifier alongside data documenting the evolution of the final set of patterns. • the evaluation procedure and test data used to evaluate the zero-shot classifier to classify transversal discourse dimensions. Statement about Rights ⊠ I certify that the author(s) of the manuscript have legitimate access to and permission to use the data used in this manuscript. Summary of Availability □All data are publicly available. ⊠Some data cannot be made publicly available. □No data can be made publicly available. Details on each Data Source Under data/ The following files are stored in data/raw: •tweet_ids.csv: A list of all tweet IDs queried with the suseurope query (see article) The following files are stored in data/processed: •tweets.csv: A table with 32 columns giving information about individual tweets per row: –tweet_id: The tweet id –created_at: date of tweet creation – SUSDIGI: The predicted score of the tweet as predicted by our classifier on whether the tweet relates to the discourse about sustainable digitalization – 29 binary variables (TRUE/FALSE) of discourse topic codes indicating whether the tweet contains a discourse topic • tweet_topic.csv: A table/ edglist with three columns storing associations of tweets with topics per row: –tweet_id: the tweet_id 1 –topic: the tweet-topic association –pattern_match: the pattern match leading to the association of the tweet with the topic Under classifiers: The following files are stored in classifiers/relevance: • train.csv: The training dataset used to train the binary classifier to identify sustainable digitalization related tweets. The dataset contains text of 4008 tweets with a binary label (0/1) where 1 indicates relevance of the tweet. • test.csv: The test dataset used to evaluate the binary classifier to identify sustainable digitalization related tweets (not used in training). The dataset contains text of 1004 tweets with a binary label (0/1) where 1 indicates relevance of the tweet. •codebook.html: The codebook used to annotate tweet texts. The following files are stored in classifiers/topics: • master_patterns.json: The set of patterns - topic label combinations used in the rule-based classifier to identify tweet topics. •topic_development/: –Four time-stamped .csv files illustrating the development of the patterns set over time. – topics_identification_log.md: Log file documenting what steps were taken during each iteration of pattern development. The following files are stored in classifiers/transversal: • annotations/: A set of five .csv files giving annotations of tweets by four different coders in terms of efficiency and economic growth orientation. • predictions/: A set of five .csv files giving the model predictions by the zero-shot classifier on efficiency and economic growth orientation. Computational requirements Software Requirements •R (code was last run with version 4.1.2) • R packages are made available in the form used in the analysis using the tools provided in the R package renv. Use renv::restore() to initialize the project on your local machine with the packages used in the analysis •The only Python dependencies are spacy(3.2) and spacy-streamlit(1.0.4). • We recommend running the R analysis in RStudio (code was last run with version 2022.07.1+554 “Spotted Wakerobin” Release (7872775ebddc40635780ca1ed238934c3345c5de, 2022-07-22) for Windows Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) QtWebEngine/5.12.8 Chrome/69.0.3497.128 Safari/537.36) Summary Approximate time needed to reproduce the analyses on a standard (2022) desktop machine: ⊠<5 minutes Details The code was last run on a 8-core Intel-based laptop with Windows 11 (64-bit). Description of programs/code •the folder classifier/ contains a packaged spacy pipeline to classify tweets under dist/. –You can install the pipeline using pip: pip install ./classifiers/relevance/en_textcat_binary_susdigi-1.0.0 2 • the script in src/visualize_model.py will allow you to interactively explore the model predictions. With spacy-streamlit installed, run the script from the repository root with streamlit run ./src/visualize_model.py • The script in src/analysis.R generates all network plots used in the main body of the article and saves them in figures/ • The script in src/evaluate_zeroshot.R reproduces the evaluation of the transversal discourse dimension classifier on annotated test sets. License for Code The code is licensed under a Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license. Instructions to Replicators •Unpack the zip file zenodo_repo.zip • Open zenodo_repo.Rproj in RStudio - this will allow here::here() to identify the correct local file paths •Run the script src/analysis.R to reproduce all figures •Optionally explore the sustainable digitalization classifier (see descriptions of programs/code) The provided code reproduces: ⊠All figures in the paper □All tables in the paper –the only table in the paper is a table of all topics used □All data preprocessing –see Data Availability 3