scieee AI-readable full text Open interactive document viewer

Interdisciplinary Colloquium on Digitalisation of Research: "Social Scientific Data Quality and Reproducibility in the AI Era: Challenges and Pathways" 04.12.2025 with Stefan Dietze

Dietze, Stefan

Abstract

Throughout the last decades, the social sciences have increasingly adopted novel forms of research data, e.g. data mined from the web and social media platforms. This together with the recent advances in artificial intelligence (AI) and related areas, e.g. natural language processing (NLP), led to a much more widespread adoption of diverse computational methods, including techniques from deep learning and, most prominently, large language models. However, increasingly complex computational methods lead to new challenges with respect to transparency, reproducibility and overall quality of research and data, further elevating an already widely recognised reproducibility crisis. This talk will, one the one hand, introduce challenges posed by the use of deep learning-based methods in social science research. On the other hand, it will show pathways to address such problems. Examples are works geared towards sharing computational (AI) methods in the social sciences in a reproducible and citable way, for understanding and tracing adoption of and relations between methods and datasets at large scale, e.g. in social science research in general (e.g. by mining scientific publications) or novel ways for providing access to sensitive research data in the social sciences (e.g. social media data) to facilitate reproducible research without violating ethical or legal constraints or principles.

Full text

Social Scientific Data Quality and Reproducibility in the AI Era: Challenges and Pathways Stefan Dietze1,2 Session moderator: Anna Jacyszyn3 Interdisciplinary Colloquium on Digitalisation of Research, DiTraRe, 4 December 2025 (1) GESIS Leibniz Institute for the Social Sciences + (2) Heinrich Heine University Düsseldorf (3) FIZ Karlsruhe - Leibniz Institute for Information Infrastructure Photos and recording Pixabay, ste_phania 2 Interdisciplinary Colloquium on Digitalisation of Research, Stefan Dietze, 4 December 2025 www.youtube.com/@DiTraRe 3 Interdisciplinary Colloquium on Digitalisation of Research, Stefan Dietze, 4 December 2025 Social scientific data quality and reproducibility in the AI era: challenges and pathways DiTraRe Interdisciplinary Colloquium on Digitalisation of Research, 4 December 2025 Stefan Dietze Social science research is changing ▪Emergence of large volumes of behavioral data (e.g. from social media) has introduced new research field (CSS), methods and data 2 Behavioral web data for the social sciences ▪Online discourse (e.g. in social media, online news) ▪Social web activity streams (posts, shares, likes, follows etc) ▪Web search behaviour, e.g. browsing, navigation or search engine interactions ▪Low-level behavioral traces (scrolling, mouse movements, gaze behavior etc) ▪General characteristics oClose to users & their personal (potentially sensitive) information oLarge and heterogeneous 3 Web data tends to be „big“ Source: Domo via PCMag 4 New kinds of data require new kinds of methods Methods widely used (e.g. for social media analysis) : ▪Time series analysis (auto-regressive models, ARIMA etc) ▪Network/graph analysis ▪Dictionary-based methods (e.g. for sentiment analysis) ▪Tailored machine learning models (trained from scratch) ▪Pretrained open source language models (e.g. BERT) ▪Pretrained proprietary LLMs (like GPT/ChatGPT) Substantial differences with respect to: ▪Scalability (ability to handle larger volumes of data) ▪Robustness (ability to handle noisy or biased data) ▪Efficiency (compute/resource requirements) ▪Transparency & interpretability ▪Reproducibility „AI“ 5 Beyond basic use of AI for data analysis: LLMs for simulating human behavior Santurkar, S., et al., Whose Opinions Do Language Models Reflect?, International Conference on Machine Learning (ICML2023) 6 Beyond just reproducibility Pineau et al., Improving reproducibility in machine learning research, Journal of Machine Learning Research 22 (2021) 1-20. 17 Beyond reproducibility: do benchmarks assess generalisable learnings? Example: Twitter bot detection Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection. ACM WebConf2023 18 „Shortcuts“ in the data Beyond reproducibility: do benchmarks assess generalisable learnings? Example: Twitter bot detection Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection. ACM WebConf2023 Take-aways ▪AI benchmark data does not represent real-world data/problems but contains shortcuts ▪Shortcut learning [Geirhos2020] is widespread and leads to poor generalisability ▪Reproducible results ≠generalisable results ▪Benchmarking, i.e. understanding what is state-of-the-art in AI/NLP is hard Geirhos, R., Jacobsen, JH., Michaelis, C. et al. Shortcut learning in deep neural networks. Nature Machine Intelligence 2, 665–673 (2020). 19 „Shortcuts“ in the data Addressing reproducibility & generalisability in CSS/AI research? 1. Empowering researchers to find state-of-the-art methods (“benchmarking / state-of-the-art crisis”) 2. Improving the interpretability of scholarly reporting (“reporting problem”) 3. Ensuring data availability & access (“access problem”) Reproducibility Replicability Robustness Generalisability 22 Overview 1. Empowering researchers to find state-of-the-art methods (“benchmarking / state-of-the-art crisis”) 2. Improving the interpretability of scholarly reporting (“reporting problem”) 3. Ensuring data availability & access (“access problem”) 23 Key challenge: how to identify high quality methods? How to find SotA methods for given task (e.g. stance detection on specific tweet sample)? •Review literature: labor-intensive, methods often poorly cited / not traceable •Code/model repositories (e.g. HuggingFace, GitHub): lack context (e.g. related research, comparisons with other methods etc) •Ad-hoc choices („I use what I know“) Benchmarking of AI/CS methods •Use of standard evaluation corpora & metrics to compare method performance / quality •In theory: benchmarks assess whether a published method is good/bad/state-of-the-art •In practice: benchmarks and benchmarking practices (eg baseline choices) are flawed, e.g. do not evaluate generalisability 24 Finding AI methods for the social sciences: GESIS Methods Hub Released in Q3 2025 Integrated into GESIS Search, MyBinder, Jupyter4NFDI •Platform for finding, sharing & using/executing data science & AI methods •Empowering social scientists with & without technical expertise to use complex state-of-the-art methods & LLMs •GESIS-curated and community-based methods and tutorials •Focus on reproducibility, quality, citability (DOIs), benchmarking, provenance https://methodshub.gesis.org 25 Benchmarking: evaluating generalisability of NLP models Feger, M., Boland, K., Dietze, S., Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments, In ACL2025. 26 Example case: argument mining in tweets/social media posts as established NLP task Feger, M., Boland, K., Dietze, S., Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments, In ACL2025. Do models actually generalise? •Train-on-one-test-on-another (dataset) experiments on 17 AM datasets •Using state-of-the-art Transformer-based language models (BERT, RoBERTa, WRAP) •Models do not generalise („do not learn to detect arguments“): performance degrades when models are tested on OOD data 27 Benchmarking: evaluating generalisability of NLP models Feger, M., Boland, K., Dietze, S., Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments, In ACL2025. •Leave-one-out cross validation: models trained on all datasets but the target dataset (rows) •Performance degradation significant (despite more diverse training data) •Performance drop particularly for datasets that seemed „easy“ to learn 28 Realistic benchmarking: evaluating generalisability of NLP models Detecting model, task and dataset mentions: model performance Otto, W., Zloch, M., Gan, L., Karmakar, S., Dietze, S. (2023). GSAP-NER: A Novel Task, Corpus, and Baseline for Scholarly Entity Extraction Focused on Machine Learning Models and Datasets. In Findings of the Association for Computational Linguistics: EMNLP 2023 36 Understanding methods and data in CSS (AAAI ICWSM publications) Tasks Methods 37 Understanding methods and data in CSS (AAAI ICWSM publications) Citations of ML models over time Citations of data sources over time 38 MethodMiner: a tool for mining task, dataset & model mentions 39 Otto, W., Upadhyaya, S., Gan, L., Silva, K. (2025), Track Machine Learning in Your Research Domain. In 2nd Conference on Research Data Infrastructure (CoRDI) Shared AI task @ ACL2025: mining data, model, software mentions https://sdproc.org/2025/somd25.html 44 Overview 1. Empowering researchers to find state-of-the-art methods (“benchmarking / state-of-the-art crisis”) 2. Improving the interpretability of scholarly reporting (“reporting problem”) 3. Ensuring data availability & access (“access problem”) 45 46 Challenge: dependencies on 3rd party gatekeepers Behavioral data is not distributed as the web but tied to platforms/gatekeepers Challenge: volatility & decay of web data •Data is not persistent •Example: deletion ratio of tweets between 25-29 % •Differs between different samples Khan, M.T., Dimitrov, D., Dietze, S., Characterization of Tweet Deletion Patterns in the Context of COVID-19 Discourse and Polarization, ACM Hypertext 2025 47 Challenge: data evolution impacts methods (quality/reproducibility) ▪Vocabulary evolves: e.g. vocabulary shift, over- /underrepresentation of topics/vocabulary in particular time periods (e.g. Twitter COVID19discourse 2020 vs prior periods) ▪PLMs/LLMs require frequent training and updates (and continuous access to data) Source: Hombaiah et al., “Dynamic Language Models for continuously evolving Content”, SIGKDD2021 48 Responsible social media archiving @ GESIS: examples X/Twitter (https://data.gesis.org/tweetskb) ▪Sampling: 1% - random sample ▪Dataset size: > 14 billion tweets ▪Time period: Feb 2013 - June 2023 Telegram (https://data.gesis.org/telescope) ▪Sampling: seed lists + snowball sampling ▪Dataset: ~120M messages from ~71K public channels and metadata for ~500K channels ▪Time period: Feb 2024 and running Fact-checked claims (https://data.gesis.org/claimskg) ▪Sampling method: 13 factchecking websites ▪Dataset: 74066 claims and 72128 claim reviews ▪Time period: claims published between 1996 –2023 4Chan ▪Sampling method: all boards ▪Dataset size: 4,676,378 threads, 264,898,231 posts ▪Time period: Nov 2023 and running ▪In preparation: BlueSky, YouTube, … https://www.gesis.org/gesis-web-data 50 https://stefandietze.net https://gesis.org/en/kts Thank you! 6 www.ditrare.de/en Thank you for joining! Stay connected ■DiTraRe ○Website: www.ditrare.de/en ○Email: ditrar[email protected] ○LinkedIn: www.linkedin.com/company/ditrare ○Mastodon: social.kit.edu/@DiTraRe ○YouTube: www.youtube.com/@DiTraRe ○Zenodo: zenodo.org/communities/ditrare ■Discussion forum: www.ditrare.de/en/forum ■Newsletter: www.ditrare.de/en/newsletter 7