scieee AI-readable full text Open interactive document viewer

Replication code availability over time and across fields: Evidence from the German Socio‐Economic Panel

Fink, Lukas,Marcus, Jan

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Fink, Lukas; Marcus, Jan Article — Published Version Replication code availability over time and across fields: Evidence from the German Socio‐Economic Panel Economic Inquiry Provided in Cooperation with: John Wiley & Sons Suggested Citation: Fink, Lukas; Marcus, Jan (2024) : Replication code availability over time and across fields: Evidence from the German Socio‐Economic Panel, Economic Inquiry, ISSN 1465-7295, Wiley, Hoboken, NJ, Vol. 63, Iss. 2, pp. 357-386, https://doi.org/10.1111/ecin.13267 This Version is available at: https://hdl.handle.net/10419/319366 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc/4.0/ Received: 31 January 2024 - Accepted: 16 October 2024 DOI: 10.1111/ecin.13267 ORIGINAL ARTICLE Replication code availability over time and across fields: Evidence from the German Socio‐Economic Panel Lukas Fink 1 |Jan Marcus 1,2,3 1 School of Business and Economics, Freie Universität Berlin, Berlin, Germany 2 Berlin School of Economics, Berlin, Germany 3 IZA Bonn, Bonn, Germany Correspondence Jan Marcus, School of Business and Economics, Freie Universität Berlin, Garystr. 21, Berlin 14195, Germany. Email: [email protected] Abstract Providing replication code is an inexpensive way to facilitate reproducibility. However, little is known about the extent of replication code provision. Therefore, we examine the availability of replication code for over 2500 peer‐ reviewed articles based on the German Socio‐Economic Panel (SOEP), one of the most widely used datasets in economics and other social sciences. We find that only 6% of SOEP‐based studies have code available, but that this proportion has increased sharply over time. We provide evidence that the increase in code provision is driven by technological advances, individual researcher initiatives, and journal policies. KEYWORDS code availability, journal policies, replication code, reproducibility, SOEP JEL CLASSIFICATION A12, C80, C88 1 | INTRODUCTION The reproducibility of scientific results is an essential part of the scientific method. It ensures that research can be independently verified and promotes trust in scientific results. Reproducibility is particularly important in economics because the results of economic studies are used to inform policy decisions and, hence, may have direct consequences for the lives of many individuals. A study is computationally reproducible, when other researchers, using the same data and code, get the same results as the original study (Pérignon et al., 2023). Hence, computational reproducibility depends on two crucial ingredients: the availability of data and the availability of code, that is, computer instructions in statistical software such as Stata, R, or Matlab. The practice of making code and data available has numerous advantages for the scientific community, in addition to allowing for an easy reproduction of the published results. First, it enables others to build upon the work, fostering further investigation and a more comprehensive understanding of the subject matter. Second, errors in the code can be Abbreviations: DIW, German Institute for Economic Research; JIF, journal impact factor; OSF, Open Science Framework; SOEP, German Socio‐ Economic Panel; TOP, Transparency and Openness Promotion. Managing Editor: Farasat Bokhari This is an open access article under the terms of the Creative Commons Attribution‐NonCommercial License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited and is not used for commercial purposes. © 2024 The Author(s). Economic Inquiry published by Wiley Periodicals LLC on behalf of Western Economic Association International. Economic Inquiry. 2025;63:357–386. wileyonlinelibrary.com/journal/ecin - 357 more easily found and corrected, strengthening the self‐correction feature of the scientific system. 1 Third, it facilitates knowledge dissemination as others can gain deeper insights into the applied methods and coding practices from the provided code. Fourth, it might serve as a deterrent against fraud and questionable research practices like p‐hacking, as researchers know that their code might be exposed to scrutiny. Fifth, it can serve as a signal that the researchers did not engage in questionable research practices. Sixth, while reproducibility is often seen as a rather low bar and “a minimum criterion for scientific credibility” (Fišar et al., 2024, p. 1343), it is an important basis for checking whether results are robust to alternative methods and analytical choices, and whether results can be replicated using different data. Therefore, increasing the transparency of research by making code publicly available could help overcome the current lack of such replication exercises (Ankel‐Peters et al., 2023). Despite the advantages, it is unclear to what extent economists and researchers from other social sciences make data and code available. We aim to contribute to filling this gap by examining the extent of code provision. Hence, we focus on the ingredient for reproducibility that is arguably easier to achieve: Basically, every empirical researcher can share code, but not everyone can share data, as researchers are frequently prohibited from sharing analyzed data due to legal and confidentiality reasons. This is especially relevant in economics, where secondary analysis of data collected by others, such as governmental organizations or general interest surveys, is common and only the primary data producer have the right to redistribute the data. Further, providing code is comparatively inexpensive in terms of both time and money for researchers, as it typically takes less time than sharing data. 2 For these reasons, we focus on a setting where data are available to researchers worldwide at no cost, making the provision of code the critical factor for reproducibility. 3 To better understand the extent of code sharing and its development over time, this metascientific study focuses on code availability for publications based on one of the most widely used datasets in economics and other social sciences, the German Socio‐Economic Panel (SOEP). For this purpose, we conducted an intensive search for code availability of all SOEP‐based publications in peer‐reviewed journals in the 1985–2021 period. We checked the journal website for each of the 2518 articles, the actual articles, the websites of 2615 distinct authors who wrote the articles, and specific online repositories. The collection of this variable is a major achievement of this paper. We use this information to address several research questions: What factors correlate with code provision? How has code provision evolved over time? Is code provision related to quality metrics, such as journal impact factors and citations? What factors drive the development of code availability over time? For the latter research question, we gathered information on journal policies on code sharing from the Transparency and Openness Promotion (TOP) framework (Nosek et al., 2015), journal websites, and a survey among editors of almost 100 journals in different disciplines. The SOEP is particularly useful for studying code provision for several reasons. First, the SOEP data are widely used by the scientific community. The SOEP has about 3500 users worldwide (Goebel et al., 2019), and more than 2500 peer‐ reviewed, SOEP‐based journal articles have been published. Second, the SOEP data are not only employed by economists but also by researchers from several other disciplines (in particular, sociology, psychology, political science, demography), allowing us to examine code‐sharing behavior in different disciplines. Third, the SOEP group provides a list of all publications based on the SOEP data, which allows us to clearly define our target population. Fourth, the SOEP makes it possible to study developments in code provision over many years as it was established 40 years ago. Fifth, there is a clear separation of the data collector (a fieldwork agency on behalf of the SOEP group at DIW Berlin) from the data analyst. This means that the raw SOEP data can be seen as exogenous to the researchers analyzing the data and that strategic decisions in the data collection process are of little relevance. Sixth, the SOEP data can be downloaded and used free of charge by researchers around the world. 4 This is a major advantage of the SOEP (and similar survey data) for studies on replicability and reproducibility, in particular compared to administrative data, which often can only be analyzed in restricted access environments and often only after payment for data access. Our empirical results show that the share of peer‐reviewed SOEP studies with code available is only 6%. Examining correlates of code availability, we find that single‐authored articles and articles written in languages other than English are less likely to provide code. In addition, code availability is positively correlated with higher journal impact factors and more citations. We also see that code provision is less common for studies in economics journals in our sample. When we control for year of publication, the correlations of code availability with article language and single authorship disappear, and the difference between economics and other disciplines becomes smaller, suggesting that parts of these observed correlations are driven by time effects. We also document that the share of SOEP studies that provide code has increased sharply since the early 2010s, reaching a share of 17.8% in 2021. We argue that the increase in code provision is driven by a mix of technological advances, individual researcher awareness, and journal policies. Technological advances, including the widespread use of individual researcher websites, the establishment of journal websites, and the creation of specialized online 358 - FINK and MARCUS repositories, have reduced the financial and logistical barriers to code sharing. We also present suggestive evidence that top‐down initiatives, such as journal policies on code sharing, have played an important role in increasing code availability over time. Moreover, the fact that a substantial fraction of articles only provide information about code availability on the authors' websites but not on the journal website or in the article suggests that individual researchers' awareness of the importance of code sharing also plays an important role. It also highlights the need to search the authors' websites to get a comprehensive overview of code availability. Our study contributes to the literature on research transparency and open science practices by providing the first analysis of the development of code provision over time and across fields. We particularly contribute to two strands of literature in this field. The first is the state of data and code‐sharing practices and factors that are associated with it. Previous studies show that many researchers favor data and code sharing (e.g., Baker, 2016; Christensen et al., 2020; Fecher et al., 2015; Ferguson et al., 2023). However, few researchers have actually shared data or code (Christensen et al., 2020). While most studies focus exclusively on the provision of data (e.g., Alsheikh‐Ali et al., 2011; Fecher et al., 2015; Gherghina & Katsanidou, 2013; Zenk‐Möltgen et al., 2018), some studies examine both data and code provision together. Colliard et al. (2023) find that less than 5% of papers published between 2010 and 2020 in one leading journal in finance, the Journal of Financial Economics, provide some data or code on the journal's website. Gertler et al. (2018) examine 203 empirical articles that were published in nine leading economics journals in the last three issues as of May 2016 and did not contain restricted data. They find that 72% of these articles contain the estimation code, but only 42% contain the data manipulation code that turns the raw data into the estimation data. Hardwicke et al. (2020) survey 156 empirical articles published between 2014 and 2017 across the social sciences and find that code is accessible for only one article, while Wallach et al. (2018) find that none of 104 empirical biomedical articles published between 2015 and 2017 mentioned the provision of code. Key (2016) studies the availability of data and code in empirical articles published in six leading political science journals in 2013/14. 58% of the 494 articles contain both data and code, and a further 2% of articles provide code but no data, giving a total of 60% of articles containing replication code, with considerable variation across journals. Key (2016) identifies journal policies as an important predictor of code and data availability. Gertler et al. (2018) examine the websites of 67 journals from different disciplines and find that most of the top political science and economics journals require authors to provide replication code and data before publication, while journals in sociology and psychology rarely do. The second strand relates to the development of research transparency practices over time. Here, a general finding in the literature is that different practices of research transparency are on the rise (for an overview on the development in economics, see Miguel, 2021). For instance, in a retrospective survey, the proportion of authors who reported using open science practices (including pre‐registration and public sharing of data, code, or study instruments) increased over time in both economics (Swanson et al., 2020) and other disciplines (Christensen et al., 2020; Ferguson et al., 2023). McCabe and Snyder (2014) and Mueller‐Langer and Watt (2018) take advantage of the fact that some journals have switched to (hybrid) open access to study the effect of (hybrid) open access on citations. Miguel (2021) presents evidence that pre‐registration and pre‐analysis plans are increasingly used in experimental economics and that more and more studies register every year at the Randomized Control Trial Registry of the American Economic Association. 5 Moreover, more and more journals in different disciplines allow for “pre‐results review,” also known as “registered reports,” where a research project is refereed based on its research plan and methodology before the research is conducted and results are obtained (Hardwicke & Ioannidis, 2018; Miguel, 2021). Further, journals have implemented various policies to foster research transparency. For instance, several journals introduced editorial statements underscoring the importance of disseminating findings with non‐significant or null results, and there is some evidence that these statements increased the share of results that are not statistically significant (Blanco‐Perez & Brodeur, 2020). Additionally, many journals implemented explicit code and data sharing policies, and several studies highlight the importance of journal policies for data and code provision as well as for reproducibility (Christensen et al., 2019; Fišar et al., 2024; Gherghina & Katsanidou, 2013; Mueller‐Langer et al., 2019; Zenk‐Möltgen et al., 2018). For these reasons, we put a special emphasis on journals' code‐sharing policies in our empirical analyses. Our study contributes to the previous literature by focusing specifically on code provision. While several studies look at data sharing, it is also important to study code sharing, the other facilitator of reproducibility, as every empirical researcher can share code but not necessarily data. Our focus on publications based on the same freely available data allows us to compare code availability across fields without the distortions that arise when the use of proprietary datasets varies across fields, and when code sharing behavior is influenced by the ability to share the data. In comparison, the existing literature primarily examines the availability of data and code within a selection of journals in a specific field (Brodeur, Cook, & Neisser, 2024; Colliard et al., 2023; Key, 2016; Rainey et al., 2024; Wallach et al., 2018). FINK and MARCUS - 359 We also contribute by applying a broader search approach as we do not only look at the journal websites for reproduction material but also at the authors' websites. Moreover, we examine a rather long time horizon spanning almost 40 years and focus not only on top journals but consider a much broader range of journals. The manuscript is organized as follows: Section 2describes our data and variables and presents descriptive statistics. Section 3reports empirical results on the state of replication code availability, its correlates, and its evolution over time, as well as how code availability relates to journal impact factors and article citations. Section 4concludes and provides an outlook for future research. 2 | DATA 2.1 | Basic data: SOEP and SOEPlit The German SOEP is one of the largest and longest‐running household panel surveys in the world: The first interviews were conducted about 40 years ago (in 1984), and currently, about 30,000 individuals in about 15,000 households are interviewed annually. The SOEP covers a wide range of topics, including questions on employment, income, health, education, demography, income, housing, life satisfaction, attitudes, values, and personality. For all these reasons, it is not surprising that the SOEP is highly regarded by economists and researchers from other social science disciplines: SOEP has more than 3500 users worldwide, and in the last decade, more than 100 peer‐reviewed articles have been published each year using SOEP data ‐ and between 250 and 300 other publications (including working papers, books, policy reports, etc.) (Goebel et al., 2019). The SOEP group at the German Institute for Economic Research (DIW Berlin) maintains a database that aims to cover all publications based on SOEP data. This database, SOEPlit (2022), is our main dataset. It contains bibliographic information about each publication, including title, type of publication, year of publication, language, journal (if applicable), digital object identifier (DOI, if available), and the names of the authors. Our version of SOEPlit is from June 2022, and we consider SOEP‐based publications up to the year 2021. We focus our analyses on peer‐reviewed articles published in journals listed in one of the Clarivate Analytics citation indices, that is, the Science Citation Index, the Social Science Citation Index, the Arts and Humanities Citation Index, and the Emerging Sources Citation Index. 6 We do not include working papers for two reasons. First, we want to avoid double counting since many working papers end up as peer‐reviewed articles. Second, authors may be more reluctant to share working paper code for fear of intellectual property theft. In total, our database consists of 2518 distinct peer‐reviewed publications. 2.2 | Data on code availability We then checked each publication for publicly available replication code. This was done in four different ways: (i) we examined the journal webpage of each article for code availability; (ii) We reviewed all 2518 articles for references to publicly available code 7 ; (iii) We searched the authors' webpages for code availability 8 ; (iv) We searched specific online repositories for code availability. 9 If code was available, we verified that the code is accessible by downloading and opening it. Our research focuses on code that is publicly available. We did not contact the authors of the articles. 10 2.3 | Additional data We integrated additional information at the article and journal level into our dataset. Article‐level information includes citations in Google Scholar, obtained through web scraping using the Python package Scholarly (Cholewiak et al., 2021). 11 Journal‐level information refers to journal impact factor (JIF) metrics as well as the primary discipline of the journal. We collected the 2‐year and 5‐year JIF metrics since 1997 from Clarivate's Master Journal List (2022). The JIF is merged by journal name and year of publication. 12 To assign the primary discipline of each journal, we use the classification of Science‐Metrix (2024), which categorizes journals into 174 mutually exclusive subfields, nested within 20 mutually exclusive fields (Archambault et al., 2011). We have grouped the subfields into five categories: Economics, Sociology, Psychology, Other Social Sciences, and Health & Other Sciences. 13 Appendix Table A1 provides, for each of our five disciplinary categories, the names of the journals with the most articles in our dataset. We merged the journal information from Science‐Metrix based on the International Standard Serial Number of the journal. 14 360 - FINK and MARCUS 2.4 | Information on journal policies Additionally, we included information about journals' policies on replication code, following the classification of the TOP framework (Nosek et al., 2015). At the lowest level (Level 0), a journal either encourages code sharing or does not specifically mention it. Moving up to Level 1, a code availability statement is required in the article, indicating whether the code is accessible and, if so, where it can be found. At Level 2, mandatory code sharing is implemented, requiring the code to be posted to a trusted repository, with any exceptions to be identified when the article is submitted. The highest level (Level 3) involves a reproducibility check, requiring that the code be posted to a trusted repository and that reported analyses be independently reproduced before publication. The TOP Factor Database (2023) provides the current level of replication code policy for 164 of the 478 journals (34%) in our sample. 15 We searched the websites of the other 315 journals for information on code‐sharing policies. 16 This search resulted in the current state of code‐sharing policies (as of November 2023). However, we are particularly interested in the historical development of these policies. We, therefore, wrote to editors of journals with a code‐ sharing policy level of at least 1 (code availability statement) to ask (i) whether our assessment was correct and (ii) when this particular policy (and previous code policies) had been adopted. This included the editors of 94 journals in our SOEP‐based sample. We received responses from 67 journals (71%) on the current status of the replication code policy, and 62 journals (67%) provided information on when the policy was adopted. 17 For the 32 journals from which we did not receive adoption dates, we tried to obtain the dates using the Wayback Machine (2024) to retrace when the policies were first on the journal's websites. In this way, we obtained dates for 27 of the 32 journals. 18 2.5 | Descriptive statistics Table 1provides descriptive statistics to get a better idea of the 2518 SOEP‐based articles in our dataset. Economics is the discipline with the most articles in our sample (45%), followed by sociology (16%), other social sciences (15%), health and other sciences (14%), and psychology (11%). While our dataset includes publications from 1985 to 2021, the number of SOEP‐based publications per year increased strongly over time, from 12 in 1990 (of which 3 in economics) to 26 in 2000 (15 in economics), 118 in 2010 (68 in economics), and 176 in 2020 (52 in economics). TABLE 1Descriptive statistics. Mean SD Min Max Economics 0.45 0.50 0.00 1.00 Sociology 0.16 0.36 0.00 1.00 Psychology 0.11 0.31 0.00 1.00 Other social sciences 0.15 0.36 0.00 1.00 Health and other sciences 0.14 0.34 0.00 1.00 Publication year 2011.91 7.38 1985.00 2021.00 Journal impact factor (JIF) 1.86 2.20 0.04 40.14 5 Year JIF 2.60 2.58 0.09 43.77 Citation count from google scholar 106.19 242.25 0.00 4359.00 Language of article is English 0.87 0.33 0.00 1.00 Single author 0.28 0.45 0.00 1.00 Two authors 0.38 0.49 0.00 1.00 Three authors 0.21 0.41 0.00 1.00 Four or more authors 0.13 0.33 0.00 1.00 Note: The table displays mean, standard deviation, minimum and maximum for different variables in our sample of 2518 SOEP‐based publications. Source: Own calculations, based on SOEPlit (2022), Google Scholar (2023), and journal impact factors from Master Journal List (2022). FINK and MARCUS - 361 While the average journal impact factor is 1.86 (2.60 for the 5‐year impact factor), this is a highly skewed variable with a maximum of 40 overall (Nature in 2016) and a maximum of 9.6 in economics (Journal of Political Economy in 2021). Another skewed variable is the number of citations, which ranges from 0 to 4359, with an average of 106. Thirty articles have more than 1000 citations, while about 430 articles have less than 10 citations (including 90 with no citations). The vast majority of articles are written in English (87%). The remaining articles are written in German (319 articles; 13%) and French (2 articles; less than 1%). All economics articles in our sample are in English. About 28% are single‐author articles, while most articles have two authors (38%). About 13% of the articles have four or more authors. The 2518 articles are published in 478 different journals, comprising 153 economics journals and 325 journals from other disciplines. Three hundred twelve journals (65%) have at most two articles in our dataset, while 14% of journals contribute 10 or more articles. The 10 most frequent journals in our dataset account for 26% of the articles in our sample. Appendix Table A1 lists the journals with the most articles in our dataset. While previous literature has focused mainly on the top journals in economics, it is evident that we consider a much broader range of journals. The most common economics journal in our dataset is “Labor Economics” (79 articles), but our dataset also includes economics journals ranked higher, including 20 articles from Top‐5 journals. As well as journals ranked lower. Similarly, a wide range of high to low‐ranked journals is included for other disciplines. Thus, our results do not primarily speak to the situation at the top, but provide a broader view of the state of replication code availability. 3 | RESULTS This section begins by looking at the current state of replication code availability and its correlates in Section 3.1. The following Section 3.2 describes trends in code availability over time and discusses possible explanations for the observed trends. Section 3.3 examines several factors related to code availability simultaneously, while Section 3.4 explores the relationship between code availability and quality metrics. Finally, Section 3.5 discusses results related to a more rigorous measure of code availability. 3.1 | Replication code availability and its correlates Across all disciplines, we found replication code for 151 of the 2518 SOEP‐based publications (6%). Figure 1shows the availability of code by discipline: We obtained replication code for 43 of the 1132 SOEP‐based articles published in economics journals (3.8%) and for 108 of the 1386 articles published in journals of other disciplines (7.8%). The share of articles with available code is lower in economics than in the other four disciplinary categories in our SOEP‐based sample. Psychology has the highest share of articles with available code (almost 10%), followed by other social sciences (9%). We also find that articles written by a single author are less likely to contain code—and this is true in both economics and other disciplines (see Appendix Figure A1). One possible reason for this observation is that in multi‐author collaborations, it is more likely that at least one author has a personal website and places importance on reproducibility. Another explanation could be that single‐author articles are less likely to be published in journals with stricter code‐ sharing policies. A further explanation is that the importance of reproducibility becomes more apparent in multi‐ author collaborations. In such collaborations, ensuring that all authors obtain identical results by running the code on their respective computers becomes a more straightforward task, underscoring the importance of reproducibility. This would also align with the observation of Herbert et al. (2021) that articles with more authors tend to have better code documentation. Furthermore, code is more likely to be available if the article is written in English (see Figure A2 in the Appendix). Section 3.3 examines whether these bivariate correlations also hold when controlling for other factors, including publication year dummies and indicators for the journal's code‐sharing policy. Next, we look more closely at the articles that provide code. The majority of articles with code provided use Stata (87%), a proprietary software, followed by R with 21% (see Appendix Figure A3). 19 SPSS, MATLAB, MPLUS, and SAS each account for less than 4% of studies with code. These proportions are roughly comparable to the figures from the 2018 SOEP user survey, which asked respondents about the statistical packages they use to analyze SOEP data: About 77% use Stata, 24% use R, 23% use SPSS, 11% use other (Britzke & Schupp, 2019). 20 Similarly, in a collection of more than 8000 economic articles with replication packages, Kranz (2023) finds that Stata is the most commonly used 362 - FINK and MARCUS software package (71.6%), followed by Matlab (24.5%) and R (9.8%). Vilhuber (2020) reports similar shares for articles in the journals of the American Economic Association. In both economics and other disciplines, most studies that provide code use the journal website to provide information about code availability (see Figure 2). 21 However, for a sizable proportion of studies, code availability is only disclosed via the authors' websites (30% in economics and 18% in other disciplines), highlighting the importance of also checking authors' websites for code availability: Looking only at journal websites would lead to a substantial underestimation of code availability. In general, only a few studies disseminate the code exclusively via online repositories, without referring to the availability of the code via the journal or the authors' websites (1 article in economics, 4 articles in other disciplines). It is also interesting to see that it seems less common in economics to disclose the availability of the code both through the journal and the authors' websites (5% vs. 32%). Economics also differs from other disciplines in terms of where the code provided is actually stored (see Figure A4 in the Appendix). While for the majority of studies in economics, code can be downloaded directly from author or journal websites, in other disciplines, 49% of studies with code rely on the repository of the Open Science Framework (OSF), compared to less than 5% in economics. 22 3.2 | Developments in code availability over time Figure 3displays how the availability of replication code has changed over the years. It starts with the first SOEP‐based peer‐reviewed publications in 1985 (1989 in economics) and goes up to 2021. The figure shows several striking patterns. First, by 2012 code is available for only five studies in economics (0.85%) and three (0.55%) in other disciplines. In economics, the first publication with available code is from 2007 (Alesina & Fuchs‐Schündeln, 2007), while overall, the earliest code we obtained is from 1995 (Rendtel et al., 1995). Here the SPSS code was printed as an appendix to the paper itself, highlighting the challenges researchers faced in sharing code before the widespread use of the Internet. The first digital code we found is from 1998 (Kohler, 1998). Second, the figure reveals that from 2012 onward, the annual share of publications with code increased strongly for both economics and other disciplines. The increase in economics and other disciplines from less than 2% in 2012 to 17% and 18% in 2021 suggests a huge change in code‐sharing behavior. In economics, the biggest growth happened after 2019, while in the other disciplines the strongest increase occurred a few years earlier. Economics was lagging behind, but is catching up in recent years. The breakdown of the non‐economics category by discipline reveals that the increase in code provision is a general phenomenon and not driven by a single discipline (see Figure A5 in the Appendix). Until around 2015, the share of FIGURE 1 Publications with publicly available code by discipline. This figure shows the share of SOEP‐based publications with publicly available replication code among all SOEP‐based publications in the respective discipline. Source: Own calculations, based on SOEPlit (2022) and own research on code availability. FINK and MARCUS - 363 publications with code was similar across disciplines. However, after 2015, the increase in sociology, psychology, and other social sciences was earlier and stronger than in economics. Next, we look at how the mode of code availability disclosure and the location of storage have changed over time. The proportion of studies we found information on code availability through the journal increased strongly over the FIGURE 2 Publications with publicly available code by mode of disclosure. This figure displays the distribution of methods used to inform others about the availability of code among all SOEP‐based publications with publicly available replication code. The left panel relates to the situation in economics, while the right panel relates to all other disciplines. Source: Own calculations, based on SOEPlit (2022) and own research on code availability. FIGURE 3 Publications with publicly available code over time. This figure shows the share of SOEP‐based publications with publicly available replication code among all SOEP‐based publications in the respective year and discipline. Source: Own calculations, based on SOEPlit (2022) and own research on code availability. 364 - FINK and MARCUS from 2007 (12 years later than the first publication with any replication code), and the first publication outside of economics that provides code according to the stricter definition is from 2012 (Kohler et al., 2012). While none of the studies providing code met the stricter definition in the period before 2007, the proportion of studies with code meeting the stricter definition remained relatively constant over time (75% in 2007–2011%, 76.7% in 2012–2016%, 75.7% in 2017–2021). This suggests that code quality, as measured by our crude measure, has not improved much over time. However, it also suggests that the increased rate of code provision over time did not come at the expense of the quality of the code provided. The Online Appendix repeats all our analyses, this time using the stricter definition of code availability. Generally, the patterns are very similar to those with our main definition of code availability. 4 | CONCLUSION In his famous article “Let's take the con out of econometrics” Edward Leamer (1983) heavily criticized the state of empirical economic research, culminating in the quote “hardly anyone takes anyone else's data analysis seriously.” If nobody takes data analysis of other researcher serious, it is a serious problem for a profession that strongly relies on empirical results and calls for evidence‐based policy‐making. However, Angrist and Pischke (2010) argue that the Leamer critique is no longer justified in empirical microeconomics due to the commitment to rigorous research designs for identifying causal effects. This is what they label the credibility revolution in their aptly titled article “The credibility revolution in empirical economics: How better research design is taking the con out of econometrics.” However, recent studies by, for example, Huntington‐Klein et al. (2021), Breznau et al. (2022), and Menkveld et al. (2024) highlight a disconcerting phenomenon: researchers drawing disparate conclusions when analyzing the same data to answer identical questions, even when they apply rigorous research designs for the identification of causal effects. This is due to the many decisions that researchers must make during data collection, preparation, and analysis, also known as the researcher's degree of freedom. It suggests that relying solely on credible research designs does not fully address the lingering concerns voiced by Leamer (1983). Making data and code used in the empirical analysis publicly available is one way to make the researcher's degree of freedom more transparent and to examine how crucial it is for a given research question. Having access to data and code would allow others to check all the details of the empirical analyses and run additional sensitivity checks. Against this backdrop, this study looks at how often researchers share the computer code they use for their empirical analyses, focusing on publications that use the SOEP dataset—a widely‐used data source in economics and other disciplines. We searched extensively to find out if the code was available for all 2518 articles using SOEP in peer‐ reviewed journals. We checked the journal websites, the articles themselves, the websites of the authors, and specific online repositories. The analysis shows that only 6% of SOEP studies have accessible replication code, a low proportion, particularly considering that code provision is crucial for the reproducibility of scientific results and associated with little costs for the researchers. However, there is a positive trend as the share of studies with available code increased strongly over time in both economics and other disciplines. We argue that this increase is driven by a mixture of three factors: Technological advances (including websites of authors and journals as well as the establishment of specific repositories), top‐down initiatives of journals, for example, mandating code sharing, and bottom‐up initiatives of individual researchers, who post code on their individual websites. Additionally, we find that studies with accessible code tend to be published in journals with higher impact factors and receive more citations, emphasizing the potential of code sharing as a quality signal in academic research. Our analysis can serve as a first step for further analyses. We sketch five of such further analyses. First, we examine whether studies provide replication code, but we do not study whether the code actually allowed the results to be reproduced. We take a first step in this direction, by applying the stricter definition of code availability where the code must load raw SOEP data. However, it is beyond the scope of this paper to analyze the reproducibility of all SOEP papers with code provided. Second, another potential avenue for future research is to go more in the direction of causality. While we document correlations between journal policies and code availability, future research could evaluate the impact of code‐sharing policies on code availability with more rigorous methods for causal inference. Similarly, future studies could look at the effects of code provision on citation metrics. 32 Third, while in the introduction we argue that code provision might work as a deterrent against fraud and questionable research practices like p‐hacking, further analyses could study whether this is actually the case. 33 Fourth, further research could identify the specific barriers that FINK and MARCUS - 371 prevent authors from adopting the habit of making code available, and ways to overcome these barriers (e.g., journals providing assistance in the form of instructions or videos on how to prepare replication packages). Fifth, future studies could also examine code availability beyond the SOEP. A possible approach would be to investigate the development of code availability in articles published in leading journals across various disciplines. However, this sampling method may not offer a comprehensive overview of code sharing as it concentrates solely on the most prestigious journals. Further, in many cases the provision of code will not be enough to allow for reproducibility as data are missing and there are substantial differences between disciplines regarding the proportion of data that are shareable with the research community. While publications based on the SOEP are not a random sample of studies in economics or other disciplines, a complimentary strategy would be to draw random samples from publications in different disciplines. However, this procedure presents several challenges. Firstly, it is necessary to verify whether the study is an empirical one and the rates of empirical studies differ between disciplines. Secondly, the disciplines vary in terms of the proportion of publications based on primary or secondary data analyses. Thirdly, it is unclear what the population is and how to construct a sampling frame. By focusing on the SOEP, we avoid many of these issues as the SOEP levels the playing field as the dataset is publicly available by researchers around the world. In this setting code availability is the most crucial ingredient for reproducibility. Providing replication code will not completely solve all problems of questionable research practices and scientific malpractice. However, it will mitigate some of these problems and make it easier for new studies to build on previous ones. In addition, code provision is a relatively inexpensive measure: Individual researchers can participate voluntarily, and journals can adopt explicit code‐sharing policies. For journals, the cost of implementing a Level‐1 code‐sharing policy (code availability statement) is relatively low. Journals simply need to update their web pages to explicitly mention the policy and ensure that each accepted paper includes a code availability statement. The cost of implementing a Level‐2 policy (mandatory code provision) is only slightly higher, as journals must enforce this policy by verifying that the code is posted in a trusted repository and that the link provided is functional. A potential downside of implementing an explicit code‐sharing policy could be that fewer researchers would submit their work to such journals. However, we believe that this is unlikely because the costs to individual researchers are relatively low for both Level‐1 and Level‐2 policies. Moreover, if the policy deters researchers who might engage in fraudulent behavior, it could be beneficial for the journal to receive fewer but higher quality submissions. From a pure reproducibility perspective, a Level‐3 policy (reproducibility check) is desirable. However, this is by far the most expensive code‐sharing policy. Verifying that the code allows for reproducibility is costly, and many journals employ data editors to perform these checks, at considerable expense. In our survey, editors from several journals indicated that while they would like to implement a Level‐3 policy, it is simply not financially feasible. While we believe that only select top‐tier journals can afford comprehensive reproducibility checks, all journals can (and should) require authors of accepted manuscripts to make their code publicly available. We conclude this paper by arguing for what we call the Credibility Revolution 2.0. This new paradigm contends that the mere focus on research designs is not enough to address the criticisms of Leamer (1983). Instead, it calls for a fundamental shift toward sharing code and, if possible, data as the new normal. In our view, the credibility lies not only in the design of the research, but also in the ability of peers to scrutinize every detail, facilitating collective validation of results through reproducibility and transparency. Providing replication code is a low‐cost measure in this new paradigm, and basically any empirical researcher can engage in it, paving the way for more trustworthy empirical research. ACKNOWLEDGMENTS We thank Toman Barsbai, Markus Jäntti, and Michael Kvasnicka for helpful comments and feedback. We are also indebted to Niema El Moufti, Ashwin Fernandez, and Alexander Grawe for outstanding research assistance. We are grateful to Uta Rahmann of the SOEP group for providing us with the SOEPlit database and for answering our questions about SOEPlit. A big thank you also goes to all the journal editors who took the time to answer our questions about the implementation of specific code‐sharing policies. Open Access funding enabled and organized by Projekt DEAL. DATA AVAILABILITY STATEMENT The code and data to reproduce all the tables and figures in the paper are openly available in the openICPSR repository at https://doi.org/10.3886/E209316V2 (Fink & Marcus, 2024). 372 - FINK and MARCUS ORCID Lukas Fink https://orcid.org/0009-0002-8786-0988 Jan Marcus https://orcid.org/0000-0001-9407-6660 ENDNOTES 1 For example, Foote and Goetz (2008) found a coding error in Donohue and Levitt (2001)'s study on legalized abortion and crime, while Herndon et al. (2013) found a coding error in an Excel spreadsheet in Reinhart and Rogoff (2010)'s original study on high public debt and low growth rates. Although these examples of coding errors received a lot of attention, (unintentional) coding errors often have only minor consequences for the conclusions of an article (Laurinavichyute et al., 2022). 2 For data sharing, researchers may need to label variables and carefully consider which variables should not be released. This may be to reduce the risk of de‐anonymizing specific observations or because they plan to publish another paper using the same data and variables that have not been used. 3 One might consider including all details of the empirical analyses in the article as an alternative to code provision. However, journals have limited space, and even the most careful researchers find it challenging to include all the important details of their analysis in the article (Krähmer et al., 2023). Moreover, our own experience in trying to reproduce results of publications using only the provided data and information provided in the article tells us that this is not only time‐consuming, but in most cases it is impossible to reproduce the results due to the complexity of most empirical analyses and the many decisions researchers have to take. 4 More specifically, researchers at universities and research institutions can use the data for research and teaching purposes by signing a data distribution contract, wherein they agree to comply with the data security and privacy requirements of the European General Data Protection Regulation. Data protection rules differ inside and outside the EU. Within the EU, users have full access to the SOEP core dataset, outside the EU, they have access to a 95% sample of this dataset. The core SOEP data contain information on the federal state. More detailed regional data are available via remote or on‐site access. See SOEP Group (2024) and Vilhuber (2023) for more information. 5 The purpose of pre‐registrations and pre‐analysis plans is to enhance confidence in a study's findings by reducing researchers' degree of freedom and the capacity to choose results selectively. While in economics, pre‐registration refers to the process of registering and describing the basic characteristics of a study (usually) before it is conducted, pre‐analysis plans additionally describe in detail how the analysis will be performed (Miguel, 2021). In economics, pre‐registration often comes without a pre‐analysis plan, while in other disciplines, pre‐registration, by definition, includes a pre‐analysis plan (Brodeur, Cook, Hartley, & Heyes, 2024). The empirical study by Brodeur, Cook, Hartley, and Heyes (2024) shows that pre‐registration reduces p‐hacking and publication bias only when accompanied by a pre‐analysis plan. 6 In our sample, based on the first category of Clarivate's Master Journal List, the Social Science Citation Index accounts for the majority of articles (86%), while the Science Citation Index makes up another 10%. 7 In particular, student assistants reviewed the acknowledgments and searched the paper for key terms such as “code,” “replication,” “syntax,” “script” and “Stata.” 8 More specifically, our student assistants checked the personal webpages or, if not found, the institutional webpages of the first four authors for code. This results in 2615 distinct authors. We only consider the first four authors as more than four authors are present in only 6% of studies. Among these, a few studies have a large number of authors, including three papers with close to 100 authors. 9 These include the GESIS data archive, Harvard Dataverse, Open Science Framework, openICPSR, and Zenodo. We used the search terms “SOEP,” “GSOEP,” “Socio‐Economic Panel,” and “Sozio‐oekonomisches Panel.” 10 Krähmer et al. (2023) took a different approach, contacting authors who had published with data from the European Social Survey and asking if they would share their code. 37.5% of successfully contacted authors provided code. 11 Citations are as of November 28, 2023. 12 Since JIF information is only available from 1997 onwards, and some journals were included only in later years, we impute the impact factor for earlier years based on the first year available (this applies to 11% of the articles in our samples). We include a missing dummy equal to one for these articles in our regressions where JIF is the dependent variable or included as a control variable. In addition, 2‐year impact factors are not available for some journals that have only recently been included in the citation indices. This applies to less than 4% of the articles in our sample. We exclude these articles in the regression analyses with JIF as the dependent variable or as a control variable. 13 Specifically, we have assigned the following subfields of “Economics & Business” to the discipline of “Economics”: “Economics,” “Econometrics,” “Economic Theory,” “Finance,” “Agricultural Economics & Policy,” “Development Studies,” and “Industrial Relations,” as well as the subfield “Health Policy & Services” of the field “Public Health & Health Services” (since in our sample this consists mainly of articles in the Journal of Health Economics and Health Economics). The subfield “Economics” accounts for more than 84% of the articles in the discipline “Economics” in our sample. We assigned to “Sociology” (discipline), the subfield “Sociology,” to “Psychology” all subfields of the field “Psychology & Cognitive Sciences” as well as the subfield “Psychiatry,” to “Other Social Sciences” the field “Social Sciences” (except the subfield ‘Sociology’) as well as all remaining subfields of the field “Economics & Business.” Journals in all other fields and subfields were assigned to the residual discipline “Health and other sciences.” We decided not to further subdivide the disciplines to have at least 10% of the observations in each category. FINK and MARCUS - 373 14 Merging was not successful for seven journals (39 articles) in our dataset. For these seven journals, we imputed the discipline based on the first category of Clarivate's Master Journal List (2022). In addition, Science‐Metrix does not provide a journal‐based classification for 15 multidisciplinary journals in our dataset (65 articles, including for example, PLOS ONE, Economics & Human Biology, PNAS, and Nature). Again, we assign these journals based on the first category of Clarivate's Master Journal List. 15 The TOP Factor Database has a community‐based approach where editors and registered users can submit a journal evaluation, which Center for Open Science staff reviews before publishing the scores (for more information, see COS, 2024). For all 164 journals, the last update of the code‐sharing policy was in 2021 or later. The coverage of the journals in the TOP Factor Database differs substantially across disciplines: The TOP Factor Database covers 69% of psychology journals of our sample, 38% of health and other sciences, 32% of sociology, 27% of economics, and 26% of other social science journals. 16 We focused on code sharing policies and considered data sharing policies only when code, syntax files, or computer programs were explicitly mentioned as part of “data.” 17 We also double checked with the dates provided in Brodeur, Cook, and Neisser (2024) and Christensen and Miguel (2018), which focused on data availability policies, for 25 and 11 economics journals respectively. In general, we took the information provided by journal websites and editors for granted and did not check the extent to which these policies are actually enforced (see also the discussion in Colliard et al., 2023). 18 As a result, we are missing the adoption dates for five journals, corresponding to 68 articles in our dataset. We exclude these 68 articles from the analyses that rely on the implementation dates of code‐sharing policies. 19 In economics, code is only provided in Stata (95%), Matlab (9%) and R (7%) in our sample. The percentages add up to more than 100% because some articles rely on more than one software program. 20 The shares are similar in the 2017 SOEP user survey with 76% Stata, 31% R, 19% SPSS, 10% other (Britzke & Schupp, 2019). 21 We also consider code found through references in the articles as having been disseminated through the journal because the articles are accessible through the journal website. 22 The OSF is a network for researchers to document their research projects, for example, by registering analysis plans and sharing their material. It also allows a DOI to be assigned to the code provided, making it possible to keep the code permanently available. 23 In line with these findings, OSF users and usage have increased exponentially since its inception in 2012 (see Nosek et al., 2022). 24 Miguel (2021, 196) highlights the importance of these online repositories, saying that they “have been so successful that it is easy to forget what an important innovation the professional curation, storage, and management of research data and code has been.” 25 The Journal of Political Economy implemented a similar Level‐2 policy in 2005 and the Review of Economic Studies in 2006. This was followed by the Brookings Papers on Economic Activity in 2007, the Canadian Journal of Economics in 2008, and the American Economic Journals in 2009. 26 While the upper panels of Figure 4show that some economic journals are rather front‐runners in terms of code‐sharing policies, the first journal with an explicit code‐sharing policy was the Stata Journal. The journal, due to its focus on software code, requires since the first issue in 2001 that software code is made available through the journal's website. 27 The Biometrical Journal was, in 2009, the first journal in our sample to implement a policy of independently reproducing the results of accepted manuscripts prior to publication (Hofner et al., 2016). However, code sharing is not mandatory: “the Editor in Chief notifies the authors …already during the review process of the journal's RR [reproducible research] policy and strongly encourages the submission of code and data to make the article reproducible” (Hofner et al., 2016, p. 418). 28 The fifth article is published in the Journal of Applied Econometrics, which also strongly emphasizes the replicability of results by other researchers. While the journal has a mandatory data sharing policy (if the data are not confidential), code sharing is only encouraged but not mandatory according to the journal website (Journal of Applied Econometrics, 2024). 29 We also exclude the 68 articles published in the five journals, for which we could not obtain the implementation dates. 30 More specifically, we estimate equations of the following form: yp¼αþx0 pβþX 2021 t¼1986 δt⋅1yearp¼t�þεp;ð1Þ where ypis a binary variable indicating whether the code is publicly available for publication p.xpis the vector of explanatory variables (e.g., indicators for the journal’s discipline); δtrepresents the coefficients of a set of dummy variables corresponding to each publication year. Finally, εpis the error term. We use heteroscedasticity‐robust standard errors throughout our analysis, as the variance of εpmay not be constant across different values of the explanatory variables. 31 Studies published in early years have more time to accumulate citations. Therefore, Figure 7controls for year of publication dummies. 32 Going in this direction, for example, Christensen et al. (2019) provide some evidence that data sharing increases citations when using changes in data sharing policies as instrumental variables. 374 - FINK and MARCUS 33 For instance, Brodeur, Cook, and Neisser (2024) find evidence that data‐sharing policies do not reduce p‐hacking and publication bias, cautioning against being too optimistic regarding the positive effects of mandatory research transparency initiatives. REFERENCES Alesina, A. & Fuchs‐Schündeln, N. (2007) Good‐bye Lenin (or not?): the effect of communism on people’s preferences. The American Economic Review, 97(4), 1507–1528. Available from: https://doi.org/10.1257/aer.97.4.1507 Alsheikh‐Ali, A.A., Qureshi, W., Al‐Mallah, M.H. & Ioannidis, J.P. (2011) Public availability of published research data in high‐impact journals. PLoS One, 6(9), e24357. Available from: https://doi.org/10.1371/journal.pone.0024357 Angrist, J.D. & Pischke, J.‐S. (2010) The credibility revolution in empirical economics: how better research design is taking the con out of econometrics. The Journal of Economic Perspectives, 24(2), 3–30. Available from: https://doi.org/10.1257/jep.24.2.3 Ankel‐Peters, J., Fiala, N. & Neubauer, F. (2023) Do economists replicate? Journal of Economic Behavior & Organization, 212, 219–232. Available from: https://doi.org/10.1016/j.jebo.2023.05.009 Archambault, É., Beauchesne, O.H. & Caruso, J. (2011) Towards a multilingual, comprehensive and open scientific journal ontology. In: Proceedings of the 13th international conference of the international society for scientometrics and informetrics, volume 13, pp. 66–77. Durban, South Africa. Baker, M. (2016) 1,500 scientists lift the lid on reproducibility. Nature, 533(26), 353–366. Available from: https://doi.org/10.1038/533452a Bernanke, B.S. (2004) Editorial statement. The American Economic Review, 94(1), 404. Blanco‐Perez, C. & Brodeur, A. (2020) Publication bias and editorial statement on negative findings. The Economic Journal, 130(629), 1226– 1247. Available from: https://doi.org/10.1093/ej/ueaa011 Breznau, N., Rinke, E.M., Wuttke, A., Nguyen, H.H., Adem, M., Adriaans, J., et al. (2022) Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty. Proceedings of the National Academy of Sciences, 119(44), e2203150119. Available from: https://doi.org/10.1073/pnas.2203150119 Britzke, J. & Schupp, J. (Eds.) (2019) SOEP Wave Report 2018. Berlin. Available from: https://www.diw.de/documents/publikationen/73/ diw_01.c.632608.de/wave_report_2018.pdf Brodeur, A., Cook, N. & Neisser, C. (2024) P‐hacking, data type and data‐sharing policy. The Economic Journal, 134(659), 985–1018. Available from: https://doi.org/10.1093/ej/uead104 Brodeur, A., Cook, N.M., Hartley, J.S. & Heyes, A. (2024) Do pre‐registration and pre‐analysis plans reduce p‐hacking and publication bias? Evidence from 15,992 test statistics and suggestions for improvement. Journal of Political Economy: Microeconomics, forthcoming, 2(3), 527–561. Available from: https://doi.org/10.1086/730455 Camerer, C.F., Dreber, A., Forsell, E., Ho, T.‐H., Huber, J., Johannesson, M., et al. (2016) Evaluating replicability of laboratory experiments in economics. Science, 351(6280), 1433–1436. Available from: https://doi.org/10.1126/science.aaf0918 Camerer, C.F., Dreber, A., Holzmeister, F., Ho, T.‐H., Huber, J., Johannesson, M., et al. (2018) Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2(9), 637–644. Available from: https://doi.org/10. 1038/s41562‐018‐0399‐z Chang, A.C. & Li, P. (2017) A preanalysis plan to replicate sixty economics research papers that worked half of the time. American Economic Review, Papers & Proceedings, 107(5), 60–64. Available from: https://doi.org/10.1257/aer.p20171034 Cholewiak, S.A., Ipeirotis, P., Silva, V. & Kannawadi, A. (2021) Scholarly: simple access to Google Scholar authors and citation using Python. Available from: https://github.com/scholarly‐python‐package/scholarly Christensen, G., Dafoe, A., Miguel, E., Moore, D.A. & Rose, A.K. (2019) A study of the impact of data sharing on article citations using journal policies as a natural experiment. PLoS One, 14(12), e0225883. Available from: https://doi.org/10.1371/journal.pone.0225883 Christensen, G. & Miguel, E. (2018) Transparency, reproducibility, and the credibility of economics research. Journal of Economic Literature, 56(3), 920–980. Available from: https://doi.org/10.1257/jel.20171350 Christensen, G., Wang, Z., Paluck, E.L., Swanson, N., Birke, D., Miguel, E., et al. (2020) Open science practices are on the rise: the state of social science (3s) survey. Working Paper Series. Center for Effective Global Action. (No. 106). Colliard, J.‐E., Hurlin, C. & Pérignon, C. (2023) The economics of research reproducibility. HEC Paris Research Paper, FIN‐2019‐1345 (December 20, 2023). COS (2024) New measure rates quality of research journals’ policies to promote transparency and reproducibility. Available from: https:// www.cos.io/about/news/new‐measure‐rates‐quality‐research‐journals‐policies‐promote‐transparency‐and‐reproducibility [Accessed 23 May 2024]. Donohue, J.J.I. & Levitt, S.D. (2001) The impact of legalized abortion on crime. Quarterly Journal of Economics, 116(2), 379–420. Available from: https://doi.org/10.1162/00335530151144050 Eubank, N. (2016) Lessons from a decade of replications at the Quarterly Journal of Political Science. PS: Political Science and Politics, 49(2), 273–276. Available from: https://doi.org/10.1017/s1049096516000196 Fecher, B., Friesike, S. & Hebing, M. (2015) What drives academic data sharing? PLoS One, 10(2), 1–25. Available from: https://doi.org/10. 1371/journal.pone.0118053 Ferguson, J., Littman, R., Christensen, G., Paluck, E.L., Swanson, N., Wang, Z., et al. (2023) Survey of open science practices and attitudes in the social sciences. Nature Communications, 14(1), 5401. Available from: https://doi.org/10.1038/s41467‐023‐41111‐1 FINK and MARCUS - 375 Fink, L. & Marcus, J. (2024) ECIN replication package for “Replication code availability over time and across fields: evidence from the German Socio‐Economic Panel”. Ann Arbor, MI: Inter‐university Consortium for Political and Social Research [distributor], 2024‐09‐25. Available from: https://doi.org/10.3886/E209316V2 Fišar, M., Greiner, B., Huber, C., Katok, E. & Ozkes, A.I. & Management Science Reproducibility Collaboration (2024) Reproducibility in management science. Management Science, 70(3), 1343–1356. Available from: https://doi.org/10.1287/mnsc.2023.03556 Foote, C.L. & Goetz, C.F. (2008) The impact of legalized abortion on crime: comment. Quarterly Journal of Economics, 123(1), 407–423. Available from: https://doi.org/10.1162/qjec.2008.123.1.407 Gertler, P., Galiani, S. & Romero, M. (2018) How to make replication the norm. Nature, 554(7693), 417–419. Available from: https://doi.org/ 10.1038/d41586‐018‐02108‐9 Gherghina, S. & Katsanidou, A. (2013) Data availability in political science journals. European Political Science, 12(3), 333–349. Available from: https://doi.org/10.1057/eps.2013.8 Goebel, J., Grabka, M.M., Liebig, S., Kroh, M., Richter, D., Schröder, C., et al. (2019) The German Socio‐Economic Panel (SOEP). Jahrbucher für Nationalokonomie und Statistik, 239(2), 345–360. Available from: https://doi.org/10.1515/jbnst‐2018‐0022 Google Scholar (2023) Google Scholar database. Available from: https://scholar.google.com [Accessed 28 November 2023]. Hardwicke, T.E. & Ioannidis, J.P. (2018) Mapping the universe of registered reports. Nature Human Behaviour, 2(11), 793–796. Available from: https://doi.org/10.1038/s41562‐018‐0444‐y Hardwicke, T.E., Wallach, J.D., Kidwell, M.C., Bendixen, T., Crüwell, S. & Ioannidis, J.P. (2020) An empirical assessment of transparency and reproducibility‐related research practices in the social sciences (2014–2017). Royal Society Open Science, 7(2), 190806. Available from: https://doi.org/10.1098/rsos.190806 Herbert, S., Kingi, H., Stanchi, F. & Vilhuber, L. (2021) The reproducibility of economics research: a case study. Banque de France Working Paper, (853). Herndon, T., Ash, M. & Pollin, R. (2013) Does high public debt consistently stifle economic growth? A critique of Reinhart and Rogoff. Cambridge Journal of Economics, 38(2), 257–279. Available from: https://doi.org/10.1093/cje/bet075 Hofner, B., Schmid, M. & Edler, L. (2016) Reproducible research in statistics: a review and guidelines for the Biometrical Journal. Biometrical Journal, 58(2), 416–427. Available from: https://doi.org/10.1002/bimj.201500156 Huntington‐Klein, N., Arenas, A., Beam, E., Bertoni, M., Bloem, J.R., Burli, P., et al. (2021) The influence of hidden researcher decisions in applied microeconomics. Economic Inquiry, 59(3), 944–960. Available from: https://doi.org/10.1111/ecin.12992 Journal of Applied Econometrics (2024) Journal of Applied Econometrics ‐ Author Guidelines. Available from: https://onlinelibrary.wiley. com/page/journal/10991255/homepage/forauthors.html [Accessed 24 May 2024]. Key, E.M. (2016) How are we doing? Data access and replication in political science. PS: Political Science & Politics, 49(2), 268–272. Available from: https://doi.org/10.1017/s1049096516000184 Kohler, U. (1998) Zur Attraktivität der Grünen bei älteren Wählern. Kölner Zeitschrift für Soziologie und Sozialpsychologie, 50, 536–559. Kohler, U., Ehlert, M., Grell, B., Heisig, J.P., Radenacker, A. & Wörz, M. (2012) Verarmungsrisiken nach kritischen Lebensereignissen in Deutschland und den USA. KZfSS Kölner Zeitschrift für Soziologie und Sozialpsychologie, 64(2), 223–245. Available from: https://doi.org/ 10.1007/s11577‐012‐0164‐7 Krähmer, D., Schächtele, L. & Schneck, A. (2023) Care to share? Experimental evidence on code sharing behavior in the social sciences. PLoS One, 18(8), e0289380. Available from: https://doi.org/10.1371/journal.pone.0289380 Kranz, S. (2023) Usage shares of programming languages in economics research. Available from: https://skranz.github.io/r/2023/12/29/ FindingEconomicArticles7.html Laurinavichyute, A., Yadav, H. & Vasishth, S. (2022) Share the code, not just the data: a case study of the reproducibility of articles published in the journal of memory and language under the open data policy. Journal of Memory and Language, 125, 104332. Available from: https://doi.org/10.1016/j.jml.2022.104332 Leamer, E.E. (1983) Let’s take the con out of econometrics. The American Economic Review, 73(1), 31–43. Master Journal List. (2022) Web of Science master journal list by Clarivate. Available from: https://mjl.clarivate.com [Accessed 10 October 2022]. McCabe, M.J. & Snyder, C.M. (2014) Identifying the effect of open access on citations using a panel of science journals. Economic Inquiry, 52(4), 1284–1300. Available from: https://doi.org/10.1111/ecin.12064 McCullough, B.D., McGeary, K.A. & Harrison, T.D. (2006) Lessons from the JMCB archive. Journal of Money, Credit, and Banking, 38(4), 1093–1107. Available from: https://doi.org/10.1353/mcb.2006.0061 Menkveld, A.J., Dreber, A., Holzmeister, F., Huber, J., Johanneson, M., Kirchler, M. et al. (2024) Non‐standard errors. The Journal of Finance, 79(3), 2339–2390. Miguel, E. (2021) Evidence on research transparency in economics. The Journal of Economic Perspectives, 35(3), 193–214. Available from: https://doi.org/10.1257/jep.35.3.193 Mueller‐Langer, F., Fecher, B., Harhoff, D. & Wagner, G.G. (2019) Replication studies in economics–how many and which papers are chosen for replication, and why? Research Policy, 48(1), 62–83. Available from: https://doi.org/10.1016/j.respol.2018.07.019 Mueller‐Langer, F. & Watt, R. (2018) How many more cites is a 3,000 open access fee buying you? Empirical evidence from a natural experiment. Economic Inquiry, 56(2), 931–954. Available from: https://doi.org/10.1111/ecin.12545 376 - FINK and MARCUS Nosek, B.A., Alter, G., Banks, G.C., Borsboom, D., Bowman, S.D., Breckler, S.J., et al. (2015) Promoting an open research culture. Science, 348(6242), 1422–1425. Available from: https://doi.org/10.1126/science.aab2374 Nosek, B.A. & Errington, T.M. (2017) Reproducibility in cancer biology: making sense of replications. Elife, 6, e23383. Available from: https://doi.org/10.7554/elife.23383 Nosek, B.A., Hardwicke, T.E., Moshontz, H., Allard, A., Corker, K.S., Dreber, A., et al. (2022) Replicability, robustness, and reproducibility in psychological science. Annual Review of Psychology, 73(1), 719–748. Available from: https://doi.org/10.1146/annurev‐psych‐020821‐ 114157 Open Science Collaboration. (2015) Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. Available from: https://doi.org/10.1126/science.aac4716 Pérignon, C., Akmansoy, O., Hurlin, C., Dreber, A., Holzmeister, F., Huber, J., et al. (2023) Computational reproducibility in finance: evidence from 1,000 tests. HEC Paris Research Paper, (FIN‐2022‐1467). Rainey, C., Roe, H., Wang, Q. & Zhou, H. (2024) Data and code availability in political science publications from 1995 to 2022. SocArXiv Preprint. Reinhart, C.M. & Rogoff, K.S. (2010) Growth in a time of debt. The American Economic Review, 100(2), 573–578. Available from: https://doi. org/10.1257/aer.100.2.573 Rendtel, U., Wagner, G.G. & Frick, J.R. (1995) Eine Strategie zur Kontrolle von Längsschnittgewichtungen in Panelerhebungen: Das Beispiel des Sozio‐Oekonomischen Panels (SOEP). Allgemeines Statistisches Archiv, 79(3), 252–277. Romero, F. (2019) Philosophy of science and the replicability crisis. Philosophy Compass, 14(11), e12633. Available from: https://doi.org/10. 1111/phc3.12633 Science‐Metrix. (2024) Science‐Metrix classification. Available from: https://science‐metrix.com/classification/ [Accessed 8 January 2024]. SOEP Group. (2024) Data access. Available from: https://www.diw.de/en/diw_01.c.601584.en/data_access.html SOEPlit. (2022) SOEPlit database. Available from: https://www.diw.de/en/diw_01.c.789503.en/publications_based_on_soep_data__soeplit. html [Accessed 1 June 2022]. Swanson, N., Christensen, G., Littman, R., Birke, D., Miguel, E., Paluck, E.L., et al. (2020) Research transparency is on the rise in economics. AEA Papers and Proceedings, 110, 61–65. Available from: https://doi.org/10.1257/pandp.20201077 TOP Factor Database. (2023) TOP Factor database. Available from: https://topfactor.org/ [Accessed 30 November 2023]. Vilhuber, L. (2020) Reproducibility and replicability in economics. Harvard Data Science Review, 2(4), 1–39. Available from: https://doi.org/ 10.1162/99608f92.4f6b9e67 Vilhuber, L. (2023) Reproducibility and transparency versus privacy and confidentiality: reflections from a data editor. Journal of Econometrics, 235(2), 2285–2294. Available from: https://doi.org/10.1016/j.jeconom.2023.05.001 Vlaeminck, S. & Herrmann, L.‐K. (2015) Data policies and data archives: a new paradigm for academic publishing in economic sciences? In Schmidt, B. & Dobreva, M. (Eds.) New Avenues for Electronic Publishing in the Age of Infinite Collections and Citizen Science. Proceedings of the 19th International Conference on Electronic Publishing. Amsterdam: IOS Press, pp. 145–155. Wallach, J.D., Boyack, K.W. & Ioannidis, J.P. (2018) Reproducible research practices, transparency, and open access data in the biomedical literature, 2015–2017. PLoS Biology, 16(11), e2006930. Available from: https://doi.org/10.1371/journal.pbio.2006930 Wayback Machine (2024) Internet Archive ‐ Wayback Machine. Available from: https://archive.org/web/ [Accessed 15 January 2024]. Zenk‐Möltgen, W., Akdeniz, E., Katsanidou, A., Naßhoven, V. & Balaban, E. (2018) Factors influencing the data sharing behavior of researchers in sociology and political science. Journal of Documentation, 74(5), 1053–1073. Available from: https://doi.org/10.1108/jd‐09‐ 2017‐0126 SUPPORTING INFORMATION Additional supporting information can be found online in the Supporting Information section at the end of this article. How to cite this article: Fink, L. & Marcus, J. (2025) Replication code availability over time and across fields: evidence from the German Socio‐Economic Panel. Economic Inquiry, 63(2), 357–386. Available from: https://doi. org/10.1111/ecin.13267 FINK and MARCUS - 377 APPENDIX A TABLE A1 Most frequent journals in our sample. Economics Labor Economics 79 Jahrbücher für Nationalökonomie und Statistik 62 Review of Income and Wealth 49 Economics Letters 48 Journal of Population Economics 44 Journal of Economic Behavior & Organization 43 Health Economics 30 German Economic Review 29 Applied Economics 26 Empirical Economics 25 Economic Journal 24 Journal of Health Economics 22 European Economic Review 21 Journal of Human Resources 21 Sociology Kölner Zeitschrift fur Soziologie und Sozialpsychologie 98 European Sociological Review 94 Zeitschrift für Soziologie 67 Soziale Welt 17 Research in Social Stratification and Mobility 16 Psychology Social Indicators Research 76 Journal of Happiness Studies 31 Journal of Personality and Social Psychology 20 Psychology and Aging 20 Journal of Research in Personality 14 Other Social Sciences Journal of Marriage and Family 29 Zeitschrift für Familienforschung 26 Journal of European Social Policy 18 Small Business Economics 18 Demography 17 Health & Other Sciences AStA‐Advances in Statistical Analysis 35 PLOS ONE 24 378 - FINK and MARCUS TABLE A1 (Continued) Gesundheitswesen 19 Social Science & Medicine 19 Journals of Gerontology Series B‐Psychological Sciences 13 Note: The table lists, for each of our five disciplinary categories, the names of the journals with the most articles in our dataset and the corresponding number of articles. Source: Own counting, based on SOEPlit (2022). FIGURE A1 Code availability: Single author versus multiple authors. This figure shows the share of SOEP‐based publications with publicly available replication code, separately for articles written by a single author and articles with multiple authors. The left panel relates to economics, while the right panel relates to all other disciplines. Source: Own calculations, based on SOEPlit (2022) and own research on code availability. FINK and MARCUS - 379 FIGURE A2 Code availability: Language of the article. This figure shows the share of SOEP‐based publications with publicly available replication code, separately for articles in English and in all other languages (mostly German). The left panel relates to economics, while the right panel relates to all other disciplines. Source: Own calculations, based on SOEPlit (2022) and own research on code availability. FIGURE A3 Software used in the replication packages. This figure shows the distribution of statistical software programs used in the available code packages in our sample. The left panel relates to economics, while the right panel relates to all other disciplines. Notethat the shares in each panel do not sum up to 1 as sometimes more than one statistical software is used. Source: Own calculations, based on SOEPlit (2022) and own research on code availability. 380 - FINK and MARCUS