scieee AI-readable full text Open interactive document viewer

Quis custodiet ipsos custodes? Despite evidence to the contrary, the American Economic Review concluded that all was well with its archive

McCullough, Bruce D.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

McCullough, Bruce D. Working Paper Quis custodiet ipsos custodes? Despite evidence to the contrary, the American Economic Review concluded that all was well with its archive Economics Discussion Papers, No. 2017-78 Provided in Cooperation with: Kiel Institute for the World Economy – Leibniz Center for Research on Global Economic Challenges Suggested Citation: McCullough, Bruce D. (2017) : Quis custodiet ipsos custodes? Despite evidence to the contrary, the American Economic Review concluded that all was well with its archive, Economics Discussion Papers, No. 2017-78, Kiel Institute for the World Economy (IfW), Kiel This Version is available at: https://hdl.handle.net/10419/169137 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/ Received September 2, 2017 Accepted as Economics Discussion Paper September 12, 2017 Published September 28, 2017 © Author(s) 2017. Licensed under the Creative Commons License - Attribution 4.0 International (CC BY 4.0) Discussion Paper No. 2017-78 | September 28, 2017 | http://www.economics-ejournal.org/economics/discussionpapers/2017-78 Quis custodiet ipsos custodes?: Despite evidence to the contrary, the American Economic Review concluded that all was well with its archive B. D. McCullough Abstract In 2011, the Annual Report of the Editor of the American Economic Review reported that the journal’s data-code archive was functioning well, and made no changes in the archive rules. This was based on an audit of the archive that the editor has commissioned. The audit was performed by a graduate student who apparently had no experience with archives, and the audit concluded that all was largely well with the archive. In point of fact, all was not well with the archive: the archive did not support the publication of reproducible research. The rules for the archive should have been changed and were not; thus the American Economic Review continued to publish articles that were not reproducible. The cause of reproducible research was set back many years. (Published in Special Issue The practice of replication) JEL B40 Keywords Replication; reproducible research Authors B. D. McCullough, Department of Decision Sciences & MIS, Drexel University, Philadelphia, PA, USA, [email protected] Citation B. D. McCullough (2017). Quis custodiet ipsos custodes?: Despite evidence to the contrary, the American Economic Review concluded that all was well with its archive. Economics Discussion Papers, No 2017-78, Kiel Institute for the World Economy. http://www.economics-ejournal.org/economics/discussionpapers/2017-78 1 A general discussion of principles about how one should do a replication Before we can discuss replication, we need to define it. The word is used in many different and sometimes conflicting ways, both within and across disciplines. The title of a recent news article from Nature (Baker 2016) describes this problem accurately: “Muddled meanings hamper efforts to fix reproducibility crisis”. This confusion harms research and retards progress. Clemens (2015) performed a yeoman’s job in classifying forty one (!) different uses of the word “replication” within economics. Clearly there is a need for a standard taxonomy. For purposes of maintaining an archive, the concept of “narrow replication” (a.k.a “reproducibility”) suffices: The data and code in the archive reproduce the published results. As an example of this need for clear thinking and precise definitions when talking about replication, consider the recent paper by Chang and Li (2017) that has received much attention. In their abstract they write (p. 2): “We successfully replicate the key qualitative result of 22 of 67 papers (33%) without contacting the authors. Excluding the 6 papers that use confidential data and the 2 papers that use software we do not possess, we replicate 29 of 59 papers (49%) with assistance from the authors.” Notice the phrase “key qualitative results”. What Chang and Li think they are doing is confirming key qualitative results, not replicating them. To see this more clearly, consider the following quote from their paper (p. 7): “We define a successful replication as when... [f]or example, if the paper estimates a fiscal multiplier for GDP of 2.0, then any multiplier greater than 1.0 would produce the same qualitative result (i.e., there is a positive multiplier effect that government spending is not merely a transfer or crowding out private investment).” However, they are not even confirming the published results because they are using the same data and code as the original author. On its face, Chang and Li’s criterion for replicability is utter nonsense. Think about it: Using the same data and same code, the original author gets 2.0 while Chang and Li get 1.0 and they think this is a successful “replication”. The number “1.0” 2 most certainly does not replicate or reproduce the number “2.0” when using the same data and code! Chang and Li (p. 2) write, “Using the authorprovided data and code replication files, we are able to replicate 22 of 67 papers (33%) independently of the authors by following the instructions in the author-provided readme files.” If Chang and Li used the same data and same code to get 1.0 when the original paper shows 2.0, then Chang and Li prove that the paper is not reproducible because the authors could not provide data and code that reproduce the published result. We can be quite confident that Chang and Li did not actually reproduce the results of 22 papers, and the actual number is probably much lower than 22. This clear distinction between reproducible and replicable is important. The recent article by Camerer et al (2016) clearly involved replication: they ran the same experiments on different subjects. In its earliest use in the physical sciences, to “replicate” an experiment meant to perform a second experiment in conditions similar to a first experiment, with the intent of confirming or disproving the result of the first experiment. With the advent of widespread computing, in about 1990 the geophysicist Claerbout coined the term “reproducible research” to refer to reproducing published results, typically using the same data and code but also allowing the coding to be done in a different language, as long as the published results are reproduced. This hair may be further split between reproducible and repeatable. Imagine taking someone’s data and code, running it on a different computer, and getting a different answer. One would say that the results might be repeatable, but they are not reproducible (Easterbrook 2014). This is what Chang and Li find when the paper’s result is 2.0 and they get 1.0 by running the same data and code: The result is repeatable, but not reproducible. It is also important to note that reproducibility does not imply correctness. For example, Donohue and Levitt’s (2001) article on abortion may have been reproducible, but it was not correct. In the course of reproducing the article using the author’s own data and code, Foote and Goetz (2008) discovered a coding error that invalidated the article’s results. The purpose of a journal’s data/code archive is to ensure that the journal’s published results are reproducible. This is a miminal standard that is easy to understand: either the results of an article can be reproduced or they cannot: it is a binary decision. To argue that “some of the results are replicable” or “the important results are replicable” is to admit that the article’s results are not reproducible. We can quibble over how many significant digits consitute reproducibility, but in the end the decision is binary. (For 3 linear procedures with moderately-sized datasets, there should be ten digit agreement, for nonlinear procedures there may be as few as four or five digits of agreement. See McCullough and Vinod (1999) for details.) Now that we have definitions established, we can discuss procedure. For computational research, it is very easy. Put the data and code in the same folder, and run the code. Barring minor accommodations such as different operating systems (e..g, the author uses Stata in Windows while the replicator uses Stata in Linux), if it fails to execute (the code doesn’t run), the person who prepared the data and code has failed to provide evidence that the article’s results are reproducible, and the article should be labeled as such. It is not the duty of the would-be replicator to spend valuable time trying to make the data and code work. To require this is to permit the original author to engage in cost-shifting; he spends less time preparing his replication files, and the replicator spends her time trying to make sense of data and code that doesn’t work. If the data and code run but do not reproduce all the published results, she does not spend her valuable time trying to fix the data and code so that they do reproduce the published results. Even if she succeeds in this effort, it remains the case that the data and code that are in the archive do not reproduce the published results. She should inform the editor that the article has failed to replicate, how it has failed to replicate, and let the editor notify the original author. If he does not swiftly respond with data and code that reproduce all the published results, the article should be flagged as not replicable. In general, no explanation of the extent of the non-replicability should be given, for this invites sloppy research. (Of course, if he used version 1.0 of the software and she used version 1.1, this is not a failure to reproduce, since the same algorithm was not applied to the data.) If the article is not computational in nature and perhaps requires human judgment for classification, then the article should enumerate protocols so that another person would arrive at the same classification. This was a part of the Hoxby/Rothstein debate. Hoxby created her controversial variable on the number of streams by looking at a map and counting “all streams that were at least 3.5 miles long and of a certain width on the map” (Hoxby 2000:1222), but she provided no further details. What was this “certain width”? Was it 1mm or 5mm in width? This lack of detail all but insured that no one else would be able to reproduce her work. As Rothstein wrote (Rothstein 2007:2033-34): “Where Hoxby reports five larger streams in Fort Lauderdale, I counted 12, and a research assistant working independently counted 15.” 4 The bottom line is that other researchers, working independently, could not get the same result she did. Her paper was not reproducible. 2 An explanation of why the candidate paper was selected for replication Recently the American Economic Review posted an advertisement for a “Data Editor”. The advertisement cited problems with its archive, the most notable being: “Posted code often does not run or does not actually replicate the results.” This contradicts the assertions made years ago by then-editor of the AER, Robert Moffitt. Moffitt had commissioned a graduate student, Philip Glandon, to conduct an audit of the AER archive. Glandon’s (2011) report, which is the subject of the present paper, was the basis on which Moffitt (2011:687) assured readers of the AER that “The vast majority of authors complied with the intent of the policy but a small fraction submitted materials that were either incomplete or that would have made replication difficult.” In response to Glandon’s report, Moffitt made no changes to the AER archive policy, implying that all was well with the archive. The fact of the matter is that all was not well with the archive, and Glandon’s report failed to make this clear; Moffitt did not make any changes to the AER archive policy and the AER continued to publish nonreproducible research. If Glandon had written his report competently, the AER would have been forced to take action to fix the archive years ago. Glandon’s report and Moffitt’s uncritical acceptance of it set the goal of reproducible economic research back by several years. The editor’s job is to ensure that the papers published in his journal can be relied upon. The editor should have known that the archive was failing. Many authors have cited Glandon’s appendix as a basis for asserting that the AER archive is fulfilling its function of ensuring the published results are reproducible. Here are some examples: 1. “[Glandon] replicated a selected sample of nine papers only from the AER.” (Chang and Li 2017) 2. “The AER conducted a self-review and found relatively good, though still incomplete, compliance with its data sharing policy (Glandon 2010).” (Christensen and Miguel 2018) 5 3. “Roughly 80% of the submissions satisfied the spirit of the AER’s data availability policy, which is to make replication and robustness studies possible independently of the author(s). The replicated results generally agreed with the published results.” (Breure and Hoogerwerf 2011) 4. “For instance, one in five articles examined from the 2006-2008 period in AER did not fully satisfy the requirement that results be reproducible from submitted data and code, leading the journal to require review by contracted grad students (Glandon 2010).” (Nylan 2015) 5. “The project on which Glandon reports covered replication of 39 articles published between 2006 and 2008 in the AER; about 80% of the submissions satisfied the spirit of the data availability policy.” (Karolyi 2011) Yet, Glandon’s (2010) report supports none of the above characterizations. In fact, Glandon’s report did not have a single successful replication! Dewald, Thursby and Anderson (1986) called into question the replicability of published economic research. They considered possible solutions to the problem, in particular they dismissed the idea of a “replication policy” that requires authors to supply data and code to would-be replicators after publication (primarily due to agency problems – once the article’s published, the author has no incentive to spend time organizing replication files). They concluded that only a mandatory data/code archive might solve the problem, provided that the data and code were deposited before publication. The above notwitstanding, then-editor of the American Economic Review Orley Ashenfelter instituted a “replication policy”. McCullough and Vinod (2003) confirmed that, as predicted, the AER replication policy did not produce replicable articles. In response, then-editor Bernanke (2003) adopted a data-code archive. Before he could implement policies to ensure that the archive would result in the AER publishing reproducible research, Bernanke left academia and resigned his post as AER editor. In his first Annual Report, Bernanke’s successor, Robert Moffit, introduced the following boilerplate that is found almost verbatim in the Editor’s Reports through the end of his term (even through that of his successor, Goldberg): In 2004, the Review began to require that authors of accepted papers who employ data in econometric exercises, simulation models, or experiments agree to post their data and programs on 6 the journal Web site unless an exemption for proprietary data is requested and granted. The policy was strengthened in 2005 with more systematic enforcement and with greater attention to searching for alternative means of data access for papers requesting exemptions. Table 8 shows the number of papers in each of the 2009 issues containing data analysis, the number of exemptions granted, the number of authors who complied on the first round (defined as supplying data after receiving the acceptance letter detailing the requirement), and the number of authors who complied after a later reminder. Full compliance was achieved for all issues. This boilerplate is the only time Moffitt mentions the archive, and no one reading the “full compliance” sentence would have any reason to think that the data and code in the archive was doing anything other than reproducing the published results. There is no other mention of the archive until his final report in 2011, in which he wrote (pp. 686-7): In the summer of 2008, the AER conducted an exercise to check the submitted files of a random set of papers to check for compliance with the policy, which requires the submission of both programs and data and an explanation of how to use them. A report prepared by Philip Glandon, included as an Appendix to this report, describes the project and the results. The vast majority of authors complied with the intent of the policy but a small fraction submitted materials that were either incomplete or that would have made replication difficult....Mr. Glandon’s report contains additional details on the project and recommendations for strengthening the AERs data posting policy. Naturally, if there had been anything seriously wrong with the archive, the Editor would have taken steps to address the problem. This lack of action implies that no action was needed, that all was more or less well with the archive. To the casual reader, Moffitt’s remark suggests that all was well with the AER archive, save perhaps the occasional glitch. The casual reader may well wonder whether the recommendations to strengthen the archive were even necessary. After all, if the archive policy needed to be strengthened, surely 7 the editor would do it. As shown above, many authors followed Moffit’s lead and reported that all was well with the AER archive. Economists familiar with replication literature were more circumspect in considering Moffit’s blandishments with respect to the archive. Dewald and Anderson (2014:208) wrote, “In 2004, AER editor Ben Bernanke adopted a mandatory data and program code archive. Compliance has been excellent, at least according to the annual reports of the editor.” [emphasis added] On the other hand, a critical reader might wonder why Moffit chose to use the word “intent”. He might then consider the difference between: (1) The vast majority of authors complied with the intent of the policy and (2) the vast majority of authors complied with the policy. A reader who merely glanced at said Appendix might not have reason to question Moffit’s assertion. Someone who read the appendix carefully, especially someone who knows something about data-code archives, would discover that the two sentences are orthogonal. The Appendix in question is called “Report on the American Economic Review Data Availability Compliance Project,” and it offers as its primary piece of “evidence” its Table 1, reproduced below. 2006 2007 Mar-08 Total Articles published 98 100 22 220 Articles subject to data policy 61 63 11 135 Articles investigated 13 24 2 39 With “readme” file 12 23 1 36 (90%) (96%) (50%) (92%) With complete submission 7 12 1 20 (54%) (50%) (50%) (51%) With proprietary data instructions 1 10 0 11 (8%) (24%) (0%) (28%) Articles investigated believed replicable 8 22 1 31 without contacting the author(s) (62%) (92%) (50%) (79%) Table 1: Glandon’s Table I: Data and Code Submissions by Year of Publication The reader’s attention is directed to the end of the last line: 79%. This is the only number that might bear on the editor’s claim that the “vast majority” of articles are compliant. Yet, even if the number is correct, it 8 Critical Finance Review (to appear) Christensen, Garret and Edward Miguel (2018), “Transparency, Reproducibility, and the Credibility of Economics Research,” Journal of Economic Literature, forthcoming Clemens, Michael A. (2015), “The Meaning of Failed Replications: A Review and Proposal,” Journal of Economic Surveys 31(1), 326-342 Dewald, William G., Jerry G. Thursby and Richard G. Anderson (1986), “Replication in Empirical Economics: The Journal of Money, Credit and Banking Project,” The American Economic Review 76(4), 587-603 Dewald, William G. and Richard G. Anderson (2014), “Replication and Reflection: A Decade at the Jouranl of Money, Credit and Banking,” Chapter 13 in Secrets of Economics Editors, Michael Szenberg and Lall Ramrattan, eds., Cambridge, MA: MIT Press Donohue, John and Steven Levitt (2001), “The Impact of Legalized Abortion on Crime,” Quarterly Journal of Economics 116, 379-420 Foote, Christopher and Christopher Goetz (2008), “The Impact of Legalized Abortion on Crime: Comment,” Quarterly Journal of Economics 123(1), 407-423 Easterbrook, Steve M. (2014), “Open Code for Open Science?” Nature Geoscience 7, 779-781 Glandon, Philip J. (2011), “Appendix to the Report of the Editor: Report on the American Economic Review Data Availability Compliance Project,” American Economic Review 101(3), 696-699 Hoxby, Caroline (2000), “Does competition among public schools benefit students and taxpayers?” American Economic Review 90(5), 1209-1238 Karolyi, G. Andrews (2011), “The Ultimate Irrelevance Proposition in Finance?” The Financial Review 46(4), 485-512 McCullough, B. D. and H. D. Vinod (2003), “Verifying the Solution from a Nonlinear Solver: A Case Study,” American Economic Review 93(3), 873-892 McCullough, B. D., Kerry Anne McGeary, and Teresa D. Harrison (2006), “Lessons from The Journal of Money, Credit and Banking Archive,” Journal of Money Credit and Banking 38(4), 10931107 15 McCullough, B. D., and H. D. Vinod (1999), “The Numerical Reliability of Econometric Software,” Journal of Economic Literature 37(2), 633-665 McCullough, B. D., and H. D. Vinod (2003), “Verifying the Solution from a Nonlinear Solver: A Case Study,” American Economic Review 93(3), 873-892 Moffitt, Robert A. (2011), “Report of the Editor,” American Economic Review 101(3), 684-689 Nylan, Brendan (2015), “Increasing the Credibility of Political Sciences Research: A Proposal for Journal Reforms,” PS: Political Science & Politics 48(S1), 78-83 16 Please note: You are most sincerely encouraged to participate in the open assessment of this discussion paper. You can do so by either recommending the paper or by posting your comments. Please go to: http://www.economics-ejournal.org/economics/discussionpapers/2017-78 The Editor © Author(s) 2017. Licensed under the Creative Commons License - Attribution 4.0 International (CC BY 4.0).