scieee AI-readable full text Open interactive document viewer

Does single-blind review encourage or discourage p-hacking?

Naguib, Costanza

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Naguib, Costanza Working Paper Does single-blind review encourage or discourage phacking? Discussion Papers, No. 25-04 Provided in Cooperation with: Department of Economics, University of Bern Suggested Citation: Naguib, Costanza (2025) : Does single-blind review encourage or discourage phacking?, Discussion Papers, No. 25-04, University of Bern, Department of Economics, Bern This Version is available at: https://hdl.handle.net/10419/324321 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Faculty of Business, Economics and Social Sciences Department of Economics Does single-blind review encourage or discourage p-hacking? Costanza Naguib 25-04 July, 2025 Schanzeneckstrasse 1 CH-3012 Bern, Switzerland http://www.vwi.unibe.ch DISCUSSION PAPERS Does single-blind review encourage or discourage p-hacking? Costanza Naguib∗ Abstract In 2011, the American Economic Association (AEA) changed its peer review policy for all their journals, shifting from a double-blind process to a single-blind peer-review process. Under this new system, referees became aware of the authors’ identities. In this paper, I explore whether this policy change influenced the prevalence of p-hacking in published papers at The American Economic Review. JEL codes: A11, A14, C13 Keywords: p-hacking, single-blind review, double-blind review, difference-in-difference 1 Introduction 1There is an ongoing debate on whether single-blind or double-blind review is more suitable for the peer review of scientific articles in economics. In a single-blind review, the author remains unaware of the referee’s identity, whereas in a double-blind review, the referee is also not informed about the author’s identity. Proponents of double-blind review argue that it helps mitigate biases from referees. They highlight two primary concerns. First, referees might unfairly reject papers from lesser-known authors or those affiliated with lower-ranked institutions, irrespective of the paper’s quality. Second, if gender or ethnic discrimination is present, papers authored by women or individuals with foreign-sounding names might face biased reviews (Blank (1991))2. Supporters of single-blind reviewing usually present the following three arguments. First, they claim that referees can often identify the author of a paper through its content or citations, so double-blind systems are seldom truly anonymous.3Second, they argue ∗University of Bern 1I thank Stefan Egli, Reto Horst and Thibaud Laurent for the excellent research assistance. 2It is however also possible that referees are stricter on more prolific authors, as they wish to leave space also to newcomers and to prevent already established scholars from publishing marginal papers, as suggested by Card and DellaVigna (2020). 3According to Blank (1991), referees were able to identify the authors of around 50% of papers at the end of the 80s. Despite the improvement in search engines, this share seems to have remained fairly stable up to now (see Cressey (2014), Hill and Provost (2003)). 1 that knowing the author’s name and institution provides valuable context that can affect how the paper is read and evaluated. Third, editors often note that double-blind reviewing involves additional administrative effort, requiring more meticulous procedures in the editorial office. In this paper, I aim to assess whether the single-blind review system encourages or discourages p-hacking. P-hacking refers to practices, such as specification searching, that researchers may use to obtain more favorable p-values. This often occurs in response to the challenges of publishing null results, as null findings are often perceived to be of lower quality (Imbens (2021), Chopra et al. (2024)). Such practices increase the prevalence of false positives in the literature, skewing the research presented to policymakers. Under a single-blind system, young researchers and those from less prestigious institutions may feel pressured to impress referees with eye-catching statistically significant results. If this is the case, it would provide yet another argument against adopting a single-blind system. In March 2011, the American Economic Association (AEA) decided to switch from a double-blind to a single-blind peer review standard.4. At the time, it was the only one among the so-called top-5 journals to use a double-blind system, whereas all the others were using single-blind reviews. This contrast provides a framework for studying a quasi-natural experiment. I aim to analyze whether the policy change by The American Economic Review was followed by a change in p-hacking practices5. It is worth noting that, as mentioned above, in the age of Google, reviewers are able to identify the authors of a paper from its text or citations with approximately 50% accuracy (Cressey (2014), Hill and Provost (2003)). This means that I actually estimate the impact on p-hacking of an increase in the probability of author identification by the referees from around 50% to 100%. 4The statement read: ”Upon a joint recommendation of the editors of the American Economic Review and the four American Economic Journals, the Executive Committee has voted to drop the ’double-blind’ refereeing process for all journals of the American Economic Association. The change to ’single-blind’ refereeing will become effective on July 1, 2011. Easy access to search engines increasingly limits the effectiveness of the double-blind process in maintaining anonymity. Further, it increases the administrative cost of the journals and makes it harder for referees to identify an author’s potential conflicts of interest arising, for example, from consulting.” Source: https://crookedtimber.org/2011/06/05/should-theamerican-economic-review-drop-double-anonymous-review/ and Goldberg (2012). In the present paper I only consider the American Economic Review and not the four American Economic Journals, as the latter only started to publish issues in 2009, i.e. two years before the change in the peer-review standard only. 5Another change of policy in the period under scrutiny is the implementation of a strict page limit on submissions in 2008 by AER alone among the top-5. While a strict limit on page count may also encourage p-hacking practices, Card and DellaVigna (2014) find that this intervention had essentially no impact on the length of published papers, and authors mostly adjusted their submission by means of purely aesthetic formatting changes in order to meet the page limit. Hence, it is likely that this intervention did not influence the extent of p-hacking. 2 First, I apply a difference-in-difference approach to determine whether a double-blind review policy is associated with a lower proportion of statistically significant test statistics being published than a single-blind one. The treatment is having a single-blind policy. Since no other Top-5 journal was adopting a double-blind policy in 2011, there is no control group, but only an always treated group and a switcher group. I am hence in the framework of ”time-reverse difference-in-difference”. As shown by Kim and Lee (2018), the estimation procedure is essentially the same as in the standard DiD framework. Second, I investigate whether a series of statistical tests can detect p-hacking under the single-blind regime, respectively under the double-blind regime in papers published at The American Economic Review. I am the first to evaluate the impact of different peer review standards on the extent of p-hacking. I find that, under a double-blind review system, authors from top institutions tend to report a higher proportion of statistically significant results, whereas authors from non-top institutions, on average, report a lower share of statistically significant findings than they do under a single-blind standard. These findings are robust to a range of sensitivity analyses, including procedures such as de-rounding and weighting. This study contributes to the growing body of research on p-hacking, building on seminal work by Brodeur et al. (2016), who documented p-hacking and publication bias in three top economics journals (The American Economic Review (AER), The Quarterly Journal of Economics (QJE), and The Journal of Political Economy (JPE)). Brodeur et al. (2020) later demonstrated that these issues vary by estimation method, with Kranz and P¨utz (2022) noting that such findings may be relevantly influenced by rounding errors. Subsequent research by Brodeur et al. (2024a, 2024b) showed that neither data availability and replication policies nor pre-registration and pre-analysis plans significantly reduce phacking. Blanco-Perez and Brodeur (2020) assessed the impact of a 2015 editorial statement from eight health economics journals encouraging the publication of statistically insignificant but economically relevant results. This intervention effectively reduced p-hacking and publication bias, lowering the proportion of tests rejecting the null hypothesis by approximately 18 percentage points. Similarly, Naguib (2024) studies the impact of the omission of significance asterisks implemented by the AEA journals in mid-2016 on the extent of p-hacking and publication bias, finding essentially no impact of the policy. Finally, McCloskey and Michaillat (2024) derive critical values for hypothesis testing that are robust to p-hacking. Studies evaluating the costs and benefits of single-blind vs double-blind peer review 3 of scientific articles in economics are scarce. The few existing papers provide suggestive evidence of editorial favoritism in the single-blind review process. Blank (1991) analyzes submissions to The American Economic Review from 1987 to 1989. In this period and experiment took place, where some papers were randomly assigned to single-blind and others to double-blind review. Blank (1991) finds that the double-blind review process results in lower acceptance rates and more critical referee comments for authors affiliated with mid-ranking top universities (ranks 6–50). However, she does not find any specific adverse effect of single-blind review on female authors. This is further confirmed by Carlsson et al. (2012), studying double and single-blind acceptance decisions for a Swedish conference held in 20086. Nevertheless, Laband and Piette (1994) analyzed more than 1,000 articles published in 28 leading economics journals in 1984 and discovered that articles from journals using double-blind review received more citations over a five-year period, indicating higher quality compared to those published in journals using single-blind review. 1.1 Potential mechanisms Under a single-blind peer review system, authors affiliated with less prestigious institutions might feel increased pressure to present statistically significant results to enhance their chances of publication. Consequently, they may be more inclined to engage in p-hacking practices compared to when they work under a double-blind review system. Conversely, in a double-blind review system, authors from reputable institutions may experience heightened pressure to present particularly compelling results, as their identities are anonymized and cannot influence reviewers. Thus, the shift from single-blind to double-blind review creates opposing incentives for different groups of authors. It is not clear, a priori, what the net effect of this transition on the prevalence of p-hacking would be. Notably, Brodeur et al. (2020) find that p-hacking practices are not relevantly influenced by authors’ experience levels or their institutions’ rankings. Regarding publication bias, this term refers to the preference exhibited by editors and referees for statistically significant results. Given that editors always know the authors’ identities, I anticipate no significant change in their attitudes towards null results between single-blind and double-blind review systems. However, referees may display greater bias against null findings submitted by authors from less prestigious institutions under a single-blind system. Conversely, referees may also 6However, Hengel (2022) finds that, under a single-blind system, women are held to higher writing standards than men. 4 demonstrate more leniency towards null results when they originate from well-established authors at top institutions in the same system. In summary, switching from a double-blind to a single-blind review policy is expected to: •Reduce p-hacking practices and publication bias among papers authored by prominent researchers from reputable institutions. •Increase p-hacking practices and publication bias among papers authored by lessknown researchers from lower-ranked institutions. The overall net impact of these two contrasting effects on p-hacking and publication bias remains unclear a priori. 2 Data Description I exploit the quasi-natural experiment of The American Economic Review (AER) passing from a double-blind to a single-blind peer review standard for their papers in July 20117. I consider data for the period 2005-2015. The analysis period stops in 2015, because in mid-2016 the AEA introduced another policy that may potentially impact the extent of phacking, e.g. the omission of significance stars from the regression tables. This intervention is described in detail in Naguib (2024). I compare the extent of p-hacking in the AER with that in comparable top-5 journals in economics such as The Quarterly Journal of Economics (QJE) and The Journal of Political Economy (JPE), which have both been adopting a single-blind review standard for the full period of analysis8. In my dataset, I exclude corrigenda, comments and replies to research papers from the analysis. I further exclude papers that do not include any estimated coefficient. Following Brodeur et al. (2020), I only collect estimates from results tables and only for the coefficients of interest, or main results, excluding regression controls, constant terms, balance and robustness checks, heterogeneity of effects, and placebo tests. I however collect coefficients drawn from multiple specifications of the same hypothesis. Moreover, I collect the estimated coefficients of interaction terms only if such terms are the variable of interest, for example in the case of the interaction between the post-treatment period and the treated dummy in a difference-in-difference setup. If the main findings of a paper are expressed by means of a Figure, e.g. impulse response functions, then I drop the paper 7The decision was announced in March 2011, and become effective on 1st July 2011. 8As mentioned above, we are hence in the framework of an ”reverse-time difference-in-difference setup”, where one group was treated (i.e. adopting a single-blind standard) for the whole period of analysis and the other group started being treated only at a certain point in time, see Kim and Lee (2018). 5 from the collection. I collect all reported decimal places. If more than one standard error per estimated coefficient is reported (e.g. obtained with different methods of clustering), I only collect the first one. There is notable overlap between my dataset and the one collected by Brodeur et al. (2016), in particular for the years 2005-2011. For those years, I only collect the main results, whereas they also collect robustness checks and similar additional results. For the years 2012-2016, I collect estimated coefficients from all the articles published in the three journals under study, whereas Brodeur et al. (2024a) only collects coefficients from a random sample of articles. In Figure 5 in the Appendix I show the distribution of zstatistics, respectively in my sample and in the sample collected by Brodeur et al. (2024a). Both histograms clearly exhibit what Brodeur et al. (2016) calls a ”two-humped camel shape”, which suggests the presence of p-hacking and/or publication bias. Further, differently from Brodeur et al. (2020), I do not restrict the analysis to articles that use one of a pre-defined set of estimation methods (DID, RDD, RTC, IV), but I collect results from all methods that produce estimated coefficients and standard errors (or t-statistics, or p-values). Notably, this means that I also include OLS estimates. However, if OLS estimates are only used to present correlations or as a sort of descriptive statistics, and hence they are not in the Results Section of the paper, but rather in the Data description, then I do not collect them. In case of IV estimations, following Brodeur et al. (2020), I only collect the coefficient(s) of the instrumented variable(s) in the second stage. Data have been coded independently by at least two of the following: the author and three research assistants. We discussed and clarified discordant cases. Finally, following Brodeur et al. (2020), since all of the test statistics in the sample relate to two-tailed tests and degrees of freedom are not always reported, I treat coefficient and standard error ratios as if they follow an asymptotically standard normal distribution. When articles report tstatistics or p-values, I transform them into equivalent z-statistics. This can be a rough approximation of significance for two reasons. First, the effective number of degrees of freedom may be modest, especially in case of results obtained with cluster-robust standard errors and a limited number of clusters. Second, different journals may adopt different rounding conventions for their reported results. If these conventions differ across journals and/or across time, this may not cancel out in my diff-in-diff setting and hence bias the results. 6 Figure 1: Percentage of tests significant, respectively at the 1% level (upper panel), at the 5% level (middle panel), and at the 10% level (bottom panel), by year of publication. 7 as in Elliot et al. (2022). Further, following Elliott et al. (2022), I interpret any p-value in Table 2 lower than 0.1 as evidence of the presence of p-hacking and publication bias. Name of test Bin. Disc. CS1 CS2B LCM N. Obs. N. articles Panel A: AER, Double-Blind Review (2005-2011) Full sample 0.201 0.994 0.647 0.137 0.959 4483 222 Theory model 0.209 0.128 0.027 0.002 1.000 2557 129 No theory model 0.434 0.570 0.050 0.024 0.994 1926 93 Single author 0.412 0.594 0.000 0.000 1.000 975 51 Not single author 0.238 0.590 0.238 0.131 0.968 3508 171 Top institution 0.314 0.890 0.000 0.000 1.000 1924 92 Not top institution 0.292 0.082 0.474 0.607 1.000 2559 130 Panel B: AER, Single-Blind Review (2012-2015) Full sample 0.939 0.748 0.102 0.070 0.568 4785 200 Theory model 0.959 0.605 0.161 0.013 0.750 3144 131 No theory model 0.678 0.608 0.000 0.000 1.000 1641 74 Single author 0.685 0.528 0.004 0.000 1.000 754 32 Not single author 0.943 0.589 0.006 0.000 0.591 4031 168 Top institution 0.849 0.207 0.026 0.003 0.947 1886 74 Not top institution 0.900 0.172 0.243 0.040 0.957 2899 126 Table 2: P-values by Subsample: AER, Double-Blind vs. Single-Blind Review From Table 2 I deduce that, under double-blind review (Panel A, 2005–2011), I can detect p-hacking in the subsamples of papers with and without a theory model, as well as single-authored and with authors coming from a top institutions (in all these cases, only hte CS1 and CS2B tests are able to reject the null of no p-hacking and no publication bias). No test rejects the null in the subsample of papers with multiple authors, as well as in the overall sample, and only the discontinuity test rejects the null among papers whose authors are not affiliated with top institutions. In Panel B (AER under a single-blind review standard), the null hypothesis is rejected by at least one test in all the subsamples as well as in the overall full sample. Similarly to panel A, only the CS1 and CS2B tests are able to detect p-hacking. It seems that it is easier to detect p-hacking under the single-blind regime. However, a word of caution is necessary, as some forms of selective reporting might not be detectable by the tests presented here and difference in significance do not correspond to significant differences (see Gelman and Stern (2006)). In Appendix D I present the results of the tests reported in Table 2 when the two above-mentioned methods for derounding are applied. My results are broadly consistent to derounding. In Table 9 I apply the derounding method used in Brodeur et al. (2016) 14 and I find that, under double-blind review, I can detect p-hacking in the subsamples of papers without a theory model, both single-authored and with multiple authors, and with authors coming from top institutions. On the other hand, under single-blind review standard I can detect p-hacking in the subsamples of papers with and without a theory model, single-authored and with authors coming from top institutions. In Table 11, where I apply the derounding method proposed by Kranz and P¨utz (2022) I find evidence of p-hacking/publication bias under double-blind review standard in the full sample of papers, as well as in the subsamples of both single-authored and multi-authored papers, of papers without a theory model and with authors coming from top institutions. Under a single-blind review standard I am able to detect the presence of p-hacking in all the six subsamples considered, but not in the full sample. 5 Conclusion This study explores the impact of single-blind versus double-blind peer review systems on the prevalence of p-hacking in economics journals, focusing on The American Economic Review’s (AER) transition to a single-blind review policy in 2011. Using a difference-indifferences approach and a series of statistical tests, the research provided suggestive hints that the change in the review system did not relevantly influence overall the extent of p-hacking in published papers at AER. However, results notably differ across subgroups. In particular, the hypothesis according to which authors coming from top institutions engage more in p-hacking-type practices and authors coming from non top-institutions engage less in them under double-blind review and vice versa under a single-blind review system is confirmed empirically. The results have broader implications for the ongoing debate between single-blind and double-blind review systems. Although proponents of double-blind reviews argue that they reduce biases and promote fairness, it might have the unintended consequence of increasing incentives for p-hacking for some groups of authors, in particular those who are single authors and come from top institutions. At the same time, a double-blind review system appears to reduce incentives for p-hacking-type practices across authors who are not affiliated with top universities and who coauthor papers with others. References 1. Blanco-Perez, C., & Brodeur, A. (2020). Publication bias and editorial statement on negative findings. The Economic Journal, 130(629), 1226-1247. 15 2. Blank, R. M. (1991). The effects of double-blind versus single-blind reviewing: Experimental evidence from the American Economic Review. The American Economic Review, 1041-1067. 3. Brodeur, A., Cook, N., & Neisser, C. (2024a). P-hacking, data type and data-sharing policy. The Economic Journal, 134(659), 985-1018. 4. Brodeur, A., Cook, N. M., Hartley, J. S., & Heyes, A. (2024b). Do Preregistration and Preanalysis Plans Reduce p-Hacking and Publication Bias? Evidence from 15,992 Test Statistics and Suggestions for Improvement. Journal of Political Economy Microeconomics, 2(3), 527-561. 5. Brodeur, A., Carrell, S., Figlio, D., & Lusher, L. (2023). Unpacking p-hacking and publication bias. American Economic Review, 113(11), 2974-3002. 6. Brodeur, A., Cook, N., & Heyes, A. (2020). Methods matter: P-hacking and publication bias in causal analysis in economics. American Economic Review, 110(11), 3634-3660. 7. Brodeur, A., L´e, M., Sangnier, M., & Zylberberg, Y. (2016). Star wars: The empirics strike back. American Economic Journal: Applied Economics, 8(1), 1-32. 8. Card, D., & DellaVigna, S. (2020). What do editors maximize? Evidence from four economics journals. Review of Economics and Statistics, 102(1), 195-217. 9. Card, D., & DellaVigna, S. (2014). Page limits on economics articles: Evidence from two journals. Journal of Economic Perspectives, 28(3), 149-168. 10. Carlsson, F., L¨ofgren, ˚ A., & Sterner, T. (2012). Discrimination in scientific review: A natural field experiment on blind versus non-blind reviews. The Scandinavian Journal of Economics, 114(2), 500-519. 11. Chopra, F., Haaland, I., Roth, C., & Stegmann, A. (2024). The null result penalty. The Economic Journal, 134(657), 193-219. 12. Cressey, D. (2014, July 14). Journals weigh up double-blind peer review. Nature. Retrieved from https://www.nature.com/news/journals-weigh-up-double-blind-peerreview-1.15564 13. Elliott, G., Kudrin, N., & W¨uthrich, K. (2022a). The Power of Tests for Detecting p-Hacking. arXiv preprint arXiv:2205.07950. 16 14. Elliott, G., Kudrin, N., & W¨uthrich, K. (2022b). Detecting p-hacking. Econometrica, 90(2), 887-906. 15. Gelman, A., & Stern, H. (2006). The difference between “significant” and “not significant” is not itself statistically significant. The American Statistician, 60(4), 328-331. 16. Goldberg, P. (2012): “Report of the Editor: American Economic Review,” American Economic Review: Papers & Proceedings, 102, 653–665. 17. Hadavand, A., Hamermesh, D. S., & Wilson, W. W. (2024). Publishing economics: How slow? Why slow? Is slow productive? How to fix slow?. Journal of Economic Literature, 62(1), 269-293. 18. Hengel, E. (2022). Publishing while female: Are women held to higher standards? Evidence from peer review. The Economic Journal, 132(648), 2951-2991. 19. Hill, S., & Provost, F. (2003). The myth of the double-blind review? Author identification using only citations. ACM SIGKDD Explorations Newsletter, 5, 179-184. 20. Imbens, G. W. (2021). Statistical significance, p-values, and the reporting of uncertainty. Journal of Economic Perspectives, 35(3), 157-174. 21. Kim, K., & Lee, M. J. (2019). Difference in differences in reverse. Empirical Economics, 57, 705-725. 22. Kranz, S., & P¨utz, P. (2022). Methods matter: P-hacking and publication bias in causal analysis in economics: Comment. American Economic Review, 112(9), 3124-3136. 23. Laband, D. N., & Piette, M. J. (1994). Does the” blindness” of peer review influence manuscript selection efficiency?. Southern Economic Journal, 896-906. 24. Erzo F.P. Luttmer, (2024). Report of the Editor American Economic Review. AEA Papers and Proceedings 2024, 114: 734–750. 25. McCloskey, A., & Michaillat, P. (2024). Critical values robust to p-hacking. Review of Economics and Statistics, 1-35. 26. Naguib, C. (2024). P-hacking and Significance Stars. Discussion paper series. University of Bern. 17 Appendix (for online publication only) A. Descriptive statistics Figure 3: Histogram (100 bins) of the z-statistic values collected from journals in the AER, QJE and JPE for the period 2012-2015, i.e. when all these journals were adopting single-blind review, by subgroups. Z-statistics larger than 10 have been trimmed in order to improve graph readability. 18 Figure 4: Histogram (100 bins) of the z-statistic values collected from journals in the AER for the period 2005-2011, i.e. under a regime of double-blind review, by subgroups. Z-statistics larger than 10 have been trimmed in order to improve graph readability. 19 % of Articles % of Tests N. Articles N. Tests AER 51.97% 50.82% 422 9268 QJE 32.88% 33.75% 267 6154 JPE 15.15% 15.43% 123 2814 Single-authored 19.58% 18.26% 159 3329 With a theoretical model 57.14% 54.05% 464 9857 Share of authors from top inst 50.25% 48.68% 408 8878 Table 3: Descriptive statistics of the treated and control group samples. Period 2005-2015. Figure 5: Histogram (125 bins) of the z-statistic values collected from QJE, JPE and AER, for the period 2005-2015 by Naguib (2024) and Brodeur et al. (2024a). Z-statistics larger than 10 have been trimmed in order to improve graph readability. In the data collected by Naguib (2024) N= 18,236, whereas in the data collected by Brodeur et al. (2024a) N= 7,658. * = 1.65, ** = 1.96, and *** = 2.58. 20 B. Additional empirical results (1) (2) (3) (4) (5) (6) (7) Overall Theory No theory Single aut No single aut Not top Top inst Double x 2005-11 0.032 -0.004 0.104 0.064 0.032 -0.055 0.107 (0.016) (0.022) (0.024) (0.042) (0.017) (0.021) (0.024) Constant YES YES YES YES YES YES YES Year FEs YES YES YES YES YES YES YES Adj R-sq 0.005 0.010 0.017 0.037 0.007 0.010 0.013 Obs 18,236 9,857 8,379 3,329 14907 9,358 8,878 Articles 812 469 351 159 653 404 408 Table 4: This table shows OLS estimates of equation (1). The dependent variable is a dummy for whether the test statistic is significant at the 1% level. Standard errors in parentheses. Top institutions are the 20 defined by Brodeur et al. (2020). The dummy is equal to one if at least half of the authors belong to a top institution at the time of the article publication. (1) (2) (3) (4) (5) (6) (7) Overall Theory No theory Single aut No single aut Not top Top inst Double x 2005-11 -0.024 -0.062 0.026 0.075 -0.034 -0.104 0.053 (0.015) (0.020) (0.023) (0.039) (0.016) (0.020) (0.022) Constant YES YES YES YES YES YES YES Year FEs YES YES YES YES YES YES YES Adj R-sq 0.006 0.011 0.016 0.038 0.009 0.010 0.018 Obs 18,236 9,857 8,379 3,329 14907 9,358 8,878 Articles 812 469 351 159 653 404 408 Table 5: This table shows OLS estimates of equation (1). The dependent variable is a dummy for whether the test statistic is significant at the 10% level. Standard errors in parentheses. Top institutions are the 20 defined by Brodeur et al. (2020). The dummy is equal to one if at least half of the authors belong to a top institution at the time of the article publication. The median lag in economics from paper submission to publication is around two years. For this reason, in this Section I present some of the baseline estimates by setting the starting date of the change in the peer-review standard to 2013, instead than to 2011. Indeed, the first papers subject to the double review standard at AER were most likely not published until 2013. I focus here on the 5% significance threshold only. 21 (1) (2) (3) (4) (5) (6) (7) Overall Theory No theory Single aut No single aut Not top Top inst Double x ’05-13 -0.003 -0.005 0.005 0.091 -0.006 -0.074 0.063 (0.016) (0.023) (0.025) (0.045) (0.018) (0.022) (0.025) Constant YES YES YES YES YES YES YES Year FEs YES YES YES YES YES YES YES Adj R-sq 0.006 0.012 0.016 0.040 0.009 0.010 0.016 Obs 18,236 9,857 8,379 3,329 14907 9,358 8,878 Articles 812 469 351 159 653 404 408 Table 6: This table shows OLS estimates of equation (1). The dependent variable is a dummy for whether the test statistic is significant at the 5% level. Standard errors in parentheses. Top institutions are the 20 defined by Brodeur et al. (2020). The dummy is equal to one if at least half of the authors belong to a top institution at the time of the article publication. 22 C. Robustness checks with weighting In this Section, I replicate the main tables and figures of the paper using article weights, in order to avoid that articles with more tests have a disproportionate influence on the results. Following Brodeur et al. (2016), to obtain the results with article weights I associate to each test statistic the inverse of the total number of tests that are reported in the same article. The result is that each article contributes in the same way to the distribution. Since I focus on the main results of each paper, for most of the papers in my sample I collect coefficients from one table only. For this reason I refrain from reporting results obtained with article and table weigths. Indeed, they would be essentially identical to the ones obtained with article weights, which are reported in the following. Similar to Brodeur et al. (2016) the ”camel shape” of z-statistics is more evident for the weighted than for the unweighted distributions. This supports the hypothesis that researchers are likely to report more estimated coefficients if their results are statistically significant and conversely they only report a few if other specifications would fail to yield statistically significant results. Weighted distributions give less weight to articles and tables in which many tests are reported. (1) (2) (3) (4) (5) (6) (7) Overall Theory No theory Single aut No single aut Not top Top inst Double x 2005-11 -0.050 -0.027 -0.081 0.060 -0.054 -0.168 0.045 (0.019) (0.025) (0.029) (0.045) (0.021) (0.026) (0.028) Constant YES YES YES YES YES YES YES Year FEs YES YES YES YES YES YES YES Adj R-sq 0.017 0.022 0.031 0.039 0.022 0.034 0.034 Obs 18,236 9,857 8,379 3,329 14907 9,358 8,878 Articles 812 469 351 159 653 404 408 Table 7: This table shows OLS estimates of equation (1). The dependent variable is a dummy for whether the test statistic is significant at the 5% level,with article weights. Standard errors in parentheses. Top institutions are the 20 defined by Brodeur et al. (2020). The dummy is equal to one if at least half of the authors belong to a top institution at the time of the article publication. 23