Learning to be overprecise
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Merkle, Christoph; Schreiber, Philipp Article — Published Version Learning to be overprecise Journal of Business Economics Provided in Cooperation with: Springer Nature Suggested Citation: Merkle, Christoph; Schreiber, Philipp (2024) : Learning to be overprecise, Journal of Business Economics, ISSN 1861-8928, Springer, Berlin, Heidelberg, Vol. 95, Iss. 2, pp. 467-497, https://doi.org/10.1007/s11573-024-01203-w This Version is available at: https://hdl.handle.net/10419/323464 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) Journal of Business Economics (2025) 95:467–497 https://doi.org/10.1007/s11573-024-01203-w ORIGINAL PAPER Learning tobe overprecise ChristophMerkle1,2 · PhilippSchreiber3 Accepted: 11 August 2024 / Published online: 10 September 2024 © The Author(s) 2024 Abstract We replicate and extend two studies on the dynamics of overconfidence among financial professionals. Using 20years of data from the ZEW Financial Market Survey with over 40,000 individual forecasts of confidence intervals, we document that participants are overprecise during the entire time period with no evidence of learning on the aggregate. We confirm that professionals update in a Bayesian manner after hits and misses by contracting or expanding their confidence intervals, respectively. However, this updating is insufficient to reach proper calibration. We cannot confirm other predictions of a Bayesian model. An explanation based on self-attribution bias fits the data better. Keywords Overconfidence· Overprecision· Miscalibration· Replication· Bayesian learning· Financial forecasting JEL Classification D03· D83· D84· G17· G41 1 Introduction Overconfidence can be detrimental to investors as it is associated with harmful behaviors such as overtrading, excessive risk taking, and underdiversification.1 However, the bias would be short-lived if frequent feedback in financial markets * Christoph Merkle [email protected] Philipp Schreiber [email protected] 1 Aarhus University BSS, Fuglesangs Allé 4, 8210Aarhus, Denmark 2 Danish Finance Institute (DFI), Frederiksberg, Denmark 3 Esslingen University, Flandernstraße 101, 73732Esslingen, Germany 1 Studies on the empirical link between overconfidence and trading behavior include Odean (1999); Barber and Odean (2000, 2001); Glaser and Weber (2007); Graham etal. (2009); Puetz and Ruenzi (2011); Anderson (2013); Merkle (2017). Additional evidence comes from experimental studies (Biais et al. 2005; Deaves etal. 2009; Glaser etal. 2013; Menkhoff etal. 2013; Fellner-Röhling and Krügel 2014).
468 C.Merkle, P.Schreiber corrected the bias, or if overconfident investors were driven out of the market due to poor performance. In contrast, the learning-to-be-overconfident hypothesis maintains that investors can become more overconfident over time, in particular, if they are successful (Gervais and Odean 2001). While important for the relevance of overconfidence on financial markets, this hypothesis has rarely been tested. In this paper, we replicate and extend two papers that test the learning-to-be-overconfident hypothesis using surveys of financial professionals. Deaves etal. (2010) (hereafter: DLS) analyze 2 years of data from the ZEW survey of financial analysts. They examine overprecision in submitted 90% confidence intervals (CIs) for DAX index estimates and find that analysts narrow confidence intervals after realizations inside the CI and widen confidence intervals after realizations outside the CI. Due to data limitations they have to work with overlapping observations and use what they themselves call an “imperfect work-around.” Their methodology extracts implied 1-month forecasts from the 6-month forecasts that survey participants submit. While this is a pragmatic solution given the short time-series, it requires strong assumptions such as participants believing that subintervals of their forecasts are i.i.d. We discuss why these assumptions are problematic. In the meantime, 20 years of data are available from the ZEW survey, which allows an extended analysis based on truly independent observations. In a first step, we replicate the results by DLS using their original methodology, but in a ten-times longer time series. We then move to 6-month forecasts that either overlap or do not overlap. The 6-month intervals correspond to the time horizon that survey participants actually forecast, and the information set they have when making these forecasts. Even with this lower time frequency, the number of observations in the replication data still exceeds the number of observations in the original study. DLS themselves state that “for a barely minimal number of independent observations, one would have to possess 10 years of data (as opposed to just over 2 years).” Our data easily exceeds this threshold. We further dispute that a symmetric response to hits and misses (their Hypothesis 2) is the correct expectation for a rational, Bayesian updater. A miss of a 90% confidence interval is a much stronger signal than a hit and should result in more pronounced updating. In this respect, we use theory introduced by Boutros etal. 2020 (hereafter: BBGHP) as a more accurate benchmark for analysts’ responses to hits and misses. We then apply their proposed empirical strategy to the ZEW survey data and analyze whether it delivers results consistent with the original approach by DLS as well as our modified and extended analysis. We then reevaluate how to interpret the results with respect to investor rationality or bias as predicted by learning-to-be-overconfident. At the same time, this analysis constitutes a replication of BBGHP in an independent sample. In their study, they use the Duke CFO survey to test their theory. They find that participants update in a Bayesian fashion, but adjust their beliefs too little, which results in a persistence of overprecision. The Duke CFO survey contains a very similar question for confidence intervals for the S&P index as the ZEW survey for the DAX. However, CIs are based on expected returns and not on index levels, and they are collected for a yearly horizon. It is known from earlier research that return forecasts and price forecasts may differ in unexpected ways
469 Learning tobe overprecise (Glaser etal. 2019). The replication further allows us to compare German financial analyst forecasts to U.S. CFO forecasts, which provides robustness across different samples of financial professionals. We further argue that the ZEW sample is the superior testing ground for the theory developed by BBGHP, as there is less fluctuation among participants. Persistence in overprecision is hard to test in a sample with high panel attrition. Replication results confirm three main findings of DLS. First, there is a strong presence of overprecision in the sample. The average hit rate in the extended sample from 2003–2023 is close to 40% irrespective of using implied 1-month forecasts or 6-month forecasts. The main reason for the low hit rate is that confidence intervals estimated by forecasters are too narrow. Confidence intervals would need to be more than twice as wide on average to be consistent with historical volatility of the DAX. While forecasters are on average also too pessimistic, the location of confidence intervals plays a minor role for hit rates. Second, forecasters revise confidence interval widths in the correct direction in response to both hits (resulting in narrower intervals) and misses (resulting in wider intervals). We replicate this finding in the original sample and in the extended sample. Third, the reaction to hits and misses is about symmetric in economic magnitude. In particular, we cannot reject symmetry in our most powerful test of overlapping 6-month forecast periods in the extended sample. Forecasters tend to expand their confidence intervals after a miss to a similar degree as they narrow the intervals after a hit. We highlight that this is not what a Bayesian forecaster would do. In contrast to DLS, the analysis of the extended sample shows little evidence of a discernible time trend in hit rates at the group level, particularly when controlling for squared realized returns. This suggests that hit rates are predominantly influenced by the extremity of realized returns, rather than substantial improvements in forecasting quality. To gain a better understanding of whether forecasters update in a Bayesian fashion, we replicate the analysis by BBGHP. The results provide some evidence in favour, but also evidence against the predictions put forward by BBGHP. Forecasters who have missed the realized value enlarge their confidence interval by 3 to 4%-points relative to those who score a hit. This adjustment is insufficient to reach proper calibration and even smaller than the one reported by BBGHP. We can thus confirm their finding of correct directional, yet insufficient adjustment. We cannot confirm that misses on the upside or on the downside of the confidence interval lead to proportionally more adjustments of the upper or lower bound of the CI, respectively. Instead, we find that participants enlarge their confidence intervals only after a miss on the downside. This is puzzling, as from a Bayesian perspective a miss challenges the prior confidence interval irrespective of whether it occurs on the upside or downside. BBGHP explain the insufficient adjustment after misses by a too high conviction in the prior and perform several tests for this assumption. First, they show that subsequent misses lead to less adjustments in confidence intervals. They argue that priors incorporate more information with each return realization, and thus subsequent return realizations receive less weight in the Bayesian updating process. The ZEW sample is ideal to test this assumption, as many forecasters participate over
470 C.Merkle, P.Schreiber many waves. The response to misses does not decrease monotonically in this sample, instead adjustments become stronger over the first few years of participation. Second, the Bayesian model predicts that forecasters with the strongest initial overprecision will adjust their forecasts the least due to their higher conviction in the prior. Our replication shows that the initial calibration is not predictive for the subsequent response to misses. We find similar responses to hits and misses independent of the initial calibration of forecasters. We can confirm a finding on path-dependence, though. The CI adjustment after a miss-hit sequence is stronger compared to a hit-miss sequence. This is consistent with the idea of a later miss having less effect on the prior. However, this analysis uses information from two independent forecasts intervals 6 months apart and ignores information revealed in between. When properly accounting for this information, path dependence can be explained by the number of accumulated hits and misses. Finally, we provide additional evidence of an equally strong and robust response to hits, which poses a significant challenge to a strict Bayesian model. We argue that a self-attribution bias accounts better for this finding, as it suggests that participants attribute prediction success to their own skill and prediction failure to external circumstances. Overconfidence is thus expected to increase after success, but not to decrease sufficiently after failure. In line with this argument, we show that a widening of confidence intervals only occurs after relatively large misses. Presumably, participants can convince themselves that they ‘almost’ got it right after narrow misses. A self-attribution bias also aligns well with the finding that adjustments are muted for misses on the upside, as high return realizations are another metric of success for financial professionals. Due to the bias, participants learn to be overprecise even if they issued a well-calibrated confidence interval in the beginning. Levels of overprecision reach an equilibrium at relatively low hit rates. 2 Data In this replication study, we use data from the monthly ZEW Financial Market Survey Panel. The survey is administered by the Leibniz Centre for European Economic Research (ZEW) and send out to financial professionals from financial institutions and financial departments of non-financial firms in Germany. A detailed and current description of the data set is provided by Brückbauer and Schröder (2023). Numerical questions for the future DAX index value and 90% confidence intervals were first introduced in 02/2003. The exact wording in the survey is “Six month ahead, I expect the DAX to stand at _____ points. With a probability of 90 per cent the DAX will then range between _____ and _____ points.” The responses form the basis for our study and have been elicited in the survey to the present day. During that time period, the survey collected on average 200 individual responses each month, which gives us a total of 48,050 observations. After pre-registered exclusion criteria,2 there remains a final sample of N=47,011. 2 Pre-registered with the OSF Registries at https:// osf. io/ d8tf6.
471 Learning tobe overprecise We further obtain realized DAX values from Thomson Reuters Eikon. This data is used to assess whether confidence intervals contain the realized value (hit) or not (miss). To determine hits and misses accurately, we use individual response dates of survey participants, as the survey is usually stretched out over 1 or more weeks.3 Given that realizations occur 6 months after the forecast and are necessary for the analysis, we only use survey waves up to 03/2023. As control variables, we obtain VDAX-NEW index data (an index of implied volatility analogous to the VIX) and also calculate realized DAX volatility. As DAX index values change significantly over the period analyzed, the point forecasts and widths of confidence intervals are easier to compare when using returns. To convert forecasts of index values into return forecasts, we follow Glaser etal. (2019) and use the DAX daily open level on the day a forecast is made.4 We summarize the return and confidence interval forecasts in Table1. The average 6-month return expectations are quite conservative at 2.4%, while the realized semiannual return matched to the forecast periods was 5.7%. Figure B.1 in the Online Appendix shows the distribution of return expectations and realizations. The average width of confidence intervals is 17.1%-points. The very stable participant pool is the main advantage of the ZEW Financial Market Survey, which results in a median of 33 completed surveys and almost 5years of participation. While BBGHP highlight that “almost two dozen CFOs have responded 30 or more times” (p. 5), in our data this is the norm rather than the exception. As Table1 further reveals, the panel consists of 785 individual forecasters. Table 1 Return expectations and realizations The table shows summary statistics for the ZEW Financial Market Survey (means, standard deviations, and percentiles). Return expectations are derived from the DAX index value expectation and the current DAX value ( DAX expectation current DAX − 1 ), as in Glaser etal. (2019). The lower and upper bound return are derived from the submitted 90% range of DAX index values in the same manner. The confidence interval width is expressed in percentage points. Realized return are the realizations matching the time period of each individual estimate. Participation frequency and length are measured on the individual forecaster level N Mean Std. dev. 5th perc. Median 95th perc. Return expectation 47,011 2.44 7.92 − 10.44 2.63 14.40 CI lower bound 47,011 − 7.86 9.57 − 24.97 − 6.92 5.26 CI upper bound 47,011 9.26 9.15 − 3.04 8.15 24.73 CI width 47,011 17.12 12.09 4.04 14.32 40.54 Realized return 47,011 5.70 13.65 − 9.60 6.93 25.47 Participation frequency 785 59.89 65.85 1.00 33.00 206.00 Participation duration (months) 785 83.40 79.72 0.00 57.00 241.00 3 This is particularly relevant for the early surveys that were not fully online. DLS follow the same approach. 4 For responses submitted on weekends or bank holidays, we use the last available DAX daily close level. DLS do not specify which value they use for their variable “current DAX.” However, this choice is unlikely to systematically affect replication results.
472 C.Merkle, P.Schreiber 3 Motivation andhypotheses The motivation for this replication study is threefold. First, the theoretical literature mostly models overconfidence in terms of investors overestimating the precision of information (Kyle and Albert Wang 1997; Benos 1998; Daniel etal. 1998; Odean 1998; Gervais and Odean 2001). This is akin to “overprecision”, which is one of several forms of overconfidence.5 In these models, investors underestimate the error variance of an information signal they receive. In reality, such signals and beliefs about signals are usually unobservable. Return expectations and confidence intervals around return expectations are perhaps the empirical measures that match the theory most closely. Secondly, the empirical literature on the dynamics of financial overconfidence is relatively scarce. A potential reason is that it requires repeated measurements of overconfidence over time among a relevant group of financial market participants. Merkle (2017) provides evidence on the dynamics of overconfidence for online brokerage investors. DLS, Ben-David etal. (2013) and BBGHP analyze data from panel surveys of financial professionals. While all studies establish that investors are strongly overconfident, the results on its dynamics are more mixed. Merkle (2017) finds that investment success leads to increased overprecision in line with learning-to-be-overconfident (Gervais and Odean 2001). DLS report evidence for some degree of rational learning and against learning-to-be-overconfident, whereas BBGHP find support for Bayesian learning with insufficient updating. Finally, the time series in existing studies is often relatively short. Merkle (2017) analyzes nine quarterly surveys over a period of 2 years. Few participants respond more than five times and, in addition, the financial crisis with its extreme return realizations is part of the sample. DLS examine a time-period of similar length. In an attempt to enhance their sample size, they generate implicit 1-month forecasts from 6-month forecasts. We provide evidence that this methodology is problematic. While BBGHP possess a long time series (16 years), the Duke CFO survey is characterized by high fluctuation, which likewise means that few CFOs participate more than ten times. We conclude that a replication within a long-run survey that features many long-term participants is warranted. It may help to distinguish which earlier findings are robust and which are chance results. In addition, BBGHP provide new theoretical predictions that can be validated in an independent sample. Besides the time series extension, the replication thus offers a cross-country validation by comparing results obtained with U.S. CFOs to those obtained with German financial professionals. Based on the previous literature, we formulate concrete hypotheses that have been pre-registered with the OSF Registries at https:// osf. io/ d8tf6. The first hypothesis concerns static overprecision: 5 In an influential study, Moore and Healy (2008) distinguish three forms of overconfidence: overestimation, overplacement, and overprecision. Overprecision hereby refers to excessive precision in one’s beliefs.
473 Learning tobe overprecise H1: Market forecasters as a group are overprecise in the sense that eventual DAX realizations fall within their 90% confidence interval much less than 90% of the time. This is the analog to DLS, Hypothesis 1. Unlike DLS we do not hypothesize that forecasters are “properly calibrated” as empirical evidence has accumulated that confidence intervals are far too narrow also in a financial context (Ben-David etal. 2013; Glaser etal. 2013; Merkle 2017). However, we still test the null of 90% calibration. It is important to note that 90% calibration cannot be expected in each survey wave, as realized returns are not independent and extreme outcomes affect all forecasters. Nevertheless, 90% should be reached in a properly calibrated population in the long run. We thus next track overprecision over time: H1a: There is no time trend in overprecision on group level. DLS report suggestive evidence that respondents become more accurate as a group over time. This may be a result of the stable market period at the end of their sample. In contrast, Ben-David etal. (2013) and BBGHP do not find a time trend in overprecision. Statistics on group level aggregate the responses of participants who experienced a hit or a miss in their previous estimate, of those who enter the panel for the first time, and of those who re-enter after a hiatus. It is thus difficult to formulate a dynamic prediction for calibration on group level as learning theories such as learning-to-be-overconfident are path dependent. We thus turn to the individual level and look into reactions to hits and misses: H2: Learning takes place in the sense that after hits confidence intervals contract and after misses confidence intervals expand. Adjustments are of the same magnitude. This hypothesis corresponds to DLS, Hypothesis 2. DLS associate this hypothesis, in particular the symmetric adjustments, with rational learning (p.406). To us, there seems little basis for this claim. With a 90% confidence interval, hits should be observed frequently and give rise to much less adjustment than misses. Moreover, a symmetric adjustment will not lead forecasters to converge to proper calibration, because at higher hit rates, they will narrow their confidence intervals too much. We propose Hypothesis 2a as an alternative to H2: H2a: Bayesian learning takes place in the sense that after a miss confidence intervals expand, but insufficient to obtain proper calibration. After a hit confidence intervals contract. BBGHP derive this prediction from a Bayesian learning model. The direction of updating is the same between H2 and H2a, but no symmetry is assumed in H2a. Insufficient updating is generated by overly high conviction in one’s prior. BBGHP argue that being overprecise already implies having too much confidence in one’s prior, which dampens adjustments in response to return realizations. The model generates a further hypothesis on how confidence intervals are adjusted:
474 C.Merkle, P.Schreiber H2b: A miss on the downside (upside) results in relatively more adjustment of the downside (upside) part of the confidence interval. DLS do not test H2b, but BBGHP confirm its prediction with the CFO data. We will further replicate a test on path dependence, which is another specific result of their Bayesian model. Our data are perhaps most powerful, though, in looking at the long-term effects of learning. DLS base their analysis on job experience, which they gather from a one-time demographic survey. Unfortunately, this data is not available for most respondents, in particular for those joining the panel later. Instead, we test for an experience effect in the spirit of BBGHP, predicting that strength of adjustments after misses decreases with survey experience (leading to persistence in overprecision): H3: Subsequent misses lead to progressively less adjustments in confidence intervals. The reasoning behind this hypothesis is that priors incorporate more information with each return realization, and thus new return realizations receive less weight. Conviction thus increases with (survey) experience. As explained above, initial overprecision can also be equated with higher conviction in one’s prior: H3a: Respondents who are most overprecise adjust the least. However, this final hypothesis is debatable, as one might interpret confidence intervals as an estimate of stock market variance (second moment of the return distribution) that is distinct from conviction in the first moment of the distribution. Taken together, testing the formulated hypotheses will provide us with a good understanding of the dynamics of overprecision and, in particular, whether and how sophisticated financial professionals learn. 4 Replication results 4.1 Static overprecision (H1) A 90% confidence interval should contain the realized value about 90% of the time. A first pass thus is whether financial analysts as a group are well calibrated. We calculate the proportion of hits (i.e., the realized DAX value falls within the specified range of a forecaster) for each survey wave. The results displayed in Fig.1 show that hit rates are well below 90% for almost all waves. The average hit rate throughout the extended sample from 2003 to 2023 is 42% (38% when using 6-month hit rates). As the average hit rate in the original sample of DLS was 50%, overprecision is even stronger in the extended sample. Reasons for the low hit rate may be that confidence intervals are too narrow or that the placement of the confidence intervals is off. We calculate the return volatility implied in the width of confidence intervals using the method of Keefer and Bodily (1983). Submitted confidence intervals are consistent with an annual volatility
481 Learning tobe overprecise covers 20 years of data. In fact, we observe 41 independent 6-month periods, which even exceeds the 26 1-month periods in the original sample. We are thus able to perform replications based on non-overlapping 6-month forecasts. However, this is not our preferred methodology as it throws away a lot of data.11 We instead advocate the use of overlapping 6-month forecasts. The major argument against overlapping forecast periods is that one “surprise” would affect multiple forecasts and thus distort hit rates. This is a valid argument when using a short time series, during which the number of extreme return realizations varies a lot by chance. However, in the long run “surprises” should not occur more frequently than anticipated by forecasters. A confidence interval that is in line with the volatility of the DAX index achieves proper calibration independent of using non-overlapping or overlapping observations. Ben-David etal. (2013) and BBGHP likewise use overlapping observations (quarterly forecasts of 1-year returns). We re-estimate regression Eq.4 over 6-month horizons. This means all variables representing changes are now defined relative to their levels in t−6 . The dummy Table 4 Change in confidence interval width (6-month periods) The table shows coefficients of linear panel regressions with p-values (for clustered standard errors) in parentheses. The dependent variable is the change in confidence intervals from month t − 6 to month t relative to the cross-sectional average width in t − 6. Independent variables are an indicator whether the previous confidence interval contained the realized value ( Hitt−6 ) and changes in DAX return and DAX expected volatility (from the 6m VDAX, from 11/2006, and the VDAX before). Columns (1) and (2) show results for the original sample and columns (3) and (4) for the extended sample, using overlapping observations. Columns (5) and (6) show results for non-overlapping observations in the extended sample (using waves [1, 7, 13, ...]). A Wald-test tests for symmetry of reactions to a hit and a miss, the table reports the two-sided p-value of this test. */**/*** stars denote statistical significance at the 10%/5%/1% level Overlapping 6m-periods Non-overlapping 6m-periods Original sample Extended sample Extended sample (1) (2) (3) (4) (5) (6) Hitt−6 − 0.155*** − 0.205*** − 0.205*** − 0.185*** − 0.142*** − 0.137*** (0.000) (0.000) (0.000) (0.000) (0.000) (0.000) Δ log(DAX) − 0.731*** 0.263*** 0.089 (0.000) (0.000) (0.215) Δ log(VDAX) − 0.066 0.391**** 0.379*** (0.300) (0.000) (0.000) Constant − 0.020** 0.070*** 0.109*** 0.099*** 0.086*** 0.089*** (0.043) (0.000) (0.000) (0.000) (0.000) (0.000) R 2 0.031 0.065 0.027 0.063 0.020 0.066 Observations 3639 3639 35,775 35,775 5980 5980 Wald-Test 𝛽1 = −2𝛼 0.000 0.037 0.176 0.195 0.003 0.001 11 In our base sample, we use the survey waves [1, 7, 13, ...], while we use the other non-overlapping samples starting in waves 2–6 for robustness.
482 C.Merkle, P.Schreiber variable for a previous hit now indicates whether the confidence interval from the survey 6 months ago contains the realized value or not. This corresponds to the information participants have when completing the current survey, as for any later survey ( t−5 to t−1 ) the final realization is not yet known. Table4 shows results for regressions using 6-month periods. Participants reduce the widths of their confidence intervals after a hit and they expand it after a miss (with the exception of one negative coefficient for the constant in column (1)). We will focus on the results for the extended sample presented in columns (3) and (4). The economic magnitude of the response to hits and misses is about twice as large as when using imputed 1-month CIs (cp. Table3, columns (5) and (6)). This suggest that 6-month realizations are the more salient outcome for participants. Participants also expand their confidence intervals in response to high returns and higher uncertainty. For the extended sample, we cannot reject the null that reactions to hits and misses are symmetric. We reject this hypothesis for the original sample and the non-overlapping sample used as baseline. In Online Appendix, Table A.2, we report robustness results for the other non-overlapping samples. The coefficient estimates are consistent across samples with mixed results for the test of symmetry. We interpret these results as supportive of Hypothesis 2. There is little doubt that participants adjust confidence intervals in the correct direction. We further conclude that this adjustment is about symmetric (in an economic sense, not always in a strict statistical sense). As argued before, we disagree with DLS that this should be taken as evidence for a rational response. Instead, the symmetric adjustment preserves the status quo of an almost constant CI width over the entire sample period well below proper calibration. We next examine whether this insufficient adjustment is consistent with a Bayesian model. 4.2.4 An alternative test BBGHP suggest an alternative test for the response to hits and misses. It is based on a Baysesian model in which a miss challenges the prior belief about a confidence interval. A forecaster thus updates by expanding the CI to a degree dictated by the weights given to the prior and the new information. BBGHP estimate the following regression: The equation looks similar to regression equation4, but there are some differences. First, the change in confidence intervals is measured in percentage points as BBGHP elicit return confidence intervals as opposed to index value confidence intervals. We transform DAX point value forecasts into return forecasts, which makes scaling by average interval width redundant. The correlation between the dependent variables in Eqs.4 and 5 is still very high (0.91). The Bayesian model focuses on reactions to misses, which is why the indicator variable is one for a miss. This should just flip the coefficients and not lead to material changes. Further, BBGHP include expected and unexpected volatility derived from historical volatility instead of option-implied volatility. Likewise, this change in control variables should not be consequential. (5) ΔCI widthi,t= 𝛼 + 𝛽 1Missi,t−1+ 𝛽 2Unexp.Volat+ 𝛽 3Exp.ΔVolat+ 𝛾 i+ 𝜔 t
483 Learning tobe overprecise Finally, BBGHP include forecaster fixed effects 𝛾i and survey-wave fixed effects 𝜔t in some regressions. Participants who missed the realized value enlarge their confidence interval by 3–4%-points relative to those who score a hit (see Table 5). However, given that the latter expand the confidence interval, the net effect is much smaller. We follow BBGHP and report the total change and percentage change in confidence interval width for the group recording a miss.12 We find that the adjustments in response to a miss are smaller in the ZEW sample relative to the Duke CFO sample. From a Bayesian perspective, they should be larger, as a miss for a 90% CI is a stronger signal than a miss for a 80% CI. Table 5 Change in confidence interval width (BBGHP regression) The table shows coefficients of linear panel regressions with clustered standard errors in parentheses. The dependent variable is the change in confidence interval width from month t − 6 to month t in return percentage points. Independent variables are an indicator whether the previous confidence interval missed the realized value ( Misst−6 ) and unexpected volatility and changes in expected volatility. The regressions contain different sets of fixed effects. The total change in CI width is the total change in percentage points for a forecaster that misses the interval. The total % change is the change in percent relative to the average prior confidence interval for a forecaster that misses the interval. Both are calculated using a linear prediction based on the estimation results. */**/*** stars denote statistical significance at the 10%/5%/1% level Δ CI width (1) (2) (3) (4) Misst−6 3.37*** 3.40*** 3.78*** 3.14*** (0.16) (0.16) (0.18) (0.15) Unexpected vol. 0.30*** (0.01) Exp. change in vol. 0.35*** (0.05) Constant − 2.88*** − 2.49*** − 17.55*** − 2.12*** (0.16) (0.10) (1.06) (0.09) R 2 0.02 0.02 0.23 0.13 Observations 35,775 35,775 35,775 35,775 Total Δ CI width 0.48 0.90 0.90 0.89 (0.14) (0.06) (0.64) (0.13) Total % Δ CI width 3.6% 6.6% 6.6% 6.5% Forecaster fixed effects N Y Y Y Time fixed effects N N Y N 12 We use predicted values instead of adding up coefficients as in BBGHP. For the simple regression in column (1), both methods give the same result. However, summing fixed effects is a flawed approach, as results depend on the omitted category. This becomes evident in the high variation in effect sizes in the later tables in BBGHP.
484 C.Merkle, P.Schreiber We find support for Hypothesis 2a. As the expansion of confidence intervals after a miss is smaller than in BBGHP, we conclude that even if participants were Bayesians, their adjustments are insufficient or, to speak in terms of the model, their conviction in their prior is too high. However, we did also confirm H2 before, which presents a challenge to a Bayesian model. A Bayesian updater, in particular one with high conviction in the prior, should adjust very little to a hit. In addition to the adjustment in CI width, BBGHP also test for the direction of the adjustment. The prediction stated in H2b is that a miss on the downside will affect the lower portion of the confidence interval more strongly than the upper portion Table 6 Response to misses on the upside and downside The table shows coefficients of linear panel regressions with clustered standard errors in parentheses. The dependent variable is the change in the upper portion of the CI ( Δ UCI) or lower portion of the CI ( Δ LCI) from month t − 6 to month t in return percentage points. Independent variables are two indicators whether the previous confidence interval missed the realized value on the upside ( Miss Hight−6 ) or downside ( Miss Lowt−6 ), and unexpected volatility and changes in expected volatility. The regressions contain different sets of fixed effects. The total change in UCI width or LCI width is the total change in percentage points for a forecaster that misses the interval (upside or downside). The total % change is the change in percent relative to the average prior confidence interval for a forecaster that misses the interval (upside or downside). Both are calculated using a linear prediction based on the estimation results. */**/*** stars denote statistical significance at the 10%/5%/1% level Δ UCI Δ LCI (1) (2) (3) (4) (5) (6) Miss Hight−6 0.55*** 1.81*** 0.91*** 0.92*** 2.49*** 1.33*** (0.08) (0.13) (0.08) (0.11) (0.16) (0.11) Miss Lowt−6 3.19*** 0.67*** 2.13*** 4.52*** 2.11*** 3.40*** (0.13) (0.15) (0.12) (0.19) (0.19) (0.17) Unexpected Vol. 0.11*** 0.11*** (0.01) (0.01) Exp. Change in Vol. 0.22*** 0.27*** (0.03) (0.04) Constant − 1.15*** − 7.69*** − 0.90*** − 1.62*** − 10.12*** − 1.32*** (0.08) (0.62) (0.05) (0.10) (0.80) (0.07) R 2 0.05 0.14 0.08 0.05 0.13 0.08 Observations 35,775 35,775 35,775 35,775 35,775 35,775 Miss High Total Δ CI width − 0.60 − 0.45 − 0.47 − 0.70 − 0.51 − 0.52 (0.06) (0.37) (0.07) (0.09) (0.49) (0.09) Total % Δ CI width − 10.3% − 7.7% − 7.9% − 8.2% − 5.9% − 6.1% Miss Low Total Δ CI width 2.04 2.13 2.16 2.90 3.06 3.07 (0.12) (0.44) (0.12) (0.18) (0.57) (0.17) Total % Δ CI width 43.6% 45.5% 45.9% 41.6% 44.0% 44.0% Forecaster Fixed Effects N Y Y N Y Y Time Fixed Effects N Y N N Y N
485 Learning tobe overprecise (and vice versa for a miss on the upside). The lower portion hereby is defined as the part of the confidence interval between the lower bound and the point forecast, while the upper portion is the part between the point forecast and the upper bound. The indicator for a miss in regression equation5 is thus replaced by two separate indicators for a miss on the upside and on the downside, respectively. Dependent variables are changes in total CI width, lower portion width, and upper portion width. Table 6 shows results for the modified regression and replicates Table 6 in BBGHP. We not only find that participants respond more strongly to misses on the downside, but also that they only expand their confidence intervals after misses on the downside. When they miss on the upside, they even narrow intervals. The results show a similar reaction for the upper portion of the confidence interval (columns (1)–(3)) as for lower portion of the confidence interval (columns (4)–(6)). BBGHP also find a stronger response to misses on the downside in their Tables5 and 6.13 However, they do also observe an asymmetry in the sense that a downside miss leads to more adjustment of the lower side of the CI and and upside miss to the upper side. We cannot replicate this result and thus cannot confirm Hypothesis 2b. When we consider these results jointly, they are inconsistent with Bayesian updating. First, a Bayesian should expand confidence intervals after a miss independent of whether the miss occurred on the downside or upside, as both challenge the prior. Second, a Bayesian should consider the side of the miss to determine which part of the confidence interval needs to be expanded. An alternative to changing the skewness of the interval is to shift the location of the interval. 4.3 Long‑term dynamics (H3) As a final verdict on the Bayesian model is still out, we replicate further tests proposed by BBGHP. Based on the assumption that every realization increases the strength of the prior, the reaction to a miss will decline for subsequent misses (H3). To test this hypothesis, BBGHP interact the miss indicator with a count variable that indicates the number of the response by the respective individual. To be clear, this is not a count of misses, but a count of the response number independent of whether previous responses were hits or misses (but the count only includes responses for which a prior outcome is available). We first replicate their analysis for up to ten responses. Figure3, Panel A, shows the results for the total change in confidence intervals after the jth response. We observe a pattern that is strikingly different from the one reported by BBGHP.14 Observing the first few outcomes, ZEW survey participants narrow their confidence intervals even after a miss, before expanding them again after later misses. A closer look reveals that forecasters during the first three survey 13 We also replicate the results for total CI width (BBGHP, Table 5) in Online Appendix Table A.3. While there is no clear prediction for the change in total CI width depending on misses on the downside vs. upside, it is interesting that the overall expansion of confidence intervals is entirely driven by misses on the downside. 14 To reassure that our methodology is correct, we replicate the BBGHP result with the Duke CFO data. We successfully replicate their Figure6 and include it in the Online Appendix Figure B.5 for comparison.
486 C.Merkle, P.Schreiber waves provide rather large confidence intervals and then narrow them regardless of the outcome (although they narrow them more after a hit). This is likely due to the special market situation around the time when the question for DAX CIs was first introduced.15 To abstract from such wave effects and to leverage the abundant data, we expand the graph to the 100th response, which is possible as there are still 88 Fig. 3 Change in confidence interval and experience. Panel A shows the total change in the predicted confidence interval in response to the jth observed outcome if this outcome was a miss. The dashed line is the baseline estimate from the second specification in Table5. Panel B extends the display to the first 100 responses 15 After the dot-com boom, the DAX had seen a long decline from above 8,000 index points to a low of 2,200 points in March 2003. Elevated confidence intervals seem to reflect this market uncertainty. In a robustness test, we exclude the first three waves of the panel and re-estimate the models presented in Table4. The coefficient estimates are very similar (see Online Appendix Table A.4). Symmetry of reac-
487 Learning tobe overprecise forecasters who record a miss in their 100th response (see Fig.3, Panel B). The second panel shows that experienced forecasters widen their CIs more in response to a miss, until after very many forecasts the response levels off to the baseline. It is possible that after initial learning, participants perceive the survey as a routine task after many years. However, the pattern might also be an artefact produced by the financial crisis that early participants encounter around response 40–60.16 A second dynamic prediction of the Bayesian model is that the most overprecise forecasters will adjust the least (H3a). We test this hypothesis following the Table 7 Initial calibration and response to a miss The table shows coefficients of linear panel regressions with clustered standard errors in parentheses. The dependent variable is the change in confidence interval width from month t−6 to month t in return percentage points. Independent variables are an indicator whether the previous confidence interval missed the realized value ( Misst−6 ) and interactions of this indicator with initial calibration group (G1– G3). The regressions in columns (1) and (2) use the regression specification with individual and time fixed effects. The total change in CI width is the total change in percentage points for a forecaster that misses the interval. The total % change is the change in percent relative to the average prior confidence interval for a forecaster that misses the interval. Both are calculated using a linear prediction based on the estimation results. */**/*** stars denote statistical significance at the 10%/5%/1% level Baseline Interactions (1) (2) Misst−6 3.43*** (0.18) Misst−6×G1 3.21*** (0.49) Misst−6×G2 3.28*** (0.27) Misst−6 × G3 3.53*** (0.22) Constant − 9.73*** − 9.73*** (0.97) (0.97) R 2 0.20 0.20 Observations 33,871 33,871 Baseline Most overprecise Least overprecise Miss ×G1 Miss ×G2 Miss ×G3 Total Δ CI width 1.10 0.97 1.20 1.07 (0.64) (0.81) (0.69) (0.64) Total % Δ CI width 8.4% 10.2% 10.1% 7.4% 16 Figure B.6, Panel A in the Online Appendix reproduces the graph including only forecasters that reach the 100th response to more cleanly track the same group of participants. The general pattern is very similar. When only considering participants who joined the survey after the financial crisis, there is no evidence of learning (see Figure B.6, Panel B). However, this figure is based on fewer observations. tions to hits and misses is rejected in one additional specification. Footnote 15 (continued)
488 C.Merkle, P.Schreiber approach by BBGHP. They use the average confidence interval from the first four responses to determine initial calibration. They then group forecasters into three groups from most to least overprecise using 10 and 20 percentage points as cutoffs for CI width. We keep these cutoffs even though they produce uneven groups in our sample (the split is about 10/30/60). The model from Eq.5 is then augmented by interaction terms of initial calibration group times the miss indicator. The estimation of the modes includes only observations that are not used to determine initial calibration. The results displayed in Table 7 show that initial calibration is not predictive for the subsequent response to misses. The coefficients for all interaction terms in regression (2) are very similar and close to the coefficient obtained in the baseline regression. The same holds for the total change in CI width and the percentage change. Contrary to the results found by BBGHP, percentage-wise the change is even largest for the most overprecise group. For robustness, we use different cutoffs to obtain more even groups and also a longer calibration period of twelve responses. Results shown in Online Appendix Table A.5 confirm that, if anything, the initially more overprecise participants make larger adjustments. These results seem intuitive as the most overprecise participants start out with the narrowest confidence intervals and face greater need for adjustment to become well calibrated. A narrow confidence interval makes it also more likely to encounter a realization that is far off the bounds of the interval and has a participant reconsider. However, the Bayesian logic presented by BBGHP is different. In their model, initial overprecision is a measure for the tightness of one’s conviction in the prior. They admit that their data does “not provide a direct measure of the tightness of CFOs’ beliefs” (p. 16), but work with the assumption that the belief about the variance expressed in confidence intervals is closely related to the tightness of conviction in the mean return forecast. This is in fact an empirical question, and we measure for each participant with at least twenty forecasts (N=468) the variance of the return forecast and the variance of CI width. These are two direct measures for conviction of beliefs (in the mean and variance of the belief distribution) as proposed in BBGHP, Table 3. In their paper, they do not measure conviction of belief directly, presumably because a too low fraction of CFOs participates long enough. Instead, they argue that CI width itself can serve as a measure of conviction, hence the sorting of participants by initial overprecision. We correlate the two variance measures with average individual CI width and find that CI width is highly correlated with the variance of CI width (0.78) but not with the variance of the return forecast (0.12). This means that in the time-series, people with a tight confidence interval shift their return expectations from period to period just as much as everyone else, but will not change their confidence interval width dramatically. While not pre-registered as a separate hypothesis, we also replicate a result on path dependence reported in BBGHP. They find a stronger widening of CIs after a miss-hit sequence compared to a hit-miss sequence. They argue that a hit-miss sequence will induce a smaller adjustment, as a hit increases the weight on the prior which then mutes the reaction to a subsequent miss. We follow their methodology and compute the change in confidence interval width from t−12 to t. A miss-hit
489 Learning tobe overprecise sequence implies a miss for the forecasting period t−12 to t−6 followed by a hit for the period t−6 to t, and vice versa for a hit-miss sequence. We can confirm that the overall response to the miss-hit sequence is stronger than to the hit-miss sequence (see Online Appendix, Table A.6). The difference is also similar in economic magnitude to the results found by BBGHP.17 In all fairness, this is a result that aligns well with the Bayesian prediction, but we do not consider it particularly strong evidence. We would agree that the result represented a (pure) sequence effect, if the predictions in t−12 and t−6 were the only predictions the forecasters made. However, this ignores predictions and realizations that are revealed in between. Forecasters with a miss for the forecast period t−12 to t−6 are likely to have observed further misses being revealed in periods t−11 to t−7 . In fact, the average hit rate of a forecaster with an initial miss is 29% over this period. We can thus interpret the adjustment of CI width as a response to accumulated misses. The average hit rate for a forecaster with a miss in the second period is 41% for the other predictions revealed during this period. This is an alternative explanation for a less strong adjustment. In the ZEW Financial Market Survey we cannot confirm H3 and H3a. The longrun data for many participants allow for a powerful test of the dynamics of overprecision. A diminishing response to misses when participating repeatedly would suggest that after hundred or more responses the reaction to misses (or hits) should approach zero. Their priors would dictate participants to converge to an immutable response. Similarly, an initial narrow confidence interval would prevent them from learning much from return realizations. However, this is not what we find, as participants make strong adjustments to CIs even with long forecasting experience and are also not constrained by their initial calibration. Table8 provides a snapshot of our results. We confirm both tested results from DLS and two out of five results from BBGHP. We will discuss these findings further in the next section to understand how they align. 5 Learning tobe overprecise Self-attribution bias suggests that overconfidence increases after success, but does not decrease to the same extent after failure. DLS report evidence of a symmetric reaction to hits and misses, which we can broadly confirm in our replication. They interpret this as “inconsistent with self-attribution bias (p.408).” BBGHP instead propose a Bayesian model with insufficient updating to explain how CFOs respond to misses. They conclude that their result “supports the broad class of models built on Bayes’ rule that study underand over-reaction to news (p.35).” They explicitly 17 However, as the table also shows, the result does not differ when splitting the sample by accumulated prediction experience. This is puzzling, as after fifty forecast the marginal effect of one additional forecast on the weight of the prior should be negligible (see columns (4) and (5)). In Figure B.7, we shed light on the path dependence of CIs. The analysis shows a selection effect, as the group with an initial miss had a narrower confidence interval in t − 12 . The following adjustments are almost symmetric, with a slightly stronger response to a hit in both sub-periods. Both paths end with a similar confidence interval in t.
490 C.Merkle, P.Schreiber cite the model by Gervais and Odean (2001), which heavily builds on biased selfattribution, as an example for such models. We can confirm some of their results, but other results, in particular for the dynamics of overconfidence, are not robust. Can these conflicting findings be reconciled? First, let us review the arguments by DLS. They suggest that self-attribution bias would predict a stronger absolute reaction to a hit than to a miss when adjusting confidence intervals. However, this is only warranted if signal strength is comparable. This is not the case for a 90% confidence interval, as a hit is the expected and more frequent outcome, while a miss should be observed only rarely. Several misses strongly point to the fact that a confidence interval might be too narrow, while several hits do not give reason for concern. DLS further suggest that learning-to-be-overprecise means that participants become ever more overprecise which is inconsistent with the rather stable hit rates observed over time. However, with rising overprecision the probability of a miss increases, which Table 8 Summary of replication results The table shows a summary of all tested hypotheses, the original results, and our replication results. It includes one not pre-registered (NPR) hypothesis # Hypothesis Source Original result Replication result H1 Market forecasters as a group are overprecise in the sense that eventual DAX realizations fall within their 90% confidence interval much less than 90% of the time. DLS Confirmed Confirmed H1a There is no time trend in overprecision on group level. DLS not tested Confirmed H2 Learning takes place in the sense that after hits confidence intervals contract and after misses confidence intervals expand. Adjustments are of the same magnitude. DLS Confirmed Broadly confirmed H2a Bayesian learning takes place in the sense that after a miss confidence intervals expand, but insufficient to obtain proper calibration. After a hit confidence intervals contract. BBGHP Confirmed Confirmed H2b A miss on the downside (upside) results in relatively more adjustment of the downside (upside) part of the confidence interval. BBGHP Confirmed Not confirmed H3 Subsequent misses lead to progressively less adjustments in confidence intervals. BBGHP Confirmed Not confirmed H3a Respondents who are most overprecise adjust the least. BBGHP Confirmed Not confirmed NPR The overall reaction of respondents will be stronger after a miss-hit sequence compared to a hit-miss sequence. BBGHP Confirmed Confirmed
497 Learning tobe overprecise Barber BM, Odean T (2000) Trading is hazardous to your wealth: the common stock investment performance of individual investors. J Financ 55(2):773–806 Barber BM, Odean T (2001) Boys will be boys: gender, overconfidence, and common stock investment. Q J Econ 116:261–292 Ben-David I, Graham JR, Harvey CR (2013) Managerial miscalibration. Q J Econ 128(4):1547–1584 Benos AV (1998) Aggressiveness and survival of overconfident traders. J Financ Mark 1(3):353–383 Biais B, Hilton D, Mazurier K, Pouget S (2005) Judgemental overconfidence, self-monitoring, and trading performance in an experimental financial market. Rev Econ Stud 72(2):287–312 Boutros M, Ben-David I, Graham JR, Harvey CR, Payne JW (2020) The persistence of miscalibration. National Bureau of Economic Research Brückbauer F, Schröder M (2023) The ZEW financial market survey panel. J Econ Stat 243(3–4):451–469 Daniel K, Hirshleifer D, Subrahmanyam A (1998) Investor psychology and security market underand overreactions. J Financ 53(6):1838–1885 Deaves R, Lüders E, Luo GY (2009) An experimental test of the impact of overconfidence and gender on trading activity. Rev Financ 13(3):575–595 Deaves R, Lüders E, Schröder M (2010) The dynamics of overconfidence: evidence from stock market forecasters. J Econ Behav Org 75:402–412 Deaves R, Lei J, Schröder M (2019) Forecaster overconfidence and market survey performance. J Behav Financ 20(2):173–194 Fellner-Röhling G, Krügel S (2014) Judgmental overconfidence and trading activity. J Econ Behav Org 107:827–842 Gervais S, Odean T (2001) Learning to be overconfident. Rev Finac Stud 14(1):1–27 Glaser M, Weber M (2007) Overconfidence and trading volume. GENEVA Risk Insur Rev 32(1):1–36 Glaser M, Langer T, Weber M (2013) True overconfidence in interval estimates: evidence based on a new measure of miscalibration. J Behav Decis Mak 26(5):405–417 Glaser M, Iliewa Z, Weber M (2019) Thinking about prices versus thinking about returns in financial markets. J Financ 74(6):2997–3039 Graham JR, Harvey CR, Huang H (2009) Investor competence, trading frequency, and home bias. Manag Sci 55(7):1094–1106 Keefer DL, Bodily SE (1983) Three-point approximations for continuous random variables. Manag Sci 29(5):595–609 Kinari Y (2016) Properties of expectation biases: optimism and overconfidence. J Behav Exp Financ 10:32–49 Kyle AS, Albert Wang F (1997) Speculation duopoly with agreement to disagree: can overconfidence survive the market test? J Financ 52(5):2073–2090 Menkhoff L, Schmeling M, Schmidt U (2013) Overconfidence, experience, and professionalism: an experimental study. J Econ Behav Org 86:92–101 Merkle C (2017) Financial overconfidence over time: foresight, hindsight, and insight of investors. J Bank Financ 84:68–87 Merkle C (2018) The curious case of negative volatility. J Financ Mark 40:92–108 Moore DA, Healy PJ (2008) The trouble with overconfidence. Psychol Rev 115(2):502–517 Odean T (1998) Volume, volatility, price, and profit when all traders are above average. J Financ 53(6):1887–1934 Odean T (1999) Do investors trade too much? Am Econ Rev 89(5):1279–1298 Puetz A, Ruenzi S (2011) Overconfidence among professional investors: evidence from mutual fund managers. J Bus Financ Account 38(5–6):684–712 Sniezek JA, Buckley T (1991) Confidence depends on level of aggregation. J Behav Decis Mak 4(4):263–272 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.