Full text
A perfect-model perspective on the signal-to-noise paradox in initialized decadal predictions Rashed Mahmood,a,b Markus G. Donat,a,c Francisco J. Doblas-Reyes,a,c Etienne Tourigny,a a Barcelona Supercomputing Center, Barcelona, Spain b National Center for Climate Research (NCKF), Danish Meteorological Institute, Copenhagen, Denmark c Catalan Institution for Research and Advanced Studies (ICREA), Barcelona, Spain Corresponding author: Rashed Mahmood, [email protected] ABSTRACT Initialized climate predictions have shown success in predicting interannual to decadal climate variations in some regions. However, the initialized predictions also suffer from different issues arising from imperfect initializations and inconsistencies between the model and the real world climate and processes. In particular, a so-called signal-to-noise paradox has been identified in recent years. The paradox implies that models can predict observations better than they predict themselves despite some physical inconsistencies between modeled and real world climate. This is often interpreted as an indicator of model deficiencies. Here we present results of a perfect-model decadal prediction experiment, where the predictions have been initialized using climate states from the model's own transient simulation. This experiment avoids issues related to model inconsistencies, initialization shock and the climate drift that affect real-world initialized climate predictions. We find that the perfect-model decadal predictions are highly skillful in predicting the near-surface air temperature and sea level pressure of the reference run on decadal timescales. Interestingly, we also find signal-to-noise issues– meaning that the perfect-model reference run is predicted with higher skill than any of the initialized prediction members. This counterintuitive result suggests that the signal-to-noise paradox may not be due just to model deficiencies in representing the observed climate in initialized predictions. We illustrate that Manuscript (non-LaTeX) 1 Early Online Release: This preliminary version has been accepted for publication in Journal of Climate, may be fully cited, and has been assigned DOI 10.1175/JCLI-D-24-0381.1. The final typeset copyedited article will replace the EOR at the above DOI when it is published. © 2025 American Meteorological Society. This is an Author Accepted Manuscript distributed under the terms of the default AMS reuse license. For information regarding reuse and general copyright information, consult the AMS Copyright Policy (www.ametsoc.org/PUBSReuseLicenses). Unauthenticated | Downloaded 05/15/25 04:01 PM UTC Addendum for Research Funded by cOAlition S Organizations CC BY 4.0 license. This is the accepted version of the following article: [https://doi.org/10.1175/JCLI-D-24-0381.1, which has been published in final form at https://journals.ametsoc.org/view/journals/clim/aop/JCLID-24-0381.1/JCLI-D-24-0381.1.xml.
2 this signal-to-noise problem , on multi-annual to decadal timescales, is related to analysis practices that concatenate time series from different discontinuous initialized simulations, which introduces inconsistencies compared to the continuous transient climate realizations and the observations. In particular, the concatenation of predictions initialized independently into a single time series breaks its auto-correlation. 1. Introduction Decadal climate predictions aim to predict the climate up to ten years into the future. These predictions consider the varying radiative forcing from changing greenhouse gases and aerosols, and are initialized towards an observational climate state with the aim to align the phasing of climate variability between the model and the real-world and also improving the forced response (e.g. Dobal-Reyes et al. 2013; Meehl et al. 2021). For this, the initialized climate predictions are performed by starting the models from reconstructed or assimilated observational data (Smith et al. 2007; Keenlyside et al. 2008; Boer et al. 2016). The individual model integrations tend to show small signal-to-noise ratios and therefore ensembles of simulations are performed by adding small perturbations to the initial conditions (Sienz et al. 2016). The initialized prediction systems have shown success in predicting interannual to decadal climate variations, with added value from initialization over the forced historical or projection simulations in some regions (e.g. Smith et al. 2007; Keenlyside et al. 2008; Doblas-Reyes et al. 2013; Kushnir et al. 2019; Athanasiadis et al. 2020; Smith et al. 2020; Delgado-Torres et al. 2022). Initialized predictions, however, also suffer from some issues that potentially reduce their skill such as model errors, imperfect initial conditions, initialization shock and the subsequent climate drift (e.g. Kharin et al. 2012; Sanchez-Gomez et al. 2016; Kröger et al. 2018; Bilbao et al. 2021). The initial shocks in the decadal predictions develop as a consequence of inconsistencies between the model attractor and the observational climate states. Thus starting the model simulations away from the model’s attractor causes fast and slow adjustments that do not represent meaningful climate trajectories due to the initialization shocks and drift. The so-called perfect-model prediction experiments, that initialize decadal predictions from climate states of a transient historical run, have been used to investigate model-specific predictability in the absence of such issues related to initialization shock and model drift (e.g. Liu et al. 2019, 2023). Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
3 A prominent issue in decadal prediction research in recent years has been the indication that the initialized climate predictions appear to be affected by the presence of a so-called ‘signal-to-noise paradox (hereafter referred to as SNP)’ at various prediction timescales (Eade et al. 2014; Dunstone et al. 2016; Scaife and Smith 2018; Smith et al. 2019). The paradox suggests that the model can predict observations with higher skill than itself, despite some physical inconsistencies between modeled and real world climate. The exact sources for the existence of the SNP in climate simulations are still being debated (Weisheimer et al. 2024) with some studies pointing towards the responsibility of fundamental model deficiencies (e.g. O’Reilly et al. 2018; Smith et al. 2019, 2020; Zhang and Kirtman 2019) and others to the uncertainties in the statistical properties of the quantities used for analysis (Weisheimer et al. 2019; Bröcker et al. 2023). Currently there is no general consensus on the origin of the SNP in the climate model simulations. This study aims to explore predictability and possible signal-to-noise issues in an idealized framework, a so-called perfect-model prediction experiment, where the model predictions are initialized from a historical climate simulation of the same model. This implies that, by definition, the predictions are initialized from climate states that are compatible with the model-specific climate attractor and therefore the predictions are not affected by shock or drift. The predictions are then evaluated against the reference run from which they were initialized (i.e. replacing the observations used in real-world predictions for initialization and evaluation by a transient simulation with the same climate model used for the predictions). This implies that the climate model used for the predictions is physically fully consistent with the climate realization it aims to predict. A similar approach was used by Liu et al. (2019) to understand limits of achievable skill in predicting climate on annual to decadal timescales albeit using a different climate model. Besides investigating the model-specific predictability, we explore possible indications of the SNP in the initialized predictions in the perfect-model experiment. As explained above, these predictions are (by definition) not affected by physical inconsistencies between the predictions and the reference, as both are from the same model and not affected by initialization shocks or climate drifts which affect the real-world climate predictions. In particular, possible signal-to-noise issues, should they exist in the perfect-model predictions, could not be explained by potential model shortcomings, as the predictions are performed with the same model against which the predictions are evaluated and initialized with perfect “observations”. The perfect-model prediction experiments provide therefore an idealized Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
4 environment to exclude or isolate specific factors, which has the potential to shed new light on model predictability and issues notoriously affecting decadal climate predictions. 2. Data and Methods The model simulations used in this study are carried out using the European Consortium’s coupled atmosphere-ocean general circulation model version 3 (EC-Earth3) with the same configurations as used for the Decadal Climate Prediction Project (DCPP) experiment (Bilbao et al. 2021) as part of the Coupled Model Intercomparison Project phase 6 (CMIP6) simulations. The EC-Earth3 model comes with different configurations including options for high, low and standard resolutions (Döscher et al. 2022). For this experiment we use here the standard resolution version of the EC-Earth3. The atmospheric component of the model is based on the European Center for Medium-Range Weather Forecasts (ECMWF) Integrated Forecast System (IFS) with ~80 km horizontal resolution and 91 vertical levels. The ocean component of the model is based on Nucleus for European Modelling of the Ocean (NEMO) version 3.6 with 1° horizontal resolution and 75 vertical levels. More details about the ECEarth3 model have been documented by e.g. Döscher et al. (2022) and Bilbao et al. (2021). First we conducted a reference historical simulation in order to generate and save the model state at the beginning of November every year, using CMIP6 historical forcings. Running a new reference historical simulation was necessary to provide initial conditions based on exactly the same version of the model that we used for the perfect model predictions. Similar to the DCPP experiments of CMIP6, also for the idealized prediction experiment the model was initialized every year from 1960 to 2014. Different to the DCPP hindcasts, we used as initial conditions the model states from the reference historical simulation instead of model reconstructions based on observations. A total of 10 ensemble members were performed by slightly (at an order of 10-5 K) perturbing the atmospheric initial conditions. The predictions were then run for up to 11 years after initialisation. An ensemble mean was computed by averaging all 10 members of the perfect model initialisation experiment (hereafter referred to as ”Perfect_Init”). For comparison we also use a ten member ensemble of DCPP initialized with observational data (hereafter referred to as “ClimPred_Init”). Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
5 While much of this study focuses on the different experiments (i.e. both Perfect_Init and ClimPred_Init hindcasts) with the EC-Earth3 model, we also analyzed a multi-model ensemble of decadal hindcasts provided within the DCPP-A component of CMIP6. We use data from a total of 8 other CMIP6 models apart from EC-Earth3 (Table S1, in supplementary material). All data sets were remapped to a uniform 5°✕5° grid before performing any analysis, following recommendations by Goddard et al. (2012). We converted all data sets to monthly anomalies using the reference climatology period of 1981 to 2010. For the initialized predictions, anomalies were computed using a lead-time dependent climatology (GarcíaSerrano and Doblas-Reyes 2012), to account for the drift in the hindcasts initialized with actual observations. We focus our analysis on forecasts years 2-9, which is similar to other studies focusing on decadal timescales (e.g. Smith et al. 2019; Scaife and Smith 2018). A total of 45 start-dates from 1960 to 2004 were used for analysis in order to use model simulations during the CMIP6 historical period. In this study, we use near-surface air temperature (tas) and mean sea level pressure (psl). The skill assessment of the Perfect_Init is carried out using the same historical simulation from which the predictions were initialized as evaluation reference, while for ClimPred_Init we use observations as evaluation reference. The observed near-surface temperature data was obtained from HadCRUT4 (Morice et al. 2012), which combines near-surface air temperature over land and sea surface temperature over ocean. For sea level pressure we used reanalysis data from ERA5 (Hersbach et al. 2020) as observational reference. For skill assessments we consider the anomaly correlation coefficient (ACC), which measures the agreement between two time series (i.e. the prediction ensemble mean and a reference) with focus on common variations rather than their magnitudes. Since the ACC can be strongly affected by the warming trend, the added value from initialisation of the climate predictions is also assessed by calculating the residual correlations. These correlate the predicted and observed residuals after removing an estimate of the forcing response following Smith et al. (2019). Here we estimate the forcing response from the ensemble mean of ten historical simulations of the EC-Earth3 model. We also analyze the ratio of predictable components (RPC) following (Scaife and Smith 2018) to explore the SNP in different ensembles. Specifically we computed the RPC by taking the square root of the ratio of squared correlations between the ensemble mean and the evaluation reference and the ensemble mean with individual members. According to Scaife and Smith (2018), the RPC can be written as: Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
6 RPC = √𝑟(𝑒𝑚,𝑜) 2 𝑟(𝑒𝑚,𝑚𝑒𝑚𝑏) 2 (1) Where r2(em,o) is the correlation between the model ensemble mean with the observation (or the reference run in the perfect-model case), and r2(em,memb) is the correlation between the model ensemble mean with a member (not included in the ensemble mean). For r2(em,memb) we use the median of the correlations between the model ensemble mean and the individual members. The initialized hindcast studies typically construct time series by concatenating the relevant forecast time averages from the different (annual) initializations. These are typically concatenated for the ensemble members with the same identifier (e.g. r1i1p1f1 from different initialization times); however the different initializations with the same identifier are in fact independent of each other. The resulting hindcast time series for a member identifier is therefore based on several discontinuous simulation chunks, and this discontinuity could potentially affect the correlation between the ensemble mean and the individual ensemble members. We will revisit this issue in section 3.2. 3. Results 3.1 Skill of the initialized predictions We first evaluate the (potential) skill of the different ensembles in simulating near-surface air temperature (Figure 1) and sea level pressure (Figure 2). We find high ACC values for the Perfect_Init over most of the globe, suggesting that the perfectly initialized ensemble is highly skillful in predicting its reference run (Figure 1a). The skill in Perfect_Init extends to most regions of the globe including some regions where the hindcasts initialized with observations with the same EC-Earth3 model do not show much skill, such as the northeast and southeast Pacific and several land regions (cf Figure 1a and 1d). Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
7 Fig. 1. Skill assessment of different ensembles in simulating near-surface air temperature for the average of forecast years 2-9 based on ACC (left column), residual correlation (middle column), and the ratio of predictable components (right column). The top row shows results for Perfect_Init and the bottom row for ClimPred_Init. The stippling represents regions where ACC (a and d) and residual correlations (b and e) are not statistically significant at the 95% confidence level. For RPC (c and f) the regions where ACC is negative are stippled. The black markers indicate the grid point used for plotting near-surface air temperature and SLP time series in Fig 3. Similarly for SLP we find that the Perfect_Init experiment shows higher and widespread skill (in predicting its reference run) over many regions of the globe compared to the skill of the real world predictions (cf Figure 2a and 2d). No skill is found in the real world predictions for SLP over the tropical and subtropical Indian and Atlantic oceans and the bordering continental regions. Perfect_Init, on the other hand, shows positive ACC values over most of these regions although at some locations ACC is not statistically significant at the 95% confidence level. These results indicate that there may be some decadal-scale predictability in the climate system of the EC-Earth3 model which is currently not being captured by the real-world (i.e. ClimPred_Init) predictions. Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
8 Fig. 2. Same as Fig. 1 but for mean sea level pressure (SLP). The skill measure based on ACC can be strongly affected by the response to external forcings in both the Perfect_Init and the ClimPred_Init predictions. To evaluate the added value from initializing the climate predictions, Figure 1b and 1e show residual correlations obtained after removing an estimate of the forcing response following Smith et al. (2019). We find more widespread added value from initialization in the Perfect_Init compared to the ClimPred_Init predictions. For example, over the Atlantic, Indian, parts of the Pacific oceans and some neighboring continental areas significant added skill in Perfect_Init is found, while the ClimPred_Init only shows added skill over the north Atlantic subpolar gyre and central north Pacific. Similarly for sea level pressure the Perfect_Init predictions show higher and more widespread positive residual correlation values compared to ClimPred_Init (Figure 2b and 2e). The real-world predictions suffer strongly from the degradation of the skill especially in the Atlantic and the Indian oceans hence no added value (and even negative residual correlations) is found in these regions in ClimPred_Init. We note that the skill (and added skill) in the perfect-model predictions does not necessarily imply predictability of the real world. While the perfect-model certainly avoids some issues that persist in affecting the real-world predictions, the deterioration of skill in the real-world predictions can also be due to inconsistencies between the model and the real-world climate, for example related to misrepresenting some key processes. Still, these results indicate the predictability that can be achieved with the EC-Earth3 model in the case that perfect initial conditions were available Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
9 (i.e. perfect knowledge of the initial state in each climate system variable) and the model being physically consistent with the reference to be predicted. 3.2 Diagnosing the signal-to-noise paradox The ratio of predictable components has been used in the past to diagnose the paradoxical behavior that climate predictions tend to predict the observations with higher skill than they predict an individual realization of the prediction model (e.g. Scaife and Smith 2018, Smith et al. 2019). This issue is indicated by RPC values larger than 1, and has often been interpreted as indication of some model shortcomings. It is therefore interesting to analyze the RPC in the perfect-model predictions, as these, by definition, are free of the model shortcomings in comparison to the perfect-model reference (because they are all part of the same model world). Both the ClimPred_Init and Perfect_Init show areas with RPC values larger than 1 for near-surface air temperature (Figure 1c and 1f) and sea level pressure (Figure 2c and 2f), implying a counterintuitive suggestion that the model can predict the observations and even the perfect-model reference better than another model member used to make the predictions (note that RPC is only meaningful in areas where the ACC is not negative). While we find RPC values greater than one in case of near-surface air temperature also, however, the paradox is relatively more prominent when predicting sea level pressure, where RPC values can reach 2 or even higher. Such large RPC is often discussed as an indication of model deficiencies (e.g. Scaife and Smith 2018; Smith et al. 2019), however in the case of perfect-model predictions the model used to make the predictions is physically consistent with the reference and therefore there cannot be a deficiency regarding the physical representation of climate. To develop an understanding of the behavior of the different predictions we first inspect the time series at one arbitrary grid point in the Pacific region as shown in Figure 3. It is immediately apparent how the individual ensemble members show very different characteristics compared to both the ensemble mean and the evaluation references - for both the Perfect_Init and the ClimPred_Init. The initialized ensemble members appear much noisier compared to both the transient reference simulation and also the observational reference time series. This is in agreement with Athanasiadis et al. (2020) who also noted that the time series generated from discontinuous initialized hindcasts are generally noisier than the observational reference. As also mentioned by Athanasiadis et al. (2020), the 8 year averages used for the evaluating references (i.e. observation and the transient climate Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
16 indicate that the large RPC values, which are often used to diagnose the SNP in the initialized climate predictions, could be a result of concatenating individual member time series from independent initializations with the same ensemble member identifier (such as ensemble member #1 from the independent simulations with initialization in 1960, 1961, 1962, etc). 4 Discussion and Conclusions We find that the perfect-model predictions show more widespread (potential) skill in predicting model-specific near-surface air temperature and sea level pressure over most regions of the globe compared to the real-world predictions in predicting their corresponding references. Although it is well understood that the higher skill of a perfect model prediction cannot be directly translated to the skill of real-world predictions, the results from these experiments suggest that the prediction can potentially be improved by providing model consistent and accurate initial-conditions (Liu et al. 2019) and improving the model quality. We also find that the added value from initialization in the perfect-model simulations extends to large regions globally while it remains very limited in the real world predictions. These results indicate that there may be predictability in the climate system on decadal timescales that is currently not being captured by the realworld prediction system, but could also point to fundamental differences between the model-specific and the real-world climate systems. New and improved initialization methods may provide more skillful global climate predictions on decadal timescales if the model and real-world climate are closer in terms of relevant processes. Similar to previous studies (e.g. Scaife and Smith 2018), we also find that the SNP persists in the latest versions of the initialized climate predictions especially on multi-annual to decadal (and longer) timescales, diagnosed by RPC values larger than 1, and suggesting that the models can predict the real world climate with higher skill than they can predict another model realization. Surprisingly, the perfect model predictions also show large areas with RPC>1, implying that the model is better in predicting its reference run than the individual members. These results suggest that the SNP is not likely dominated by the issues arising from imperfect initialization of the real-world predictions, but also that the SNP is not necessarily an indicator of model deficiencies if it also occurs in a perfect-model experiment. Indeed, we show here that the concatenation of predictions corresponding to initializations in consecutive and somehow independent years into a single time series breaks the autocorrelation of the resulting time series. The measures that consider such individual member Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
17 trajectories, such as the RPC, can therefore be misleading – and should not be interpreted as an indicator for potential model deficiencies. In the RPC case, the changes in auto-correlation result in lower values for the ensemble mean correlation with an individual member (which is affected by the broken auto-correlation after concatenating predictions initialized at different point in time and assigned to the same ensemble member identifier), compared to its correlation with the reference (which is not affected by the discontinuity introduced by the concatenation). We also find that the unrealistic changes in the auto-correlation of individual members of an initialized prediction system is not limited to the EC-Earth3 model. We find similarly lower auto-correlations in other decadal climate predictions of the CMIP6 decadal prediction ensembles, confirming that the practice of concatenating simulations with different initial start dates into a time series assigned to an ensemble member leads to lower correlation of the ensemble mean with individual members. Therefore, our results provide an alternative explanation for the SNP in initialized climate predictions compared to previous studies that suggested differences in initial ensemble and observation spreads (Mayer et al. 2021), overall model deficiencies (e.g. Smith et al. 2019; Scaife and Smith 2018; Strommen and Palmer 2018), differences in statistical properties of the quantities used for analysis (e.g. Weisheimer et al. 2019; Bröcker et al. 2023), and limited sample size (Weisheimer et al. 2019). It is likely that a combination of factors play a role in causing the SNP. Our study highlights the importance of artificially reducing the auto-correlation of individual ensemble members when combining data from different discontinuous initializations (e.g. with the same ensemble member identifier) into a single time series. While this issue of constructing time series from discontinuous simulations also affects seasonal hindcasts (which are initialized at specific dates every year), the described effects on the auto-correlation of the time series are larger for multi-annual to decadal (or longer) predictions than for e.g. seasonal predictions (Figure S5 and S6). However, these effects become relevant as soon as we average over multiple years (as shown for two-year predictions in Figures S7, S8). That is because when calculating multi-annual averages in continuous time series (e.g. observations or our transient reference simulation), these moving windows have several data points in common, which leads to high auto-correlations. But when calculating multi-year averages from discontinuous initialised simulations, even temporally overlapping averages may not have data points in common as they come from different independent simulations. In the specific example of the 8-year averages presented in Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
18 this study, the reference time series include seven annual data points which are common among two consecutive 8-year averages. This is not the case when calculating these multiannual averages from the discontinuous simulations initialized independently in each year, for example. Our finding that these auto-correlation issues affect predictions averaging over several forecast years may explain that the signal-to-noise issues (as quantified e.g. by RPC>1) are typically larger in decadal than in seasonal predictions (Eade et al. 2014; Smith et al. 2019). Athanasiadis et al. (2020) pointed out a similar issue of time series from initialized hindcasts being noisier than their observational references and they applied an additional temporal smoothing using seven-point running averages across the start dates (i.e. in addition to the smoothing introduced by the 8-year average predictions) to address this symptom. This approach has not been used here because averaging predictions across the start dates makes the nature of the prediction and the observational reference very different. This additional averaging leads to a forecast product that, although smoother, might not have more skill because older predictions are included in it. Smith et al. (2020) constructed very large ensembles by combining the initialized simulations from several (lagged) start dates. While the primary rationale for this approach was to substantially increase the ensemble size, a collateral effect of this approach could be to increase the temporal smoothness of the time series (for the ensemble mean at least), due to using the information from several consecutive initializations. The time series constructed from concatenating simulations with different start dates into a specific ensemble member does not represent physically meaningful continuous climate trajectories, contrary to the reference time series used in the validation. This implies that diagnostic metrics that consider e.g. the correlation of an ensemble mean with the individual ensemble member time series, such as the RPC, will be affected by this artifact from the concatenation of different (discontinuous) initialized simulations. As a consequence, large RPC values (or low correlations of a prediction ensemble average with individual initialized ensemble member time series) should not necessarily be interpreted as an indicator of potential model deficiencies. While there is large evidence that the initialized climate predictions underestimate the magnitude of the prediction signal, in particular for atmospheric circulation (e.g. Scaife and Smith 2018; Smith et al. 2020), our results suggest that the so-called paradox (that the realworld climate is predicted with higher skill than an individual ensemble member) may be an analysis-related artifact. Our study shows that this paradoxical effect is largely a consequence Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
19 of artificially changing the statistical characteristics of the individual initialized ensemble members by combining the information from different independent initializations. While the so-called paradox may be a consequence of how we process and analyze the data, the problem of the small predictable signals relative to noise remains (Siegert et al. 2016; Smith et al. 2020). This signal-to-noise issue may be rooted in models not capturing relevant processes, and the underestimation of the signal magnitude is not paradoxical as such. Future climate model developments and improvements are hoped to resolve this issue of small signal-to-noise ratios. Acknowledgments. This research contributed to the Spanish Ministry for Science and Innovation projects PRECEDE (grant no. EUR2022-134059) and PATHFINDER (grant no. PID2021127943NB-I00). We are also grateful for partial support by the Horizon Europe projects ASPECT (grant number 101081460) and Impetus4Change (grant number 101081555) and support by the Departament de Recerca i Universitats de la Generalitat de Catalunya for the Climate Variability and Change (CVC) Research Group (Reference: 2021 SGR 00786). High-performance computing resources used to perform some of the experiments were obtained from the ECMWF special project spesiccf-2024 “Understanding inter-annual to decadal predictability in the EC-Earth3 model and the potential benefits from perfect initialisation” as well as PRACE (HiRes-NTCP, project 3: grant no. 2017174177) and the Red Española de Supercomputación (AECT-2019-2-0003 and AECT-2019-3-0006 projects). Data Availability Statement. All model simulation data (except Perfect model simulations) used in this study are freely available from ESGF nodes (e.g. https://esgf-node.ipsl.upmc.fr/) and observational and reanalysis data are available from their respective sources. The perfect model simulation and the corresponding reference run data can be made available upon request. REFERENCES Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
20 Athanasiadis, P. J., Yeager, S., Kwon, Y.-O., Bellucci, A., Smith, D. W., & Tibaldi, S. (2020). Decadal predictability of North Atlantic blocking and the NAO. Npj Climate and Atmospheric Science, 3(1), 20. https://doi.org/10.1038/s41612-020-0120-6 Boer, G. J., Smith, D. M., Cassou, C., Doblas-Reyes, F., Danabasoglu, G., Kirtman, B., Kushnir, Y., Kimoto, M., Meehl, G. A., Msadek, R., Mueller, W. A., Taylor, K. E., Zwiers, F., Rixen, M., Ruprich-Robert, Y., & Eade, R. (2016). The Decadal Climate Prediction Project (DCPP) contribution to CMIP6. Geoscientific Model Development, 9(10), 3751–3777. https://doi.org/10.5194/gmd-9-3751-2016 Bröcker, J., Charlton–Perez, A. J., & Weisheimer, A. (2023). A statistical perspective on the signal‐to‐noise paradox. Quarterly Journal of the Royal Meteorological Society, 149(752), 911–923. https://doi.org/10.1002/qj.4440 Doblas-Reyes, F. J., Andreu-Burillo, I., Chikamoto, Y., García-Serrano, J., Guemas, V., Kimoto, M., Mochizuki, T., Rodrigues, L. R. L., & Van Oldenborgh, G. J. (2013). Initialized near-term regional climate change prediction. Nature Communications, 4(1), 1715. https://doi.org/10.1038/ncomms2704 Döscher, R., Acosta, M., Alessandri, A., Anthoni, P., Arsouze, T., Bergman, T., Bernardello, R., Boussetta, S., Caron, L.-P., Carver, G., Castrillo, M., Catalano, F., Cvijanovic, I., Davini, P., Dekker, E., Doblas-Reyes, F. J., Docquier, D., Echevarria, P., Fladrich, U., … Zhang, Q. (2022). The EC-Earth3 Earth system model for the Coupled Model Intercomparison Project 6. Geoscientific Model Development, 15(7), 2973–3020. https://doi.org/10.5194/gmd-15-2973-2022 Dunstone, N., Smith, D., Scaife, A., Hermanson, L., Eade, R., Robinson, N., Andrews, M., & Knight, J. (2016). Skilful predictions of the winter North Atlantic Oscillation one year ahead. Nature Geoscience, 9(11), 809–814. https://doi.org/10.1038/ngeo2824 Eade, R., Smith, D., Scaife, A., Wallace, E., Dunstone, N., Hermanson, L., & Robinson, N. (2014). Do seasonal‐to‐decadal climate predictions underestimate the predictability of the real world? Geophysical Research Letters, 41(15), 5620–5628. https://doi.org/10.1002/2014GL061146 García-Serrano, J., & Doblas-Reyes, F. J. (2012). On the assessment of near-surface global temperature and North Atlantic multi-decadal variability in the ENSEMBLES decadal Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
21 hindcast. Climate Dynamics, 39(7–8), 2025–2040. https://doi.org/10.1007/s00382-0121413-1 Hawkins, E., & Sutton, R. (2009). The Potential to Narrow Uncertainty in Regional Climate Predictions. Bulletin of the American Meteorological Society, 90(8), 1095–1108. https://doi.org/10.1175/2009BAMS2607.1 Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A., Muñoz‐Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., … Thépaut, J. (2020). The ERA5 global reanalysis. Quarterly Journal of the Royal Meteorological Society, 146(730), 1999–2049. https://doi.org/10.1002/qj.3803 Keenlyside, N. S., Latif, M., Jungclaus, J., Kornblueh, L., & Roeckner, E. (2008). Advancing decadal-scale climate prediction in the North Atlantic sector. Nature, 453(7191), 84–88. https://doi.org/10.1038/nature06921 Kharin, V. V., Boer, G. J., Merryfield, W. J., Scinocca, J. F., & Lee, W. ‐S. (2012). Statistical adjustment of decadal predictions in a changing climate. Geophysical Research Letters, 39(19), 2012GL052647. https://doi.org/10.1029/2012GL052647 Kröger, J., Pohlmann, H., Sienz, F., Marotzke, J., Baehr, J., Köhl, A., Modali, K., Polkova, I., Stammer, D., Vamborg, F. S. E., & Müller, W. A. (2018). Full-field initialized decadal predictions with the MPI earth system model: An initial shock in the North Atlantic. Climate Dynamics, 51(7–8), 2593–2608. https://doi.org/10.1007/s00382-017-4030-1 Liu, Y., Donat, M. G., Taschetto, A. S., Doblas‐Reyes, F. J., Alexander, L. V., & England, M. H. (2019). A Framework to Determine the Limits of Achievable Skill for Interannual to Decadal Climate Predictions. Journal of Geophysical Research: Atmospheres, 124(6), 2882–2896. https://doi.org/10.1029/2018JD029541 Liu, Y., Donat, M. G., England, M. H., Alexander, L. V., Hirsch, A. L., & Delgado-Torres, C. (2023). Enhanced multi-year predictability after El Niño and La Niña events. Nature Communications, 14(1), 6387. https://doi.org/10.1038/s41467-023-42113-9 Mayer, B., Düsterhus, A., & Baehr, J. (2021). When Does the Lorenz 1963 Model Exhibit the Signal‐To‐Noise Paradox? Geophysical Research Letters, 48(4), e2020GL089283. https://doi.org/10.1029/2020GL089283 Morice, C. P., Kennedy, J. J., Rayner, N. A., & Jones, P. D. (2012). Quantifying uncertainties in global and regional temperature change using an ensemble of observational estimates: Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
22 The HadCRUT4 data set. Journal of Geophysical Research: Atmospheres, 117(D8), 2011JD017187. https://doi.org/10.1029/2011JD017187 O’Reilly, C. H., Weisheimer, A., Woollings, T., Gray, L. J., & MacLeod, D. (2019). The importance of stratospheric initial conditions for winter North Atlantic Oscillation predictability and implications for the signal‐to‐noise paradox. Quarterly Journal of the Royal Meteorological Society, 145(718), 131–146. https://doi.org/10.1002/qj.3413 Sanchez-Gomez, E., Cassou, C., Ruprich-Robert, Y., Fernandez, E., & Terray, L. (2016). Drift dynamics in a coupled model initialized for decadal forecasts. Climate Dynamics, 46(5–6), 1819–1840. https://doi.org/10.1007/s00382-015-2678-y Scaife, A. A., & Smith, D. (2018). A signal-to-noise paradox in climate science. Npj Climate and Atmospheric Science, 1(1), 28. https://doi.org/10.1038/s41612-018-0038-4 Sienz, F., Müller, W. A., & Pohlmann, H. (2016). Ensemble size impact on the decadal predictive skill assessment. Meteorologische Zeitschrift, 25(6), 645–655. https://doi.org/10.1127/metz/2016/0670 Siegert, S., Stephenson, D. B., Sansom, P. G., Scaife, A. A., Eade, R., & Arribas, A. (2016). A Bayesian Framework for Verification and Recalibration of Ensemble Forecasts: How Uncertain is NAO Predictability? Journal of Climate, 29(3), 995–1012. https://doi.org/10.1175/JCLI-D-15-0196.1 Smith, D. M., Scaife, A. A., Eade, R., Athanasiadis, P., Bellucci, A., Bethke, I., Bilbao, R., Borchert, L. F., Caron, L.-P., Counillon, F., Danabasoglu, G., Delworth, T., DoblasReyes, F. J., Dunstone, N. J., Estella-Perez, V., Flavoni, S., Hermanson, L., Keenlyside, N., Kharin, V., … Zhang, L. (2020). North Atlantic climate far more predictable than models imply. Nature, 583(7818), 796–800. https://doi.org/10.1038/s41586-020-2525-0 Smith, D. M., R. Eade, A. A. Scaife, L. P. Caron, G. Danabasoglu, T. M. DelSole, T. Delworth, F. J. Doblas-Reyes, N. J. Dunstone, L. Hermanson, V. Kharin, M. Kimoto, W. J. Merryfield, T. Mochizuki, W. A. Müller, H. Pohlmann, S. Yeager, and X. Yang. 2019. “Robust Skill of Decadal Climate Predictions.” Npj Climate and Atmospheric Science 2(1):13. doi: 10.1038/s41612-019-0071-y Smith, D. M., Cusack, S., Colman, A. W., Folland, C. K., Harris, G. R., & Murphy, J. M. (2007). Improved Surface Temperature Prediction for the Coming Decade from a Global Climate Model. Science, 317(5839), 796–799. https://doi.org/10.1126/science.1139540 Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC
23 Strommen, K., & Palmer, T. N. (2019). Signal and noise in regime systems: A hypothesis on the predictability of the North Atlantic Oscillation. Quarterly Journal of the Royal Meteorological Society, 145(718), 147–163. https://doi.org/10.1002/qj.3414 Weisheimer, A., Baker, L. H., Bröcker, J., Garfinkel, C. I., Hardiman, S. C., Hodson, D. L. R., Palmer, T. N., Robson, J. I., Scaife, A. A., Screen, J. A., Shepherd, T. G., Smith, D. M., & Sutton, R. T. (2024). The Signal-to-Noise Paradox in Climate Forecasts: Revisiting Our Understanding and Identifying Future Priorities. Bulletin of the American Meteorological Society, 105(3), E651–E659. https://doi.org/10.1175/BAMS-D-24-0019.1 Weisheimer, A., Decremer, D., MacLeod, D., O’Reilly, C., Stockdale, T. N., Johnson, S., & Palmer, T. N. (2019). How confident are predictability estimates of the winter North Atlantic Oscillation? Quarterly Journal of the Royal Meteorological Society, 145(S1), 140–159. https://doi.org/10.1002/qj.3446 Zhang, W., & Kirtman, B. (2019). Understanding the Signal‐to‐Noise Paradox with a Simple Markov Model. Geophysical Research Letters, 46(22), 13308–13317. https://doi.org/10.1029/2019GL085159 Accepted for publication in Journal of Climate. DOI 10.1175/JCLI-D-24-0381.1. Unauthenticated | Downloaded 05/15/25 04:01 PM UTC