scieee AI-readable full text Open interactive document viewer

Forecasting Turkish Industrial Production Growth With Static Factor Models

Günay, Mahmut

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Günay, Mahmut Article Forecasting Turkish Industrial Production Growth With Static Factor Models International Econometric Review (IER) Provided in Cooperation with: Econometric Research Association (ERA), Ankara Suggested Citation: Günay, Mahmut (2015) : Forecasting Turkish Industrial Production Growth With Static Factor Models, International Econometric Review (IER), ISSN 1308-8815, Econometric Research Association (ERA), Ankara, Vol. 7, Iss. 2, pp. 64-78, https://doi.org/10.33818/ier.278041 This Version is available at: https://hdl.handle.net/10419/238817 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/ Günay-Forecasting Turkish Industrial Production Growth With Static Factor Models 64 Forecasting Turkish Industrial Production Growth With Static Factor Models Mahmut Günay Yıldırım Beyazıt University and The Central Bank of the Republic of Turkey ABSTRACT In this paper, we forecast industrial production growth for the Turkish economy using static factor models. We evaluate how the performance of the models change based on the number of factors we extract from our data as well as the level of aggregation for the series in the data set. We consider two evaluation samples for the out-of-sample forecasting exercise to assess the stability of the forecasting performance. We find that the effect of the data set size on the forecasting performance is not independent from the number of factors extracted from this data set. Rankings of the models change in different evaluation samples. We conclude that using a dynamic approach to evaluate models from different dimensions is important in the forecasting process. Key words: Forecasting, Factor Models, Principal Components JEL Classifications: E37, C32, C33 1. INTRODUCTION A quote attributed to the Nobel laureate Niels Bohr states that “prediction is very difficult, especially if it is about future”. However, forecasts of the key macro variables are vital for real time policy making due to lags in the transmission mechanism. Since we are living in a stochastic world, in general, realizations will be different from predictions and time to time by a high margin. Hence, over an evaluation period, it would be unrealistic to expect zero forecast errors from a forecasting model. In this respect, efficiency of forecasts is as important as accuracy. Inefficiency can occur due to various reasons, such as not using an indicator that has adequate forecasting power in the prediction process, not using a modelling technique that is known at the time of forecasting, or not considering the appropriate parameters in the models. Hence, it is important for forecasters to check whether all information in the economy is utilized to the greatest extent possible and in an efficient way. In this paper, we approach the issue from two perspectives: how to utilize the wide range of available data and understand the effect of model specification on forecasting performance. There are a lot of candidate indicators that can be used in the forecasting, and this number is increasing with the advances in information technology. Due to increasing connectedness within the global economy, considering international data in addition to domestic indicators in the forecasting of local variables may be necessary. However, one can use only a limited number of variables in an OLS or VAR type forecasting model due to the degrees of freedom problem. Stock and Watson (2002a:147) state that some variable selection procedures may be  Mahmut Günay, Yıldırım Beyazıt University and The Central Bank of the Republic of Turkey, (email: [email protected]). This work is part of the PhD thesis that is going on at the Department of Economics in the Yıldırım Beyazıt University under the supervision of Asst. Prof. Sıdıka Başçı. The views attributed in this study are those of the author and cannot necessarily be attributed to the Yıldırım Beyazıt University or the Central Bank of the Republic of Turkey. International Econometric Review (IER) 65 used for determining the forecasting model, but the performance rests on the few variables chosen. Hence, forecasters need techniques that enable them to use large amounts of data in the forecasting model. Factor models became popular in the last decade for dealing with large data. In factor models, information in a large data set is summarized with a few underlying factors and then these factors are used in the forecasting equation (Stock and Watson, 2002a and 2002b). Factor models enable us to incorporate as many series as we want in the forecasting process, but, there may not be a linear relation between forecasting performance and the number of series we use for extracting factors. Also, the number of factors we extract from a given data set may affect the forecasting performance. Hence, analyzing the effect of modelling decisions in factor models on forecast performance may provide valuable information to the forecasters. Factor model approach is a tool that enables us summarize information in a, possibly large, data set with few underlying factors. The basic rationale of factor models is presented in Equation 3.1. We decompose each series (X) into a part that is explained by the factors (in the jargon of factor models, common part) and to a part that is specific to the series (in the factor model jargon, idiosyncratic component). ittiit eFX     (3.1) where X is a stationary series; F, is a vector of factors and  , lambda is factor loadings. Equation 3.1 is a theoretical representation of a factor model but we have several issues to address when applying these models in practice. First of all, we do not observe factors. There are methods to extract factors, but they require some parameters as input. Another issue is that we need to construct a data set. With the advance of information technology, the cost of accessing information has decreased considerably, and we can gather large amounts of data relatively easily. Increasing the size of data set, however, may not always improve forecasting power. Thus, we need to analyze the effect of the data set structure on the performance of the models. Once we decide the data set, we need to choose how many factors to extract from this set. In a meta-analysis where the results of papers on factor models are analyzed, Eickmeier and Ziegler (2008) find that the relative performance of the factor models depends on several things such as the variable that is forecast, the country of the study, and the size of the data set. Hence, although factor models let us use large amounts of data in the forecasting, there is no guarantee for obtaining a better forecast than when using simpler methods. In this respect, we think careful analysis of the sensitivity of the forecasting performance to the modeling choices is necessary. Boivin and Ng (2005) is an example for an effort in this direction. They analyze direct and indirect approach and different factor extraction methods. They show that factor model specification indeed affects the forecasting performance. In this paper, we forecast industrial production growth for the Turkish economy using a large number of indicators from different blocks of data with factor models. Our data set covers indicators from production, foreign trade, financial variables, confidence indicators, interest rates, commodity prices, and international variables. We set up three different data sets from these categories to see how the forecasting performance changes with data set size. The number of factors from these data sets is obtained with the different criteria suggested in the literature. Günay-Forecasting Turkish Industrial Production Growth With Static Factor Models 66 We find that modelling choices such as the number of factors and size of the data set affect the forecasting performance of factor models. More importantly, the effect of the data set size on the forecasting performance is not independent from the number of factors used in the forecasting. Another finding is that the evaluation sample of the models may play a considerable role on the relative performance of the models. In this respect, the forecasting models to be used in the future must be selected with caution based on past performance. All in all, our results point out the importance of considering, continuously, all the dimensions of modelling for efficient forecasting. In the next sections we describe the data and summarize the methodology we use in the paper; we then present our results and conclude. 2. DATA A critical issue that a forecaster needs to address before setting up the forecasting model is the composition of the data set. This choice is even more important in the case of factor models since we can use as many series as we can collect for extracting the factors. Yet, there is no consensus on the ideal number of series or on the distribution of indicators from different blocks in the data set from which the factors are extracted. For example, Rünstler et al. (2009) forecast GDP growth using large data sets for several European economies. The number of series used for different countries in Rünstler et al. (2009) ranges from 76 to 393. Moreover, the distribution of the data in different blocks changes considerably. For instance, they do not use any price variable for Euro Area but use 42 price series for Belgium. Boivin and Ng (2006) note that adding more data may not always be useful for forecasting. They find that factors extracted from 40 pre-selected variables may yield better forecasting performance than using 147 series for factor extraction. Hence, the composition of the data set may have some effect on the forecasting performance. Deciding whether to use aggregated or disaggregated data and determining the level of detail for the disaggregation is another key issue that a forecaster faces when constructing a data set. For example, we have data on industrial production as headline index; in MIGS (Main Industrial Groupings) we see industrial production as the sum of intermediate goods, consumer goods, investment goods, and energy. In another classification, we see a more detailed picture of industrial production, such as production of food, textile, and so on for about 20 different sectors. A similar picture arises for soft data. We can use consumer confidence as the headline index, or we can also consider subcomponents, which are questions about the recent state of the economy as well as about expectations. Angelini et al. (2010) use series from different detail levels in the same data set. On the other hand, Barhouimi et al. (2010) use different data sets depending on the level of detail. Barhoumi et al. (2010) find that Stock and Watson (2002b)’s static approach with a small data set, which uses headline series rather than subcomponents, led to competitive results. We follow the same approach as Barhoumi et al. (2010) and construct three data sets with different aggregation levels resulting in three data sets: small (22 series,), medium (63 series), and large (167 series). In the small data set, we use only headline growth in industrial production. In the medium data set we adopt the definition of industrial production as the sum of five categories defined by the MIGS classification. Hence, we use the growth rate of the each of these five items in the data set. In the large data set, we use a more detailed disaggregated sectoral classification for industrial production. Table 2.1 demonstrates the increasing level of detail described above. International Econometric Review (IER) 67 We include data about industrial production, foreign trade, consumer and business confidence, interest rates, exchange rates, European Union industrial production and confidence indicators, commodity prices, stock exchange, and global risk perception indicators. Details about these sets are provided in the Appendix (Table A.4 to Table A.6). The series are transformed by taking logs, if appropriate, and first differenced to ensure stationarity. For the series that exhibit seasonality, we use seasonally adjusted series. In the pseudo out-of-sample forecasting exercise, we standardize data at each point before extracting factors. Small Data Set Medium Data Set Large Data Set Industrial Production Intermediate Mining Capital Food Non-durable Beverage Durable Tobacco Energy Textile Apparel Leather Wood Paper Media Refined petroleum Chemical Pharmaceutical Rubber Other Mineral Basic Metal Fabricated Metal Electronic and Optical Electrical Equipment Machinery and Equipment Motor Vehicles Other Transport Furniture Other manufacturing Repair of mach-eq Electricity, gas and steam Table 2.1 Example of Increasing Detail: Case of Industrial Production. Notes: We show an example of increasing detail level of the data set. In the small data set we use headline, in the medium data set we use MIGS classification and in the large data set we use a more disaggregated sectoral detail. 3. METHODOLOGY In this paper, we use factor models for forecasting Turkish industrial production (cumulative) growth for 3 and 12 month-ahead with three types of data sets; small, medium, and large. In this section, we provide details about our modelling strategies. 3.1. Factor Extraction We extract factors with principal components as suggested by Stock and Watson (2002b). Factors can be obtained by using the formula in Equation 3.2. We obtain the eigenvectors of X'X, and using the eigenvectors corresponding to the largest r eigenvalues (we will discuss how we set r below in more detail) we can obtain factors. Günay-Forecasting Turkish Industrial Production Growth With Static Factor Models 68 N x Ft   ˆ ˆ ˆ (3.2)  ˆ = eigenvectors of X'X corresponding tor largest eigenvalues where r is the number of factors. Figure 3.1 Principal Components a. First Principal Components b. Second Principal Components c. Third Principal Components d. Fourth Principal Components Notes: We show the principal components that we obtain from three different data sets namely, small, medium and large. As we discussed in the data section, we work with three different data sets. In Figures 3.1.a to 3.1.d, we plot the first four principal components from these data sets. These principal components are the factors that we will use in the forecasting. Figure 3.1.a shows the first principal component from each of the three data sets. We see that they show similar patterns over the sample. Second, principal components are also fairly similar for medium and small data sets. When we go to the fourth principal components, the series become less similar. These observations suggest that, the number of factors we use in the forecasting may affect -20 -15 -10 -5 0 5 10 2005M02 2005M07 2005M12 2006M05 2006M10 2007M03 2007M08 2008M01 2008M06 2008M11 2009M04 2009M09 2010M02 2010M07 2010M12 2011M05 2011M10 2012M03 2012M08 2013M01 2013M06 2013M11 2014M04 2014M09 Factor_Large_1 Factor_Medium_1 Factor_Small_1 -25 -20 -15 -10 -5 0 5 10 15 2005M02 2005M07 2005M12 2006M05 2006M10 2007M03 2007M08 2008M01 2008M06 2008M11 2009M04 2009M09 2010M02 2010M07 2010M12 2011M05 2011M10 2012M03 2012M08 2013M01 2013M06 2013M11 2014M04 2014M09 Factor_Large_2 Factor_Medium_2 Factor_Small_2 -15 -10 -5 0 5 10 2005M02 2005M07 2005M12 2006M05 2006M10 2007M03 2007M08 2008M01 2008M06 2008M11 2009M04 2009M09 2010M02 2010M07 2010M12 2011M05 2011M10 2012M03 2012M08 2013M01 2013M06 2013M11 2014M04 2014M09 Factor_Large_3 Factor_Medium_3 Factor_Small_3 -10 -5 0 5 10 2005M02 2005M07 2005M12 2006M05 2006M10 2007M03 2007M08 2008M01 2008M06 2008M11 2009M04 2009M09 2010M02 2010M07 2010M12 2011M05 2011M10 2012M03 2012M08 2013M01 2013M06 2013M11 2014M04 2014M09 Factor_Large_4 Factor_Medium_4 Factor_Small_4 International Econometric Review (IER) 69 the conclusion about the effect of the size of the dataset. In particular, if we use only one factor, forecasts may be quite similar for three data sets. But if we use more than 2 factors, we may get different forecasts. In this respect, in the next section we discuss the choice of the number of factors. Figure 3.2 Number of Factors Obtained from Different Information Criteria a. Number of Factors Obtained from Different Information Criteria from Medium Data Set b. Number of Factors Obtained from Different Information Criteria from Large Data Set Notes: In the paper, we do recursive out–of-sample forecasting exercise. In the evaluation sample, at each month we get number of factors proposed by different criteria from Bai and Ng (2002). X axis shows the date where data set that we extract factors from ends. For small data set we set the maximum number of factors (required for PC1, PC2 and PC 3) as four, for medium data set as seven and for large data set as nine. 3.2. Number of Factors In the Stock and Watson (2002b) approach to factor extraction, there is only one parameter that we need to set before we get the factors: number of factors. Bai and Ng (2002) note that if we know the true number of factors, we can use the Bayesian Information Criteria (BIC) to determine this number. When the factors are unknown and have to be estimated, however, the BIC will not always consistently estimate the true number of factors. Bai and Ng (2002) offered seven criteria to determine the number of factors. They find that PC1, PC2, IC1, and IC2 seem to perform better than PC3 and IC3 (for formulas see, Bai and Ng, 2002:201). In the presence of cross-section correlations, BIC3 has very good properties (see Bai and Ng, 2002:202 and 207). This criterion can be used despite not fulfilling all the conditions of Theorem 2 in their paper. Figure 3.2.a and Figure 3.2.b show the number of factors that we get by recursively expanding our medium and large data sets. As evident in the figures below, the seven criteria of the Bai and Ng (2002) give diverging results from each other in terms of the number of factors, and this number may change as we add more observations through time. Some authors use one of these criteria and do not always check the role of using a certain information criterion on forecasting performance. For instance, Barhoumi et al. (2013) analyze the effect of the number of factors on forecasting performance. Although they are specifically interested in the effect of the number of factors on forecasting performance, they only employ IC1 among the Bai and Ng (2002) criteria. Gupta and Kabundi (2011) forecast South African variables with factor models. They find that PC1 and PC2 suggest seven factors, while IC1 and IC2 suggest five for their data set. They do not consider the BIC3 criterion for selecting the number of factors. They state that they use five factors. However, since the number of factors changes slightly over time and substantially depending on the choice of the criterion, it is still an empirical question to check whether using other criteria changes the forecast’s performance. In this respect, we consider all of the seven criteria suggested by Bai and Ng (2002). 0 2 4 6 8 01.2011 06.2011 11.2011 04.2012 09.2012 02.2013 07.2013 12.2013 05.2014 PC1 PC2 PC3 IC1 IC2 IC3 BIC3 0 2 4 6 8 10 01.2011 06.2011 11.2011 04.2012 09.2012 02.2013 07.2013 12.2013 05.2014 PC1 PC2 PC3 IC1 IC2 IC3 BIC3 Günay-Forecasting Turkish Industrial Production Growth With Static Factor Models 70 3.3. Forecast Equation In this paper we present forecasting results for 3 and 12 month-ahead forecasts. Though other horizons for the analysis exist, we focus our attention for the sake of clarity in presentation. The 3 month-ahead forecast performance is expected to be informative about the short run performance of models, while the 12 month-ahead forecast is thought to be informative about the longer run. An important question emerges when we forecast more than one period ahead. Consider the month-on-month growth rate of industrial production as presented in Figure 3.3. We can define the 3 month-ahead forecast as the month-on-month growth from three months from now. For example, in a case where we have the January figures as the last data point, we can forecast what would be the monthly growth rate in April, which is three months from January. However, this is a highly volatile series, which would be very hard to forecast. Also, the monthly growth rate in April will depend on the monthly growth rate in March. Hence, the month-on-month growth rate from 3 or 12 month from now may not be very interesting from a policy maker’s perspective. Rather, policy-makers may be interested in the over-all growth during these periods. Figure 3.3 Month-on-Month Growth of Industrial Production (Annualized) Figure 3.4 Three and Twelve Month Cumulative Growth of Industrial Production (Annualized) -80 -60 -40 -20 0 20 40 60 80 Şub.05 Tem.05 Ara.05 May.06 Eki.06 Mar.07 Ağu.07 Oca.08 Haz.08 Kas.08 Nis.09 Eyl.09 Şub.10 Tem.10 Ara.10 May.11 Eki.11 Mar.12 Ağu.12 Oca.13 Haz.13 Kas.13 Nis.14 Eyl.14 Month-on-Month Growth (annualized) -60 -50 -40 -30 -20 -10 0 10 20 30 40 Şub.05 Tem.05 Ara.05 May.06 Eki.06 Mar.07 Ağu.07 Oca.08 Haz.08 Kas.08 Nis.09 Eyl.09 Şub.10 Tem.10 Ara.10 May.11 Eki.11 Mar.12 Ağu.12 Oca.13 Haz.13 Kas.13 Nis.14 Eyl.14 3 Month Ahead Cumulative Growth (annualized) 12 Month Ahead Cumulative Growth (annualized) International Econometric Review (IER) 71 In this respect, we follow Stock and Watson (2002a) and forecast the cumulative growth rate for 3 and 12 month-ahead periods. In this approach, for the case that we can access January data, we forecast the growth rate in April relative to the level in January; in other words, we work with the cumulative growth in the horizon of interest. The 3 and 12 month-ahead cumulative growth rates in Figure 3.4 show that, as expected, the 12 month-ahead cumulative growth rates are relatively more stable than the 3 month-ahead rates. We also observe that volatility of the three month growth increased after around mid-2011. Equation 3.3 shows the forecasting model where we obtain coefficients with OLS. In this equation, the dependent variable is the cumulative growth rate from time t to time t + h so that we are forecasting h-period ahead. We use month-on-month change of industrial production and the estimated factors as the independent variables. Using different letters in the notation of Equation 3.3 for the lag length (namely m and p) indicates that we can allow different number of lags for the lag of dependent variable and for the factors. We use a cap on the F which shows that we are working with estimated factors since we cannot observe actual factors. 1 1 1 1 ˆ ˆ ˆ ˆ ˆ          jt p jhjjt m jhjh h t h tYFY  (3.3) where Y is the variable that we want to forecast. In the direct forecasting approach, and change for each horizon. Subscript "h" in the dependent variable indicates that we define cumulative growth for each forecast horizon, h. 3.4. Forecast Evaluation Our evaluation criterion for comparing the models is the Root Mean Squared Error (RMSE) that we get from a pseudo out-of-sample forecasting exercise. Stock and Watson (2003) note that the relative performance of the models may change in different samples. They divide their evaluation sample into two and compare the relative performance of selected indicators for forecasting output relative to a benchmark. They find that only 10 percent of the indicators beat the benchmark in both periods, while around 20 percent of the indicators beat the benchmark in only one of the evaluation periods. Altug and Uluceviz (2013) analyze the forecasting performance of selected indicators for Turkish industrial production. Their results show that the forecast performance relative to an AR model changes depending on the evaluation sample. They find that recently it gets harder to beat the AR model. We estimate models starting from February 2005 and do the evaluation for two samples to see whether the forecast performance is stable or not. In the first evaluation sample, the out-of- sample recursion starts in January 2010 and ends in September 2011. For the second evaluation sample, the recursion starts in October 2011 and ends in September 2013. We have data up until September 2014, and the longest horizon that we are interest in is 12 monthahead. So, September 2013 is the last point in the recursion that we can compare our 12 month-ahead forecast with a realization. In other words, we will produce a forecast for the growth rate between September 2014 and September 2013. We will then compare this forecast with the realization. At each step we get the factors, determine lag lengths in Equation 3.3, estimate the appropriate equation for h step-ahead forecasting, and derive the forecasts. We estimate two versions of Equation 3.3. In the first version, we use lags of the explanatory variables, as per the DI-AR Lag specification in Stock and Watson (2002b:149). The second specification is the DI of Stock and Watson (2002b), where we use only contemporaneous values of the Günay-Forecasting Turkish Industrial Production Growth With Static Factor Models 78 Bai, J. and S. Ng (2002). Determining the Number of Factors in Approximate Factor Models. Econometrica, 70, 191-221. Barhoumi, K., O. Darne and L. Ferrara (2010). Are Disaggregate Data Useful for Factor Analysis in Forecasting French GDP? Journal of Forecasting, 29, 132-144. Barhoumi, K., O. Darne and L. Ferrara (2013). Testing the Number of Factors: An Empirical Assessment for a Forecasting Purpose. Oxford Bulletin of Economics and Statistics, 75, 64-79. Boivin, J. and S. Ng (2005). Understanding and Comparing Factor Based Forecasts. International Journal of Central Banking, 1, 117-151. Boivin, J. and S. Ng (2006). Are More Data Always Better for Factor Analysis? Journal of Econometrics, 132, 169-194. Eickmeier, S. and C. Ziegler (2008). How Successful are Dynamic Factor Models at Forecasting Output and Inflation? A Meta-Analytic Approach. Journal of Forecasting, 27, 237-265. Gupta, R. and A. Kabundi (2011). A Large Factor Model for Forecasting Macroeconomic Variables in South Africa. International Journal of Forecasting, 27, 1076-1088. Rünstler, G., K. Barhoumi, S. Benk, R. Cristadoro, A. Den Reijer, A. Jakaitiene, P. Jelonek, A. Rua, K. Ruth and C. Van Nieuwenhuyze (2009). Short-term Forecasting of GDP Using Large Datasets: A Pseudo Real-time Forecast Evaluation Exercise. Journal of Forecasting, 28, 595-611. Stock, J. and M. Watson (2002a). Macroeconomic Forecasting Using Diffusion Indexes. Journal of Business and Economic Statistics, 20 (2), 147-162. Stock, J. and M. Watson (2002b). Forecasting with Principal Components from a Large Number of Predictors. Journal of American Statistical Association, 97, 1167-1179. Stock, J. and M. Watson (2003). Forecasting Output and Inflation: The Role of Asset Prices. Journal of Economic Literature, 16, 788-829.