scieee AI-readable full text Open interactive document viewer

Imputation of large gaps in tourism management time series

Caridad Ocerin, José M.; Marmolejo Martín, Juan Antonio; Fiala, Petr; Caridad López del Río, Lorena

Abstract

The tourist industry relies every day in the capability to foresee the near future: in fixing prices offered to potential customers, adapting the number of employees to demand, making available a certain number of rooms, and so on. The statistics used are both internal and from the Statistical Institutes and specialized companies. But it is usual to find gaps in the time series data, sometimes quite large. And, to analyze temporal data, it is convenient to avoid missing data, and outliers. We propose a strategy for dealing with large gaps in series related to this sector, based on obtaining back and forward forecast to minimize the forecasting errors. In a case study, using the average daily rate (ADR) for the Spanish hotels in a long period, the methodology proved to produce smaller RMSE, and get reliable forecasts.

Full text

— 197 — IMPUTATION OF LARGE GAPS IN TOURISM MANAGEMENT TIME SERIES José M. Caridad Ocerin iManagement&Tourism. Seville. Spain [email protected] Juan A. Marmolejo Martín. Universidad de Granada. Melilla. Spain Dpt. of Statistics and O.R. Petr Fiala Prague University of Economics and Business [email protected] Dpt. of Econometrics Lorena Caridad y López del Río. University of Seville. Spain. [email protected] Dpt. of Financial Economics and O.M. Abstract The tourist industry relies every day in the capability to foresee the near future: in fixing prices offered to potential customers, adapting the number of employees to demand, making available a certain number of rooms, and so on. The statistics used are both internal and from the Statistical Institutes and specialized companies. But it is usual to find gaps in the time series data, sometimes quite large. And, to analyze temporal data, it is convenient to avoid missing data, and outliers. We propose a strategy for dealing with large gaps in series related to this sector, based on obtaining back and forward forecast to minimize the forecasting errors. In a case study, using the average daily rate (ADR) for the Spanish hotels in a long period, the methodology proved to produce smaller RMSE, and get reliable forecasts. Keywords: hotel management, ADR, imputation, large gaps, time series 1. INTRODUCTION Statistics in the tourism sector compile information from usual variables about the number of visitors, length of the stay, types of accommodation, number of places, and so on. Regarding the management of establishments, several financial data are collected on many countries on a monthly basis; for example in the hotel sector the revenue per available room (RevPAR), the occupancy rate (OR), the average daily rate (ADR) per occupied rooms, the number of rooms on offer (Rooms), the gross operating profit per available room, total revenue per available room, including all sources of revenue, cost per occupied room, net promoter score, market penetration index, average length of stay, and so on. These metrics are essential for strategic decision-making in the hospitality industry (Binesh & Belamino, 2021, and Corgel & Sturman, 2011). They help managers to optimize operations, set competitive pricing, and improve guest experiences, ultimately leading to increased profitability and market share. For hotels, these variables are recorded usually on a monthly basis, but, it is not unusual to have periods where these data are missing, forming a large gap. For example, if the hotel is closed for repairs, or when changing the furniture, or, at it happened in 2020, for a pandemic induced closure. In the last case, a closure can be followed by a recover period, were data in a time series may be labelled as outliers, widening the gap. The imputation of outliers or o missing data is the most common situation in time series analysis, but wide gaps’ imputation present a different challenge, as most of current methods may fail and distort the time series data, which is the opposite aim of the imputation procedures. This is the problem consider her, mainly for time series used in hotel management. Even in financial published series (Caridad, L., 2020) separate data are common. Forecasting in hotel management is a crucial practice that helps hoteliers predicts future business conditions, enabling them to make informed decisions about pricing, staffing, marketing, and other operational aspects. Effective forecasting can significantly enhance revenue management — 198 — strategies, optimize resource allocation, and improve overall hotel performance (Enz et al., 2001). Key Aspects of Forecasting in hotel management can be classified in several topics: - Demand forecasting: its purpose is to predict the number of guests who may book rooms on any given day. The common approach is to uses historical booking data, current booking pace, and external factors like seasonality, holidays, local events, and economic conditions. And its expected outcome leads to helps in setting room rates, planning promotions, and managing inventory effectively. - Rate optimization: it aims to determine the optimal pricing of hotel rooms to maximize revenue. It is often implemented through revenue management systems that use sophisticated algorithms to analyze past and current booking trends, competitor pricing, and market demand. The results help in dynamic pricing strategies that adjust rates based on forecasted demand to ensure competitive pricing and high occupancy rates. - Financial forecasting would obtain estimates of future revenues, costs, and profits. It is based on historical financial data, market trends, and operational forecasts, including projected occupancy and room rates. It is essential for budgeting and financial planning, helping hoteliers ensure profitability and manage cash flow. - Staffing and resource allocation is needed to predict staffing needs based on expected hotel occupancy, and it is based in analyzing past staffing patterns, guest influx, and special events to forecast staffing requirements. With these, it can ensure optimal staff levels to maintain service quality without incurring unnecessary labour costs. Some techniques and tools used in forecasting in hotel management are the following: - Time Series Analysis: it involves methods like moving averages, exponential smoothing, and ARIMA models, which are used to predict future data points based on past data trends. - Econometric causal modelling: An approach to estimate relationships among variables. It's used to understand how dependent variables such as bookings are influenced by independent variables like price changes, seasons, or economic indicators. - Machine Learning: these models can analyze large datasets to detect patterns and predict future outcomes more accurately. Techniques such as neural networks and decision trees are employed to forecast demand and pricing. - Simulation Models: they can allow hotel managers to understand how changes in certain variables could affect outcomes. This is particularly useful for scenario planning and risk management. - Revenue Management Systems (RMS): these are specialized software solutions that integrate historical data, real-time market data, and forecasting tools to help optimize pricing and inventory decisions. Forecasting is a challenging task, and it depends on many factors: - Data quality and availability: reliable forecasting depends on high-quality, accurate historical data. Incomplete or incorrect data can lead to poor predictions. - External factors: many external variables, such as changes in economic conditions, new competition, or unforeseen events like natural disasters, can significantly impact the accuracy of forecasts. - Adapting to market changes: the hospitality industry is dynamic, and consumer preferences can shift rapidly. Forecasting models need to be adaptable and responsive to these changes. — 199 — 2. IMPUTATION IN TIME SERIES Introduction In the realm of time series analysis, dealing with missing data is an inevitable challenge that can significantly affect the accuracy and reliability of forecasts and other statistical insights. Imputation plays a crucial role in addressing this challenge by providing methodologies to estimate missing values in a time series. Here are explored the importance of imputation, various imputation methods, and their applications in different fields, highlighting both the advantages and limitations of these approaches. The Importance of Imputation in time series data, characterized by sequential measurements over time, is prevalent across various domains such as finance, meteorology, economics, and healthcare. The integrity of this data is often compromised by gaps due to non-recording, errors in data collection, or other disruptions. Missing data can lead to biased estimates and reduce the statistical power of time series models. Imputation helps mitigate these issues by filling in missing values, thus enabling more accurate and comprehensive analysis. Imputation techniques for time series data range from simple methods for handling missing data to more sophisticated approaches that consider the time-dependent nature of the data, (Ahn et al. 2022, and Moritz et al. 2015): - Mean or Median Imputation: This method involves replacing missing values with the mean or median of available data points. It is suitable for data with little variability and no strong trend or seasonal components. It is limited as it can reduce the variability of the dataset and does not account for time series dynamics. - Last Observation Carried Forward (LOCF) and Next Observation Carried Backward (NOCB): LOCF imputes missing values using the last observed value, while NOCB uses the next available value. It is effective for data with minimal time gaps and when the data points are not highly volatile, but, it does not reflect changes between periods, potentially leading to misleading analysis in volatile series. - Linear Interpolation: it estimates missing values by linearly interpolating between neighbouring data points, and it is useful in datasets with gradual trends and no abrupt jumps between observations. It assumes a linear relationship between points, which may not hold in all cases, particularly in nonlinear time series. - Time Series Decomposition into trend, seasonal, and residual components, and then imputing missing values based on these components. It is ideal for time series with clear trends and seasonal patterns, but, requires a relatively long time series to accurately estimate the components. - Dynamic Models (e.g., ARIMA, State Space Models, Tramo-Seats, X13, Neural Networks): these models use the underlying patterns and structures in the data, such as autocorrelations, to estimate missing values. They are best suited for complex time series with significant autocorrelations and systematic patterns, but need qualified personnel to use them (Wolf et al. 2021, and Fruet-Cardozo et al. 2022). Imputing missing values in time series data where there are large gaps presents a unique set of challenges. Such gaps can severely distort the time series analysis, affecting everything from trend detection to forecasting. The task becomes even more demanding when these gaps are not randomly distributed but occur in large contiguous blocks, which might be due to systemic issues, significant disruptions, or long periods of data collection failure. In these types of series, among the usual challenges, are the following: - Loss of Information: Large gaps mean a significant loss of information, which can lead to poor estimation of the time series components (trend, seasonality, and cyclical components). — 200 — - Biased Estimations: Standard imputation methods might introduce bias if used without adjustments for such extensive missing data, as they typically assume that the missing points are more sporadically placed. - Decreased Reliability: The reliability of predictions and analyses decreases as the gap widens, because the assumptions about data continuity and pattern stability are weakened. When dealing with large gaps in time series data, more sophisticated imputation techniques are often necessary to maintain the integrity of the data analysis. Here are some effective strategies (Ribeiro & Castro, 2022): - Model-Based Imputation such as State Space Models and Kalman Filtering: This approach is wellsuited for handling missing data in time series with large gaps. State space models, used in conjunction with Kalman filtering techniques, can effectively estimate missing values by considering the observations as part of a dynamic system. The Kalman filter excels in scenarios where the data points are part of a linear dynamic system influenced by Gaussian noise. - ARIMA Models: For time series that exhibit strong autocorrelations, AutoRegressive Integrated Moving Average (ARIMA) models can be adjusted to forecast missing values. These models can be particularly useful if the structure of the data outside the gaps can be well modelled by ARIMA processes. - Multiple Imputation: it involves creating several different plausible imputations for missing values. By treating the imputation process as a simulation, this method acknowledges the uncertainty inherent in the prediction of missing values, especially with large gaps. The final imputation can then be obtained by averaging across different simulated datasets, thereby providing a robust estimate that includes a measure of variability. - Machine Learning Techniques such as Long Short-Term Memory Networks (LSTMs) or Gaussian Process Regression (GPR). LSTM are a type of recurrent neural network that is capable of learning order dependence in sequence prediction problems. They can be particularly effective in handling large gaps if the time series exhibits complex non-linear patterns that simpler methods cannot capture. GPR is a non-parametric kernel-based probabilistic model can be an excellent tool for imputing values in time series with large gaps, especially if the series can be assumed to follow a smooth underlying process. - Interpolation with External Data: when internal data are insufficient due to large gaps, external datasets that correlate well with the time series in question can be used for imputation. For instance, using economic indicators to estimate missing financial market data, or regional weather patterns to estimate local hotel’ rooms demand. Understanding data patterns, before choosing an imputation method, it is critical to understand the underlying structure and characteristics of the time series. This includes seasonality, trend components, and volatility. Whatever method is used, it’s important to validate the imputations against known data (if available) to assess the accuracy and reliability of the imputation technique. Especially with large gaps, it's important to account for increased uncertainty in the imputed values. Techniques like multiple imputation can help quantify this uncertainty. We propose, for hotel management variables large gaps of missing information, the development of multiple forecasts, and a combination method based on the forecasting errors in each procedure. — 201 — 3. METHODOLOGY The imputation with just one large gap not at the end of the series can be represented as follows: let’s consider the series y1, y2, …, yn, divided in three sectors, corresponding to the full period of n data,, but with g cases missing after the instant t = m; the missing data are, thus, ym + 1, ym + 2, …, ym + g; then come the last available cases ym + g + 1, ym + g + 2, …. , yn. Subset I is formed by the first m data, being m the last date available before the gap. This period starts at t = m + 1 and ends at instant m + g; after this, the last n – m – g data are also available. The objective is to forecast the gap sector, that is to obtain the g estimates ŷm + 1, ŷm + 2, …, ŷm + g. If there are several large gaps, the strategy would be to estimate first the gap with more information available at both ends, and then proceed with the second gap using the same criteria (and with the series imputed from the missing data of last gap); and so on. If the series has classical components that can be considered additive (using, if necessary, a Box-Cox transformation), a classical dynamic model would be suitable to obtain a first estimate of the missing data with its forecast errors standard deviation. We would see about this in the case study considered, in the next section. These first estimates ŷm + 1, ŷm + 2, …, ŷm + g could be used, to fill part of the gap, if the corresponding estimates are considered suitable, reducing the length of the remaining gap. Usually, the filled part would be at one or both sides of the gap. With this procedure, the gap would be smaller and more prone to be imputed with side models, both using forward forecasts and backward forecasts. The former would be based on the first sector of the series, that is using the interval from t = 1 to t = m. If with the first set of forecasts, r of the first observations after t = m, having been imputed, we consider as New value m = Original m + r. With some time series methods, and data from the (new) first sector (before the actual gap), the forecasted values and its error’s standard deviations are (ŷm + 1I, sm + 1I), (ŷm + 2I, sm + 2I), ….., (ŷm + gI, sm + gI ) If some initial forecasts are used to fill part of the end of the gap, for example r* missing data, we will consider the new m + g + 1 as the first case of the third sector (that is, New m + g + 1 = Original m + g + 1 – r*) ; the gap would have also shorten by r* cases at its end (and by r cases in the other end). The new g will be then equal to g – r – r*. Another set of forecasts for the same time interval will be obtained using the third available set of data, ym + g + 1, ym + g + 2, …. , yn, with its time order inverted, originating the time series y1* = yn, y2* = yn - 1, …, yn – m – g* = ym + g + 1; using this series, g (new) forecasts ŷn – m – g + 1*, ŷn – m – g + 2*, …, ŷn – m – g + g*, with its error’s standard deviations would be obtained. These backward forecasts should be put in increasing time order inverting the order of the y*s, leading to the second set of g forecasts (ŷm + 1II, sm + 1II), (ŷm + 2II, sm + 2II), ….., (ŷm + gII, sm + gII ) The rationale of using forecasts obtained from the data after both ends of the gap (forward and backward forecasts), is that it is usual that the first forecasts, after the last available data of a series, tend to be more precise than the following forecasts. Thus the first forecasts of type I and the last included in the set III ought to be (generally) more precise than the rest. Combining both types of forecast is it possible to obtain more precise forecasts (Amstrong, 2001). As we have considered two sets of forecasts, one criteria to merge them is used a weighted average using weights proportional to the inverse of the corresponding standard deviations of the corresponding observations, that is, the final set of forecasts would be I II I II I II I II // ˆ ˆ ˆ / / / / m t m t m t m t m t m t m t m t m t ss y y y s s s s ++ + + + + + + + =+ ++ 11 1 1 1 1 — 202 — Of course, if we use also the original forecasts, the final set of forecasts would be a weighted average or the three forecasts obtained. And this could be extended if some of the forecasts with different sets of data were obtained with different methods. The increase of precision in the forecasts using these weights has the advantage of reducing the variability of the errors that are not near the ends of the gap. Combining forecasts for data imputation in time series is an effective approach to handle missing data by leveraging the strengths of different forecasting models. Here are some key methods and considerations for combining forecasts for this purpose: - Simple Averaging: it involves taking the average of the forecasts produced by different models. This approach is straightforward and often improves forecast accuracy by balancing out the individual model errors. As the main advantage is that it is easy to implement, but it may not always produce the best results if the individual models have vastly different accuracies. - Weighted Averaging assigns different weights to each model based on their past performance or some other criteria. Models that perform better receive higher weights. It is a more flexible method than simple averaging, and can improve accuracy by giving more importance to better-performing models. The problem is to determine appropriate weights. - Regression-based methods involve using regression techniques to combine forecasts. For example, a linear regression model can be trained where the individual forecasts are the independent variables, and the actual observed values are the dependent variables. It can capture complex relationships between the forecasts, but requires a certain amount of historical data to train the regression model. - Machine Learning approaches such as neural networks, random forests, or support vector machines can be used to combine forecasts. These models can learn complex patterns and relationships between the forecasts. It is highly flexible and can model complex interactions, but it requires significant computational resources and large datasets for training. - Bayesian Model Averaging involves weighting the forecasts by their posterior probabilities, which are updated as more data becomes available. It provides a principled way to combine forecasts, accounting for model uncertainty. The disadvantage is how to compute a priori probabilities and define a set of computable a posteriori distributions. The steps required to combine forecasts for Data Imputation if structured in several phases: 1. Identify Missing Data: determine the missing values in your time series data and if there are wide gaps of missing data. 2. Generate forecasts using different forecasting models (e.g., ARIMA, Exponential Smoothing, Neural Networks). 3. Combine forecasts, applying one of the combination methods (e.g., simple averaging, weighted averaging, regression-based methods) to combine the forecasts from different models. 4. Impute Missing Data, using the combined forecast to impute the missing values in the time series. Ensure that the models being combined are diverse and not highly correlated to gain the benefits of combination. Validate the results using cross-validation to assess the performance of the combined forecast method. Computational resources should be available: some methods, especially machine learning-based, may require significant computational resources. — 203 — Combining forecasts for data imputation can significantly enhance the accuracy and reliability of time series data, especially when dealing with complex and uncertain environments. A firm such as an hotel, can have scarce resources, and for sure it lacks of personnel with expertise in many econometric methods. Thus, here is proposed a combination of forecasts using a weighted average, as it is easy to implement, and in the hotel industry, usual time series models are realistic to implement. When imputing gaps of several missing observations, one has to be aware that the larger the gap is (and the presence of several large gaps in the series) can jeopardize the results. Also, if the forecasting methods are biased in the same direction, the precision of the forecast will be affected. A subjective judgement, beside the ‘correct’ specification of the time series used at both sides of a gap, is an aim desirable. 4. A CASE STUDY IN THE HOTEL INDUSTRY Several metrics are usual in the hotel management: the revenue per available room (RevPAR), the occupancy rate (OR), the average daily rate (ADR) and some others. To check the methods proposed the Spanish hotels ADR series from 2008 to April 2024 (Lozano et al., 2020) will be used. In this series there is a large gap corresponding to the pandemic years In the pandemic period, starting in March 2020, and we can consider it lasted until the end of 2021, we have a quite large gap of outliers, during a run of 22 months. To be able to estimate these data to ‘linearize’ de series, some non ordinary imputation method should be used. But, if we want to test the proposed method, it would be feasible to eliminate one set of 12 data, for example the whole year 2015, and estimate the missing observations using several methods, and then compare these with the real data. After this, the pandemic months could be imputed considering this period as missing data. 0 20 40 60 80 100 120 140 2008 2010 2012 2014 2016 2018 2020 2022 2024 ADR Spanish Hotels Figure 1. ADR for Spanish Hotels. Source: INE (2024) — 204 — Series y is the ADR series without the data corresponding to 2015, and several forecasts will be considered to estimate this missing year, and data till 2020 will be used in this imputation process yt = 71.268 – 0.1208t + 0.001914t2 – 2.844Jan – 1.138Feb – 3202Mar – 3.546Apr – – 5.214May – 0.835Jun + 8.809Jul + 15.066Aug + 1.593Sep – 3.477Oct – – 3.182Nov – 2.028Dec + et being significant both the trend and the seasonal component wit very low p-values, and R2 = 0.932; the month dummy variables are binary. This model (Caridad et al., 2023, EViews, 2023) is used to obtain the initial forecasts for the year 2015. Using the data from 2008 to 2014 (sector I) an Arima model is estimated 12yt = (1 – 0.407B)at from which the forecasts ytI for 2015 with sector I data are obtained. To generate the back-forecasts for 2015 with sector II data (2016 to 2019), the series in these years should be inverted. This series yt*is adjusted to the Arima model (1 + 0.837B12)12yt* = at and the corresponding forecasts for 2015 are derived; once put in the right order, the back forecasts ytII for 2015 are obtained. Table I. 2015 values, forecasts I, II, combined and classical. Weights I and II. yt ŷtI ŷtII Combined ŷt Classical ŷt wtI wtII January 71.76 71.47 70.62 71.27 71.99 0.767 0.233 February 70.95 73.64 72.87 73.43 73.9 0.730 0.270 March 74.02 69.95 70.47 70.11 72.04 0.697 0.303 April 71.07 70.66 72.76 71.36 71.92 0.667 0.333 May 71.49 68.95 71.99 70.06 70.46 0.635 0.365 June 76.05 72.58 79.32 75.25 75.07 0.604 0.396 July 87.68 82.96 90.42 86.17 84.94 0.570 0.430 August 95.30 90.54 97.70 93.87 91.42 0.535 0.465 September 79.99 77.43 81.93 79.70 78.18 0.495 0.505 October 74.85 71.25 74.43 73.01 73.35 0.447 0.553 November 75.16 72.08 72.20 72.15 73.89 0.387 0.613 December 75.64 74.17 74.97 74.73 75.28 0.301 0.699 To assess the different forecast for 2015, the RMSE is obtained Table 2. Root Mean Square Error for forecasts with different methods ŷtI ŷtII Combined ŷt Classical ŷt RMSE 9.8772 4.8184 3.5118 3.7687 — 205 — As expected the combined forecasts using both forward and backward forecasts shows smaller RMSE's. The forward forecasts tend to be more precise in the first months of 2015, while the backward forecasts are preferable in the last months of the gap period. The classical model (estimated from all data in the whole interval 2008-2019, but the gap year 2015, is more stable along the gap, as the ADR series (yt) follows a quite defined pattern all along this period. 68 72 76 80 84 88 92 96 100 M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 2015 Forward (I) Backward (II) y y combined y = ADR (EUR/day) Figure 2. ADR forecasts and real values If we combine, with the same procedure and weights proportional to the inverse of the standard errors of the three forecasts, included the one obtained with the classical model, the precision increases slightly, being RMSE = 3.3276. Of course, this can be done in this case, as the third forecast was quite precise. In a real gap we could not calculate RMSEs for different forecasts, as in the pandemic years. As we have observed the behaviour of the ADR series, a sound strategy to impute a 22 months gap could be to use a classical method to impute data in the months of recovery, that is for 2021, and to obtain forward and backward forecasts for the ten months of 2020 (Caridad et al., 2021). The procedure could then be repeated to obtain a second estimation for the last year of the gap period. 5. CONCLUSIONS Effective forecasting in hotel management is about more than just filling rooms; it's about strategically managing a hotel's resources to maximize profitability and efficiency. As technology advances, the tools and methods for forecasting become more sophisticated, providing hoteliers with deeper insights and more accurate predictions. By leveraging these technologies and continually refining their forecasting models, hotels can better anticipate guest needs, optimize their operations, and enhance their competitive edge in the market. Imputation in time series is a critical statistical technique that enhances the quality and usability of datasets across various fields. By understanding and applying appropriate imputation methods, analysts can overcome the challenges posed by missing data, leading to more reliable and insightful outcomes. However, the choice of imputation method should be carefully considered based on the specific characteristics of the time series and the analysis objectives. As data continues to grow