Predicting wholesale edible oil prices through Gaussian process regressions tuned with Bayesian optimization and cross-validation
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Jin, Bingzi; Xu, Xiaojie Article Predicting wholesale edible oil prices through Gaussian process regressions tuned with Bayesian optimization and cross-validation Asian Journal of Economics and Banking (AJEB) Provided in Cooperation with: Ho Chi Minh University of Banking (HUB), Ho Chi Minh City Suggested Citation: Jin, Bingzi; Xu, Xiaojie (2025) : Predicting wholesale edible oil prices through Gaussian process regressions tuned with Bayesian optimization and cross-validation, Asian Journal of Economics and Banking (AJEB), ISSN 2633-7991, Emerald, Leeds, Vol. 9, Iss. 1, pp. 64-82, https://doi.org/10.1108/AJEB-06-2024-0070 This Version is available at: https://hdl.handle.net/10419/334137 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Predicting wholesale edible oil prices through Gaussian process regressions tuned with Bayesian optimization and cross-validation Bingzi Jin Advanced Micro Devices (China) Co., Ltd, Shanghai, China, and Xiaojie Xu North Carolina State University at Raleigh, Raleigh, North Carolina, USA Abstract Purpose –Developing price forecasts for various agricultural commodities has long been a significant undertaking for a variety of agricultural market players. The weekly wholesale price of edible oil in the Chinese market over a ten-year period, from January 1, 2010 to January 3, 2020, is the forecasting issue we explore. Design/methodology/approach –Using Bayesian optimisations and cross-validation, we study Gaussian process (GP) regressions for our forecasting needs. Findings –The produced models delivered precise price predictions for the one-year period between January 4, 2019 and January 3, 2020, with an out-of-sample relative root mean square error of 5.0812%, a root mean square error (RMSEA) of 4.7324 and a mean absolute error (MAE) of 2.9382. Originality/value –The projection’s output may be utilised as stand-alone technical predictions or in combination with other projections for policy research that involves making assessment. Keywords Wholesale edible oil, Price forecasting, Gaussian process regression, Bayesian optimization, Cross-validation, Chinese market Paper type Research paper 1. Introduction Food price projections from the agriculture sector are crucial for a range of market participants, such as processors, speculators, hedgers and policymakers (Raihan et al., 2023). Producers, for example, often need price forecast data to set sales prices before production begins, exporters and processors to fulfil their contractual duties, speculators to profit, hedgers to control risks, and policymakers to develop, track and evaluate strategic plans and policies (Dacha et al., 2021). China’s significant agricultural market share (Ji et al., 2022), close linkages to the energy sector (Ranguwal et al., 2023), and the influence of financial markets and macroeconomic factors (Raihan, 2023) make edible oil price forecasting there no exception. Because of the physical limitations on the food supply—such as land, agricultural technology, environmental sustainability and climate change—policymakers see food prices as strategic issues. This is especially true, given China’s massive population and expanding economy. A number of macroeconomic and financial factors, such as interest AJEB 9,1 64 JEL Classification — C22, C53, C63, Q11, Q13 © Bingzi Jin and Xiaojie Xu. Published in Asian Journal of Economics and Banking. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at http://creativecommons.org/licences/by/4.0/ legalcode Conflict of interest: There is no conflict of interest. Data availability statement: Data are available upon request. The current issue and full text archive of this journal is available on Emerald Insight at: https://www.emerald.com/insight/2615-9821.htm Received 13 June 2024 Revised 16 September 2024 16 October 2024 2 November 2024 Accepted 4 November 2024 Asian Journal of Economics and Banking Vol. 9 No. 1, 2025 pp. 64-82 Emerald Publishing Limited e-ISSN: 2633-7991 p-ISSN: 2615-9821 DOI 10.1108/AJEB-06-2024-0070 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
rates, stock prices, exchange rates and financialisation levels, as well as changes in the energy markets, such as the price of oil, ethanol and the demand for biofuels, could put food price stability at risk. Price forecasting may not need to be motivated a lot since agricultural commodity prices often show erratic volatility patterns (Yeasin et al., 2020), have a significant impact on market participants’ decisions (Rahoveanu et al., 2018), and ultimately affect resource allocations and overall economic wellbeing (Wulandari et al., 2021). Many research studies have been conducted in the literature on a variety of time series approaches for price prediction (Jin and Xu, 2024a). In these early studies described below, models like vector autoregressive (VAR) models, autoregressive integrated moving average models, vector error correction models and many variants of these models are often mentioned. The autoregressive integrated moving average (ARIMA), for example, has been shown in previous studies to be a highly preferred choice for a variety of time series forecasting applications. It was shown that ARIMA performs noticeably better than expert views and structural model-based forecasts for the US hog and cattle markets. The accuracy of hog price projections will only be slightly improved by moving from the ARIMA to models that include more data from the sow’s farrowing price, according to another research. A number of distinctions exist between this empirical data and the wheat price data, which demonstrated that the ARIMA model’s forecast accuracy may be enhanced with adding exchange rate series data. Prior studies have shown that combining the ARIMA with various model types may increase prediction accuracy more than depending just on one data source. One well-liked econometric method for price series forecasts is the VAR approach, which emphasises the connections between various economic factors. The comparison’s conclusions demonstrate that the VAR predicts US cotton prices better than structural models at times of normal price volatility. In distinguishing the predictive content of a set of wheat futures prices from various countries and in differentiating between the prices of soybeans and soy in various US regions, the VAR was shown to be useful. Long-term relationships between economic variables are also taken into consideration by the vector error correction model (VECM), which is closely related to the VAR, via cointegration. It may be particularly helpful for long-term price forecasts. According to studies, for example, the VECM often performs better than the VAR in predicting global wheat prices. In price predicting studies, the previously outlined econometric models have proven useful, especially in edible oil studies. Karia et al. (2016), for example, used the auto-regressive fractionally integrated moving average (ARFIMA) and ARIMA for palm, rapeseed, soybean, linseed and sunflower oil to study price forecasts. Resolving the over-differencing issue had little influence on the prediction accuracy of any model, and they found contradictory results about the efficacy of several models. According to Priyanga et al. (2019), the ARIMA might be useful for forecasting the price of coconut oil in Kerala. The ARIMA model was used by Darekar and Reddy (2017) to forecast the pricing of an oil seed, namely groundnuts, in India. They found that farmers, legislators and marketers may benefit from the projections. There are minor differences between the model’s projected and actual values, according to Meena et al. (2014)’s analysis of oil and mustard seed prices in India using the ARIMA model. A multivariate ARIMA, which combines the ARIMA with an econometric equation predetermined for Malaysian palm oil price estimates, was proposed by Shamsudin and Arshad (2000) to improve forecasts based on the ARIMA. Khin et al. (2011) found similar empirical evidence to back their Malaysian palm oil price predictions. The ARIMA, exponential generalised auto-regressive conditional heteroscedastic (GARCH) model (EGARCH) and GARCH model were studied by Lama et al. (2015) in order to forecast edible oil prices both domestically and internationally. Because it better captures the volatility pattern, they found that the EGARCH model outperforms the other two. In order to anticipate the prices of soybean and rapeseed oil, Wang et al. (2013) demonstrated the effectiveness of the seasonal VECM in China. In order to demonstrate the availability of possible predictive information from crude oil prices to those of edible oil, Hasanov et al. (2016) used the GARCH-in-mean model and volatility impulse response function analysis. Their findings Asian Journal of Economics and Banking 65 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
showed that crude oil prices might aid in forecasting edible oil prices. This empirical evidence may vary depending on the historical periods in question. Using the VECM and directed acyclic graph approach, Yu et al. (2006) examined the price correlations between crude oil and various food oils and concluded that there was no discernible impact of crude oil prices on edible oil prices. Researchers have lately shown a great lot of interest in examining the uses of machine learning algorithms for agricultural commodity price forecasts because of the ease with which computer resources and technology are now accessible (Alade et al., 2021;Jin and Xu, 2024b). Therefore, a variety of commodities—such as soybeans, sugar, corn, wheat, soybean oil, coffee, cotton, green beans, canola, edible oil and peanut oil—have been the subject of research using neural networks, genetic programming, deep learning, support vector regressions, random forests, K-nearest neighbours, multivariate adaptive regression splines, decision trees, ensembles, and boosting. Neural networks may be the most popular machine learning model for predicting the price of agricultural commodities, according to these and other findings from earlier studies, however this is by no means an exhaustive analysis. Furthermore, previous empirical studies demonstrating the effectiveness of machine learning methods for financial and economic forecasting are often in agreement with these evaluations. The study demonstrates that machine learning techniques are being used more and more to forecast edible oil prices. In their investigation of price projections for Malaysian palm, soybean, coconut, olive, rapeseed and sunflower oils, for example, Kanchymalay et al. (2017) found that the sequential minimum approach to support vector regression (SVR) improves forecast accuracy. The random forest may be useful in forecasting Myanmar’s edible oil price, claim Myat and Tun (2019). In key Indian markets, Singh (2021) and Jha and Sinha (2014) examined neural networks with ARIMA for price forecasts of mustard, groundnut, rapeseed, and soybean oil. They found that, on average, neural networks are more accurate than ARIMA. Mishra and Singh (2013) focused on predicting groundnut oil prices in Delhi using neural networks and ARIMA; they found mixed findings on the two models’ efficacy. Lama et al. (2016) suggest that combining the neural network with GARCH might improve the forecasts of the individual model for edible oil prices in both domestic and international economies. The price forecasting challenges for maize and palm oil were examined by Jaiswal et al. (2022) using the ARIMA, deep long short-term memory neural network and time-delay neural network. They discovered that the most accurate predictions were made by the deep long shortterm memory neural network. According to Jaiswal et al. (2023), a nonlinear autoregressive neural network containing exogenous variables could be a helpful tool for soybean oil price forecasting. Silalahi (2013) discovered that the neural network model that was optimised by the genetic algorithm could provide a reasonable level of accuracy in price forecasting for soybean and palm oil. Salman et al. (2018) proposed tuning the backpropagation neural network for palm oil price forecasting using particle swarm optimisation, which increases prediction accuracy compared to the conventional backpropagation neural network. Amal (2021) discovered that the long short-term memory neural network model could be tuned using adaptive moment estimate optimisation to provide a palm oil price prediction with a high degree of accuracy. For time-series data gathered on edible oil prices, however, the predictions generated by the Gaussian process (GP) regression have not gotten much attention. A novel regression technique is put forward, drawing on Neal’s work on Bayesian learning for neural networks (Neal, 2012). Since the technique relies on priors over functions of Gaussian processes, simulating noisy data makes sense (Jin and Xu, 2024c). The convergence of several neural network-based Bayesian regression models to Gaussian processes close to the edge of an infinite network was demonstrated (Neal, 2012). Research has shown the effectiveness of GP regressions in replicating data that is either noisy (Williams and Rasmussen, 1995) or noisefree (Neal, 1997). In their study of Gaussian processes using radial basis function neural networks for forecasting problems involving stationary time-series data, Brahim-Belhouari and Vesin discovered that Bayesian learning yields superior prediction outcomes AJEB 9,1 66 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
(Brahim-Belhouari and Vesin, 2001). According to the results of Brahim-Belhouari and Bermak (2004), this study looked at a wide range of covariance functions. GP prediction techniques are also helpful for forecasting problems with non-stationary time-series data. Based on the work of Brahim-Belhouari and Bermak, GP regressions are more effective than radial basis function neural networks overall (Brahim-Belhouari and Bermak, 2004). The exact matrix operations required to integrate prior and noisy models are also the reason for the GP formulation’s effectiveness and success (Brahim-Belhouari and Bermak, 2004). Additionally, Brahim-Belhouari and Bermak proposed the use of GP predictors in the multi-model forecasting technique (Brahim-Belhouari and Bermak, 2004), which function similarly to how we would use model averaging. Over a ten-year period, from January 1, 2010 to January 3, 2020, we use the weekly wholesale edible oil price to conduct our inquiry. The forecast model we use is GP regression. The edible oil price index for the Chinese market has a significant underlying economic relevance as it is meant to reflect the wholesale market trend for all types of edible oil throughout the country. This price index may provide useful projections for policymakers and other market participants. To the best of our knowledge, none of the few earlier studies—including the ones mentioned above—have examined the pertinent prediction issue for this price index. Our prediction exercise is based on this price index, and we follow the literature on commodity price forecasts in order to fill this research gap. Therefore, our results serve to provide useful forecast information of the important price index to different forecast consumers. We examine how well models trained using crossvalidation and Bayesian optimisations provide forecasts. As it turns out, our models are rather straightforward and provide accurate and reliable forecasts. As far as we are aware, and considering the aforementioned studies, this is the first study to predict China’s wholesale edible oil pricing using the machine learning technique of GP regression. Model prediction performance may be improved by using Bayesian optimisation to provide GP regressions more flexibility, especially when dealing with underlying data that exhibits nonlinear properties. On the one hand, since it doesn’t take a lot of computing time to implement, this prediction framework is rather efficient. However, this approach produces forecasts that are quite accurate. There is no doubting the importance of accurate and timely agricultural commodity price projections for both policymakers and market players. Risk management, market evaluations and portfolio adjustments may benefit from such forecasts. This research assists decision-makers in making timely judgements by looking at difficulties related to projecting agricultural commodity prices and utilising weekly data, which are rather high-frequency for the wholesale market. Our results may be utilised as independent technical price projections, on the one hand. Nonetheless, they might be used in combination with other (basic) prediction findings for policy research and the creation of hypotheses about pricing trends. This method may be helpful for both market participants and policymakers as it facilitates the extension of the framework to potential commodity price projections in other business sectors. 2. Data Over the course of ten years, from January 1, 2010, to January 3, 2020, we examine weekly wholesale edible oil costs for the Chinese market. These prices are taken from China’s National Wholesale Price Information System. Figure 1’s top panel displays the price series plot on the left, the quantile-quantile plot against the standard normal distribution on the right, and the 50-bin histogram with kernel estimates in the centre. For the price series’ differences, the relevant data visualisation is shown in the bottom panel of Figure 1. For the pricing data, Table 1 offers crucial summary information. The price series definitely deviates from normal distributions, as shown by the results of the Jarque-Bera and Anderson–Darling tests. This may not come as a surprise considering the nature of economic data (Jin and Xu, 2024d). In particular, the price series exhibits platykurtic behaviour and right skew. As can be seen Asian Journal of Economics and Banking 67 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
Figure 1. Visualization of weekly wholesale edible oil prices and their first differences during the period of January 1, 2010–January 3, 2020, together with associated 50-bin histograms with kernel estimates and quantile-quantile plots against the standard normal distribution AJEB 9,1 68 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
Table 1. Data summary of weekly wholesale edible oil prices during the period of January 1, 2010–January 3, 2020 Series Minimum 1st percentile 5th percentile Mean Median Standard deviation 95th percentile 99th percentile Maximum Skewness Kurtosis JarqueBera AndersonDarling Price 65.4700 87.6898 90.1130 107.4965 102.3900 16.1876 133.6985 138.1588 140.1600 0.4619 1.7983 <0.001 <0.0005 First difference �36.2600 �6.7388 �3.4080 �0.02600 �0.1000 3.6374 3.8600 7.3028 34.9400 �0.0096 46.2291 <0.001 <0.0005 Source(s): Table by authors Asian Journal of Economics and Banking 69 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
from Figure 1, the price series has four notable surges. Keep in mind that the base period price, which is 100, is derived from the average weekly price for June 1994. Previous studies on the emergence of nonlinear features at higher moments over a broad range of time-series data have been published in significant numbers by the financial and economic domains (Yang et al., 2008). To find any possible nonlinear trends, the price series is put through the Brock–Dechert–Scheinkman (BDS) test (Brock et al., 1996). As can be seen, the test yielded almost zero p-values. These results imply the presence of nonlinearities in the price series. The purpose of this study is to predict the price series using GP regressions while taking these particulars into consideration. 3. Method The primary focus of this work is on the forecasting approach of GP regressions, a kind of probabilistic kernel model that has shown predictive ability in predicting a range of nonlinear patterns across several scientific fields (Jin and Xu, 2024e). The training data used to estimate model parameters with an uncertain distribution is represented by the model as follows: xi;yi ð Þ;i¼1;2;...;Tf g. These are the expressions for the d�dimension predictors: xi∈Rd, and the reflection of the target occurs via yi∈R. By use of twelve-lag prices as predictors, the price estimate is produced. Prices from the previous twelve weeks, for instance, will be used as predictors to determine the price for the thirteenth week. The approach below may be used to express a linear regression: y5x T βþε, with ε∼N0;σ2 ð Þreporting the error item. However, in GP regressions, an explicit basis and latent variables are used to define the target variable. l(x i ) represents the latent variables from Gaussian processes that satisfy the requirements for a joint Gaussian distribution, whereas b represents the basis function. The basis function’s purpose is to project different predictors onto the feature space; the objective’s smoothness will be shown by the latent variables’ covariance function (Zhang and Xu, 2020;LI et al., 2015). Usually, a Gaussian process (GP) is described by two metrics: mean and covariance. We would adopt k x;x0 ð Þ ¼ Cov lðxÞ;l x0 ð Þ½ �to report the covariance and m(x)5E(l(x)) to report the mean. It would be reported that y5b(x) T βþl(x) express the GP regression, where lðxÞ∼GP 0;k x;x0 ð Þð Þand bðxÞ∈Rp. Via a hyper-parameter named θ,k x;x0jθð Þwould be able to be parameterized. A GP regression is often trained using the following variables, which are often calculated using a particular approach: σ 2 ,θ, and β. Furthermore, we would provide the basis functions (called b’s) and kernels (called k’s) that would be used throughout the model’s training processes. This paper would examine two categories of kernels: one is the nonisotropic kernel (also known as automatic relevance determination kernel) and the other is the isotropic kernel. To explore both isotropic kernels and nonisotropic kernels, five distinct kernels of each category would be applied. Equations (1) through (10) provide the details for each kernel under discussion. To indicate the scale-mixture parameter, we would use the notation α> 0. We would utilize σ l for showing the characteristic length scale associated with isotropic kernels. We would use σ f for showing the standard deviation associated with the signal. r¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi xi−xj � �0xi−xj � � q. Applying θ5(θ 1 ,θ 2 )5(log σ l , log σ f ) suggests an approach for the purpose of guaranteeing that σ l and σ f are above zero. The length scale associated with each predictor that corresponds to a nonisotropic kernel would be unique and would be reflected thorough the use of σ m (m51, 2, . . .,d). Consequently, θwould be reflected via the use of θ5(θ 1 ,θ 2 ,...,θ d ,θ dþ1 )5(log σ 1 , log σ 2 ,..., log σ d , log σ f ). Isotropic Exponential:k xi;xj��θ � �¼σ2 fe−r σl(1) Isotropic Squared Exponential:k xi;xj��θ � �¼σ2 fe −1 2 xi�xj ð ÞTxi−xj ð Þ σ2 l(2) AJEB 9,1 70 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
Isotropic Matern 5=2: k xi;xj��θ � �¼σ2 f1þffiffiffi5 pr σlþ5r2 3σ2 l !e−ffiffi5 pr σl(3) Isotropic Rational Quadratic:k xi;xj��θ � �¼σ2 f1þr2 2ασ2 l � �−α (4) Isotropic Matern 3=2: k xi;xj��θ � �¼σ2 f1þffiffiffi3 pr σl � �e−ffiffi3 pr σl(5) Nonisotropic Exponential:k xi;xj��θ � �¼σ2 fe −ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi X d m¼1 xim�xjm ð Þ2 σ2 m s(6) Nonisotropic Squared Exponential:k xi;xj��θ � �¼σ2 fe −1 2X d m¼1 xim�xjm ð Þ2 σ2 m(7) Nonisotropic Matern 5=2: k xi;xj��θ � �¼σ2 f1þffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 5X d m¼1 xim �xjm � �2 σ2 m v u u tþ5 3X d m¼1 xim �xjm � �2 σ2 m 0 @1 Ae −ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 5X d m¼1 xim�xjm ð Þ2 σ2 m s (8) Nonisotropic Rational Quadratic:k xi;xj��θ � �¼σ2 f1þ1 2αX d m¼1 xim �xjm � �2 σ2 m !−α (9) Nonisotropic Matern 3=2: k xi;xj��θ � �¼σ2 f1þffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 3X d m¼1 xim �xjm � �2 σ2 m v u u t 0 @1 Ae −ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 3X d m¼1 xim�xjm ð Þ2 σ2 m s (10) Analyses that are close to those of various kinds of kernels of four different possible basis functions (indicted in Equations (11)–(14)) would be carried out in our work as well. From Equations (11–14),X¼x1;x2;. . . ;xn ð Þ0,X2¼ x2 11 x2 12 ��� x2 1d x2 21 x2 22 ��� x2 2d . . .. . .. . .. . . x2 T1x2 T2��� x2 Td 1 C C C C C C C C A 0 B B B B B B B B @ ,B¼b x1 ð Þ;ð b x2 ð Þ;... ;b xn ð ÞÞ0, and an empty matrix is referring to a matrix who has at least one of its dimensions that is zero. Asian Journal of Economics and Banking 71 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
precise estimates of commodity prices. Price projections are required for policymakers to know in order to conduct out market evaluations, create policies, put such policies into operation and make continual changes. Proper forecasting is crucial for a number of reasons, including maintaining a positive corporate environment and averting market collapse. To the best of their knowledge, econometric models—in particular, time-series econometric models—are the basis of the forecasting approach that the government and many investors use when a significant proportion of commodity prices are at risk. In the meanwhile, price estimates are nevertheless often based on expert opinions. This is supported by the possibility that developing, implementing, and maintaining econometric models and expert judgements will be comparatively simple. And because many of these models have been extensively used for decades by a broad range of forecast clients, it is likely that many of them have a respectable prediction accuracy. Given the growing affordability of computer resources and the solid foundation for expected nonlinear properties in price series of many commodities, machine learning models are widely recognised to have promise and worth more investigation. It may be difficult for some investors and policymakers to properly examine these models, however, since some decision-makers may still see them as too complicated forecasting tools. Actually, advanced investors and some governments have been increasingly interested in machine learning technologies in recent years. The study conducted here is a component of a broader approach that explores the possible uses of GP regression as a machine learning technique to solve the edible oil forecasting issue. The findings presented suggest that machine learning models may be worth looking into, maybe for a wider range of commodities, given the suggested method of building such a model and the shown strong forecast accuracy and stabilities. 7. Conclusion Commodity price forecasts are important to a number of agricultural industry stakeholders. In this research, we use weekly wholesale edible oil prices over a ten-year period, from January 1, 2010, to January 3, 2020, to predict the Chinese market. Not enough attention has been paid to the projections for this significant price series in the literature. As a forecasting tool, the GP regression is investigated using cross-validation and Bayesian optimisations, yielding accurate and dependable findings. More precisely, for the period spanning from January 4, 2019 to January 3, 2020, the generated models were capable of generating predicting results of the prices with an out-of-sample RRMSE of 5.0812%, RMSE of 4.7324 and MAE of 2.9382. The benefit of the generated models is shown by our benchmark study, which compares the GP regression models with a number of other time-series and machine learning models. This data may be used for independent technical predictions or combined with other estimations when doing policy research that calls for price trend views. The modelling methodology that underlies these predictions may also be used to forecast problems of a similar kind in other economic areas. The simplicity and convenience of use of this framework may be essential for a number of decision-making processes. The technique presented here may be limited by the absence of potentially useful predictive data from other economic aspects, which might improve prediction performance. Examining regression models using Gaussian processes and exogenous inputs might help resolve this potential limitation. This is an important area for further study if data on other economic factors could be gathered. Future research on additional Bayesian optimisation procedures may be worthwhile, since the current study focusses on the EI per second plus method for Bayesian optimisations. Given that the time period under examination concludes in January 2020, it may also be a valuable direction for future research that considers more current time periods for analysis. We use the Bayesian optimisation technique to guide our model development procedure. The method determines the best GPR model based on the training sample by experimenting with different kernels, basis functions, and matching parameters. Following the construction of the optimum model, out-of-sample forecasting is performed. A promising direction for future research is the notion of building AJEB 9,1 78 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
several models and selecting significant ones based on the model confidence set. The price projection problem for China’s wholesale edible oil market has been investigated in this paper. Investigating the possibilities of this empirical forecasting framework based on other commodity prices would be interesting. Abbreviation ARIMA – Autoregressive Integrated Moving Average VAR – Vector Autoregressive VECM – Vector Error Correction Model BDS – Brock–Dechert–Scheinkman GP – Gaussian Process EIPSP – Expected Improvement Per Second Plus EI – Expected Improvement EIPS – Expected Improvement Per Second RRMSE – Relative Root Mean Square Error RMSE – Root Mean Square Error MAE – Mean Absolute Error CV – Cross Validation AR – Autoregressive GARCH – Generalized Autoregressive Conditional Heteroskedasticity SVR – Support Vector Regression RT – Regression Tree MSE – Mean Square Error MDM – Modified Diebold-Mariano GPR – Gaussian Process Regression CART – Classification Analysis and Regression Tree References Alade, I.O., Zhang, Y. and Xu, X. (2021), “Modeling and prediction of lattice parameters of binary spinel compounds (am 2 x 4 ) using support vector regression with Bayesian optimization”, New Journal of Chemistry, Vol. 45 No. 34, pp. 15255-15266, doi: 10.1039/d1nj01523k. Amal, I. (2021), “Crude palm oil price prediction using multilayer perceptron and long short-term memory”, Journal of Mathematical and Computational Science, Vol. 11, pp. 8034-8045, doi: 10.28919/jmcs/6680. Brahim-Belhouari, S. and Bermak, A. (2004), “Gaussian process for nonstationary time series prediction”, Computational Statistics and Data Analysis, Vol. 47 No. 4, pp. 705-712, doi: 10.1016/j.csda.2004.02.006. Brahim-Belhouari, S. and Vesin, J.-M. (2001), “Bayesian learning using Gaussian process for time series prediction”, Proceedings of the 11th IEEE Signal Processing Workshop on Statistical Signal Processing (Cat. No. 01TH8563), IEEE, pp. 433-436, doi: 10.1109/SSP.2001.955315. Breiman, L. (2017), Classification and Regression Trees, Routledge. Brock, W.A., Scheinkman, J.A., Dechert, W.D. and LeBaron, B. (1996), “A test for independence based on the correlation dimension”, Econometric Reviews, Vol. 15 No. 3, pp. 197-235, doi: 10.1080/07474939608800353. Bull, A.D. (2011), “Convergence rates of efficient global optimization algorithms”, Journal of Machine Learning Research, Vol. 12. Dacha, K., Cherukupalli, R. and Sinha, A. (2021), “Food index forecasting”, in Applied Advanced Analytics, Springer, pp. 125-134, doi: 10.1007/978-981-33-6656-5_11. Asian Journal of Economics and Banking 79 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
Darekar, A. and Reddy, A. (2017), “Forecasting oilseeds prices in India: case of groundnut, forecasting oilseeds prices in India: case of groundnut (december 14, 2017)”, Journal of Oilseeds Research, Vol. 34, pp. 235-240, doi: 10.2139/ssrn.3237483. Diebold, F.X. and Mariano, R.S. (2002), “Comparing predictive accuracy”, Journal of Business and Economic Statistics, Vol. 20 No. 3, pp. 134-144, doi: 10.2307/1392185. Harvey, D., Leybourne, S. and Newbold, P. (1997), “Testing the equality of prediction mean squared errors”, International Journal of Forecasting, Vol. 13 No. 2, pp. 281-291, doi: 10.1016/S01692070(96)00719-4. Hasanov, A.S., Do, H.X. and Shaiban, M.S. (2016), “Fossil fuel price uncertainty and feedstock edible oil prices: evidence from mgarch-m and virf analysis”, Energy Economics, Vol. 57, pp. 16-27, doi: 10.1016/j.eneco.2016.04.015. Jaiswal, R., Jha, G.K., Kumar, R.R. and Choudhary, K. (2022), “Deep long short-term memory based model for agricultural price forecasting”, Neural Computing and Applications, Vol. 34 No. 6, pp. 4661-4676, doi: 10.1007/s00521-021-06621-3. Jaiswal, R., Jha, G.K., Kumar, R.R. and Lama, A. (2023), “Agricultural price forecasting using narx model for soybean oil”, Current Science, pp. 79-84. Jamieson, P., Porter, J. and Wilson, D. (1991), “A test of the computer simulation model arcwheat1 on wheat crops grown in New Zealand”, Field Crops Research, Vol. 27 No. 4, pp. 337-350, doi: 10.1016/0378-4290(91)90040-3. Jha, G.K. and Sinha, K. (2014), “Time-delay neural networks for time series prediction: an application to the monthly wholesale price of oilseeds in India”, Neural Computing and Applications, Vol. 24 Nos 3-4, pp. 563-571, doi: 10.1007/s00521-012-1264-z. Ji, M., Liu, P., Deng, Z. and Wu, Q. (2022), “Prediction of national agricultural products wholesale price index in China using deep learning”, Progress in Artificial Intelligence, Vol. 11 No. 1, pp. 121-129, doi: 10.1007/s13748-021-00264-0. Jin, B. and Xu, X. (2024a), “Contemporaneous causality among price indices of ten major steel products”, Ironmaking and Steelmaking, Vol. 51 No. 6, pp. 515-526, doi: 10.1177/ 03019233241249361. Jin, B. and Xu, X. (2024b), “Forecasts of coking coal futures price indices through Gaussian process regressions”, Mineral Economics. doi: 10.1007/s13563-024-00472-9. Jin, B. and Xu, X. (2024c), “Price forecasting through neural networks for crude oil, heating oil, and natural gas”, Measurement: Energy, Vol. 1, 100001, doi: 10.1016/j.meaene.2024.100001. Jin, B. and Xu, X. (2024d), “Forecasting wholesale prices of yellow corn through the Gaussian process regression”, Neural Computing and Applications, Vol. 36 No. 15, pp. 8693-8710, doi: 10.1007/ s00521-024-09531-2. Jin, B. and Xu, X. (2024e), “Wholesale price forecasts of green grams using the neural network”, Asian Journal of Economics and Banking. doi: 10.1108/AJEB-01-2024-0007. Kanchymalay, K., Salim, N., Sukprasert, A., Krishnan, R. and Hashim, U.R. (2017), “Multivariate time series forecasting of crude palm oil price using machine learning techniques”, IOP Conference Series: Materials Science and Engineering, Vol. 226, IOP Publishing. doi: 10.1088/1757-899X/ 226/1/012117. Karia, A.A., Abd Hakim, T. and Bujang, I. (2016), “World edible oil prices prediction: evidence from mix effect of ever difference on box-jerkins approach”, Journal of Business and Retail Management Research, Vol. 10. Khin, A.A., Mohamed, Z. and Malarvizhi, C. (2011), Forecasting Methods of Spot Palm Oil Prices: Comparative Techniques. Lama, A., Jha, G.K., Paul, R.K. and Gurung, B. (2015), “Modelling and forecasting of price volatility: an application of garch and egarch models”, Agricultural Economics Research Review, Vol. 28 No. 1, pp. 73-82, doi: 10.5958/0974-0279.2015.00005.1. Lama, A., Jha, G.K., Gurung, B., Paul, R.K., Bharadwaj, A. and Parsad, R. (2016), “A comparative study on time-delay neural network and garch models for forecasting agricultural commodity price volatility”, Journal of the Indian Society of Agricultural Statistics, Vol. 70, pp. 7-18. AJEB 9,1 80 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
Li, F., Gao, F. and Kou, P. (2015), “Integrating piecewise linear representation and Gaussian process classification for stock turning points prediction”, Journal of Computer Applications, Vol. 35, p. 2397, doi: 10.11772/j.issn.1001-9081.2015.08.2397. Meena, D.C., Singh, O. and Singh, R. (2014), “Forecasting mustard seed and oil prices in India using arima model”, Annals of Agri Bio Research, Vol. 19, pp. 183-189. Mishra, G. and Singh, A. (2013), “A study on forecasting prices of groundnut oil in Delhi by arima methodology and artificial neural networks”, Agris on-line Papers in Economics and Informatics, Vol. 5, pp. 25-34, doi: 10.22004/ag.econ.157527. Myat, A.K. and Tun, M.T.Z. (2019), “Predicting palm oil price direction using random forest”, 2019 17th International Conference on ICT and Knowledge Engineering (ICT&KE), IEEE, pp. 1-6, doi: 10.1109/ICTKE47035.2019.8966799. Neal, R.M. (1997), “Monte Carlo implementation of Gaussian process models for Bayesian regression and classification”, arXiv preprint physics/9701026. Neal, R.M. (2012), “Bayesian learning for neural networks”, Springer Science and Business Media, Vol. 118. Priyanga, V., Lazarus, T.P., Mathew, S. and Joseph, B. (2019), “Forecasting coconut oil price using auto regressive integrated moving average (arima) model”, Journal of Pharmacognosy and Phytochemistry, Vol. 8, pp. 2164-2169. Rahoveanu, A.T., Rahoveanu, M.M.T. and Ion, R.A. (2018), “Energy crops, the edible oil processing industry and land use paradigms in Romania–an economic analysis”, Land Use Policy, Vol. 71, pp. 261-270, doi: 10.1016/j.landusepol.2017.12.004. Raihan, A. (2023), “The dynamic nexus between economic growth, renewable energy use, urbanization, industrialization, tourism, agricultural productivity, forest area, and carbon dioxide emissions in the Philippines”, Energy Nexus, Vol. 9, 100180, doi: 10.1016/j.nexus.2023.100180. Raihan, A., Muhtasim, D.A., Farhana, S., Hasan, M.A.U., Pavel, M.I., Faruk, O., Rahman, M. and Mahmood, A. (2023), “An econometric analysis of greenhouse gas emissions from different agricultural factors in Bangladesh”, Energy Nexus, Vol. 9, 100179, doi: 10.1016/j. nexus.2023.100179. Ranguwal, S., Sidana, B.K., Singh, J., Sachdeva, J., Kumar, S., Sharma, R.K. and Dhillon, J. (2023), “Quantifying the energy use efficiency and greenhouse gas emissions in Punjab (India) agriculture”, Energy Nexus, Vol. 11, 100238, doi: 10.1016/j.nexus.2023.100238. Salman, N., Lawi, A. and Syarif, S. (2018), “Artificial neural network backpropagation with particle swarm optimization for crude palm oil price prediction”, Journal of Physics: Conference Series, Vol. 1114, 012088, doi: 10.1088/1742-6596/1114/1/012088. Shahriari, B., Swersky, K., Wang, Z., Adams, R.P. and De Freitas, N. (2015), “Taking the human out of the loop: a review of Bayesian optimization”, Proceedings of the IEEE, Vol. 104, pp. 148-175, doi: 10.1109/jproc.2015.2494218. Shamsudin, M.N. and Arshad, F.M. (2000), “Short term forecasting of malaysian crude palm oil prices”, URL: econ1, available at: upm.edu.my/fatimah/pipoc.html Silalahi, D.D. (2013), “Application of neural network model with genetic algorithm to predict the international price of crude palm oil (cpo) and soybean oil (sbo)”, in 12th National Convention on Statistics (NCS), pp. 1-2, Mandaluyong City, Philippine, October. Singh, A. (2021), “Comparison of artificial neural networks and statistical methods for forecasting prices of different edible oils in indian markets”, International Research Journal of Modernization in Engineering Technology and Science, Vol. 3, pp. 1044-1050. Wang, J., Dharmasena, S. and Bessler, D.A. (2013), “Price dynamics and forecasts of world and China vegetable oil markets”. doi: 10.22004/ag.econ.151150. Williams, C. and Rasmussen, C. (1995), “Gaussian processes for regression”, Advances in Neural Information Processing Systems, Vol. 8. Wulandari, R., Surarso, B., Irawanto, B. and Farikhin, F. (2021), “The forecasting of palm oil based on fuzzy time series-two factor”, Journal of Soft Computing Exploration, Vol. 2, pp. 11-16. Asian Journal of Economics and Banking 81 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
Yang, J., Su, X. and Kolari, J.W. (2008), “Do euro exchange rates follow a martingale? Some out-ofsample evidence”, Journal of Banking and Finance, Vol. 32 No. 5, pp. 729-740, doi: 10.1016/j. jbankfin.2007.05.009. Yeasin, M., Singh, K., Lama, A. and Paul, R.K. (2020), “Modelling volatility influenced by exogenous factors using an improved garch-x model”, Journal of the Indian Society of Agricultural Statistics, Vol. 74, pp. 209-216. Yu, T.-H.E., Bessler, D.A. and Fuller, S.W. (2006), “Cointegration and causality analysis of world vegetable oil and crude oil prices”. doi: 10.22004/ag.econ.21439. Zhang, Y. and Xu, X. (2020), “Curie temperature modeling of magnetocaloric lanthanum manganites using Gaussian process regression”, Journal of Magnetism and Magnetic Materials, Vol. 512, 166998, doi: 10.1016/j.jmmm.2020.166998. Corresponding author Xiaojie Xu can be contacted at: [email protected] For instructions on how to order reprints of this article, please visit our website: www.emeraldgrouppublishing.com/licensing/reprints.htm Or contact us for further details: [email protected] AJEB 9,1 82 Downloaded from http://www.emerald.com/ajeb/article-pdf/9/1/64/9700262/ajeb-06-2024-0070.pdf by ZBW German National Library of Economics user on 16 December 2025
