Dynamic pricing using flexible heterogeneous sales response models
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Aschersleben, Philipp; Steiner, Winfried J. Article — Published Version Dynamic pricing using flexible heterogeneous sales response models OR Spectrum Provided in Cooperation with: Springer Nature Suggested Citation: Aschersleben, Philipp; Steiner, Winfried J. (2024) : Dynamic pricing using flexible heterogeneous sales response models, OR Spectrum, ISSN 1436-6304, Springer, Berlin, Heidelberg, Vol. 46, Iss. 1, pp. 29-72, https://doi.org/10.1007/s00291-024-00756-0 This Version is available at: https://hdl.handle.net/10419/313823 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) OR Spectrum (2024) 46:29–72 https://doi.org/10.1007/s00291-024-00756-0 1 3 ORIGINAL ARTICLE Dynamic pricing using flexible heterogeneous sales response models PhilippAschersleben1 · WinfriedJ.Steiner1 Received: 31 March 2022 / Accepted: 12 February 2024 / Published online: 29 March 2024 © The Author(s) 2024 Abstract We combine nonparametric price response modeling and dynamic pricing. In particular, we model sales response for fast-moving consumer goods sold by a physical retailer using a Bayesian semiparametric approach and incorporate the price of the previous period as well as further time-dependent covariates. All nonlinear effects including the one-period lagged price dynamics are modeled via P-splines, and embedding the semiparametric model into a Hierarchical Bayesian framework enables the estimation of nonlinear heterogeneous (i.e., store-specific) immediate and lagged price effects. The nonlinear heterogeneous model specification is used for price optimization and allows the derivation of optimal price paths of brands for individual stores of retailers. In an empirical study, we demonstrate that our proposed model can provide higher expected profits compared to competing benchmark models, while at the same time not seriously suffering from boundary problems for optimized prices and sales quantities. Optimal price policies for brands are determined by a discrete dynamic programming algorithm. Keywords Sales response models· Functional flexibility· Store heterogeneity· Price dynamics· Price optimization· Discrete dynamic programming 1 Introduction Estimating price response functions to support pricing decisions is a highly relevant topic in the marketing literature and in marketing research practice. A price response function or, more generally speaking, a sales response function relates the sales of a brand to own and competitive brand prices, to accompanying marketing activities * Philipp Aschersleben [email protected] Winfried J. Steiner [email protected] 1 Department ofMarketing, Clausthal University ofTechnology, Julius-Albert-Str. 2, 38678Clausthal-Zellerfeld, Germany
30 P.Aschersleben, W.J.Steiner 1 3 (like feature or display advertising), and/or to further covariates accounting for timeor store-specific effects. In this paper, we focus on sales response modeling for fastmoving consumer goods that are sold by physical retailers and for which scanner data are nowadays widely available. Publications in this research area have particularly focused on the following dimensions: First, the specification of the ‘correct’ functional form for the relationship between sales and prices was dominated by strictly parametric modeling until the 2000s, starting with simple linear regression (lin-lin) models and followed by nonlinear parametric models, especially multiplicative (log-log) and exponential (loglin) models. To overcome the problem that parametric models can largely fail to approximate the true functional form inherent to real data (Härdle 1990; VanHeerde 1999; Leeflang etal. 2000), researchers started to apply more flexible nonparametric techniques, which are able to explore the functional shape directly from data instead of assuming a predefined parametric functional form (Hanssens etal. 2002). There is very clear evidence from many studies that these nonparametric methods can (greatly) improve the predictive model performance as well as expected profits over parametric modeling.1 According to VanHeerde etal. (2002), managers should rely on models which provide the most accurate predictions. Besides, more flexible nonparametric estimation methods have also become established in choice modeling, see, e.g., Abe (1991), Abe (1995), Abe etal. (2004), or Boztuğ etal. (2014). Second, it is well-known that the aggregation of store-level data across stores of a retail chain leads to biased estimates of the effects of marketing activities if these effects differ for individual stores (as one would expect) but even if marketing effects are homogeneous across stores (e.g., Krishnamurthi etal. 1990; Christen etal. 1997). Further, if marketing effects were different across individual stores of a retailer, pooling store-level data into a homogeneous model should also cause a bias in the effect estimates. Although empirical findings are not unequivocal, heterogeneous store-level sales response models that enable the estimation of store-specific effects often provided more or at least not less accurate sales predictions compared to their homogeneous (pooled) counterparts. More importantly, only heterogeneous sales response models allow for store-specific optimal pricing, also known as micromarketing pricing (Montgomery 1997). There is further empirical evidence that accounting for (unobserved) heterogeneity can also increase expected chain profits (e.g., Montgomery 1997; Lang etal. 2015). Third, it is reasonable to assume that past prices of brands or expected prices in future periods might affect a brand’s sales or price response in the current period. Therefore, the accommodation of price dynamics in sales response models has also been of a long tradition, be it via considering lagged or lead price covariates (e.g., VanHeerde etal. 2000, 2004), allowing for time-varying parameters (e.g., Foekens et al. 1999; Kopalle et al. 1999), or incorporating market-level reference price variables (e.g., Greenleaf 1995; Fibich etal. 2003; Aschersleben 1 We use the term nonparametric function to model the relationship between a specific metric predictor (like price) and a metric criterion variable (like sales) using flexible regression techniques, while the term semiparametric approach will describe a model where one or more predictor effects are captured parametrically, while others nonparametrically.
31 1 3 Dynamic pricing using flexible heterogeneous sales response… and Steiner 2022). Almost all of these dynamic approaches have been directed at either explaining structural effects of sales promotions (for example the lack of postpromotion dips in store data), or showing that dynamic pricing can increase profits over static models. Importantly, dynamic pricing allows the computation of optimal price paths, i.e., optimal prices that vary over time due to the consideration of time-dependent effects on price or sales response. From the previous explanations, it is clear that accommodating functional flexibility in price response, store heterogeneity in marketing effects, and price dynamics can improve both sales predictions and expected chain profits, which is why it seems promising to consider these features jointly in a store-level sales response model. However, not a single approach has yet accounted for all three dimensions simultaneously. Therefore, our first research question is whether a sales response model that combines the three components can provide higher expected chain profits compared to existing simpler (nested) models. In particular, our proposed approach will allow for the derivation of optimal price paths by modeling heterogeneous nonlinear price effects via nonparametric price functions (including the one-period lagged price effect), and we will assess whether this higher model complexity pays off compared to simpler models. As indicated above, some researchers went beyond the descriptive modeling stage and used an estimated sales or price response model as basis to subsequently derive optimal prices or price paths. Most of these optimization models assume linear price effects, i.e., price effects estimated by strictly parametric response modeling, and almost all of these parametric approaches also included price dynamics. The optimization results for many of these parametric approaches provide evidence or at least suggest the existence of corner solutions, i.e., that optimized prices hit the upper boundary of observed prices, which would make the price optimization exercise seem less useful. At this point, it is important to note that none of the dynamic optimization approaches has modeled price response more flexibly, like we propose in our model. Our second research question therefore relates to this boundary pricing problem, making our proposed approach recommendable only if optimal prices do not, or not always or not very frequently hit the boundaries of observed prices (and as well do not induce boundary solutions for related sales quantities). More explicitly, we want to analyze how prone our proposed model is to boundary price hits. If the boundary pricing problem is not an issue, we are probably the first to enable the calculation of optimal price paths for each of the stores of a retail chain via nonparametric heterogeneous price effects modeling. The rest of the paper is organized as follows: In Sect.2, we provide a review of the relevant literature for our approach. In Sect.3, we introduce our new dynamic semiparametric Hierarchical Bayesian sales response model together with nested model versions. We will use the latter for model comparison. We further propose a discrete dynamic programming approach for calculating optimal price paths. Section4 starts with a short description of the store-level scanner data used in our empirical study and provides details about model estimation and validation procedures. Moreover, optimal prices, optimized profits, and expected losses when not
32 P.Aschersleben, W.J.Steiner 1 3 using the proposed model are analyzed. In Sect.5, we conclude with a summary of findings and discuss limitations of our research. 2 Literature review Table 1 provides a summary of relevant articles related to sales response modeling in the field of fast-moving consumer goods that have combined at least two of the dimensions discussed in the introduction (i.e., functional flexibility in price response, store heterogeneity in marketing effects, price dynamics, price optimization) in their empirical applications. As mentioned before, nonlinear parametric models were state-of-the-art for a long time to capture price or more generally sales effects for fast-moving consumer goods (e.g., Hruschka 1997; Montgomery 1997; Foekens etal. 1999; Kopalle etal. 1999; VanHeerde etal. 2000, 2002; Hruschka 2006b; Andrews etal. 2008), and they still serve as benchmarks for the more sophisticated nonparametric approaches observed nowadays in many publications. Possible nonparametric operationalizations have ranged from kernel regression (e.g., VanHeerde etal. 2001) and neural nets (e.g., Hruschka 2006a, 2007) to splines such as stochastic cubic splines (Kalyanam and Shively 1998), B-splines (Hruschka 2000; Haupt and Kagerer 2012; Haupt etal. 2014) or Bayesian P-Splines (Steiner etal. 2007; Brezger and Steiner 2008; Weber and Steiner 2012; Lang etal. 2015; Weber etal. 2017). It is clear from the existing studies in this research area that more flexible estimation techniques are more powerful to uncover complex nonlinearities in sales response and, if these are at work, may lead to much better sales predictions and higher expected profits. The consideration of (unobserved) store heterogeneity in sales response modeling has become even more established, see the column ‘Store heterogeneity’ in Table1. Some researchers have only included store intercepts in order to account for heterogeneity in baseline sales across stores. The more advanced approaches to also capture store heterogeneity in marketing effects have all been embedded into Hierarchical Bayesian (HB) estimation frameworks,2 where some of them have accommodated both heterogeneity and functional flexibility in price effects (Hruschka 2006a, 2007; Lang etal. 2015; Weber etal. 2017) and the latter three of them price optimization in addition. Interestingly, some researchers found that accounting for store heterogeneity in a parametric sales response model did not necessarily improve the predictive model performance (Andrews etal. 2008; Weber and Steiner 2012, 2021), whereas Weber et al. (2017) did report major advantages from addressing store heterogeneity as soon as functional flexibility in price response was taken into account. Leeflang etal. (2000) already very early called for more intense research with regard to the comparison of heterogeneous versus homogeneous store-level sales response models, and we pick up their suggestion for model comparison in our empirical study later on. 2 While the single normal distribution is typically used in these HB models to represent store heterogeneity on the upper (population) level, Wedel and Zhang (2004) used a Dirichlet process prior to capture heterogeneity in price effects across stores more flexibly.
33 1 3 Dynamic pricing using flexible heterogeneous sales response… Table 1 Overview of empirical studies in the context of (store-level) sales response modeling considering at least two of the following dimensions: functional flexibility in price response, store heterogeneity in marketing effects, price dynamics, price optimization Study Price dynamics Functional flexibility Store heterogeneity Price optimization Greenleaf (1995) Reference prices – – Dynamic programming Montgomery (1997) – – Hierarchical Bayes (HB) Sequ. quadr. programming Foekens etal. (1999) Time-varying param. – Store intercepts – Kopalle etal. (1999) Time-varying param. – Store intercepts Dynamic programming VanHeerde etal. (2001) – Kernel regression Store intercepts – Fibich etal. (2003) Reference prices – – Analytical solution VanHeerde etal. (2004) Lead and lag prices Local polynom. regression – – Fok etal. (2006) Error correction model – Hierarchical Bayes (HB) – Hruschka (2006a) – Neural nets Hierarchical Bayes (HB) – Hruschka (2006b) – – Hierarchical Bayes (HB) Grid search Hruschka (2007) – Neural nets Store clusters, HB Improving hit-and-run Steiner etal. (2007) – Bayesian P-splines Store intercepts – Brezger and Steiner (2008) – Bayesian P-splines Store intercepts – Haupt and Kagerer (2012) – B-splines Store intercepts – Horváth and Fok (2013) Lag prices – Hierarchical Bayes (HB) – Haupt etal. (2014) – B-splines Store intercepts – Lang etal. (2015) – Bayesian P-splines Hierarchical Bayes (HB) Grid search Weber etal. (2017) – Bayesian P-splines General heterogeneity SUR model, HB Evolutionary algorithm Aschersleben and Steiner (2022) Reference prices Bayesian P-splines – – This study Lag prices Bayesian P-splines Hierarchical Bayes (HB) Dynamic programming
34 P.Aschersleben, W.J.Steiner 1 3 Price dynamics were only seldom considered in models with nonlinear price effects, VanHeerde etal. (2004) and more recently Aschersleben and Steiner (2022) are the exceptions with their empirical studies. These two approaches, however, did not account for store-specific marketing effects. Fok etal. (2006) and Horváth and Fok (2013) proposed Hierarchical Bayesian vector autoregression models to analyze brand-/category-specific differences in dynamic price effects on sales. But the focus of these two studies was more on the investigation of brandand category-specific characteristics as moderators for differences in ownor cross-price effects between brands and product categories rather than on within-brand store heterogeneity in price response (especially as Fok etal. 2006 used data from only one single store). The two approaches that combined price dynamics and store heterogeneity (Foekens etal. 1999; Kopalle etal. 1999) only included store intercepts to consider differences in baseline sales across stores. Similar to Leeflang etal. (2000), we here see the need for a more intense research to assess the effect of price dynamics in storelevel sales response models that as well account for store-specific effects and ideally also for functional flexibility in price response. Generally, we observe quite different strategies to accommodate price dynamics in sales response models: via lagged and lead prices (VanHeerde etal. 2004; also, e.g., VanHeerde etal. 2000), via time-varying parameter models (Foekens etal. 1999; Kopalle etal. 1999; also, e.g., Ataman etal. 2010), via vector-autoregressive specifications (Fok etal. 2006; Horváth and Fok 2013; also, e.g., Nijs etal. 2001), as well as via market-level reference prices (Greenleaf 1995; Fibich etal. 2003; Aschersleben and Steiner 2022). Only the minority of the papers collected in Table1 provided normative implications with respect to expected profits or toexpected losses from not using the sales response model with the best predictive performance. Like the milestone article on micro-marketing pricing strategies by Montgomery (1997), all more recent optimization approaches used a Hierarchical Bayesian estimation framework to allow for store-specific (or at least clusterwise) optimal pricing (Hruschka 2006a, b, 2007; Weber etal. 2017; Lang etal. 2015). None of them, however, incorporated price dynamics for price optimization. On the other hand, the three remaining approaches that considered dynamic price effects on sales response (Greenleaf 1995; Kopalle etal. 1999; Fibich etal. 2003) did not allow for price optimization by modeling nonlinear and/or heterogeneous store-level pricing effects. But, the findings from the latter three studies still indicated that accommodating price dynamics can increase profits over static price optimization. As another important issue, several of the price optimization studies seem to provide evidence of corner solutions for optimized prices, i.e., that optimized prices would have hit the upper boundary of the observed price range if this upper boundary had been imposed as a constraint in the model.3 Interestingly, this applies to all three papers that incorporated price dynamics for price optimization (Greenleaf 1995; Kopalle etal. 1999; Fibich etal. 2003), and these three dynamic optimization approaches have further in common that they used linear pricing effects (i.e., price effects estimated by a parametric sales response model). The corner solution problem seems as well evident in the paper of Montgomery (1997), who did not consider 3 We thank one referee for pointing us to this boundary price phenomenon.
35 1 3 Dynamic pricing using flexible heterogeneous sales response… price dynamics but also used a parametric sales response model (an exponential one) to determine optimal prices. Three optimization approaches were based on semiparametric sales response models, where nonlinear price effects were estimated using neural nets or splines in the descriptive model stage (Hruschka 2007; Lang etal. 2015; Weber etal. 2017). Indeed, the findings reported in Hruschka (2007) and Weber etal. (2017) suggest that flexible estimation methods might be less affected by the boundary pricing problem.4 Note that we found no evidence that optimized prices hit the lower bound of observed prices. And, neither paper further allows conclusions about boundary effects for predicted sales quantities based on optimized prices. At this point, it is very important to mention that the findings from our explorative research on boundary price effects should be treated with caution since the corner solution problem is not explicitly addressed in any of the papers and information on boundary effects is sparse. In general, optimal prices depend on price elasticities and on variable costs (in our case the wholesale prices of a retailer). Still, the price elasticity in parametric models is much more “rigid”, as it depends on one or only few price parameter estimates which globally determine the shape of the response function over the entire observed price range. Nonparametric estimation techniques like splines, on the other hand, fit the data locally, entailing greater flexibility for uncovering complex price response patterns and allowing locally varying price elasticities. This locally fitting property can make flexible models less prone to corner solutions for optimal prices. Note that none of the dynamic optimization approaches has modeled price response more flexibly, like we propose in our dynamic model. And different from all previous studies in this field, we will put a special focus on the boundary pricing problem in our empirical study. Based on the literature review, the central contribution of this paper is that our proposed model will allow for price optimization by modeling heterogeneous nonlinear pricing effects through nonparametric functions for own-, cross-, and lagged price response. In other words, we are probably the first to enable the computation of truly dynamic price paths5 for each store of a retail chain due to heterogeneous nonlinear pricing effects embedded in a semiparametric sales response model. None of the previously proposed sales response models has combined the three features functional flexibility, store heterogeneity, and price dynamics to derive optimal pricing strategies for fast-moving consumer goods. 4 Hruschka (2007) determined an eight-cluster solution for the stores of a retail chain with equal optimal brand prices for all stores within a cluster, and optimized prices largely varied across clusters for most brands considered. On the one hand, the large variation of optimal prices might indicate less boundary problems. It remained unclear, however, how often the optimal cluster prices hit the upper boundary of observed prices at the individual store level. Weber etal. (2017) provided a plot of optimal store-specific brand prices for one brand as example, suggesting a fairly low number of boundary effects for their flexible model. 5 Note that time-varying optimal prices can also result from static price optimization if variable costs are not constant over time.
36 P.Aschersleben, W.J.Steiner 1 3 3 Model specification andoptimization approach In the following, we introduce a semiparametric, heterogeneous, and dynamic storelevel sales response model. We use Bayesian P-splines as nonparametric method to estimate immediate and lagged price effects flexibly, allowing us to uncover possibly exceptional pricing effects (like distinct threshold or saturation effects) directly from the data. An advantage of P-splines is that they can easily be constrained to provide monotonic shapes of price response (Brezger and Steiner 2008), which is reasonable from an economic point of view for fast-moving consumer goods. Following Lang etal. (2015), heterogeneity in price response across stores is captured via store-specific scaling factors, which can be as well easily embedded as additional parameters into the Gibbs sampling procedure of our Hierarchial Bayesian estimation framework. Since the shape of the price response is pooled over stores by this heterogeneity specification (i.e., functional flexibility and heterogeneity in price response are not decoupled), we further provide robustness checks by varying the number of knots and the degree of the underlying B-spline basis functions. Dynamics are accommodated via the one-period lagged own-item price since more time lags were not supported by the data for almost all brands considered. We will also check the assumptions required for a proper estimation of the models (multicollinearity, heteroskedasticity, autocorrelation). We apply a discrete dynamic programming algorithm to derive optimal price paths. 3.1 Semiparametric heterogeneous dynamic model We use the following additive sales response model with smooth nonparametric and multiplicative random price effects.6 Accordingly, the (log) unit sales of a brand in a specific store and week are assumed to depend on ownand competitive price effects (where the latter are captured at the quality tier level), as well as a dynamic, one-week lagged own-price effect. We further include promotional activities for the brand (use of a display, odd pricing) and accommodate seasonality effects via a smooth monthly trend and unobserved store-specific effects. Models are estimated for each brand separately, thus brand indices are omitted. where f0 is an unknown smooth nonlinear time trend of the calendar month ( mt ); f1 is an unknown smooth nonlinear decreasing function of the brand’s own price ( pst ); f2 is an unknown smooth nonlinear increasing function of the brand’s oneweek lagged own price ( ps;t−1 ); fci are unknown smooth nonlinear increasing functions for cross-price effects ( pci st ) captured at the level of the price-quality tiers ci , (1) log(q st )=𝜂 st +𝜀 st =f0(mt)+(1+𝛼s1)f1(pst)+(1+𝛼s2)f2(ps;t−1 ) + ∑i (1+𝛼s,c i )fc i (pci st)+v� st 𝜸s+𝛽s+𝜀st, 6 Effects of marketing instruments other than prices as well as store intercepts are captured parametrically, which is why the model is called a semiparametric model, also see footnote 1.
43 1 3 Dynamic pricing using flexible heterogeneous sales response… version to compare probabilistic and deterministic outcomes. Then, it can be used to measure the squared error between a probabilistic forecast and a point measure, the former relating to the draw-based predictions for a brand’s unit sales and the latter relating to the corresponding observed sales of a brand’s unit sales in a certain store and week in our context. In this variant, the CRPS can be seen as a continuous version of the well-known Brier score that is defined as the mean squared error between a discrete outcome y and a corresponding probabilistic forecast based on the cdf F (see, e.g., Gneiting and Raftery 2007; Jordan 2016,pp.37–38): In empirical applications, the underlying true distribution F is commonly unknown. Using an empirical cdf Fecdf n , based on n observations, as an approximation of F , the CRPS score can be computed by the following consistent modification of Eq. (11) (Jordan 2016, Sect.6; Krüger etal. 2021): Figure1 illustrates the concept behind the CRPS for a better understanding: as the score is calculated as the integral over squared differences between two distributions, with one of them being degenerated to a point measure, the CRPS represents the size of the gray-shaded area between the cumulative distribution functions. In other words, the distribution function for the predictive distribution is compared to the ideal distribution function of a point mass in a newly observed value by forming the squared distance and integrating over it. The more the two distribution functions overlap, i.e., the more they coincide, the lower is the distance between them. In our case, the draws saved from the Markov chain after convergence are samples from an (unknown) posterior distribution Fpost , and using this scoring rule we again account for parameter uncertainty. Note that we obtain one CRPS value as “counterpart” to every observation qst , and averaging the CRPS values again across all weeks and stores provides us with a mean CRPS value ( MCRPS ) for each holdout: (11) CRPS (F,y)= �ℝ (F(z)−1z≥y)2dz . (12) CRPS ( Fecdf n,y)= 2 n2 n ∑ i=1 (X(i)−y) ( n1y<X(i)−i+1 2 ). Fig. 1 Example figure for the CRPS scoring rule: the gray-shaded area represents the integrated squared difference between the empirical cumulative distribution function Fecdf n of a sample of F and the singlepoint empirical cdf (ecdf) of observation y
44 P.Aschersleben, W.J.Steiner 1 3 Finally, averaging the individual MCRPS values over the C folds results in the AMCRPS value as our second predictive accuracy measure corresponding to the ARMSE . We use the R package scoringRules (Jordan etal. 2019) to compute the CRPS values according to Eq. (12). 4.3 Estimation results andpredictive model performance 4.3.1 Estimation results Estimation results for the most complex DynFlexHet model as described in Eq. (1) are illustrated for the brand “Minute Maid” as example in Fig.2. Remember that this model accounts for heterogeneous effects across stores, functional flexibility in price effects, and price dynamics represented by the one-period lagged own-price effect. Depicted are the heterogeneous spline estimates together with the partial residuals as well as violin plots for the display and 9or 99-ending price effects. (13) MCRPS =1 S S ∑ s=1 1 Ts T s ∑ t=1 CRPS(Fpost st ,qst) , (14) ⇒ AMCRPS =1 C C ∑ c=1 1 S S ∑ s=1 1 T s T (c) s ∑ t=1 CRPS(Fpost;(c) st ,q(c) st ) . Fig. 2 Estimation results for the DynFlexHet model using the brand "Minute Maid" as example: estimated effects and partial residuals for own, lagged, and competitive price variables (accounting for heterogeneity via scaled spline slopes) as well as estimated effects for display and odd price endings (heterogeneous effects displayed by violin plots)
45 1 3 Dynamic pricing using flexible heterogeneous sales response… First, estimates and the corresponding effect sizes suggest face validity. The strongest effect is observed for the own-price effect, and it is also the one of the price effects showing a larger amount of heterogeneity across stores as is represented by the relatively large bandwidth of the scalings of the spline functions. More specifically, 21% of (a total of 3240) pairwise 80% credible intervals for the estimated scaling factors do not overlap between stores, confirming that own-price effects significantly differ between many stores. We will further discuss below in more detail that, as a rule, ownprice effects for all eight brands are much more heterogeneous across stores than both cross-price effects and lagged price effects (see Fig.4). The overall slope of the ownprice effect reveals a threshold effect near 1.75$ beyond which sales of “Minute Maid” strongly increase. The lagged price effect shows a steeper slope for small price levels than for medium and high price levels (where the lagged price effect is rather flat), and in addition a less distinct threshold effect at 2.50$. This suggests that customers respond less to high prices of “Minute Maid” in the previous period, in terms of buying more of the brand in the current period, than to a very low previous price in terms of buying less of “Minute Maid” in the current period. In other words, less purchases seem to be postponed if the previous price is very high while a certain level of stockpiling appears to occur in case of a very low one, albeit these conclusions are difficult to validate with aggregate data. Some purchase deceleration however obviously exists for prices larger than 2.50$. Compared to the own-price effects (including the lagged price effect), cross-price effects turn out rather flat with sales of “Minute Maid” being least (most) affected by the private label brand (premium brands). Interestingly, the sales effect of a 9-ending price (excluding a price ending in 99) is frequently (much) larger than the one for a 99-ending price (excluding other price endings in 9) and turns out significantly positive for 93.8 % of the stores ( D99 : 59.3 %). The display effect is not significant for all stores but one. Figure 3 displays the lagged price effects in the four dynamic models for the brand “Minute Maid”, see the homogeneous and heterogeneous parametric models (top left and right panels) and the corresponding flexible model versions (bottom panels). The advantage of using a flexible approach instead of modeling the effect in a parametric way becomes obvious again, as was already visible in Fig.2: nonlinearities with piecewise steeper or flatter slopes can be modeled with Bayesian P-splines7 but not with parametric functions. Nevertheless, a closer look at the bottom left panel reveals wide confidence bands at the lowest price levels, where only a few observations are available.8 Note that there is virtually no difference in the estimated own-price effects between the respective static and dynamic model variants, as illustrated in Fig.9 in the Appendix. This implies that the immediate own-price effect remains highly robust even if the models are extended to capture price dynamics. Plots of the estimated lagged price effects for all eight brands obtained by the DynFlexHet models 7 Note that we use the centered sampling method for spline estimation provided in BayesX such that the sum of the spline function over all observations equals zero. 8 We did not include the confidence bands in the lower right panel for the flexible heterogeneous dynamic model to prevent clutter. The corresponding confidence bands at the lower bound of the price range look highly similar.
46 P.Aschersleben, W.J.Steiner 1 3 are provided in Fig.10 in the Appendix, showing very different shapes across brands and hence confirming the benefits of nonparametric estimation as well for the dynamic price effects. The complete estimation results for all brands are available from the authors upon request. Figure4 shows density plots of the estimated store-specific scaling or random effects parameters for the DynParHet and DynFlexHet models in order to get a deeper understanding about how much heterogeneity is inherent to our data and how Fig. 3 Estimated effect of the lagged price on the sales of the brand “Minute Maid” for the different dynamic model variants Fig. 4 Density plots of store-specific scaling or random effect parameter estimates (centered around their mean) for the DynParHet and DynFlexHet models
47 1 3 Dynamic pricing using flexible heterogeneous sales response… this heterogeneity is handled by these two model variants9. Some points from Fig.4 are in particular noteworthy. First, as already illustrated in Fig.1 for the brand “Minute Maid”, heterogeneity across stores is especially distinct for the own-price effect; this applies more or less to all eight brands independent whether the own-price effect is captured flexibly or parametrically. Second, heterogeneity between stores does not seem to be an issue for both the lagged price effect nor for cross-price effects if price effects are modeled parametrically ( DynParHet ). However, once price response is modeled flexibly both lagged price effects and cross-price effects do reveal a moderate amount of store heterogeneity for all brands. Lang etal. (2015) have previously provided a possible explanation for this phenomenon, albeit not in the context of price dynamics. And third, there is also some heterogeneity across stores in the effects of using odd prices and a display for all brands and, like for the own-price, the effects are rather stable across the two models ( DynParHet , DynFlexHet ). 4.3.2 Predictive performance We evaluate the predictive performance of the competing models in terms of the Average Root Mean Squared Sales Prediction Error ( ARMSE ) and the Average Mean Continuous Ranked Probability Score ( AMCRPS ), which were introduced in Sect.4.2. We use a 9-fold subsampling, where 8/9 of the data (randomly drawn without replacement) are used each time for model estimation and the remaining 1/9 of the data serves as holdout sample. The predictive validity results for the two performance measures are reported in Table3 for the ARMSE measure and in Table4 for the AMCRPS measure, together with relative improvements or deteriorations of the different models in predictive accuracy compared to the proposed DynFlexHet model (for each brand and aggregated at the median) and frequencies in how many cases each model provided the best performance (row ‘# best’). From Table3, we at first observe that different models perform best for different brands in terms of the ARMSE measure, i.e., there seems to be no clear favorite model under this error measure at first glance. For seven out of eight brands, a flexible model version provides the best predictive accuracy, as does a dynamic model version for six of the brands, while for three brands the best model is a heterogeneous one. On the other hand, when aggregated across brands, the proposed DynFlex - Het model is only slightly outperformed by the DynFlexHom model and provides the second-best predictions, see the median in relative changes in the ARMSE measure at the bottom of the table. In Table4, the picture is completely different and very clear. Using the CRPS measure to assess the predictive model performance, the most complex model ( DynFlexHet ) always provides the most accurate predictions, and concerning the three model dimensions functional flexibility, store heterogeneity, and price dynamics the more complex model variant outperforms its simpler counterpart (i.e., 9 We here abstain from displaying the corresponding density plots for the static heterogeneous counterpart models ( StatParHet and StatFlexHet ) because there is virtually no difference between the static and dynamic model variants except that the lagged price is not included in the static models.
48 P.Aschersleben, W.J.Steiner 1 3 heterogeneous models perform better than homogeneous ones, flexible models better than parametric ones, and dynamic models better than static ones). To validate our conjecture that the ARMSE measure is more prone to extreme observations (e.g., very high sales at very low prices) than the AMCRPS measure, we computed the Average Root Median Squared Error ( ARMedSE ) as further performance measure, which is known to be much more robust against extreme observations compared to the ARMSE (Franses and Ghijsels 1999). According to the ARMedSE measure, the DynFlexHom model performs best for all brands, and dynamic and flexible models almost always perform better than their static and parametric counterparts – just as when using the AMCRPS . Different from the AMCRPS results, homogeneous models are preferred to heterogeneous ones in Table 3 Out-of-sample predictive performance of the competing models evaluated by the Average Root Mean Squared Sales Prediction Error ( ARMSE ) in holdout samples compared to the proposed DynFlex - Het model. Best models per brand are marked in bold The bottom lines show the median relative improvement / deterioration of each model (i.e., across brands) compared to the DynFlexHet model and the number of times a considered model outperformed the other models Stat Dyn Par Flex Par Flex Hom Het Hom Het Hom Het Hom Het Flor. Natrl. 26.5 26.8 25.4 28.5 23.8 23.824.6 26.6 (−0.2%) (+1.1%) (−4.5%) (+7.5%) (−10%) (−10%) (−7.2%) – Tropic. Pure 57.4 57.1 53.0 52.1 55.9 55.1 52.2 51.1 (+12%) (+12%) (+3.8%) (+1.9%) (+9.4%) (+7.8%) (+2.1%) – Citrus Hill 120.7 122.1 101.1 135.0 104.7 107.1 86.0107.4 (+12%) (+14%) (−5.8%) (+26%) (−2.5%) (−0.3%) (−20%) – Flor. Gold 57.5 57.7 55.6 55.9 55.5 55.8 53.453.8 (+ 6.8 %) (+ 7.3 %) (+ 3.4 %) (+ 3.9 %) (+ 3.1 %) (+ 3.7 %) (− 0.8 %) – Min. Maid 53.8 53.3 52.5 52.054.2 53.6 53.2 52.8 (+1.8%) (+0.9%) (−0.6%) (−1.6%) (+2.6%) (+1.4%) (+0.7%) – Tree Fresh 83.6 83.1 64.483.7 81.1 80.6 64.5 79.7 (+4.8%) (+4.2%) (−19%) (+5.0%) (+1.7%) (+1.1%) (−19%) – Tropicana 110.4 112.3 105.4 108.1 100.9 101.9 93.794.6 (+17%) (+19%) (+11%) (+14%) (+6.7%) (+7.7%) (−1.0%) – Dominick’s 311.7 316.1 307.5 313.8 300.9 306.6 278.5281.1 (+11%) (+12%) (+9.4%) (+12%) (+7.0%) (+9.1%) (−0.9%) – median +8.9% +9.5% +1.4% +6.2% +2.8% +2.5% −1.0% – # best – – 1 1 – 1 4 1
49 1 3 Dynamic pricing using flexible heterogeneous sales response… nearly all cases.10 Based on the overall consideration of the results for all three predictive validity measures ( ARMSE , AMCRPS , ARMedSE ) we can recommend the use of either the DynFlexHet or DynFlexHom for sales predictions. 4.3.3 Further tests androbustness checks We also conducted tests for all models (by brand) to assess the assumptions required for a proper estimation of them (multicollinearity, heteroskedasticity, Table 4 Out-of-sample predictive performance of the competing models evaluated by the Average Mean Continuous Ranked Probability Score ( AMCRPS ) in holdout samples compared to the proposed Dyn - FlexHet model. Best models per brand are marked in bold The bottom lines show the median relative improvement / deterioration of each model (i.e., across brands) compared to the DynFlexHet model and the number of times a considered model outperformed the other models Stat Dyn Par Flex Par Flex Hom Het Hom Het Hom Het Hom Het Flor. Natrl. 10.1 9.8 9.5 8.6 9.2 8.8 9.0 8.1 (+24%) (+20%) (+17%) (+6.0%) (+14%) (+8.1%) (+11%) – Tropic. Pure 24.4 23.2 22.8 20.8 23.3 21.8 22.1 19.9 (+23%) (+17%) (+15%) (+4.5%) (+17%) (+9.6%) (+11%) – Citrus Hill 28.9 27.8 24.7 21.9 24.8 23.6 21.5 18.8 (+53%) (+47%) (+31%) (+16%) (+32%) (+25%) (+14%) – Flor. Gold 20.4 19.9 19.5 18.8 19.6 19.2 18.6 18.0 (+ 14 %) (+ 11 %) (+ 8.5 %) (+ 4.4 %) (+ 8.8 %) (+ 6.4 %) (+ 3.5 %) – Min. Maid 21.8 20.6 20.9 19.3 20.8 19.5 20.1 18.6 (+17%) (+11%) (+12%) (+3.9%) (+12%) (+4.8%) (+8.1%) – Tree Fresh 19.2 18.5 15.9 13.7 18.8 18.1 15.8 13.5 (+42%) (+37%) (+17%) (+0.8%) (+39%) (+33%) (+17%) – Tropicana 55.0 53.3 52.0 49.8 50.9 49.4 46.7 44.2 (+25%) (+21%) (+18%) (+13%) (+15%) (+12%) (+5.8%) – Dominick’s 150.9 148.3 145.0 140.2 144.4 141.4 134.6 128.2 (+18%) (+16%) (+13%) (+9.4%) (+13%) (+10%) (+5.0%) – median +24% +19% +16% +5.2% +14% +9.9% +9.5% – # best – – – – – – – 8 10 The ARMedSE results can be obtained from the authors upon request. We further conducted a small simulation study to test the sensitivity of the CRPS in comparison to the RMSE against outliers. For this, we first generated 100 observations from the standard normal and added in a second step one or two outliers to the observations. The RMSE increased considerably, while the CRPS remained highly robust every time. This means that the CRPS weighs central observations more strongly than more extreme observations compared to the RMSE .
50 P.Aschersleben, W.J.Steiner 1 3 autocorrelation), and briefly summarize results for the proposed DynFlexHet model in the following.11 Multicollinearity was not a problem in any case: variance inflation factors (VIF) were never higher than 2.5 across brands and thus far from being critical, where as a rule the highest VIFs could be observed for the own-item price covariates (immediate and lagged prices). Heteroskedasticity was evaluated using the Brown-Forsythe test, which is robust against deviations from normality (Brown and Forsythe 1974). We applied the test to compare the variance of the residuals against the fitted values, and results were more mixed here. The Brown-Forsythe test was not significant for four brands, indicating that heteroskedasticity is not an issue there (e.g., with p=0.82 for “Minute Maid”), while it turned out significant for the other four brands ( p<0.01 ). We further used a modified version of the Durbin-Watson test for panel data to assess autocorrelation of the residuals, which is as well robust against deviations from the normal distribution (Bhargava etal. 1982; Ali and Sharma 1993). Given the number of covariates and sample size per brand (between 6624 and 6827 observations), the lower critical value around 1.85 is slightly undercut for two brands ( 𝛼=0.05 ), indicating positive autocorrelation only for these two brands. We found no evidence for negative autocorrelation across brands. Heteroskedasticity and autocorrelation lead to biased estimates of standard errors and can affect the validity of statistical tests, with the result of a possibly misleading statistical inference. However, as our focus lies on the predictive model performance and related profit implications, violations of these assumptions seem less severe. We further performed robustness checks starting with modifications for the specification of the P-splines. By default, we used 20 knots and B-spline basis functions of degree 3 to estimate price effects flexibly. In order to assess the sensitivity of the predictive power ( AMCRPS , ARMSE ) of the DynFlexHet and DynFlexHom models depending on the number of knots and/or the degree of the B-spline basis functions, we further estimated them with a lower degree of the B-splines (degree 1) and/or more knots (40 knots, corresponding to the suggested upper bound by Eilers and Marx 1996) for the immediate and lagged own-price effects. The CRPS scores turned out highly robust for all brands and, except for one brand, the RMSE measure as well, indicating that 20 knots are sufficient and B-splines of degree 1 would perform comparable. Details are provided in Tables7 and 8 in the Appendix. Finally, we considered only one time lag in the dynamic model versions to consider price dynamics. By definition, this implies that only short-term price-change or short-lived post-promotion effects can be analyzed, while longer persisting marketing effects cannot be accommodated. To test if the data support higher order autoregressive structures, we added two more lags for the own price to the Dyn - FlexHet and DynFlexHom models (with the default settings of 20 knots and degree 3 for the P-splines), capturing the higher-order lag dynamics (lag 2, lag 3) parametrically to keep the model complexity manageable. Again, the CRPS scores remained extremely robust in both models except for the store brand, where adding a second price lag somewhat improved the predictive performance. In terms of ARMSE , 11 The complete results for all tests and models can be obtained from the authors upon request. Note that multicollinearity measures only differ across brands and not across models, because all models share the same covariates per brand.
51 1 3 Dynamic pricing using flexible heterogeneous sales response… adding more lags did not pay off for all brands in the DynFlexHom model and for six brands (including the store brand) in the DynFlexHet model. Details about these dynamic model extensions are provided in Tables9 and 10 in the Appendix. 4.4 Optimization 4.4.1 Settings andconstraints In order to find optimal (i.e., profit maximizing) store-specific prices for each brand m , we use the dynamic program (P) introduced in Sect.3.3. As the action space for determining a brand’s optimal price in period t is independent of the current state (i.e., the brand’s price in the previous period), the set of generally possible prices for the brand is constant over time, hence in each period the same with Xt(st)=X . And since st+1=g(xt,st)=xt , the state space equals the action space ( S=X ). For each brand and store, we set individual state and action spaces as to contain a grid of possible prices in 0.01$ steps between the minimum and maximum observed price levels, denoted as [ p obs m,i,min ; p obs m,i,max ], for the considered brand m in store i : The restriction of optimized prices to the range of observed price levels is chosen to preserve the retailer’s price image perceived by customers. It further helps to obtain realistic predictions, especially when using flexible functions that fit the underlying data locally such that extrapolations could be still more problematic compared to parametric functions. To start the recursive process and solve the dynamic optimization problem, the initial state s1=p0 has to be fixed. As starting value, we decided to use the price in the middle of the observed price range of a considered brand and store. By conducting some sensitivity analyses, we checked that optimal price paths were fairly robust against variations of this starting value. We further set the discount rate r to 0 without loss of generality. As a second constraint, we limit the predicted sales q of a brand in a certain store to the largest number of sales qobs m,i,max observed for this brand in this store in our data. Using this second constraint lets us stay conservative and prevents us from an overestimation of a brand’s unit sales beyond the upper bound of observed unit sales per week and store for (very) low prices. This guarantees realistic sales scenarios as observed in the data on the one hand, but may trim the greater flexibility of the spline estimates on the other hand. In other words, if the sales forecast for a brand in a store from one of our estimated models is higher compared to the maximum observed sales, we truncate the prediction to the maximum of observed sales: (15) S =Smi = { pobs m,i,min,pobs m,i,min +0.01, …,pobs m,i,max } =Xmi =X . (16) q =min { q,qobs m,i,max }.
52 P.Aschersleben, W.J.Steiner 1 3 Note that predictions q are again made at the drawlevel as already described in Eq. (8), i.e., we integrate the optimization criterion over the posterior distribution of the parameters by averaging the optimization criterion across draws, before optimizing it. To further speed up optimization, we use every tenth draw from the saved draws of the Markov chain after convergence. 4.4.2 Optimal price paths We use forward recursion to solve the optimization problems and generate 5184 optimal price paths in total (81 stores times 8 brands times 8 sales response models). It seems reasonable to assume that optimal price paths might vary depending on the store-specific scaling factors (1+𝛼s1) . For illustration, we continue with the brand “Minute Maid” as example and choose for this brand three out of the 81 stores with a low vs.medium vs.high scaling factor for the own-price effect estimated by the DynFlexHet model. Figure5 shows optimized price paths for the brand “Minute Maid” for the three selected stores and all heterogeneous model variants. Further depicted are variable costs (also provided in the data) and observed prices as benchmarks for optimized prices (panels in the top row). Observed prices show a very high variation during the first 25 weeks, switching back and forth mostly between only two different price levels with no more than three consecutive weeks at the high price level and no more than two consecutive weeks at the lower price level (note that the high price level is highest in the store with the high scaling parameter, 3.17$, and lowest for the store with the low scaling parameter, 2.62$.) This clear hi-lo price pattern dilutes after the first 25 weeks, but some price levels still show up more often than others (e.g., 1.99$ in all three selected stores), and price cuts of varying depths continue to occur. Also note that the general price level of “Minute Maid” decreases with decreasing variable costs (wholesale prices), as expected. The optimized prices resulting from the StatParHet model (red lines in the middle row) “follow” the costs very closely, i.e., increased costs lead to increased prices and vice versa, and the correlation between observed costs and optimized prices (0.92) is considerably higher than the correlation between observed costs and observed prices (0.60). It is further reasonable that optimized prices end in 9 or 99 since odd pricing was shown to have positive effects on expected sales (see Sect.4.3.1). This optimized price pattern in parts also holds for the StatFlexHet model (middle row; blue lines), but with a more clear hi-lo pricing scheme and a smaller number of different price levels. Completely different optimal price paths are obtained for the dynamic models. For the store with the low scaling parameter for the own-price effect (bottom left panel), optimized prices resulting from the DynParHet model vary in a much smaller bandwidth. Small price cuts are observed primarily in the second half of the data time window, and the high price level is constant over time. Optimized prices obtained from the DynFlexHet model hardly show any variation (bottom row; blue lines). Deep price cuts are observed in only three weeks, in all other weeks the optimal price is set consistently high at the same unique price level as suggested by the DynParHet model (bottom row; red lines).
59 1 3 Dynamic pricing using flexible heterogeneous sales response… different predictive validity measures used. Nevertheless, more complexity always provided a better forecasting accuracy according to the AMCRPS measure, suggesting that the DynFlexHet model was the model of choice here, while the ARMedSE measure unequivocally tended to the DynFlexHom model instead. Regarding the ARMSE measure, at least one of these two models ( DynFlexHet , DynFlexHom ) also performed best for five brands as well as not dramatically worse for the other three brands. Therefore, without loss of generality, we choose the DynFlexHet model to compute expected losses next. In particular, we take the optimal prices for a brand determined by the DynFlex - Het model and calculate the total (store) profit for this brand (i.e., summing up over all weeks) under the DynFlexHet model. Expected losses in total (store) profit can then be computed when using optimal prices determined by a different model but inserted in the DynFlexHet model for expected profit calculations (cf., e.g., Lang etal. 2015). Note that expected store profits from the DynFlexHet model based on optimal prices of a different model can turn out higher than expected profits from the DynFlexHet model based on its true (i.e., own) optimal prices in particular weeks but not when summed up over all weeks. Therefore, expected losses in total (store) profits are negative by definition compared to the assumed “true” model. Figure7 shows box plots for expected losses by stores for the different heterogeneous model alternatives (homogeneous models are omitted in the figure as they provided patterns highly similar to those of their heterogeneous counterparts). One important finding here is that expected losses relative to the DynFlexHet model are generally not larger than 10% for the DynParHet model across all brands (although even a loss of “only” 10% can translate into a huge loss expressed in monetary units). The least complex heterogeneous model ( StatParHet ) either leads to the highest expected losses for six brands (with a maximum median expected loss of 26 % and largest individual store losses of up to −40 % across brands) or expected losses similarly high as for its flexible counterpart ( StatFlexHet ). The DynFlex - Hom model on the other hand provides the most inconsistent results here: it shows (much) higher expected losses than the DynParHet model for five brands but performs excellently for the other three brands, where it comes up with only very small or almost negligible expected losses. Overall, the DynParHet seems to be the most robust model variant to reduce expected losses relative to the proposed Dyn - FlexHet model. Therefore, at least for the data at hand, the use of a heterogeneous dynamic model for pricing decisions is strongly suggested. The main (preliminary) conclusion resulting from our optimization exercise is that nonlinear heterogeneous dynamic pricing provides better pricing decisions and higher expected chain profits, and can prevent losses due to suboptimal pricing. That’s why we recommend to prefer the proposed DynFlexHet model. Finally, it is important to note that optimizing profits or evaluating expected losses based on an econometric model like we do is prone to the Lucas critique, at first glance. Transferred to our context, Lucas (1976) states that since the structure of an econometric model involves optimal decision rules of brand managers and since these rules change systematically with the structure of the time series data relevant to the retailer’s policy, any change in the retailer’s policy will alter the structure of the econometric model (Lucas 1976,p.41). However, the focus of this paper does
60 P.Aschersleben, W.J.Steiner 1 3 not lie on the evaluation of the retailer’s pricing policy but on the comparison of the statistical and normative capabilities of different variants of econometric models and in particular on assessing the benefits of including functional flexibility in price response, store heterogeneity, and/or price dynamics in a sales response model. Fig. 7 Box plots of expected losses by store (aggregated over weeks) for the heterogeneous model variants using the DynFlexHet model as “true” model Fig. 8 Distributions of store-specific shares of optimized prices hitting the upper bound of the observed price range for the dynamic model variants DynParHet , DynFlexHom , and DynFlexHet
61 1 3 Dynamic pricing using flexible heterogeneous sales response… 4.4.5 Boundary pricing effects All optimization results discussed above would be of less value if optimized prices always or frequently hit the boundaries of the observed store-specific price ranges. In the following, we illustrate the boundary pricing issue for the following three models: the DynFlexHet , which provided the highest expected profits for most brands (and as well on average across brands) and the most accurate sales predictions based on the CRPS score; the DynFlexHom , which as well showed an excellent forecasting accuracy based on all three predictive validity statistics ( CRPS , RMSE , RMedSE ); and the DynParHet , which consistently (i.e., across brands) led to the least expected losses compared to the DynFlexHet . We included the latter model also because there seemed to be evidence from our literature search that dynamic parametric models were especially prone to upper bound corner solutions, compare Sect.2. Figure8 displays the distributions of store-specific shares of optimized prices hitting the upper bound of the observed price range, both by brand and across all brands. The plots indicate that boundary pricing is not a general problem by brand or model (in the sense that the boundary is always or mostly hit for certain brands or models) and that the occurrence of corner solutions is not critical for most brands. For our example brand “Minute Maid”, optimized prices almost never hit the upper bound in most stores and in the worst case in 12% of the weeks in a store. For two brands (“Citrus Hill”, “Tree Fresh”), however, the upper price bound is hit (much) more often in a larger number of stores. On the other hand, we observe highly different shares of boundary hits across the stores for these two brands (especially for the heterogeneous models), including stores with zero share (no upper boundary hits) as well. Across brands and the three dynamic models shown, 50% (75%) of the stores show upper boundary price hits in less than 10% (35%) of the weeks (not displayed in the figure). As it is further obvious, the upper price bound patterns differ more by brands than by models, which consistently applies for the other models not shown, too. Our findings contradict the assumption that parametric models or models with linear price effects might be generally prone to boundary pricing effects. However, we consistently find more stores with higher shares of upper boundary price hits for the dynamic models compared to their static counterparts. As a rule, we also see a (much) greater bandwidth of stores with different shares of upper boundary price hits for heterogeneous models compared to their homogeneous counterparts across brands, as one might have expected. No such systematic can be derived for flexible versus parametric models. In contrast to upper boundary price hits, lower boundary price hits as well as upper boundary sales hits were hardly observed. The share of lower bound price hits averaged across stores is not larger than 1% by both brand and type of model. The number of upper boundary sales hits never exceeds 5 times in a store across brands and models, corresponding to a maximum share of 6%. The detailed results on boundary effects by brand for all models can be obtained from the authors upon request.
62 P.Aschersleben, W.J.Steiner 1 3 5 Conclusions andlimitations In this paper, we proposed a Hierarchical Bayesian semiparametric store sales model with price dynamics for optimal pricing of fast-moving consumer goods offered in physical stores of a retailer and maximizing related expected brand profits. Using Bayesian P-splines as a nonparametric technique to estimate price effects allowed us to dispense with the specification of a specific functional form for ownand crossprice response a priori and to identify possible “irregular” pricing effects (e.g., kinks and steps in price response) that are difficult to capture with parametric functions. Heterogeneity of price effects across stores was accommodated via scaling factors for the P-splines, which serve as random effect parameters to scale the price functions upor downwards for individual stores while preserving their overall shape. Accommodating heterogeneity has become state-of-the-art if not a must in econometric marketing models, far beyond the present context of store-level sales response models. Price dynamics were accounted for via one-week lagged own prices, which allowed us to implicitly address stockpiling or customer-holdover effects, although such dynamic effects are more difficult to explore based on aggregate sales data. For this reason, we further provided robustness checks for the dynamic variants of the flexible nonlinear models by including more time lags for the own-item price, indicating that higher order autoregressive structures were (mostly) not supported by the data. Optimal price paths for brands were determined by a discrete dynamic programming algorithm. We further imposed additional constraints for the price optimization step on price ranges and upper bounds for brand sales to preserve a realistic scenario for profit implications. To the best of our knowledge, addressing the three dimensions functional flexibility, store heterogeneity, and price dynamics simultaneously in one approach (both for sales response modeling and subsequent price optimization) has not been proposed previously. Therefore, this approach allows us for the first time to disentangle the effects of these three dimensions on sales, pricing and profits. In addition, we introduced the Continuous Ranked Probability Score as a new measure to assess the predictive model performance, a measure that has not been used before in the sales response modeling field. We applied our proposed approach in an empirical study to data for refrigerated orange juice brands offered in physical stores of a large retail chain and compared it to several benchmark models (which ignore heterogeneity, functional flexibility, and/or price dynamics) in terms of forecasting accuracy, optimal pricing or price paths, optimized profits, and expected losses. Based on three different predictive performance measures, we found that accommodating both functional flexibility in price response and price dynamics provided the best sales predictions. Moreover, sales predictions from all models were almost fairly well dimensioned with regard to observed brand sales, i.e., the constraints imposed on the price optimization step worked well and protected our models from overor underestimation biases in sales predictions and as a result in expected profit calculations, too. In addition, optimized expected profits were highest under the proposed flexible, heterogeneous, and dynamic sales response model for five brands and by either again a flexible or a
63 1 3 Dynamic pricing using flexible heterogeneous sales response… dynamic model for the other three brands. Importantly, optimized expected profits were highest for six brands based on models including price dynamics, suggesting that it is very important for managers to accommodate price dynamics for pricing strategies as well, at least for our data at hand. Not least, the best models by brand were all heterogeneous ones, and the proposed flexible, heterogeneous, and dynamic sales response model provided the highest expected profits on average across brands. As our first research question was whether a sales response model that combines nonlinear pricing, store heterogeneity in price (and other marketing) effects, and price dynamics can provide higher expected chain profits, the answer to this question is yes. The benefits from accommodating price dynamics for optimal pricing decisions were also clearly visible in our analyses on expected losses. Furthermore, and this is a very important point, we could show that upper boundary pricing effects were usually moderate across brands and never critical. In addition, the storespecific shares of price hits at the upper bound turned out highly heterogeneous for some brands. And lower boundary price hits as well as upper boundary sales hits occurred only in very few cases. As such, our second research question can also be positively evaluated in the sense that boundary pricing effects are not a critical issue for the proposed model, at least not for the data at hand. This finding together with the other ones summarized above let us recommend the use of the new model for price optimization. Some limitations of our study should also be addressed. First, we took the perspective of a brand manager, who should be interested in analyzing the consequences of suboptimal pricing strategies from using a “wrong” sales response model for her/ his brand, assuming fixed pricing patterns for substitute brands and the possibility to adjust the price of the own brand in the retailer’s assortment. This brand perspective is also an option for product category managers of retailers to analyze how price changes for one brand in their assortment would affect category sales and profits. Nevertheless, a next step could then be an extension of our model for optimal category pricing, using a seemingly unrelated regression (SUR) approach for model estimation as proposed, e.g., by Weber etal. (2017). Second, other options for accommodating price dynamics could be considered, for example by allowing for time-varying price effects instead of using simple lagged prices (leaning on reparametrization approaches such as proposed by Foekens etal. 1999 or Kopalle etal. 1999). Also, lead price effects could be added to account for anticipatory responses of consumers as a result of price promotion announcements (following, e.g., VanHeerde etal. 2000, 2004). Third, our approach could be further extended to allow for synergy effects between marketing instruments. For example, Ataman etal. (2010) proposed a dynamic model that relates brand sales to the long-term marketing strategy represented by advertising, product, and place in addition to price. Finally, we did not treat the issue of price endogeneity, which, if ignored, could lead to biased effect estimates. A good instrument to account for endogeneity in sales response models is the lagged price. However lagged prices are not exogenous in our dynamic model(s) and therefore couldnot be taken as instrument there. Wholesale prices, in our case
64 P.Aschersleben, W.J.Steiner 1 3 used as costs in our profit functions, represent another candidate for an instrumental variable. As objected by Rossi (2014), however, retail prices often show a much higher variation than wholesale prices. This also holds for many stores in our data, as was visible from Fig.2 (top row) for the brand “Minute Maid” as example. Furthermore, two-step estimation as best known instrumental variable technique is restricted to linear models and cannot be applied to our nonlinear models (see, e.g., Hruschka 2017). Moreover, Ebbes etal. (2011) showed that models which account for endogeneity via instrumental variables do not necessarily provide better predictions. In the absence of adequate instruments one could use instrument-free techniques to accommodate endogeneity. Hruschka (2017) has recently provided an overview of the various options here, emphasizing, however, that currently only the copula-based method by Park and Gupta (2012) could be extended to nonlinear models like the ones proposed in this article (cf. Hruschka 2017,p.28). To the best of our knowledge, such an extension of the copula-based method has not yet been developed for the flexible scaling models we consider. We therefore leave the issue of endogeneity as well as the other limitations mentioned above, for future research. Appendix Additional figures See Figs.9, 10 and 11. Fig. 9 Comparison of estimated own-price effects between static and dynamic model variants for the brand “Minute Maid”
65 1 3 Dynamic pricing using flexible heterogeneous sales response… Fig. 10 Estimated lagged price effects resulting from the DynFlexHet model for all eight brands in the refrigerated orange juice category
66 P.Aschersleben, W.J.Steiner 1 3 Fig. 11 For each brand, the numbers in the upper triangular part of the panels refer to the relative frequencies with which each two different models share the same optimal price levels in a particular store and week. Additionally, the diagonal elements indicate the relative frequencies where the optimized price and the observed price per store and week coincide for a particular model. Highlighted by colored borders are the pairwise comparisons between models that are more similar to each other, i.e., that differ in only one of the three model dimensions (functional flexibility, store heterogeneity, price dynamics). For example, the comparison between the DynFlexHom model and the StatFlexHom model is framed in red as both models are flexible and homogeneous and just differ in the dynamic component
67 1 3 Dynamic pricing using flexible heterogeneous sales response… Additional tables See Tables6, 7, 8, 9 and 10. Table 6 Descriptive statistics for weekly prices, market shares, and unit sales per store *The share of price changes represents the share of weeks with a different store-specific price compared to the previous week ( pt≠pt−1 ) Brand Retail price ($)/Price changes (%)* M. share (%) Unit sales Range Mean SD pt≠ p t−1 Mean SD Mean SD Premium brands Florida Natural [1.54, 3.35] 2.85 0.33 38.6 4.7 6.5 27.2 46.2 Tropicana Pure [1.29, 3.87] 2.96 0.57 46.7 12.3 13.6 74.8 98.2 National brands Citrus Hill [0.99, 3.07] 2.31 0.35 42.8 8.0 12.7 53.5 157.4 Florida Gold [0.99, 3.08] 2.18 0.39 43.3 5.2 7.9 33.4 64.0 Minute Maid [0.99, 3.17] 2.22 0.42 55.2 10.2 13.8 52.0 76.8 Tree Fresh [0.99, 2.69] 2.15 0.31 42.3 7.7 8.6 48.8 92.5 Tropicana [1.41, 2.99] 2.20 0.38 56.1 18.4 21.1 112.8 159.0 Private brand Dominick’s [0.99, 2.69] 1.76 0.42 47.3 34.0 25.5 308.4 538.8 Table 7 Out-of-sample predictive performance of the DynFlexHom and DynFlexHet models for different numbers of knots and different degrees of the B-spline basis functions, evaluated by the Average Root Mean Squared Error ( ARMSE ) in holdout samples (relative improvements/deteriorations inARMSEvalues refer to the default specification of the P-splines with 20 knots and degree 3) Brand DynFlexHom DynFlexHet 20/1 20/3 40/1 40/3 20/1 20/3 40/1 40/3 Flor. Natrl. 24.9 25.0 24.9 25.0 27.0 27.2 27.0 26.9 ( − 0.2%) – ( − 0.3%) ( − 0.1%) ( − 0.7%) – ( − 0.4%) ( − 0.8%) Tropic. Pure 52.8 52.9 52.9 53.0 51.5 51.6 51.7 51.7 ( − 0.1%) – ( + 0.1%) ( + 0.2%) ( − 0.3%) – ( + 0.1%) ( + 0.2%) Citrus Hill 84.6 87.3 84.0 86.3 107.2 112.0 107.2 110.4 ( − 3.1%) – ( − 3.8%) ( − 1.1%) ( − 4.2%) – ( − 4.3%) ( − 1.4%) Flor. Gold 54.0 54.1 54.0 54.1 54.5 54.6 54.5 54.5 ( − 0.2%) – ( − 0.1%) (±0.0%) ( − 0.1%) – ( − 0.1%) ( − 0.1%) Min. Maid 53.8 54.1 53.7 54.1 53.7 53.8 53.5 54.2 ( − 0.5%) – ( − 0.7%) ( − 0.1%) ( − 0.2%) – ( − 0.5%) ( + 0.8%) Tree Fresh 65.8 65.7 65.8 65.8 76.5 78.4 76.9 78.5 ( + 0.1%) – (±0.0%) ( + 0.1%) ( − 2.4%) – ( − 1.9%) ( + 0.1%) Tropicana 91.5 91.8 91.7 91.8 92.7 93.0 93.0 93.1 ( − 0.3%) – ( − 0.1%) (±0.0%) ( − 0.3%) – ( − 0.1%) (±0.0%) Dominick’s 282.9 283.6 283.5 284.5 285.1 286.2 286.3 286.9 ( − 0.3%) – (±0.0%) ( + 0.3%) ( − 0.4%) – (±0.0%) ( + 0.3%)
68 P.Aschersleben, W.J.Steiner 1 3 Table 8 Out-of-sample predictive performance of the DynFlexHom and DynFlexHet models for different numbers of knots and different degrees of the B-spline basis functions, evaluated by the Average Mean Continuous Ranked Probability Score ( AMCRPS ) in holdout samples(relative improvements/deteriorations in AMCRPSvalues refer to the default specification of the P-splines with 20 knots and degree 3) Brand DynFlexHom DynFlexHet 20/1 20/3 40/1 40/3 20/1 20/3 40/1 40/3 Flor. Natrl. 8.9 9.0 8.9 8.9 7.9 8.0 7.9 8.0 ( − 0.7%) – ( − 0.7%) ( − 0.4%) ( − 0.9%) – ( − 0.8%) ( − 0.5%) Tropic. Pure 22.3 22.4 22.3 22.3 20.1 20.1 20.1 20.1 ( − 0.1%) – ( − 0.2%) ( − 0.2%) ( − 0.2%) – ( − 0.2%) ( − 0.1%) Citrus Hill 21.5 21.8 21.5 21.7 18.8 19.1 18.9 19.0 ( − 1.4%) – ( − 1.4%) ( − 0.5%) ( − 1.7%) – ( − 1.2%) ( − 0.5%) Flor. Gold 18.9 19.0 18.9 18.9 18.3 18.4 18.3 18.3 ( − 0.2%) – ( − 0.1%) ( − 0.1%) ( − 0.3%) – ( − 0.3%) ( − 0.2%) Min. Maid 20.3 20.4 20.3 20.4 18.9 19.0 18.8 18.8 ( − 0.6%) – ( − 0.7%) ( − 0.4%) ( − 0.4%) – ( − 1.0%) ( − 0.7%) Tree Fresh 16.1 16.0 16.0 16.0 13.6 13.7 13.7 13.8 ( + 0.1%) – ( − 0.2%) (±0.0%) ( − 0.3%) – ( + 0.2%) ( + 0.5%) Tropicana 45.3 45.5 45.3 45.4 43.3 43.4 43.4 43.5 ( − 0.5%) – ( − 0.4%) ( − 0.2%) ( − 0.4%) – (±0.0%) ( + 0.1%) Dominick’s 137.7 138.4 138.0 138.5 131.6 132.2 132.0 132.3 ( − 0.6%) – ( − 0.3%) (±0.0%) ( − 0.5%) – ( − 0.1%) ( + 0.1%)