Discrete choice in marketing through the lens of rational inattention
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Turlo, Sergey; Fina, Matteo; Kasinger, Johannes; Laghaie, Arash; Otter, Thomas Article — Published Version Discrete choice in marketing through the lens of rational inattention Quantitative Marketing and Economics Provided in Cooperation with: Springer Nature Suggested Citation: Turlo, Sergey; Fina, Matteo; Kasinger, Johannes; Laghaie, Arash; Otter, Thomas (2025) : Discrete choice in marketing through the lens of rational inattention, Quantitative Marketing and Economics, ISSN 1573-711X, Springer US, New York, NY, Vol. 23, Iss. 1, pp. 45-104, https://doi.org/10.1007/s11129-025-09292-9 This Version is available at: https://hdl.handle.net/10419/323528 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Quantitative Marketing and Economics (2025) 23:45–104 https://doi.org/10.1007/s11129-025-09292-9 Discrete choice in marketing through the lens of rational inattention Sergey Turlo1·Matteo Fina1·Johannes Kasinger2,3 ·Arash Laghaie4· Thomas Otter1,4 Received: 10 August 2023 / Accepted: 24 December 2024 / Published online: 6 February 2025 BSergey Turlo [email protected] Matteo Fina [email protected] Johannes Kasinger [email protected] Arash Laghaie [email protected] Thomas Otter [email protected] 1Johann Wolfgang Goethe University Frankfurt, Frankfurt, Germany 2Tilburg School of Economics and Management, Tilburg University, Tiburg, Netherlands 3Leibniz Institute for Financial Research (SAFE), Frankfurt, Germany 4Nova School of Business and Economics, Carcavelos, Portugal 123 © The Author(s) 2025 Abstract Models derived from random utility theory represent the workhorse methods to learn about consumer preferences from discrete choice data. However, a large body of literature documents various behavioral patterns that cannot be captured by basic random utility models and require different non-unified adjustments to accommodate these patterns. In this article, we discuss strategies how to apply rational inattention theory—which explains a large variety of such departures—to the analysis of discrete choice among multiple alternatives described along multiple attributes. We first review existing applications that make restrictive belief assumptions to obtain choice probabilities in closed multinomial logit form. We then propose a model that allows for general consumer beliefs and demonstrate its empirical identification. Further, we illustrate how this model naturally motivates stylized empirical results that are hard to reconcile from a random utility perspective. Keywords Choice modeling ·Rational inattention ·Conjoint analysis · Discrete choice experiments JEL Classification C00 ·C35 ·D83 ·M31
46 S. Turlo et al. 1 Introduction Discrete choice models based on rational inattention (RI) theory are becoming popular in economics and marketing for studying consumer choices and preferences when information processing is limited but adaptive. This article demonstrates how to apply these models in various multi-attribute, multi-alternative (MAMA) contexts. Unlike the traditional random utility model (RUM), RI-based choice models provide a unified framework to explain phenomena like consideration sets, stake sensitivity, and brand-specific price responses, which previously required ad-hoc adjustments to the RUM. This article aims to support applied researchers by: first, raising awareness of the potential of discrete choice models under RI; second, outlining the steps for implementing RI for empirical research; and third, clarifying the strengths and limitations of different implementations. Based on foundational ideas from psychology (e.g., Simon & Newell, 1971), rational inattention theory, introduced by Sims (2003) in macroeconomics, suggests that decision-makers (DMs) face cognitive limitations and, therefore, do not take in all available information when making choices. DMs recognize their cognitive limitations and strategically decide how much and what type of costly information to process in each decision scenario. RI suggests that DMs adjust their processing efforts based on prominent, accessible aspects of a choice task, which shape their prior beliefs about unknown factors affecting utility. The adaptive and partial information processing implied by RI motivates a rich set of behaviors, even with standard additively separable utility. Under RI, probabilistic choice follows from costly and thus imperfect processing of information, i.e., from DMs’ residual uncertainty about what the utility maximizing choice alternative is. In contrast, the RUM by McFadden (1974) derives choice probabilities by assuming that the DM acts on a larger information set than observed by the analyst. Unlike RUMs, RI choice models can explain choices from MAMA sets without assuming that only the decision-maker observes certain utility factors. This aligns with earlier research that explained randomness in choice through cognitive processes (e.g., Thurstone, 1927; Quandt, 1956; Louviere et al., 1999). RI adds to this literature by offering a micro-foundation for probabilistic choice based on economic optimization. The general discrete choice problem under RI lacks a closed form solution. To utilize standard estimation methods, current empirical RI discrete choice models (RIDCMs) make particular—and potentially unrealistic—assumptions about consumers’ prior beliefs, resulting in choice probabilities that follow a closed-form multinomial logit (MNL) function of the underlying utility index. While these models can motivate some deviations from the full information RUM, they face similar conceptual issues as logit models, such as unrealistic substitution patters and exogenous consideration sets. To address these limitations, we demonstrate how to estimate a RI-DCM under general prior belief assumptions, including rational expectations, and consumer heterogeneity. In this model, alternatives’ payoffs are represented by linear utility indices based on preferences and attributes, which may be simple or complex. DMs can process simple attributes at no cost, whereas processing or integrating utility from complex 123
Discrete choice in marketing through... 47 attributes requires costly effort. DMs have prior beliefs about the complex attributes, assuming rational expectations. In choice experiments, these expectations are shaped by the experimental design. We show that both preference parameters and the distinction between simple and complex attributes can be likelihood identified in this model. The primary advantages of this model over existing empirical RI-DCMs are that i) it can explain a broader range of phenomena that deviate from the RUM-DCM framework and ii) it enables more flexible counterfactual analysis. The drawback of this approach is that it requires a numerical solution to the formal RI problem. Table 1 provides a comparison of existing methods for applying RI to MAMA data. Using the RI-DCM with general beliefs, we demonstrate how different phenomena contradicting the microeconomic foundation of RUMs in MAMA settings endogenously follow from the optimal deployment of limited cognitive resources. Examples include brand-specific price coefficients (e.g., Carmone & Green, 1981; Sawtooth Software, 1996; Kalra & Goodstein, 1998), separate coefficients for different aspects of price such as, e.g., a coefficient for regular price and one for a price discount or a tax (Guadagni & Little, 1983; Blattberg & Neslin, 1989; Chetty et al., 2009), as well as consideration sets and attribute non-attendance. Table 1 Comparison of different empirical RI-DCMs Strategy Paper Prior beliefs Implied choice probabilities Comments RI-DCM with choice in closed MNL form Brown and Jeon (2024) Beliefs over index of unknown attributes follows Cardell distribution Equivalent to multinomial logit (additive separability over alternative characteristics, fully compensatory) Assumed belief distribution has full support on the real line which may be unreasonable, e.g., with prices; does not reproduce certain RI features like consideration sets Joo (2023) Beliefs over all utility components are a function of non-utility components, implicitly defined such that resulting choice probabilities are equivalent to multinomial logit in closed form Equivalent to multinomial logit (additive separability over alternative characteristics, fully compensatory) Approach motivates the inclusion of nonutility attributes in the logit index; implied prior distribution is not accessible by the analyst; counterfactuals with respect to beliefs are available only in a restricted fashion; does not reproduce certain RI features like consideration sets RI-DCM with general beliefs This paper Any prior belief distribution over a discrete state space, e.g., rational expectations over choice tasks in a DCE non-compensatory, not additively separable Resulting model reproduces all qualitative features of discrete choice under RI, e.g., consideration sets or attribute interactions 123
48 S. Turlo et al. In the RI-DCM with general beliefs, predictions of, e.g., consideration sets and attribute non-attendance become implicit functions of the composition of a choice set, reflecting adaptations to prominent features of the set. This flexibility is absent in current empirical RI-DCMs which invoke very specific assumptions about consumer beliefs (e.g., Joo, 2023; Brown & Jeon, 2024). Similarly, this RI-DCM predicts effects of attribute range, number of attribute levels, and the size of choice sets on choice behavior. RUMs typically do not account for these common properties of choice data, at least not comprehensively in one unified model. Different from extant models of consumer search and learning,1RI does not restrict the structure of informative signals that the DM uses. This feature makes RI distinct and more generally applicable than models where all uncertainty is resolved upon search. For example, in the context of discrete choice experiments (DCEs) all relevant pieces of information, i.e., attribute information for all alternatives, are readily presented to DMs. Thus, the distinction between more and less processing of the available information is qualitatively different from the distinction between knowing or not knowing certain product attribute values as typical of search models. This idea resembles the distinction between evaluation costs and search costs in Guo (2021) and Gu and Wang (2022). Consequently, RI-DCMs are particularly important and useful departures from extant models when evaluating and integrating information for an updated overall understanding of a choice situation is decisive and effortful. In contrast, search models are arguably more adequate when knowing or not knowing a particular attribute makes the difference. Finally, search models that allow for partial learning about the value of available alternatives are more closely related to RI (e.g., Ursu et al., 2020). However, sequential search/learning models applied to MAMA choice likely will be computationally intractable without observing the search/learning sequence. Alas, this sequence that involves mental operations beyond reading attribute information may well be fundamentally unobservable. In addition, the standard assumption of normally distributed prior beliefs and signals in these models does not correspond with the empirical distribution of attributes and their bounded support, especially in DCEs. In this article, we provide an accessible presentation of the RI theory for studying discrete choice in MAMA settings common in economics and marketing. Our contributions are threefold. First, we outline two general strategies for applying RI to discrete choice MAMA data and discuss their respective advantages and disadvantages. The first strategy, typical for existing RI-DCMs, invokes specific assumptions about consumer beliefs (continuous with full support following a Cardell distribution) that yield choice probabilities in closed MNL form. The full support assumption contradicts observable distributions of attributes. The second strategy can accommodate general consumer beliefs and requires a numerical solver. A researcher choosing between these strategies faces a trade-off between conceptual realism and computational burden. Our second contribution is that this article is, to our knowledge, the first to detail the necessary steps for implementing the second empirical strategy (“RI-DCM with general prior beliefs") while accounting for consumer heterogeneity. Third, we demonstrate that a RI-DCM with general beliefs can effectively reproduce various behavioral 1For a recent overview, see Honka et al. (2019). 123
Discrete choice in marketing through... 49 patterns inconsistent with standard RUM models in MAMA settings parsimoniously. In contrast, existing RI-DCMs that yield choice probabilities in closed MNL form imply the same constraints as MNL derived from RU theory, such as independence of irrelevant alternatives and strictly positive choice probabilities for all alternatives in a set.2 The remainder of this paper is organized as follows. Section 2derives discrete choice among multiple alternatives described along multiple attributes under the RI framework. Section 3illustrates estimation and empirical identification of the model. Section 4discusses key features of the RI model and provides illustrative simulations. Section 5concludes with a discussion and an agenda for future research. 2 A rational inattention model of discrete choice The basic idea behind RI theory is that DMs face an abundant amount of information and cannot process all of it. However, they are aware of this limitation and decide how to process the available information optimally, trading off costs and benefits of being better informed. This idea was suggested by Sims (2003) to provide a unifying framework for different frictions in macroeconomics. While the original model was developed for continuous action spaces, Matˇejka and McKay (2015) extend this theory to discrete choices. Our presentation builds on the discrete choice version of the RI model (Matˇejka & McKay, 2015; Caplin et al., 2019) and tailors it to the typical MAMA setting. Introducing the model, we first present the RI choice problem and discuss how its various components translate into the MAMA setting. Then, we turn to the problem’s solution and cover how the various primitives affect the resulting choice behavior. In particular, this will illustrate the impact of the complexity of a choice task and of the incentives to process information. To ease the exposition of the various components of the RI framework, we will refer as an example to a DCE where a DM has to choose between a car and an outside option. In this example, the final price paid by the DM consists of two components: i) a list price that is easily evaluated by the DM, and ii) a discount that applies only to specific cars (thus encouraging the purchase of such vehicles). While both components have the same impact on final utility, we assume it is more effortful to find out if and what discount applies to a particular car. This simple example may align well with existing traditional search models if determining the eligibility of a specific car is a simple search task, e.g., checking whether the discount in monetary terms applies to a specific car model. However, we posit a scenario wherein the eligibility of a specific car hinges upon (a combination of) diverse characteristics. In such a situation, judging the applicability of the discount requires the consumer to collect information from various sources and integrate it to 2Matˇejka and McKay (2015) already pointed out that a random utility model cannot generally capture behavior implied by RI agents. 123
50 S. Turlo et al. determine the overall value of the discount. Consequently, it is possible that the DM processes only some parts of the information and, therefore, may arrive at a faulty evaluation of the final price, which in turn leads to choice errors.3 2.1 Formal problem and its translation into MAMA settings We closely follow Caplin et al. (2019) in defining the problem faced by the rationally inattentive DM. There is a finite number of states the DM can learn about.4An action ais a mapping from states to utilities. Adenotes the set of all possible actions. The mapping u:A×→Rdescribes the utility from any action in each state. The problem faced by the DM is non-trivial because, typically, different actions are optimal in different states, and the DM is uncertain about the true state. However, as we will explain later, the DM can costly learn about the true state. The general nature of RI theory provides room for different translations of the framework into the typical MAMA setting in marketing. We naturally impose that actions correspond to different alternatives from which the DM chooses and define states as representing different choice sets characterized by the specific attribute compositions of alternatives available to the DM. Accordingly, in this setting, corresponds to the set of all attainable choice sets in a given choice environment. Payoffs Similar to the distinction between directly observable attributes and attributes that need to be searched in search models (e.g., Honka et al., 2019; Gardete & Hunter, 2020) or the distinction between attributes that guide consideration and attributes that are only processed upon consideration in two-stage models of choice (e.g., Aribarg et al., 2018), we assume that the subjective value of alternatives is derived from attributes that fall into two categories. The first category consists of simple attributes xswhose joint valuation is immediate to the DM. The second category comprises complex attributes xcwhose joint valuation and integration with simple attributes requires cognitive effort and time.5 We assume additive separability such that the subjective utility of an alternative is given by u(a,ω)=x a,s(ω)βs+x a,c(ω)βc,(1) 3There are many more things about a car that are likely payoff relevant to a DM, and it may or may not be effortful to evaluate and integrate them into an overall evaluation. However, this minimal example will help develop basic principles. 4An alternative formulation with a continuous state space is given in Matˇejka and McKay (2015). 5Studies that explore the choice process using eye traces have documented an “orientation phase” where the DM acquires partial information about the products, which guide her subsequent information acquisition (e.g., Russo & Leclerc, 1994; Musalem et al., 2021). This orientation phase and subsequent behavior involves both bottom-up and top-down processing (see Corbetta & Shulman, 2002 for a review). 123
Discrete choice in marketing through... 51 where βsand βcare the respective part-worths of simple and complex attributes.6 The dependence of xa,sand xa,con ωabove highlights that attributes of alternatives change from choice set to choice set.7 Before learning, the DM has some beliefs about the value of complex attributes xa,c, which become more precise as the DM processes information. Note that any uncertainty is due to complex attributes. We refer to the portion of utility derived from simple attributes as the “simple utility component” while the portion derived from complex attributes is termed the “complex utility component”. In our example, the list price of a car is a simple attribute, and the discount is a complex attribute that requires time and cognitive effort to process and integrate with the simple list price to arrive at a final price and an assessment of utility. To further illustrate the challenges associated with integrating information, consider the following examples involving two price components: (i) prices are given in currency units and discount values in percentages given as numbers, e.g., 9.90 Euro and a 16% discount, (ii) discount values in currency units, e.g., 9.90 Euro and a 1.58 Euro discount, (iii) discount values mentioned in some way together with the final price of, in this example, 8.32 Euro. Cases (i) and (ii) require the DM to integrate the discount information with the price. However, this exercise arguably is more involved in case (i) than in case (ii) because the DM has to calculate the discount value as part of the integration exercise. The main point, however, is that case (ii)—where both the price and the discount are presented in currency units—is harder than case (iii), where the final price is displayed. This added difficulty stems from the need to integrate two pieces of information by subtracting the discount from the price to determine the overall value. Prior beliefs The DM’s problem is given by a pair (μ, A). Here, μ∈() is her prior belief over the states of the world, with () being the set of distributions over , and A⊂Ais the set of actions she can choose from. In our illustrative example, a state ωcorresponds to a specific choice set characterized by a particular combination of attribute realizations. Since simple attributes are processed at no cost by the DM, each combination of simple attribute realizations xsinduces a different prior belief distribution μs∈() over possible choice sets ω∈. In general, these prior beliefs, conditional on costless information, will differ from the unconditional distribution over choice sets. In particular, the DM obtains prior beliefs μsby conditioning the distribution over all choice sets on the simple attribute realizations faced in a specific choice set ω. Formally, prior beliefs are given by μs(ω) =Pr ω|{(xa,s(ω))}a∈A where the distribution over choice sets, Pr(ω), is determined by the choice environment. 6Additive separability is by no means a necessary but often a natural assumption when, e.g., different price components add up to a total price. 7With the present notation, any state or choice set is defined by the configuration of the alternatives: ω={(xa,s(ω), xa,c(ω))}a∈A. 123
52 S. Turlo et al. Table 2 Set of choice sets with respective payoffs and prior beliefs Choice set ω1ω2ω3ω4 Probability of ωiγp(1−γd)γ pγd(1−γp)(1−γd)(1−γp)γd Payoffs: Inside alternative uIβb−pLβb−pL+Dβb−pHβb−pH+D Outside alternative uO00 0 0 Information set p=pL,p=pL,p=pH,p=pH, d∈{0,D}d∈{0,D}d∈{0,D}d∈{0,D} Prior beliefs μs: Pr(uI=βb−pL)1−γd1−γd00 Pr(uI=βb−pL+D)γ dγd00 Pr(uI=βb−pH)00 1−γd1−γd Pr(uI=βb−pH+D)00 γdγd Each column represents a different choice set ωi. In addition to the payoffs, the objective probability of each choice set, the information set of the DM before any learning takes place, as well as the resulting prior beliefs are displayed. Note that this fully characterizes the DM’s prior since for all choice sets Pr(uO=0|ωi)=1 Intuitively, one can think of a sequentially updating DM who is aware of the choice environment (e.g., an experimental design) and the implied distribution of attributes over all possible choice sets. Once she observes the realized simple attribute values (xs) in a specific choice set, she forms conditional prior beliefs μs, which in turn determine how she processes complex attributes xc. Returning to our exemplary DCE, suppose the DM chooses only between a single car brand and an outside option of not buying. The utility of the inside alternative, i.e., the car, is given by uI=βb−p+dwith βbbeing the brand coefficient, pbeing a simple price, and dbeing a complex discount. For the sake of a minimal example, we assume that brand is a simple attribute as well and that whatever (complex) criteria qualify the car for the discount only contribute to utility through the discount, such that cars with and without the discount have the same brand coefficient βb.8We further assume that the DM has to commit to purchasing the car at a price p, i.e., pay pand only after is reimbursed d, depending on the eligibility of the car. The utility of the outside option is normalized to uO=0. Further, the experimental design is such that p∈{pL,pH} with Pr(p=pL)=γpand d∈{0,D},D>0 with Pr(d=D)=γd>0.9In Table 2, the columns represent the possible choice sets of this design. The DM can face four different choice sets as there are two attributes with two levels each, so ||=4. The objective probabilities of each choice set ω(prior to the realization of the simple attributes) are displayed in Table 2. Since there are only two possible realizations of simple attributes, characterized by the two levels of the simple price pLand pH, there 8In the notation of Eq. 1, we have the following decomposition into simple and complex contributions to utility: [1,p]s[βb,−1]+dc·1 with a shared price coefficient equal to 1. 9In our example, the conditional distribution of the complex attribute equals the marginal distribution due to the orthogonal design. In designs with built-in correlations, e.g., with conditional pricing, the realized values of simple attributes will predict the distribution of complex attributes values. 123
Discrete choice in marketing through... 59 conditional on a specific choice set for an alternative aread P(a|ω) =exp{(x a,s(Cβs+βs)+x a,cβc)/λ} bexp{(x b,s(Cβs+βs)+x b,cβc)/λ}. Defining βs≡Cβs+βs, one can see how this derivation can motivate different coefficients for, e.g., a simple and a complex component of price. However, the resulting model otherwise is indistinguishable from a standard logit RUM. Prior beliefs consistent with RI choice probabilities in closed MNL form In the application by Joo (2023), each alternative a∈Ais characterized by a set of observable consideration shifters da, e.g., advertising or shelf placing, that solely affect prior beliefs but not consumption utility.15 All product attributes are assumed to be complex so that u(a,ω)=x c,aβc. Given prior beliefs μ, that depend solely on the informational shifters {da}a∈A, and information costs λthe DM learns and chooses following the RI framework. Joo (2023) shows that for any combination of alternative specific payoffs {u(a,ω)}a, information costs λ, and strictly positive unconditional choice probabilities P(a), there are prior beliefs μthat are consistent with choice behavior of a rationally inattentive DM. However, the actual parameterization Joo (2023) brings to the data, i.e., a logit form with an additively separable index of alternative specific attributes xc,aand information shifters da, requires very specific beliefs about the index from alternative specific attributes xc,a. These beliefs are not derived from the objective distribution of this index in the marketplace and are substantially different from this distribution, as already implied by the full support assumption.16 This limits the formulation proposed in Joo (2023) as a model of rationally inattentive DMs that acquire knowledge about the distribution of alternative specific attributes xc,ain the marketplace over longer time horizons. 3.2 RI-DCM with general prior beliefs The RI-DCM with general prior beliefs does not have a closed-form solution, and we need to solve for choice probabilities numerically to compute the likelihood in this model. However, over and above incorporating more realistic assumptions about beliefs, this generalization yields the qualitatively distinctive features of RI discrete choice we previewed in Section 1and will elaborate on in Section 4. In applications, prior beliefs can be determined, for instance, by assuming rational expectations where εaare identically and independently T1-EV distributed. The Cardell distributed prior beliefs over complex utility components xa,c, together with εa, follow the T1-EV distribution. This feature is key for obtaining unconditional choice probabilities in closed form. For a detailed derivation, see Supplemental Appendix A-1 of Brown and Jeon (2024) and Appendices A.1 and A.2 in Bertoli et al. (2020). 15 Natan (2021) employs a similar identification strategy. In this paper, however, information processing costs λexplicitly depend on the size of the choice set. 16 While Cardell beliefs result in additive separability as demonstrated by Brown and Jeon (2024), the model in Joo (2023) is not immediately consistent with Cardell beliefs because of the assumption that the outside good payoff is known with certainty. 123
60 S. Turlo et al. Fig. 1 Flowchart of the Blahut-Arimoto algorithm to compute the likelihood in RI-DCM about the experimental design (in case of DCE data) or the empirical distribution of attributes (with observational data). However, the model could readily accommodate directly elicited, potentially heterogeneous beliefs about the (conditional) distributions of complex attributes and implied choice sets. Likelihood computation Conditional on βsand βcthe utility of the DM, u(a,ω), isgivenbyEq.1.Givenu(a,ω),μs(ω), and λ, the likelihood can be obtained by computing the optimal conditional choice probabilities, P(a|ω),inEq.3.AsP(a|ω) is a function of the endogenous unconditional choice probabilities Ps(a)in Eq. 3, we solve for both P(a|ω) and Ps(a)using the Blahut-Arimoto algorithm (Cover and Thomas, 2006). The algorithm starts by initializing the Ps(a)and iterates between updating Ps(a|ω) and Ps(a)until convergence. The first step of each iteration tuses the optimality condition in Eq. 3to compute Pt+1 s(a|ω) given Pt s(a)obtained in the past iteration. The second step computes Pt+1 s(a)by integrating Pt+1 s(a|ω) over the (conditional) prior belief distribution: Pt+1 s(a)= ω∈ μs(ω)Pt+1 s(a|ω) =Pt s(a) ω∈ μs(ω) exp{u(a,ω)/λ} b∈APt s(b)exp{u(b,ω)/λ}.(5) We calculate the distance between unconditional distributions obtained in subsequent steps, Pt+1 sand Pt s, via the Bhattacharyya (1946) distance D(Pt+1 s,Pt s).The algorithm stops when D(Pt+1 s,Pt s)<ξor when the maximum number of iterations itermax is reached. The converged conditional choice probabilities give us the individual level likelihood L(μs(ω), βs,βc,λ)=P∗ s(a|ω).17 Note that parameters ξand itermax govern the precision of the numerical solution, and care must be taken in setting their value. Figure 1shows the flowchart of the algorithm. Inference based on solutions from the Blahut-Arimoto algorithm, and specifically Markov Chain Monte Carlo (MCMC) estimation, is computationally costly as we need solutions for every unique combination of a choice set indexed by simple attribute realizations with model parameters visited by the MCMC. To speed up computations, without giving up on required precision, we employ two optimization strategies. First, we optimize the initialization of the starting values (Pinit(a)in Fig. 1) by leveraging (saved) solutions from the current state of the MCMC. At each MCMC draw r,wesetPinit(a)=(1−γ)P∗ s(a)|r−1+γ/|A|, 17 Convergence of the algorithm has been proven by Csiszár (1974). 123
Discrete choice in marketing through... 61 where P∗ s(a)|r−1denotes the converged unconditional choice probabilities from the last MCMC draw, γ∈(0,1]is a weight parameter, and |A|is the number of alternatives in the choice set.18 For r=1 we use uniform choice probabilities as initial values. Second, we implement a parallelized version of the Blahut-Arimoto algorithm, which simultaneously computes likelihoods for multiple choice observations. Identification As only the ratio of costs and preferences can be identified from choice data, preferences βsand βccan be identified by fixing λand leveraging variations in simple and complex attributes across alternatives and choice sets presented to the DM.19 However, there may be situations without variation in simple attributes. For example, one could think of “brand” as the only simple attribute. The RI model is still identified in this case, subject to variation in a complex attribute, e.g., total price, in the same way as a simple RUM with alternative specific constants. However, with only one configuration of simple attributes, there is only one optimal processing strategy (see expression 2). If this strategy results in a full consideration set, the model is empirically indistinguishable from a standard RUM. However, note that depending on unobserved heterogeneous preferences, consumers may have different consideration sets that are endogenous to their brand preferences and price sensitivity. However, if the environment is such that not all the brands are available all the time, there will be variation in the conditioning argument to the optimal processing strategy under RI. Finally, it is important to note that point identification of preferences may not be achieved, since deterministic choice is a possible endogenous outcome under RI. This occurs, for instance, at extreme values of λ. Intuitively, a DM with a very high information acquisition cost does not react to variations in the complex attributes, and one with very low information acquisition costs fully processes the complex attributes, leading to deterministic choices (see Section 4). Both of these cases result in set-identification of preference parameters.20 In this context, our Bayesian inference framework, illustrated next, will be useful, as the likelihood surface then exhibits flat regions, complicating maximum likelihood estimation. 3.3 Empirical identification with DCE data under preference heterogeneity We now illustrate the Bayesian estimation and empirical identification of the RIDCM with general beliefs proposed in Section 3.2, using simulated data, in a “small T, large N" setting, typical of DCEs in marketing. We use standard weakly informative subjective prior settings for the parameters indexing hierarchical prior distributions (see Appendix A.2 for details). 18 The transition kernel of our MCMC naturally results in unconditional distributions (implied by a proposal for a new parameter value) that often are close to those at the current state. The initialization guarantees non-degenerate starting values, i.e., a vector of probabilities with all entries larger than zero and smaller than one. In our simulations, we found γ=0.2 to work well. 19 Appendix A.1 illustrates how a difference in processing costs λ can be empirically identified if a DM with invariant preferences makes decisions in different environments. 20 We illustrate set-identification of preferences in Appendix A.3. 123
62 S. Turlo et al. Table 3 Posterior means of preference distributions for different model specifications Model Brand Price Discount |βp/βd||βp/βb|LMD Data generation 2.50 -1.00 ≡−βp1.00 0.40 RI-DCM 2.50 -1.00 ≡−βp1.00 0.40 -2,530.29 (0.05) (0.02) RU logit separate 27.61 -9.16 4.44 2.06 0.34 -2,620.36 (0.69) (0.22) (0.13) RU logit joint 9.81 -3.95 ≡−βp1.00 0.40 -4,558.35 (0.19) (0.07) We report data generating parameters as well as the estimated posterior means of the RI-DCM and the benchmark logit models with and without the constraint βp=βd. Standard errors are in parentheses. |βp/βd|and |βp/βb|are the ratios of mean coefficients We simulate data from the following hierarchical setup. A sample of rationally inattentive DMs (N=1,000) face T=20 choices between an inside and an outside good each. The utility of the inside good to DM jin choice task tis given by uj,t= βb,j+βp,j(pt−dt)where βb,jis the brand coefficient, βp,jthe price coefficient, ptis the price, and dtis the discount. The utility of the outside option is normalized to zero: uO=0. In our simulation, brand and price are simple attributes and thus perceived and processed, i.e., integrated to an overall utility, immediately and at no cost, while the discount requires costly processing. We first illustrate the case without heterogeneity in processing costs. Here, all individuals have the same processing costs of λ=0.25. DMs differ in their structural utility parameters. Preference coefficients are generated from the following distributions: βb∼N(2.5,0.25),βp∼N(−1,0.04). Prices pand discounts dare drawn uniformly and independently from the following sets: p∈{2.5,3,3.5,4,4.5}and d∈{0,0.5,1,1.5,2}such that the resulting design is orthogonal. Individuals know the value of the price p, and for all prices and for any discount level dthe prior beliefs are given by Pr(d=d|p)=1/5. With the simulated data, we fit the RI-DCM and two RU logit specifications: one that allows for separate price and discount coefficients (“RU logit separate”) and one with only one coefficient measuring the utility of money (“RU logit joint”), as in the data-generating process. We rely on Rossi’s bayesm-package for the estimation of the hierarchical RU logit (Rossi et al., 2005). The estimation of the hierarchical RI-DCM employs Metropolis-Hastings steps to update individual-level preference parameters and relies on standard results for updating parameters indexing the hierarchical prior distributions (e.g., Rossi et al., 2005). We obtain the likelihood by solving the problem Eq. 2for given parameters numerically with the Blahut-Arimoto algorithm, as described in Section 3.2. Without loss of generality, we fix the value of λto be equal to its true value in estimation. We will revisit this point below. Table 3summarizes posterior means and Table 4reports posterior variances. We see that the estimated RI-DCM nicely recovers data-generating parameters. We use log marginal density (LMD) estimates throughout the paper to compare model fits 123
Discrete choice in marketing through... 63 Table 4 Posterior variances of preference distributions for different model specifications Model Brand Price Discount Data generation 0.25 0.04 ≡Var (βp) RI-DCM 0.33 0.09 ≡Var (βp) (0.05) (0.01) RU logit separate 16.00 1.71 1.93 (4.02) (0.39) (0.31) RU logit joint 4.03 0.71 ≡Var (βp) (0.71) (0.11) We report data generating parameters as well as the estimated posterior variances of the RI-DCM and the benchmark logit models with and without the constraint βp=βd. Standard errors are in parentheses (see Rossi et al., 2005). The inferior fits of the RU logit models, the (still) current benchmark, testify to the empirical identifiability of the RI-DCM (see the last column in Table 3). The RU logit with separate parameters for price and discount fits the data much better than the RU logit with only one price coefficient, i.e., suggesting that different “sources of money” are valued differently. However, here the larger magnitude of the price coefficient, relative to the discount coefficient, simply reflects that rationally inattentive DMs react to the realized discount value adaptively, both as a function of the realized (simple) price that varies across choice sets and as a function of heterogeneous preference coefficients that vary across DMs. As a consequence, the RU logit also struggles with measuring heterogeneity in preference parameters. For example, the RU logit dramatically overestimates the heterogeneity in the price coefficient (and the discount coefficient, where separately specified). This observation is important given that existing applications of RI to discrete choice with observational data present reinterpretations of RU logit choice conditioned on additively separable indices. Finally, Table 5reports elasticities for all models given changes in different discount and price levels. Elasticities are calculated using posterior expected (changes in) choice probabilities and, therefore, incorporate all posterior uncertainty. The columns Discount A and Price A report elasticities resulting from a discount decrease from d=1tod=0.5, and a price increase from p=4top=4.5, respectively. Similarly, columns Discount B and Price B denote scenarios where discounts decrease from Table 5 Elasticities for different price and discount levels Model Discount A Price A Discount B Price B RI-DCM 0.28 0.52 0.21 0.52 RU logit separate 0.25 0.53 0.29 0.58 RU logit joint 0.38 0.38 0.38 0.38 Columns report elasticities from changes in discount (Discount A and Discount B) and changes in price (Price A and Price B). The elasticities in the first two columns are calculated for an inside good with p=4 and d=1 and in the last two columns for an inside good with p=5andd=2 123
64 S. Turlo et al. d=2tod=1.5, and prices rise from p=5top=5.5. In all instances, absolute and relative changes in total price are the same. Under RI, price elasticities depend both on the source of the price variation and on the composition of the inside good. The RI-DCM yields that the price elasticity is much smaller (larger) when the change in total price comes through the complex discount (the simple price). The RI-DCM also yields that price elasticity from changing the discount further decreases at the higher simple price (compare columns Discount A and Discount B in Table 5). The reasons are the information friction associated with processing and integrating the complex discount and the stronger prior against the inside alternative when the simple price is larger. The RU logit separate correctly picks up that price changes from the complex discount result in smaller elasticities than simple price changes. However, this model wrongly suggests that the elasticity from changing the price through the discount increases at the higher simple price. Finally, the RU logit that does not differentiate between price changes through the simple price and the complex discount necessarily overestimates (underestimates) elasticities from changing the discount (changing the simple price). 3.4 Heterogeneous preferences, heterogeneous information processing costs In general, both processing costs and preferences likely vary across DMs. Hence, it may be theoretically appealing and efficient in a hierarchical model to structure heterogeneity in βas residual heterogeneity after considering heterogeneity in λ. Such a decomposition in the context of a hierarchical RI-DCM can be viewed as a microfounded version of the idea behind Fiebig et al. (2010)’s generalized multinomial logit model that, in addition to preference heterogeneity, captures the heterogeneity in the scale of the error term in a hierarchical RU logit. As already mentioned, the objective function in the RI framework is homogeneous of degree one with respect to payoffs and costs, and only their ratio is directly identifiable from choice data. However, from a statistical point of view, the combination of continuous heterogeneity in preferences with continuous heterogeneity in information Table 6 Posterior means of preference distributions for the RI-DCM with and without scale mixture component Model Brand Price Discount σ2 log(λ) LMD Data generation 2.50 -1.00 ≡−βp0.09 RI-DCM 2.40 -0.95 ≡−βp0 -2,984.55 (0.04) (0.02) RI-DCM scale mixture 2.48 -0.99 ≡−βp0.08 -2,884.90 (0.06) (0.02) (0.01) We report data generating parameters as well as the estimated posterior means of the RI-DCM with homogenous and heterogeneous information processing costs. Standard errors are in parentheses. σ2 log(λ) is the variance of log information processing costs in the population 123
Discrete choice in marketing through... 65 Table 7 Posterior variances of preference distributions for the RI-DCM with and without scale mixture component Model Brand Price Discount Data generation 0.25 0.04 ≡Var (βp) RI-DCM 0.41 0.07 ≡Var (βp) (0.05) (0.01) RI-DCM scale mixture 0.31 0.05 ≡Va r (βp) (0.06) (0.01) We report data generating parameters as well as the estimated posterior variances of the RI-DCM with homogenous and heterogeneous information processing costs. Standard errors are in parentheses costs gives rise to a scale mixture distribution. For a well-known example, the Student t-distribution can be derived as a scale mixture of normal distributions. Because scale mixtures can be identified and distinguished from their non-mixed counterparts (see, e.g., Choy & Smith, 1997), one can identify heterogeneity in the unit information costs and preferences, as long as one is willing to assume a continuous distribution of preferences that is not a scale mixture. For example, popular semiparametric distributions such as a mixture of normals are strictly continuous and not scale mixtures. However, again because the objective function defining the RI problem is homogeneous of degree one, the joint distribution of preferences and unit information costs is only identified up to the first moment of the latter distribution. For illustration, we extend the simulation from Section 3.3 to include heterogeneity in information processing costs with log(λj)∼N(μlog(λ) =−1.4,σ2 log(λ) =0.09)in the data generating mechanism. As noted previously, the mean of this distribution is not jointly identified with the mean of the distribution of preference parameters. Hence, we fix μlog(λ) to the data generating value in estimation and without loss of generality.21 Tables 6and 7document that we recover the joint distribution of preference parameters and information costs subject to fixing the first moment of the latter. Moreover, we see that by not accounting for heterogeneity in information processing costs, one overestimates the preference heterogeneity, here reflected in brand and price variance estimates in the RI-DCM with homogeneous information costs. Complementing the illustrative simulations here, we conduct a simulation study that varies the number of inside goods (one versus two), simulates data with and without heterogeneity in processing costs (in addition to preferences heterogeneity), and adds a two-stage choice model with consideration sets from screening on price (see, e.g., Gilbride & Allenby, 2004; Pachali et al., 2023) to the model comparison. We simulate 50 data sets in the four data-generating settings. The benchmark models are estimated once with a single price parameter and once with separate parameters for the (simple) price and the (complex) discounts. We summarize results in Tables 15,16,17, and 18 and discuss additional details in Appendix A.2. Not surprisingly, we find that only the RI-DCM recovers data-generating parameters. We also find that (i) we can reliably distinguish the data generating RI-DCMs from the benchmark models, (ii) the benchmark models fare relatively much worse in the 21 Normalization of the information processing costs is analogous to that of the error variance in RU logit models in that it does not affect market share or welfare computations. 123
66 S. Turlo et al. larger choice set because of the corresponding increase in number of optimal RI information strategies, and (iii), slightly worse when processing costs are heterogeneous (in addition to preferences heterogeneity). 3.5 On the distinction between simple and complex aspects of a choice task Different from extant search models, RI motivates information frictions even in situations in which all attribute information is essentially equally accessible, and the DM’s challengeisnot to resolvevaluesof unknownattributes buttointegrate accessible information to overall utility. Hence, another practical challenge of the proposed framework is the identification of simple and complex attributes. In some cases, prior knowledge may be sufficient to classify attributes, possibly as a function of the specifics of a product category under study or the experimental design. In other cases, we envision that the distinction between simple and complex attributes must be empirical. This distinction is greatly facilitated whenever theory constrains coefficients in the utility function to be equal, such as in the example of different price components. In this case, descriptive models, or even just marginal summaries of the data, can reveal that choice probabilities react more strongly to changes, say, in price component A than in price component B. It follows that price component A is simple relative to price component B, and price component B is complex relative to component A (e.g., Brown & Jeon, 2024). Obviously, this argument fails when theory allows for different utility coefficients for different attributes. Next, we illustrate by simulation that the distinction between simple and complex attributes is likelihood identified, even in this case. Whereas theoretical results imply that any combination of rationally inattentive behavior and information processing costs can be rationalized with some state-contingent payoffs (Lipnowski and Ravid, 2022), we illustrate that additive linear separability in utility contributions can suffice to distinguish between simple and complex attributes empirically, given Shannon costs and a non-degenerate distribution over states.22 One inside good We simulate 2,000 choice tasks, each involving a DM choosing between an inside and an outside good. The inside good is characterized by two linear attributes, xsand xc, that additively combine into overall utility. Attribute xsis simple, and the attribute xcis complex. Each of the two linear attributes is represented by three levels in the experimental design: xs∈{2,2.25,5}and xc∈{1,1.5,3}. Preferences are given by βs=−βc=−1 and information processing cost is set to λ=0.5. In this design, the generated data comprises probabilistic and deterministic choices. Attribute combinations defining inside goods are drawn from independent uniform distributions over the discrete attribute support. With the simulated data, we estimate different structural RI-DCMs. In the first specification, the distinction between simple and complex attributes follows that of the data-generating process. In the second model, we (falsely) reverse what 22 In the context of search, Abaluck et al. (2022) propose how to identify limited information about an attribute under the assumptions that another attribute is known to be fully processed and all-or-nothing learning. 123
Discrete choice in marketing through... 67 Table 8 Identification of simple and complex attributes in a design with one inside alternative and one outside alternative Model min 25% 50% 75% max ML Correct - Linear -380.52 -371.39 -370.68 -370.25 -369.94 -369.93 Misspecified - Linear -1076.47 -1067.22 -1066.28 -1065.71 -1065.11 -1039.15 Misspecified - Linear Out -417.41 -412.02 -411.30 -410.95 -410.78 -410.70 Misspecified - Categorical -382.63 -375.21 -374.04 -373.13 -370.54 -369.23 Quartiles as well as the minimum and the maximum of the log-likelihood MCMC draws are reported for the correct linear, the misspecified linear, the misspecified linear with an additional outside parameter, and the misspecified categorical model, respectively. The last column reports the maximum of the likelihood (ML) under an improper prior, i.e., the “frequentist maximum” is simple and complex in estimation. The third model adds a coefficient for the outside good (equal to zero in the data-generating process). This model isolates the linearity of utility differences between inside goods in attributes as a source of identification. Finally, we estimate a model with completely flexible utility within attributes by coding the two attributes as categorical while misspecifying which attribute is simple and which is complex in estimation. As subjective prior distribution for preference parameters we use β∼N(0,100I). We also report the maximum of the log-likelihood under an improper prior, to safeguard against an undue influence of this subjective prior setting. Table 8presents quantiles of the log-likelihoods from Markov-Chain-Monte-Carlo (MCMC) estimation and the numerical maxima of the likelihoods implied by the different models. We find that, subject to constraints on the utility function, there is scope for empirical identification of what is simple and complex in a choice task (comparing the first three lines of Table 8). However, once we give up on linearity in attributes (see the last line in Table 8), we can no longer distinguish between simple and complex in this minimal example.23 Two inside goods A basic constraint from utility theory, namely that of no crosseffects between alternatives in a model of perfect substitution, does not come into play when there is only one outside good. To showcase identification from this constraint, we extend the simulation described above and include a second inside alternative. The DM chooses between two inside goods and an outside good. The inside goods have two attributes, one simple and one complex with three levels each. The major difference to the one inside good case is that simple attribute realizations of two inside goods now effectively interact in determining the optimal processing strategy. For example, particular realizations of simple attributes may lead to considering both, 23 In Appendix A.3, we additionally show that we can no longer distinguish between simple and complex attributes based on model fit once processing costs are either small enough or large enough such that (essentially) deterministic choices ensue. When λbecomes sufficiently small, all information is fully processed, and the conceptual and empirical distinction between simple and complex vanishes. When λbecomes sufficiently large, the information in complex attributes is never integrated into the overall evaluation of alternatives, and a model with extreme coefficients for the simple attribute (misspecified as complex) will approach a perfect fit to the data. 123
68 S. Turlo et al. Table 9 Identification of simple and complex attributes in a design with two inside and one outside alternative Model min 25% 50% 75% max ML Correct - Linear -715.7 -709.2 -708.6 -707.6 -707.3 -707.1 Misspecified - Categorical -1589 -1585 -1583 -1582 -1581 -1580.3 Quartiles as well as the minimum and the maximum of the log-likelihood MCMC draws are reported for the correct linear and the misspecified categorical model, respectively. The last column reports the maximum of the likelihood (ML) under an improper prior, i.e., the “frequentist maximum” only one, or none of the two inside alternatives. Misspecifying the simple attribute that drives the DM’s choice of information strategy as complex fails to capture these possibilities,evenif we drop the linearity constraint and code all attributescategorically (see Table 9). 4 Features of discrete choice in MAMA contexts under RI In this section, we illustrate the implications of RI theory for discrete choice in MAMA settings, as common in marketing. We use the RI-DCM with general beliefs and show through simulations how several well-documented phenomena in the discrete choice literature that are difficult to justify in a RU framework naturally follow from RI. We further demonstrate that estimating a standard RU logit can yield misleading conclusions about the behavior of RI agents. This approach follows a common research strategy in that literature. Often, various logit models are estimated and compared in different contexts, e.g., with varying numbers of inside alternatives. This comparison then allows us to identify the moderating effect of context, thereby revealing deviations from the standard RU logit model. While certain implications have been discussed in prior (theoretical) RI literature, we add to this body of work by exploring additional implications due to the MAMA structure. In general, the presented implications arise from how RI agents translate the MAMA choice environment into the unconditional choice probabilities outlined above. Existing RI-DCMs impose simplifying assumptions on this very process to facilitate estimation at the expense of not capturing the features presented here. In our illustrations, we distinguish between i) endogenous features of RI (Section 4.1) and ii) context effects (Section 4.2). Endogenous features arise as a result of the optimal allocation of limited cognitive resources in the RI-DCM. These features can explain empirical phenomena in MAMA settings that cannot be captured by basic RUMs. Previously, researchers have applied diverse, non-unified adjustments to RUMs to address these phenomena. Table 10 outlines important endogenous features of RI and includes examples from studies that have made non-unified modifications to RUMs to accommodate the related phenomena. In Section 4.1, we discuss the following endogenous features of discrete choice under RI: 123
Discrete choice in marketing through... 75 Fig. 3 Iso-choice-probability sets for the RI-DCM and the RU logit. This Figure displays sets of pricediscount combinations of the inside good that result in the same conditional choice probabilities for the RI-DCM (left panel) and the RU logit (right panel). The thin solid line is the 45◦line. Under the RU logit, the iso-choice probability sets are linear and individual sets are parallel to each other. In contrast, under the RI-DCM, the substitution rate varies depending on the composition of the inside good so that it is non-linear and the individual sets are not parallel. The dots in the left panel indicate actual attribute combinations of the inside good RI, it matters whether the source of utility is a simple or a complex attribute. As shown in the left panel of Fig. 3, the ratio of changes in the discount and price that keep choice probabilities constant is smaller than one, i.e., an increase of the discount by one unit offsets an increase in the price that is strictly smaller than one under RI even though both price and discount have the same impact on utility. This is explained by the information friction present in the RI-DCM, which makes choice probabilities a function not only of an alternative’s utility value but also of the source of that utility. Under the RI-DCM, there will be prices that are sufficiently low (high) so that the DM chooses deterministically (given a simple price) in the limit. Consequently, with a fixed discount distribution, the iso-choice sets become lower (upper) contour sets with a boundary that is flat in the discount. In contrast, under the RU logit, there are no combinations of finite discounts and prices so that the DM chooses deterministically. For a final illustration, Fig. 4displays the conditional choice probability of the inside good for different combinations of the price and the discount for a fixed utility. All points displayed are associated with the same net utility equal to one, uI=1 where uI=βb−βpp+βddwith βb=6 and βp=−βd=−1. However, going from left to right, we increase the discount and the price simultaneously by the same amount so that p−5=dwithin the same experimental design. Under the RU logit, the choice probability of the inside good remains constant. In contrast, choice probabilities decrease weakly as the price increases under the RI-DCM. Even though the discount is just a negative price in utility terms, the distinction matters under RI if price is simple and the discount is more complex to process and integrate. 123
76 S. Turlo et al. Fig. 4 Conditional choice probabilities for different compositions of the inside alternative under RIDCM and RU logit. This Figure depicts the conditional choice probability for the inside good in the discount under the constraint that the price increase equals the discount increase with p−5=dso that the net utility of the inside good is constant. While choice probability is constant under the RU logit, it is weakly decreasing when both the price and the discount are increasing in the considered variation of the inside good composition. The dots are the conditional choice probabilities of the specific inside good compositions 4.1.3 Inattention to alternatives A central feature of discrete choice under RI is that consideration sets form endogenously. These sets include only the alternatives that have a strictly positive probability of being chosen. Consideration sets arise as a result of the DM’s optimal information strategy. Here, we demonstrate how attributes—specifically the configuration of simple attributes—determine which alternatives are considered.26 Figure 5depicts choice probabilities in a choice set with four inside alternatives i=1, ..., 4 that provide utilities of ui=βb,i−pi+diwith βb,i∈{2,1.75,1.5,1}, pi=2 and a discount diwith Pr(di=0)=Pr(di=2)=0.5. The specific values for βb,iare chosen for illustration purposes. While βb,4represents the least preferred brand alternative, we demonstrate how the DM reacts qualitatively differently to this brand as the composition of the choice set changes. The top-left panel of Fig. 5 26 For related illustrative examples of endogenous consideration set formation that do not differentiate between attributes, see Caplin et al. (2019). 123
Discrete choice in marketing through... 77 Fig. 5 Consideration set formation. The above panels present choice probabilities of a rationally inattentive DM facing up to four alternatives {a1,a2,a3,a4}ordered from highest to lowest simple brand βb,i with identical prices piand identically distributed complex discounts di. The top-left panel shows the unconditional choice probabilities. The top-right depicts conditional choice probabilities given a choice set where alternative a4provides the highest utility. Due to information frictions, even in such a case, alternative a4is never chosen. The bottom-left panel exhibits the updated unconditional choice probabilities in a reduced choice set that drops alternative a1. Finally, the bottom-right panel shows that, as a consequence of updating unconditional choice probabilities to the smaller set, alternative a4has the highest conditional choice probability in the state where it delivers the highest payoff illustrates unconditional choice probabilities that represent the DM’s beliefs about how she will choose before any processing of the complex discount has taken place. The top-right panel shows the choice probabilities conditioned on a specific choice set, i.e., realized values of the complex discount (from the analyst’s perspective, as the DM will not necessarily learn the exact choice set because of processing costs). The complex discounts in the specific choice set are d1=d2=d3=0 and d4=2, implying that a4provides the highest utility in the specific choice set (u4=1>uj for j= 4). Yet, as apparent from the figures, the DM does not consider a4because excluding a4from her endogenous consideration set was (a priori) optimal for the DM given her processing costs and the small prior probability that a4is, in fact, optimal. However, in contrast to extant two-stage models (e.g., Gilbride & Allenby 2004; Goeree, 2008; Terui et al., 2011), RI predicts that a4will be considered once a1is no longer available to the DM (see the unconditional choice probabilities in the bottomleft panel of Fig. 5). Due to updating unconditional choice probabilities to the smaller 123
78 S. Turlo et al. set, alternative a4has the highest conditional choice probability in the state where it delivers the highest pay-off (bottom-right panel). Accordingly, RI implies deterministic consideration sets conditional on λ,prior beliefs μsand a utility function, while choice conditional on such a consideration set is stochastic. However, because prior beliefs μsare a function of realized simple attribute values, consideration sets will generally change from choice set to choice set in ways that cannot be captured by an alternative specific index or decision rule. If the simple information in a choice set does not vary, RI still implies consideration sets that are endogenous to consumers’ preferences, priors over complex attributes, and information processing costs. An attractive feature of characterizing consideration sets this way is that the exclusion of, e.g., a brand from consideration in a particular choice set, does not have to be motivated by persistent extreme tastes or screening rules. The latter may not generalize to changes in the set of available brands or the priors over complex attributes. 4.2 Effects of choice environment variations under RI 4.2.1 Impact of information processing costs and incentives A large body of experimental evidence documents that the complexity of choice tasks and incentives affect choice behavior (e.g., Swait & Adamowicz, 2001; Ding et al., 2005). In RI, both aspects influence attention allocation and, thus, observable choice. Recall that the costs of information processing in RI are the product of mutual information and the strictly positive unit information cost λ>0, see expression Eq. 2. Structurally, characteristics of the (expected) choice task as well as characteristics of the DM relate to λ(e.g., Regier et al., 2014). To continue with our example of a complex discount, processing the eligibility requirements will be affected by the number of criteria that must be checked, or even the font size used to describe the discount. Intuitively, more criteria that need checking or a smaller font size will increase λ. Similarly, a less constrained DM, or more experience with the product or the eligibility criteria, will be reflected in a relatively smaller λ. Finally, it is possible to cast λas a function of the incentives offered in a DCE.27 For example, if an incentivized DCE instructs participants facing Nchoice tasks that one of the Nchoices will become an actual transaction, the realization probability of a specific choice is ρ=1/N. With probability 1 −ρthe DM’s choice is hypothetical, i.e., she does not actually obtain the chosen alternative.28 From the perspective of the 27 The implicit assumption here is that a change in incentives does not affect the payoff function or the (subjective) prior distribution over states ω. 28 Recall that in an incentive-aligned DCE, the DM is endowed with a budget. When the DM chooses the outside option, she retains the endowed budget. 123
Discrete choice in marketing through... 79 DM, the resulting RI objective function for each choice task is given by ρ ω∈ μs(ω) a∈A P(a|ω)u(a,ω) −λ ρ ω∈ μs(ω) a∈A P(a|ω) ln P(a|ω) − a∈A P(a)ln P(a). This formulation reveals that a higher realization probability has the same impact on choice behavior as a decrease in the information processing costs. For example, a 1% increase in information processing costs λwill be offset by a 1% increase in the realization probability ρ. Thus, one way of interpreting the patterns we illustrate next is through the lens of changing incentives in a DCE. Consider the case when the DM chooses between an inside alternative, characterized by a simple brand valued at βb=1.2 at a simple price p=2 and a complex discount dwith Pr(d=0)=Pr(d=2)=0.5. Based on expected utility, the DM thus prefers the inside good. Figure 6illustrates how λaffects attention and choice in this example. The top left panel shows that when λincreases, the processing of complex attributes decreases until the DM learns nothing beyond the known distribution of the complex discount attribute (at λ=λ), i.e., conditional choice probabilities equal the unconditional choice probabilities. At λ≥λ, the DM deterministically chooses the inside option based on prior expectations (bottom left and right panel), which of course implies that choice probabilities no longer change as a function of the complex discount (top right panel). At λ=λ=0, the DM perfectly learns the complex discount and deterministically chooses the alternative with the highest utility, and hence maximally reacts to changes in the complex discount value. Note that changes in λcan impact what is revealed about the DM’s preferences. The bottom-left panel of Fig. 6shows that as λincreases, there is a choice reversal as the DM switches from choosing the outside good (based on learning that the complex discount does not apply) to choosing the inside good (based on prior expectations and without learning the true choice set ω). Finally, the bottom-right panel of Fig. 6shows that choice probabilities, and here specifically the probability of making a choice error (from the point of view of the analyst who has all information about alternative specific payoffs), can be non-monotonic in the amount of cognitive processing (topleft panel). The non-monotonic relationship here derives from the prior pointing to the payoff maximizing choice in the absence of processing complex information. To showcase what an analyst taking a RU perspective when analyzing RI choice data may find in this example, we fit logit models to RI choices conditional on different values of λ. Figure 7summarizes point estimates of logit coefficients for brand, price, and discount, across different simulated RI choice data sets with varying λ. We see that absolute values of the brand and price coefficients, i.e., the coefficients associated with the simple attributes, first decrease and then increase, while that of the complex discount decreases in λ. The latter effect is immediate, since higher 123
80 S. Turlo et al. Fig. 6 Impact of information processing costs on attention, discount effect, and choice. The upper left panel shows mutual information as a function of information costs. The upper right panel displays the impact of a fixed increase of the complex discount (from 0 to 2) on conditional choice probability for different levels of λ. The panels in the bottom row show conditional choice probabilities of the inside alternative in choice sets where the discount is 0 (left) and 2 (right). Thus, the inside good provides a lower (higher) payoff in the left (right) panel than the outside option. Note that choice is deterministic for information costs λequal to zero and larger than λ processing costs dampen the effect of the discount, as discussed previously. The rationale for the former pattern is that at small values of λ, the DM processes in most choice sets all available information, and choice becomes nearly deterministic. For intermediate levels of λ, some choice sets (characterized by different simple prices) will motivate more, and some less information processing, causing a higher level of overall stochasticity in the data that is reflected in absolutely smaller brand and price coefficients (from the viewpoint of a RU logit). Fig. 7 Logit approximation for different levels of information processing costs. Each panel shows logit estimates of the respective coefficients for varying levels of information processing costs λ. Note that each point is the result of an estimation from simulated data with T=1,000 choice tasks each. The data generating parameters are βp=−βd=−1, βb=2, pis uniformly drawn from [2,4],anddis distributed with Pr(d=0)=Pr(d=2)=0.5 123
Discrete choice in marketing through... 81 Fig. 8 Choice consistency is non-monotonic in the information processing costs λ. Choice consistency is measured by McFadden’s pseudo R-squared. For details, see Domencich and McFadden (1975) Eventually, as λincreases, realized discount values are ignored, as it becomes too costly to process the corresponding complex eligibility requirements and the discount coefficient approaches zero in the logit fit. However, as the amount of processing of complex information decreases beyond some level, so does the level of stochasticity in the data. Eventually, RI choices are based on prior information only, conditioned on the simple attributes brand and price here, and deterministic. This is reflected in absolutely increasing brand and price coefficients in the logit fits, summarized in Fig. 7. It is common in the choice modeling literature to report the estimated error term variance as a measure of choice consistency (e.g., DeShazo & Fermo, 2002. Figure 8 plots McFadden’s pseudo R-squared in relation to λ. It illustrates that RI choices are more deterministic at very low and very high values of λ, and less deterministic at intermediate values. Figure 9extends the illustration of RI choice as a function of λto the case of three alternatives.29 Alternative ai,i=1, ..., 3, yields utility ui=βb,i−pi+diwith βb,1= 3.5, βb,2=3.25, βb,3=3, pi=4, and diare independently distributed according to Pr(di=0)=Pr(di=2)=0.5. Based on prior information, alternative a1is the best and a2is the second best. Figure 9displays conditional choice probabilities for a choice set ωwhere alternative a3provides the highest payoff based on realized values of the complex discount attribute, illustrating how information costs λimpact the formation of consideration sets. As λincreases, the number of alternatives chosen with strictly positive unconditional probability first decreases from three to two (at λ), and eventually results in deterministic choice of a1based on prior considerations only (to the right of λ). As the information costs increase, the costs of resolving uncertainty about a priori less attractive alternatives outweigh the (expected) benefits. As a consequence, it becomes optimal to ignore such alternatives even if there exist choice sets ωin which the ignored 29 For ease of exposition and without loss of generality, there is no outside option in this example. 123
82 S. Turlo et al. Fig. 9 Impact of information costs on consideration set size. Conditional choice probabilities for three alternatives are displayed as a function of information processing costs λin the choice set where a3is the best alternative. As information costs increase, the DM rationally chooses to ignore alternatives in the choice set. λ,λ,andλ indicate threshold values for which the consideration set size changes. For costs λ=0 choice is deterministic, and the best alternative a3is always chosen, however, after considering all alternatives. For costs larger than λ, choice becomes deterministic again, however now because the DM ignores alternatives a2and a3, regardless of realized complex discount levels alternatives provide the highest payoff (as depicted in Fig. 9).30 Together, the illustrations in this section suggest that RI provides a useful basis for bridging across choices under different incentive or difficulty levels. 4.2.2 Attribute range/dispersion and levels effects Next, we show how the attribute range, typically measured as the difference between the highest and the lowest level of an attribute, or more generally, the dispersion of complex attributes, moderates the impact of a one-unit increase in that attribute on choice. The underlying mechanism is that as the range of the complex attribute increases, the expected gain from identifying its realized value also increases, making processing information more valuable. This ultimately increases the impact of the complex attribute on choice and in contrast to what one would expect when taking a RU perspective. Figure 10 illustrates this mechanism in our leading car example. We set βb=6 and recall that βp=−βd=−1. Here, we study the impact of an increase in the discount from 2 to 3 on conditional choice probabilities for different discount ranges. Both lines in Fig. 10 depict how conditional choice probabilities change when the complex discount increases by one unit for different values of the simple price. The dashed blue line represents the case where the complex discount is drawn from the set {1,2,3,4}, while in the second case (solid red line) the discount takes values in 30 This way, RI can motivate a positive probability of choosing an alternative that is dominated a posteriori, i.e., after processing (some) complex information (cf. Ruan et al., 2008 who propose a sequential sampling model to model dominance as a form of similarity). 123
Discrete choice in marketing through... 83 Fig. 10 Range effects of complex attributes. This Figure displays conditional choice probability differences of the inside alternative in reaction to an increase of the discount from 2 to 3 given a large range (solid red) and a small range (dashed blue) of the complex discount as a function of the simple price. Specifically, complex attribute levels are {1,2,3,4}when the range is small, and they are {0,2,3,5}when the range is large {0,2,3,5}. In both cases, the DM’s beliefs are uniform over the respective support. Figure 10 shows that as the range of the complex discount attribute increases, the impact of a one-unit increase of the discount also increases. Technically, an increase in the attribute range spreads the range of possible payoffs from choosing the inside good further, motivating larger (costly) departures of conditional choice probabilities from their unconditional counterparts as a result of the optimal processing strategy. To showcase what an analyst taking a RU perspective when analyzing RI choice data may find in this example, we fit logit models to simulated RI choices conditional on different ranges of the complex discount in the experimental design. We generate data sets as follows: βb=5, βp=−βd=−1, p∈[8,10], and dis drawn with equal probability from the binary set {4−x,4+x}with x∈[0.25,4]. Figures 11 and 12 summarize the logit estimates as well as McFadden’s pseudo R-squared values as a function of the range of the complex discount in a particular experimental design. When the range of the complex discount is small, the estimated brand and price coefficients are absolutely large, and the discount coefficient is, relatively, much smaller. As the range increases, brand and price coefficients become smaller in absolute value, and the inferred discount coefficients tend to increase. 123
84 S. Turlo et al. Fig. 11 Logit estimates for different levels of the discount range. This Figure displays estimated coefficients of brand, price, and discount for different ranges of the complex attribute. Each point is the result of fitting a logit model with simulated data from T=1,000 choices with data generating parameters λ=0.5, βb=5, βp=−βd=−1, and pbeing uniformly drawn from [8,10]. The discount range, given as the difference between the two discount levels, is varied from 0.5to8 Figure 12 exhibits a U-shaped relationship between choice consistency and the range of the complex discount. When the range is very small, the DM makes less costly mistakes when choosing based on prior beliefs, conditioned on the simple attributes brand and price in this example. As a consequence, the DM pays little attention to the realized discount levels and chooses rather consistently based on simple attributes as well as the expected discount level only. As the range increases, the mistakes when choosing based on prior beliefs can become rather costly. However, there still are prices at which learning the realized complex discount in addition does not add much value. In the RU logit approximation, the relative importance of the discount increases, however, choice consistency decreases. Finally,whenthediscountrangebecomessolarge that fewerand fewersimple prices within the support of the design translate into a good enough choice (in expectation) without knowing the realized discount level, the DM will process the complex discount consistently. This again results in less stochastic data. Fig. 12 Non-monotonic effect of the discount range on choice consistency. Choice consistency is measured as McFadden’s pseudo R-squared. For details, see Domencich and McFadden (1975) 123
Discrete choice in marketing through... 91 key question is thus how to assess and simulate (likely) beliefs in the market setting. Choosing belief distributions based on tractability in a closed-form logit framework with additively separable indices, as currently standard in empirical applications (see Section 3.1), does not seem to be a satisfactory solution. In light of a growing empirical RI literature that relies on belief assumptions motivated by analytical convenience, we feel that research into the sensitivity with respect to different assumptions about beliefs will be useful, as well as an integration of methods to study potentially heterogeneous market beliefs empirically. The characterization of an RI-DCM that allows for general prior beliefs in this paper paves the way for this line of research. Other frameworks modeling limited information, such as consumer search models or learning models, face a similar challenge of dealing with typically unobservable (consumer) beliefs as a key building block to empirical analysis and counterfactual computations and are often quite sensitive with respect to assumptions about beliefs (Chintagunta & Nair, 2011). Notably, some of these contributions have explored the added benefit of eliciting beliefs through auxiliary information rather than relying on purely theoretical assumptions, such as rational expectations. Successful approaches that have been employed in other areas to learn about beliefs use survey-based belief elicitation methods (e.g., Cavallo et al., 2017; Coibion et al., 2018; Armona et al., 2019) or observational data such as clickstream data (e.g., Hu et al., 2019), or eye-tracking data (e.g., Ursu et al., 2024). For a recent overview in marketing, see the literature section in Jindal and Aribarg (2021). 5.5 Alternative information cost functions There is an ongoing discussion in the economics literature on the question of which attention cost function is appropriate for which choice circumstances. Most of the experimental studies, either explicitly or implicitly, estimate the cost function, which is then used in subsequent analyses. Consequently, much of the experimental literature dealing with RI (implicitly) tests the applicability of different cost functions in a variety of stylized settings. As in the present paper, many results in RI theory have been derived under the assumption of costs that are linear in Shannon mutual information.35 For instance, experimental evidence suggests that non-linearities of information costs exist, where subjects pay too little attention to high rewards compared to low rewards (Caplin & Dean, 2013; Dean & Neligh, 2019). Perhaps more important for analyzing MAMA choice, a linear Shannon entropy cost function precludes the concept of perceptual distance, where some states are harder to distinguish than others (Dean & Neligh, 2019; Hébert & Woodford, 2021). As a result of this critique, some papers generalize the linear Shannon cost function or apply different cost functions. For instance, Hébert and Woodford (2021) introduce a neighborhood cost function to 35 Ma´ckowiak et al. (2023) discuss what features of RI are robust with respect to the functional form of information processing costs. 123
92 S. Turlo et al. address the problem that some states are harder to distinguish than others.36 A detailed discussion of this mostly theoretical literature is beyond the scope of this paper.37 We believe that the choice of cost function could become important for the applications envisioned in this paper as well. For instance, the model formulation based on Shannon costs implies that any two possible choice sets (states ω) within an experimental design are equally hard to distinguish. However, it is not unlikely that the distinction between two choice sets that pose many complex trade-offs may be harder than that between two choice sets that pose fewer trade-offs each. We thus conjecture that neighborhood-based costs that allow to impose that certain sets of choice sets are harder to distinguish than others may eventually become useful. We note that this is related to a finer distinction of levels of processing difficulty than the distinction between simple and complex attributes proposed in this paper. 5.6 Conceptualization of free information The history of research on choice in marketing, economics, and psychology is rich in empirical and theoretical results about what may be simple and more complex about a particular choice task or set of such tasks. For example, it may be worthwhile to reconsider the literature on heuristics in choice as information (for the analyst) about simple aspects of choice tasks that, in the context of a RI-DCM, may give rise to prior beliefs that guide the amount of processing of more complex aspects of a choice task.38 In our suggested specification of the RI-DCM, some attributes are considered simple so that processing them does not require (significant) cognitive effort. Thus, it is these attributes xsthat determine (conditional) prior beliefs μs, which then form a key ingredient to how much processing of complex attributes should occur (see Section 2). This formulation imposes a specific mapping from choice sets into prior beliefs that is not given by RI theory but assumed by the analyst (even if, as we showed, an empirical distinction between simple and complex attributes is possible, in general). In principle, any aspect of a choice set and the attribute configuration presented therein could be simple information and thus determine prior beliefs. Consider, for example, a DCE design where the discount has values d∈{50, ..., 90,100}.ADM may easily recognize that the discount equals 100 due to the substantially different visual stimulus (two vs. three digits). In this case, the DM may assign a positive belief to all choice sets where d=100 and zero to all other. This raises the conceptual question of which pieces of information in any choice task can be considered free and, consequently, how to map choice sets into (conditional) prior beliefs. Our discussion regarding the empirical identification of simple vs. complex attributes in Section 3.5 should be viewed as a special case of this broader 36 A related generalization of Shannon costs that allows for alternatives to have different information costs is suggested by Huettner et al. (2019). 37 For an overview of different information cost functions applied in the behavioral inattention literature, see Gabaix (2019); for a more detailed discussion on entropy-based cost functions, see Dean and Neligh (2019) as well as Ma´ckowiak et al. (2023). 38 See also Ma´ckowiak et al. (2023) who propose the idea that RI provides a model for the formation of heuristics. 123
Discrete choice in marketing through... 93 consideration. Due to the high dimensionality of this question, it will typically not be viable to give purely empirical answers, and an appeal to theory and previous empirical results is required. While this is beyond the scope of this paper, it suggests that the RI framework may be able to fruitfully integrate conjectured and empirically demonstrated choice simplification strategies with fully rational behavior that is guided by priors formed on the basis of whatever a DM may easily and immediately process about a choice task. With an eye towards industry-grade applications with many attributes and many alternatives, this is an important part of future research. A Appendix A.1 Identification of 1 across decision environments under stable preferences Next, we present estimation results to show that a relative change in information processing costs under stable preferences can be identified (see Table 14). We simulate 2,000 choices for one individual, with one inside and one outside option. The first half of the 2,000 choices are made in an “easy” environment with lower information processing costs λlow, and the second half in a more difficult environment subject to higher information processing costs λhigh. The inside good consists of three attributes: two simple and one complex. One simple attribute is a brand intercept, and the remaining attributes have the following levels: simple price p∈{2.5,3,3.5,4,4.5} and complex discount d∈{0,0.5,1,1.5,2}. The preference vector is given by β=(βb,βp,β d)=(2.5,−1,1)and information processing costs are λlow=0.1 and λhigh =0.3. In estimation, we fix λhigh to the data generating value and jointly estimate λlowand β. The third row in Table 14 shows that we can recover the datagenerating parameters, and the fifth row illustrates the bias from ignoring the difference in processing costs, as well as the poorer fit from doing so. A.2 Simulation study The simulation study builds on the set-up used for illustrative simulations in Section 3.3 of the paper. We vary the number of inside goods (one versus two), simulate Table 14 Posterior means of preferences and information processing costs Model Brand Price Discount λlowλhigh LMD Data generation 2.50 -1.00 ≡−βp0.10 0.30 RI-DCM 2.59 -1.03 ≡−βp0.10 0.30 -253.38 (0.10) (0.04) (0.01) RI-DCM with fixed λ2.28 -0.91 ≡−βp0.20 0.20 -298.98 (0.11) (0.04) In the first model, λhigh is fixed to data generating value in estimation. For the second model, the information processing costs are fixed to 0.2 123
94 S. Turlo et al. Table 15 Means of posterior means and variances of preference distributions for different model specifications using 50 simulations with one inside and one outside good each and homogenous information processing costs Posterior Means Model Fits Model Brand Price Discount LMD LMD Data generation 2.50 -1.00 ≡−βp RI-DCM 2.51 -1.00 ≡−βp-2,465.49 (0.02) (0.01) (63.88) RU logit separate 26.41 -8.71 4.41 -2,630.29 162.17 (0.84) (0.32) (0.19) (75.46) (26.47) RU logit joint 9.94 -4.01 ≡−βp-4,602.05 2018.19 (0.14) (0.09) (236.44) (50.34) BC separate 27.43 -9.08 4.51 -2,608.36 136.66 (0.89) (0.51) (0.25) (73.41) (25.72) BC joint 11.48 -4.13 ≡−βp-3,160.25 685.37 (0.99) (0.16) (82.37) (49.87) Posterior Variances Model Brand Price Discount Data generation 0.25 0.04 ≡Var(βp) RI-DCM 0.29 0.05 ≡Var(βp) (0.04) (0.01) RU logit separate 14.01 1.82 1.12 (0.84) (0.32) (0.19) RU logit joint 5.31 0.70 ≡Var(βp) (1.01) (0.30) BC separate 20.11 1.73 1.57 (1.06) (0.57) (0.45) BC joint 9.73 0.44 ≡Var(βp) (0.83) (0.20) Standard deviations are in parentheses. For the BC separate and BC joint model, the means of posterior means for the price threshold are 10.91 and 4.67 with standard deviations of 5.02 and 0.81, respectively. The means of the posterior variances for the price screening thresholds are 8.41 for the BC separate and 1.01 for the BC joint model data with and without heterogeneity in processing costs (in addition to preferences heterogeneity), and add a two-stage choice model with consideration sets from screening on price (see e.g., Gilbride & Allenby, 2004; Pachali et al., 2023) to the model comparison. We use the following (standard) weakly informative subjective prior parameter settings for the one and the two-inside good cases, respectively: {¯ β∼N(0,100I), Vβ∼IW(4,1.5I)}and {¯ β∼N(0,100I),Vβ∼IW(5,2I)}. The slightly more informative subjective prior for the hierarchical prior variance of {βi},Vβis owed to the increased dimensionality of the estimation problem. When we estimate the hierarchical prior variance of {log(λi)},weuseσ2 log(λ) ∼IG(3,1). 123
Discrete choice in marketing through... 95 We simulate 50 data sets in the four data-generating settings and estimate five models for each: the RI-DCM, two RU logit models, and two models with price-based consideration sets (BC separate and BC joint). The benchmark models are estimated once with a single price parameter and once with separate parameters for the (simple) price and the (complex) discount. Each data replication in this simulation study resamples N×Tchoice sets from the design base defined by prices pand discounts d, drawn uniformly and independently from the following sets: p∈{2.5,3,3.5,4,4.5} and d∈{0,0.5,1,1.5,2}and re-samples preferences (and processing costs when Table 16 Means of posterior means and variances of preference distributions for different model specifications over 50 simulations with two inside goods and one outside good each and homogenous information processing costs Posterior Means Model Fits Model Brand 1 Brand 2 Price Discount LMD LMD Data generation 2.50 2.50 -1.00 ≡−βp RI-DCM 2.54 2.53 -1.03 ≡−βp-3,807.55 (0.05) (0.05) (0.03) (56.43) RU logit separate 24.89 24.55 -8.06 4.33 -4,160.46 272.48 (0.76) (0.69) (0.31) (0.19) (102.34) (54.03) RU logit joint 10.01 9.97 -3.99 ≡−βp-6,838.41 2940.38 (0.25) (0.24) (0.05) (107.61) (60.37) BC separate 25.19 24.91 -8.41 4.44 -4,196.47 228.54 (0.69) (0.62) (0.33) (0.10) (94.31) (48.01) BC joint 11.15 11.67 -4.01 ≡−βp-5,431.81 1549.33 (0.29) (0.37) (0.09) (81.69) (50.67) Posterior Variances Model Brand 1 Brand 2 Price Discount Data generation 0.25 0.25 0.04 ≡Var(βp) RI-DCM 0.28 0.30 0.05 ≡Var(βp) (0.05) (0.06) (0.01) RU logit separate 12.01 11.19 1.97 1.22 (0.84) (0.72) (0.21) (0.11) RU logit joint 6.31 6.23 0.60 ≡Var(βp) (0.41) (0.44) (0.17) BC separate 17.41 16.33 1.70 1.27 (3.16) (3.55) (0.39) (0.21) BC joint 8.73 8.56 1.04 ≡Var(βp) (0.83) (0.81) (0.20) Standard deviations are in parentheses. For the BC separate and BC joint model, the means of posterior means for the price threshold are 12.01 and 4.26 with standard deviations 5.04 and 0.71, respectively. The means of the posterior variances for the price screening thresholds are 10.39 for the BC separate and 0.91 for the BC joint model 123
96 S. Turlo et al. Table 17 Means of posterior means and variances of preference distributions for different model specifications over 50 simulations with one inside and one outside good each and heterogeneous information processing costs Posterior Means Model Fits Model Brand Price Discount LMD LMD Data generation 2.50 -1.00 ≡−βp RI-DCM scale mixture 2.51 -1.01 ≡−βp-2,444.54 (0.04) (0.03) (53.41) RU logit separate 26.97 -8.92 4.43 -2,616.04 185.58 (1.51) (0.47) (0.10) (62.66) (17.39) RU logit joint 10.07 -4.02 ≡−βp-4,530.19 2083.55 (0.69) (0.21) (95.38) (61.37) BC separate 27.50 -8.84 4.23 -2,603.48 156.47 (2.31) (1.04) (0.68) (74.31) (21.71) BC joint 11.31 -4.40 ≡−βp-3,171.05 663.49 (1.62) (0.60) (67.46) (82.94) Posterior Variances Model Brand Price Discount σ2 log(λ) Data generation 0.25 0.04 ≡Var(βp)0.09 RI-DCM scale mixture 0.30 0.05 ≡Var(βp)0.08 (0.05) (0.01) (0.02) RU logit separate 15.11 1.62 1.73 (0.94) (0.32) (0.29) RU logit joint 6.21 0.73 ≡Var(βp) (1.54) (0.14) BC separate 14.41 1.84 1.90 (2.13) (0.41) (0.45) BC joint 10.15 0.39 ≡Var(βp) (1.23) (0.13) Standard deviations are in parentheses. For the BC separate and BC joint model, the means of posterior means for the price threshold are 10.38 and 4.13 with standard deviations 5.26 and 0.80 respectively. The means of the posterior variances for the price screening thresholds are 13.94 for the BC separate and 0.89 for the BC joint model. σ2 log(λ) is the variance of log information processing costs in the population heterogeneous) from population parameters. In each replication, a (fresh) sample of rationally inattentive DMs (N=1,000) face T=20 choices sets. The utility of inside good ito DM jin choice task tis given by uj,i,t=βb,j,i+ βp,j(pj,i,t−dj,i,t)whereβb,j,iis the brand coefficient, βp,jtheprice coefficient, pj,i,t is the price, and dj,i,tis the discount. The utility of the outside option is normalized to zero: uO=0. As in the illustrative simulations in Section 3.3, brand and price are simple attributes, while the discount requires costly processing. When processing costs are homogenous, they are set to λ=0.25. When heterogeneous, they are generated from log(λj)∼N(−1.4,0.09). Preference coefficients are independently generated 123
Discrete choice in marketing through... 97 Table 18 Means of posterior means of preference distributions for different model specifications over 50 simulations with two inside goods and one outside good each and heterogeneous information processing costs Posterior Means Model Fits Model Brand 1 Brand 2 Price Discount LMD LMD Data generation 2.50 2.50 -1.00 ≡−βp RI-DCM scale mixture 2.52 2.52 -1.02 ≡−βp-3,800.39 (0.04) (0.04) (0.02) (56.51) RU logit separate 27.33 27.19 -7.34 4.71 -4,172.86 363.87 (4.56) (5.31) (0.97) (0.53) (75.37) (44.63) RU logit joint 10.94 10.88 -5.63 ≡−βp-6,873.34 3,014.79 (1.34) (1.57) (0.81) (79.31) (84.28) BC separate 26.39 26.01 -8.11 5.01 -4,114.34 305.67 (5.11) (4.89) (1.02) (0.64) (69.37) (50.31) BC joint 10.70 10.17 -5.21 ≡−βp-5,404.29 1,603.48 (1.31) (1.29) (0.93) (79.14) (85.34) Posterior Variances Model Brand 1 Brand 2 Price Discount σ2 log λ Data generation 0.25 0.25 0.04 ≡Var(βp)0.09 RI-DCM scale mixture 0.29 0.28 0.05 ≡Var(βp)0.09 (0.06) (0.06) (0.01) (0.01) RU logit separate 13.11 9.59 1.62 1.17 (0.83) (0.68) (0.25) (0.12) RU logit joint 6.95 6.77 0.68 ≡Var(βp) (0.48) (0.48) (0.19) BC separate 16.44 15.98 1.65 1.31 (3.50) (3.31) (0.40) (0.25) BC joint 9.02 8.74 0.64 ≡Var(βp) (0.88) (0.80) (0.23) Standard deviations are in parentheses. For the BC separate and BC joint model, the means of posterior means for the price threshold are 13.31 and 4.43 with standard deviations 6.84 and 0.83 respectively. The means of the posterior variances for the price screening thresholds are 16.73 for the BC separate and 0.96 for the BC joint model. σ2 log(λ) is the variance of log information processing costs in the population from the following distributions: βb,j,i∼N(2.5,0.25)and βp,j∼N(−1,0.04), i.e., the two brands are symmetric in the population in the case of two inside brands. Tables 15,16,17, and 18 report distributions of preference estimates and model fit across data replications in our four data generating settings for each of the five models fit. The last column in the upper tables, LMD, shows log-marginal density differences relative to the RI-DCM across simulations. Positive values indicate that the RI-DCM achieves better model fit. Not surprisingly, we find that only the RI-DCM recovers data-generating parameters. We also find that (i) we can reliably distinguish the data generating RI-DCMs from all the benchmark models (see columns LMD and LMD in the respective tables), (ii) 123
98 S. Turlo et al. Table 19 Quartiles as well as the minimum and the maximum of the log-likelihood MCMC draws for different λlevels for the correct and wrong specification respectively Model min 25% 50% 75% max λ=0.01 Correct - Linear -1.42 0.00 0.00 0.00 0.00 Misspecified - Linear -1.86 0.00 0.00 0.00 0.00 λ=5 Correct - Linear -4.85 -3.83×10−7-9.82×10−9-1.71×10−10 0.00 Wrong - Linear -1.45 -3.73×10−7-1.07×10−8-2.36×10−10 0.00 the approximating benchmark models fare relatively much worse in the larger choice set because of the corresponding increase in the number of optimal RI information strategies39 (comparing LMD in Tables 15 and 17 to that in Tables 16 and 18, respectively), and (iii), slightly worse when processing costs are heterogeneous, in addition to preferences heterogeneity (comparing LMD in Tables 15 and 16 to that in Table 17 and 18, respectively). Finally, we can see that benchmark models substantially benefit from including separate coefficients for price and discount and that modeling screening based on price further improves the fit of benchmark models. However, even the combination of separate coefficients for price and discount with price screening fits reliably worse than the RI-DCM. In the case of one inside good only, what is missing from this approximation is the complex interaction between price and discount implied by optimal processing under RI. In the case of two inside goods, consideration of one inside good also depends on simple features of the other inside good. Similarly, the processed contribution of one brand’s discount depends on the simple attributes of this brand and that of the other brand. A.3 Identification in the case of extreme information processing costs Distinction of simple and complex utility aspects of a choice task Table 19 illustrates that the distinction between simple and complex attributes becomes mute once the data become (essentially) deterministic at very low or very high processing costs λ.We again simulate 2,000 choice tasks with the same design as outlined in Section 3.5 but now with information processing costs λ=0.01, making processing complex information essentially free, and λ=5, making information processing infeasible. When λis sufficiently small, all information is fully processed, and the conceptual and 39 Recall from expression Eq. 2that optimal information strategies are conditional on realized levels of simple attributes. With one inside brand and five levels of simple price, there are five different optimal processing strategies, depending on the realized simple price in a choice set. With two inside brands and five levels of simple price, there already are 52=25 different configurations of simple attributes, giving rise to different optimal processing strategies. 123
Discrete choice in marketing through... 99 Fig. 15 Histogram of MCMC draws of the complex attribute parameter under different information processing costs λ. This Figure shows histograms of the posterior distribution of complex attribute coefficient based on 100,000 draws from the implied marginal posterior for correctly specified simple and complex attributes. The solid vertical line indicates the data-generating value empirical distinction between simple and complex vanishes. When λis sufficiently large, the information in complex attributes is not integrated into the overall evaluation of alternatives. If, in this case, the analyst falsely specifies complex attributes as simple and simple attributes as complex, the estimator will infer extreme utility parameters for the latter relative to the former such that deterministic choice based on simple attributes (falsely assumed to be complex) ensues. Set identification of preferences Figure 15 illustrates that utility coefficients are only set-identified once choices become (essentially) deterministic, using the example of the complex attribute coefficient in the correctly specified model. Acknowledgements We thank Ana Martinovici, Dan Bartels, Xiaolin Li, Tamer Boyaci, Robert Zeithammer, Daniel Ackerberg, Jean-Pierre Dubé, Jeremy Fox, Wes Hartmann, Günter Hitsch, Carl Mela, and Stephan Seiler for helpful comments and suggestions. We also benefitted from the comments of seminar participants at the 17th Annual Bass FORMS Conference in Dallas, EMAC 2022 in Budapest, the INFORMS Marketing Science Conference 2022, as well as the Kleinwalsertal Seminars in 2020-2022. We gratefully acknowledge financial support by the German Research Foundation (DFG Grant ID: OT 447/21). Arash Laghaie acknowledges funding by Fundação para a CiênciaeaTecnologia(UIDB/00124/2020, UIDP/00124/2020, UID/00124, Nova School of Business and Economics and Social Sciences DataLab - PINFRA/22209/2016), POR Lisboa and POR Norte (Social Sciences DataLab, PINFRA/22209/2016). 123
100 S. Turlo et al. Funding Open Access funding enabled and organized by Projekt DEAL. Data Availability Code for estimation and replication of tables and figures in the paper are available at: https://gitfront.io/r/user-3843061/SqrbQZ5sXtbH/Rational-Inattention-Discrete-Choice/ Declarations Conflicts of Interest This work was supported by the German Research Foundation (Grant ID: OT 447/2-1). The authors have no further competing interests to declare that are relevant to the content of this article. 123 Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. References Abaluck, J., Compiani, G., & Zhang, F. (2022). A method to estimate discrete choice models that is robust to consumer search. Unpublished manuscript, Yale School of Management, University of Chicago, and University of California. Anderson, N. H. (1981). Foundations of Information Integration Theory. Academic Press. Anderson, N. H. (1982). Methods of Information Integration Theory. Academic Press. Andrews, R. L., & Srinivasan, T. (1995). Studying consideration effects in empirical choice models using scanner panel data. Journal of Marketing Research, 32(1), 30–41. Aribarg, A., Otter, T., Zantedeschi, D., Allenby, G. M., Bentley, T., Curry, D. J., Dotson, M., Henderson, T., Honka, E., Kohli, R., et al. (2018). Advancing non-compensatory choice models in marketing. Customer Needs and Solutions, 5(1), 82–92. Armenter, R., Müller-Itten, M., & Stangebye, Z. R. (2024). Geometric methods for finite rational inattention. Quantitative Economics, 15(1), 115–144. Armona, L., Fuster, A., & Zafar, B. (2019). Home price expectations and behaviour: Evidence from a randomized information experiment. The Review of Economic Studies, 86(4), 1371–1410. Bertoli, S., Moraga, J.F.-H., & Guichard, L. (2020). Rational inattention and migration decisions. Journal of International Economics, 126, Bestard, A. B., & Font, A. R. (2021). Attribute range effects: Preference anomaly or unexplained variance? Journal of Choice Modelling, 41, 100321. Bhattacharyya, A. (1946). On a measure of divergence between two multinomial populations. Sankhy¯a: The Indian Journal of Statistics, pp. 401–406. Blattberg, R. C., Briesch, R., & Fox, E. J. (1995). How promotions work. Marketing Science, 14(3), G122– G132. Blattberg, R., & Neslin, S. (1989). Sales promotion: The long and the short of it. Marketing Letters, 1, 81–97. Boyacı, T., & Akçay, Y. (2017). Pricing when customers have limited attention. Management Science, 64(7), 2995–3014. Bronnenberg, B. J., & Vanhonacker, W. R. (1996). Limited choice sets, local price response and implied measures of price competition. Journal of Marketing Research, pp. 163–173. Brown, Z. Y., & Jeon, J. (2024). Endogenous information and simplifying insurance choice. Econometrica, 92(3), 881–911. Cao, X., & Zhang, J. (2021). Preference learning and demand forecast. Marketing Science, 40(1), 62–79. Caplin, A., & Dean, M. (2013). Rational inattention and state dependent stochastic choice. Unpublished manuscript, New York University.
