scieee AI-readable full text Open interactive document viewer

Differential attention to attributes in utility-theoretic choice models

Cameron, Trudy Ann,DeShazo, J. R.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Cameron, Trudy Ann; DeShazo, J. R. Article Differential attention to attributes in utility-theoretic choice models Journal of Choice Modelling Provided in Cooperation with: Journal of Choice Modelling Suggested Citation: Cameron, Trudy Ann; DeShazo, J. R. (2010) : Differential attention to attributes in utility-theoretic choice models, Journal of Choice Modelling, ISSN 1755-5345, University of Leeds, Institute for Transport Studies, Leeds, Vol. 3, Iss. 3, pp. 73-115 This Version is available at: https://hdl.handle.net/10419/66812 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc/2.0/uk/ Journal of Choice Modelling, 3(3), pp 73-115 www.jocm.org.uk Differential Attention to Attributes in Utility-Theoretic Choice Models Trudy Ann Cameron1,* J. R. DeShazo2,† 1Department of Economics, 435 PLC, 1285 University of Oregon, Eugene, OR, 97403-1285 USA 2Department of Public Policy and Institute of the Environment, 3250 School of Public Affairs, University of California, Los Angeles, CA 90095-1656 USA Received 20 October 2008, revised version received 10 March 2010, accepted 1 November 2010 Abstract We show in a theoretical model that the benefit from additional attention to the marginal attribute within a choice set depends upon the expected utility loss from making a suboptimal choice if it is ignored. Guided by this analysis, we then develop an empirical method to measure an individual’s propensity to attend to attributes. As a proof of concept, we offer an empirical example of our method using a conjoint analysis of demand for programs to reduce health risks. Our results suggest that respondents differentially allocate attention across attributes as a function of the mix of attribute levels in a choice set. This behaviour can cause researchers who fail to model attention allocation to estimate incorrectly the marginal utilities derived from selected attributes. This illustrative example is a first attempt to implement an attention-corrected choice model with a sample of field data from a conjoint choice experiment. Keywords: Attention to Attributes, Allocation of Attention, Conjoint Choice, Choice Set Design, Bounded Rationality, Choice Heuristics * Corresponding author, T: +01-5413461242, F: + 01-5413461243, [email protected] †T: +01-3105931198, F: + 01-3102060337, [email protected] Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 1 Introduction Simple empirical choice models assume that the investigator knows exactly what information the individual uses to make a given choice—i.e. that the individual fully attends to, and costlessly processes, all the information available within a choice scenario. Economists, and choice modellers more generally, now recognize that the constituent elements of attention, including cognition and time, are scarce resources which rational individuals should allocate optimally (Simon 1955; March 1978; Heiner 1983, 1985; de Palma et al. 1994; Conlisk 1996; Gabiax and Laibson 2000). The optimal allocation of attention across attributes of alternatives will depend upon both the expected marginal benefits and marginal costs of further information processing. As a consequence, prior to making a choice, the individual may rationally attend to some attributes of the alternatives more than others. However, this process of optimally allocating attention over the array of information in a choice set may create a profound problem for discrete choice researchers. Suppose an individual does value the level of a particular attribute, but because of some resource constraint, overlooks differences in its levels across alternatives when making their choice. A random utility empirical model, based on perfect and costless information, will imply that the marginal utility associated with that attribute is zero. Similarly, incomplete attention to any particular attribute, could lead to biased estimates of the effect of variations in the level of that attribute on the individual’s choice. In particular, the apparent marginal utility in this case would be an attenuated estimate of the true marginal utility under perfect and costless information. Further complications may arise if the individual’s allocation of attention differs systematically across particular types of attributes. For example, consider the case of estimating demand for goods in economic applications. In revealed preference contexts, some individuals might allocate a disproportionately greater level of attention to prices as opposed to other attributes. This might raise the estimated relative marginal utility of net income and thereby lower the estimated willingness to pay (WTP). In stated preference contexts, however, researchers are often worried that people will instead pay too little attention to an alternative’s price, thus lowering the estimated relative marginal utility of net income and artificially inflating estimated WTP. In this paper, we derive results based on pre-choice optimization behaviour which lead to guidelines for empirical specifications. We develop a theoretical model that motivates our methodological approach, which we then illustrate with an empirical example. We argue that the benefits from additional attention allocated to the evaluation of an incremental attribute stem from the expected value of the avoided lost utility associated with a suboptimal choice (made as a result of ignoring that incremental attribute). The expected magnitude of this utility loss depends upon two components. The first component is the “other-attribute utility dissimilarity.” This component represents how close the alternatives are in utility space—given the other attributes evaluated thus far. The second component is the “own-attribute utility dissimilarity.” This component captures how much of a difference it might make, to overall utility from each alternative, if the incremental attribute is taken into account. Our theoretical model leads us to develop a practical implementation of its insights, so that empirical choice specifications can accommodate the individual’s “propensity to attend” to each attribute in a choice scenario. Conceptually, this propensity to attend to an incremental attribute is identified based on individual-specific measures of otherand ownattribute utility dissimilarities as well as the cognitive costs of attribute evaluation. We introduce a multiplicative propensity-to-attend parameter for each attribute which can be 74 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 viewed equivalently as affecting either the apparent marginal utility associated with the marginal attribute, or the perceived difference in the level of this attribute across the two alternatives. As a proof of concept, we offer an empirical example of our method using a conjoint analysis of demand for programs to reduce health risks. Our results suggest that the combination of other-attribute dissimilarity and own-attribute dissimilarity causes respondents to differentially allocate attention across attributes. More specifically, this process appears to result in a tendency for the researcher to overestimate the marginal utility derived from net income, overestimate the marginal disutility of sick-years and lost-lifeyears, but perhaps to underestimate the marginal disutility of recovered-years, on average. Although there is certainly a great deal of room to expand the theoretical scope of our model, the empirical version of this simple attention-corrected model is important for several reasons. First, the attention-corrected model has the potential to identify and eliminate distortions in the estimated parameters of discrete choice models which are estimated using either revealed or stated preference data. Second, it provides a clear measurement framework to test emerging hypotheses from the behavioural literature about the determinants of an individual’s allocation of attention. Third, when the prospective real consequences of a choice vary from context to context, we might also expect the individual’s budgeted attention to vary as well. Our attention-corrected model might also be used to explain some types of observed differences across data generation methods, such as revealed preference (RP) versus stated preference (SP) information. Fourth, our model opens up the possibility of beginning to measure the effectiveness of agents (such as marketers, salespeople, politicians, etc.) who strategically seek to direct the individual’s attention towards some attributes and away from others as they design choice sets. The extent, and implications, of such strategic behaviour on choice outcomes, and therefore upon individual welfare, cannot be assessed adequately without the benefit of a framework with features similar to those of the model presented here. The scope of the theoretical model is modest; it emphasizes the role of expected marginal benefits in determining the allocation of attention, although we are careful to outline how costs could also be incorporated. An illustration that includes the role of cost information awaits richer data. The present paper also focuses only on the allocation of attention to the marginal attribute. We leave a similar analysis of optimal attention to the marginal alternative for a subsequent paper.1 1.1 Related literature Economists have long recognized that individuals face various resource constraints as they acquire and deploy information in their decision-making processes (Simon 1955; March 1978; Heiner 1983, 1985; de Palma, et al. 1994; Conlisk 1996). Theoretical models that predict how individuals might optimally respond to such constraints have only recently begun to emerge (Gabiax and Laibson 2000) and be tested in experimental settings (Gabaix et al. 2006). DellaVigna (2009) reviews the state of the literature on psychology and economics and inventories a small literature wherein inattention is assumed to vary inversely with the salience of “opaque” information and directly with the number of competing stimuli, but this literature appears to emphasize the identification of types of opaque information to which decision makers are not fully attentive. While many advances 1 The notion of individual-specific “consideration sets” of relevant alternatives has been addressed in Haab and Hicks (1999), Chakravarti and Janiszewski (2003), Paulssen and Bagozzi (2005) and Jedidi and Kohli (2005). 75 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 have recently been introduced that inform the design of attribute-based field studies of demand, no general methods exist for directly modelling the effects of the allocation of attention, within a choice task, on the estimated marginal utilities (or part-worths) for these attributes.2 In developing a directed cognition model, Gabaix and Laibson (2000; 2005) approximate the economic value of additional attention using two ideas stemming from the analysis of option values. First, the option value of continued consideration declines as one alternative gains a large edge over other available alternatives. In the context of the present paper, this corresponds to the case where alternatives come to be perceived as less similar in terms of the utility they generate (i.e. when one alternative clearly dominates the other(s) in terms of the current information set). Second, the option value also declines when continued consideration yields little new information. In the context of the present paper, this corresponds to the case where additional attributes are more similar across alternatives or as units of these additional attributes provide minimal marginal utility. Overall, the experimental results in Gabaix et al. (2006) are consistent with two implications of their option value framework, which defines how many search operations the subject should pursue. Translating their discussion into the terminology of our paper, the value of [attribute] exploration [for a given alternative] decreases the larger the gap between the active [alternative] and the next best [alternative]. Second, the value of [attribute] exploration increases with the variability of the information that will be obtained (p. 1053). When extended to their yoked case, where an additional attribute is revealed simultaneously for all alternatives, these insights appear to be essentially equivalent to the ones we derive analytically in this paper, in the context of a conventional empirical random utility choice model. The Gabaix et al. (2006) experimental design, couched exclusively in monetary amounts, eliminates the need to measure physical quantities of an attribute, to estimate the marginal utility associated with that attribute, or to infer the marginal WTP for units of each attribute by considering marginal rates of substitution between that particular attribute and money. Money-denominated attributes cannot differ in their salience across individuals, since utility is implicitly considered to map directly into the total number of cents paid. However, real choice situations in field settings are confounded by heterogeneity in marginal utilities across attributes, differences in attribute metrics, as well as differences in individuals’ cognitive abilities and the opportunity costs of their time.3 Many advances have recently occurred with respect to the design of attribute-based field studies of demand but they fall short of the goal of modelling the allocation of attention.4 Hensher and co-authors have initiated several intriguing explorations of attention within a standard multivariate discrete choice setting. Hensher et al. (2005) use a specific follow-up question about which attributes the respondent did not use in making his or her choices. Hensher et al. (2007) also uses the same follow-up question to identify nine distinct attribute processing rules in the same. Respondent adherence to these rules is modelled as stochastic. The authors then use a modified mixed logit model which conditions each 2 However, Swait and Adamowicz (2001b) take to task the empirical choice modelling community for its persistence in assuming a “utility-maximizing, omniscient, indefatigable consumer.” 3 The choice experiments in Fischer et al. (2000b) use the same Mouselab software as in Gabaix et al. (2006), but the choice data from their study is not utilized in an econometric random utility model. 4 Psychologists have explored how various within-choice-set conditions may increase the cognitive costs of attribute evaluation and comparison (Bettman et al. 1993, 1998; Dellaert et al. 1999; Fisher et al. 2000a,b; Luce et al. 2003; Johnson 2008). Similarly, marketing scholars and others have explored the effects of task complexity on choice outcomes (Shugan 1980; Malhotra 1982; Mazzotta and Opaluch 1995). 76 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 parameter on whether a respondent included or excluded an attribute in their information processing strategy. In their conclusions, these authors acknowledge that there may be differences “between what people say they think and what they really think” (p. 216), and they question whether the “simply conscious statements” made by survey respondents represent an adequate measure of information processing. They emphasize that individuals’ information processing strategies “should be built into the estimation of choice data from stated choice studies” (p. 214). This is precisely what we endeavour to accomplish in the present paper. Employing a similar research method, Hensher et al. (2006a) find that that probability of a respondent considering more attributes decreases as the attributes used in their survey are drawn from distributions with narrower ranges. However, attribute ranges in this study appear to be varied simultaneously across all attributes and the dependent variable is available only at the level of the individual, not the choice set. In contrast, our models lead us to consider differences in the ranges of attributes within a single choice set as additional potential determinants of attention, and therefore of apparent marginal utilities and ultimately our estimates of WTP. Very relevant to our study is Hensher’s finding that individuals’ processing strategies depend on the nature of the attribute information in the choice set, not just the quantity of such information (i.e. the number of attributes). Another related study is Puckett and Hensher (2008). This study seeks to integrate the attribute processing strategies (APSs) reported by respondents into the analysis of choices. The relevant APSs involve certain attributes being ignored, or aggregated to a different extent, by different respondents. This work builds on Hensher et al. (2006a) in that it considers the effects of APSs utilized by respondents for every alternative in every choice set, including across choice tasks faced by a given respondent. This approach can accommodate cases where attribute level mixes are outside of the acceptable choice bounds for the individual. The wording of their debriefing question for each choice was: “Is any of the information shown not relevant when you make your choice? If an attribute did not matter to your decision, please click on the label of the attribute below. If any particular attributes for a given alternative did not matter to your decision, please click on the specific attribute.” Subjective all-or-nothing attention to different attributes is thus elicited directly from each respondent, rather than being inferred from choice behavior. Finally, there is large literature that explores how the design of choice sets affects individuals’ choice consistency and willingness to pay. Early versions of these include Mazzotta and Opaluch (1995) and DeShazo and Fermo (2002), leading up to the very ambitious “design-of-designs” studies by Hensher (2006b). Much of this work, however, is motivated by a concern with how the cognitive costs of information processing vary with choice set design. Our focus in the present paper concerns an optimizing model for how the expected benefit from additional information drives the allocation of attention across attributes. Several of these other papers emphasize how, through deliberate manipulation of choice set design, the researcher can alter the estimated parameters. We contend that even if all of the survey instruments in a study employ the identical choice set design (in terms of numbers of attributes, alternatives, choice sets, attribute levels, and ranges) there can still be artefacts of the researcher’s design decisions—with respect to the mix of attributes in any given choice set—that can unintentionally or intentionally affect the recovered utility parameter estimates. Furthermore, these effects can vary across individuals. 77 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 2 A Theoretical Model for Attention to Attributes Suppose subjects in a stated preference (SP) choice experiment (or in revealed preference (RP) choice data) actually do care about the level of a particular attribute. However, for some reason, they fail to evaluate its levels across alternatives when making their choices. In this situation, a random utility empirical model, based on perfect and costless information, will imply that the marginal utility associated with that overlooked attribute is zero. Likewise, simply incomplete attention to any particular attribute, as opposed to zero attention, could be expected to result in a lesser-than-expected effect of variations in the level of that attribute on people’s choices. The apparent marginal utility in this case would be an attenuated estimate of the “true” marginal utility under perfect and costless information. If the subject’s cognitive resources are limited, but this “inattention effect” is uniform across attributes, then perhaps all of the indirect utility parameters may be proportionally attenuated. This is observationally equivalent to the case where the “scale factor” in a discrete-choice model is smaller (i.e. the error variance is larger). If the propensity to attend to attributes is scaled down equally for all attributes (including net income) when an individual is paying less attention, we would expect to see no bias created in the implied point estimate of marginal WTP for that attribute. The ratio of the marginal utility associated with any attribute, relative to the marginal utility of income, is typically all that matters in simple models.5 However, it is possible that inattention to attributes differs across attributes. In particular, when decision resources are limited, there may remain a disproportionate level of attention devoted to an alternative’s cost as opposed to other attributes. In this case, distortions in WTP could be expected. Relatively less attenuation in the estimated marginal utility of income will inflate the denominator relative to the numerator in the usual WTP calculation, so that the implied WTP for every alternative could be biased toward zero. In contrast, in the context of SP models, there is great concern that because of the hypothetical nature of the choices involved, the subject will fail to pay sufficient attention to the cost variable. The emphasis on providing a “cheap talk” script as part of the survey is designed to draw the respondent’s attention specifically to the cost variable and its implications. Disproportionately greater attention to the implications of the cost variable may serve to amplify attention to this attribute relative to other attributes when the expected utility loss is otherwise rather low because of the less directly consequential nature of many SP choices. If other attributes of the offered alternatives are not similarly emphasized, the practice of offering only a cheap talk script—generally intended to increase the apparent marginal utility of income—can be expected to produce a downward bias in estimated willingness to pay (relative to a scenario that worked equally hard to draw respondents’ scarce attention toward all attributes).6 In this paper, we derive some results based on optimization behaviour which lead to guidelines for empirical specifications. Our models concern how the individual’s optimal 5 Conlon et al. (2001) use response time as a proxy for consumer effort devoted to a choice, where effort is regressed on choice set characteristics and involvement measures. One choice set characteristic is the expected utility difference across alternatives, based on a preliminary multinomial logit choice model. 6 An anonymous referee has suggested that when attention to attributes is not scaled back proportionally for all attributes, the effect on a conventional choice model might be analogous to the introduction of an attribute-specific error term. However, we concentrate upon a possible structural interpretation involving the marginal benefits of attention, rather than relegation of this phenomenon to the stochastic structure of the model. 78 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 amount of attention to a particular attribute might be determined by the nature of the decision context and how this context interacts with the preferences of that individual. 2.1 A Two-Alternative Case Consider first a familiar binary choice model (with alternatives indexed by 0 and 1) where the underlying indirect utility function is linear and additively separable in net income (i.e. () j ii YT− , 1 ki Xk= = income minus the cost of option j) as well as several other attributes, . ,..., K () () 11 1 12 00 0 12 K iii kki k K iii kki k VYT X VYT X 1 0 i i ββ ε ββ ε = = =−+ + =−+ + ∑ ∑ (1) The utility-difference expression driving the choice between alternatives 1 and 0 can thus be written as: () ( )( 10 01 1 0 10 12 12 K ii ii kkiki ii k K ikkii k VV TT X X tx ) ββ εε ββε = = −= − + − +− =− + + ∑ ∑ (2) where is often distinguished from the other attribute differences because of the role of its coefficient, ()( ) 01 1 0 11 1ii i i i TT X X x t−=−==− i 1 β , in the calculation of WTP. In general, lower-case variable names will be used to denote differences in attribute levels between alternative 1 and alternative 0 (i.e. “net” levels of attributes in this two-alternative context). In this simple linear specification, WTP is calculated by setting the utility difference to zero and solving for the level of which creates this indifference between alternatives. The implied WTP and E[WTP] functions are * i t () [] 2 * 1 2 11 K kki i k ii K kki ki i x WTP t x EWTP E E βε β βε ββ = = + == ⎡⎤ ⎡ ⎤ ⎢⎥ =+ ⎢ ⎥ ⎢⎥ ⎣ ⎦ ⎣⎦ ∑ ∑ (3) 2.1.1 Marginal Benefit of Attention to an Additional Attribute The marginal benefit to the individual of paying attention to an additional attribute can be equated to the avoided expected utility loss from making an incorrect choice as a result of ignoring in the choice process. There are two ways that the subject can experience a loss from failing to consider . First, she might choose alternative 1, when in fact (i.e. on the basis of the full set of attributes) her utility would actually be higher under alternative 0. Or, she might choose alternative 0 when her utility would actually be higher under alternative 1, k X k X 79 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 if all attributes were taken into account. Her expected utility loss from failing to consider the level of will be: k X [ ] [ ] () [] () 01 10 1 | 0 Pr 0 |1 ii ii E U L chosen optimal V V chosen optimal V V =− − Pr + oss x (4) Let the full (true) utility-difference function be 10 1 ' ' K ii kkii k ii ki k ki k i VV x xx βε βε ββ ε = −− + =+ =+ ∑ −= (5) + 0 where each attribute-difference term 1 ki ki ki x XX=− concerns a single attribute k and its associated single indirect utility-difference coefficient, k β . In a linear and additively separable model, this coefficient will be the same as the marginal utility of . k X The second line of equation (5) illustrates our convention for referring to the complete inner product, ' i x β , of all attribute differences and their associated coefficients that actually enter into the systematic portion of the individual’s utility function. The third line of the equation shows how we decompose this inner product into two terms, one being the inner product of all attribute differences other than ki x and their corresponding parameters, denoted ' ki xk β −− , and the other being the attribute difference and its own coefficient, th k ki k x β . The probability that alternative 1 or alternative 0 is truly optimal for the individual (based on a full consideration of all attributes and their differences) would be given by () () '' ' Pr 1 Pr 0 Pr Pr 0 Pr ii ii ii al x x al x β εε β εβ ⎡⎤⎡ =+>=< ⎣⎦⎣ ⎡⎤ => ⎣⎦ ⎤ ⎦ (6) optim optim In contrast, if the individual completely ignores attribute k x (either because he or she does not think or bother to consider it, or believes incorrectly that it confers zero marginal utility), the probabilities of the observed choices will depend only on the levels of the other attributes, the vector k x −: () () '' ' Pr 1 0 Pr Pr 0 Pr ki k i i ki k ikik x x c n x Prchosen hose β εε β εβ −− −− −− ⎡⎤⎡ =+>=< ⎣⎦⎣ ⎡⎤ => ⎣⎦ ⎤ ⎦ (7) There are thus two ways for the individual to make a “mistake.” The subject could choose alternative 1 when alternative 0 is optimal, or choose 0 when 1 is optimal. 80 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 interpreted as “complete attention.”9 The third formulation treats the propensity to attend to each attribute as simply a non-negative factor that scales the true marginal utility associated with an attribute either up or down as the value of this factor is greater or less than one. The sign of the underlying true marginal utility, k β , is thereby preserved. This assumption may be the most empirically hospitable one when a sign restriction is desired.10,11 3.1.1 Implementation Our theory section has suggested specific information that should be included among the ki Z variables that determine the subject’s propensity to attend to the attribute of the alternatives in the choice set. These factors contribute to the expected benefits (or the costs) of paying attention to a particular attribute. The propensity-to-attend measure, , should be some explicit function of how different the alternatives are in terms of the utility based on all other attributes, which we will denote by the construct that measures this in the twoalternative case, th k ki a ' ki k x β −− . The propensity-to-attend measure will also depend on the difference across alternatives in utility derived from this attribute, denoted by ki xk β for the two-alternative case. Finally, it will depend on any available variables which capture the marginal cost of attention to this attribute (which may or may not differ across attributes k). Taking the first specification in equation (12) as our example, we now differentiate among the generic coefficients in ' ki ki k Z1a ⎡ ⎤ γ =+ ⎣ ⎦, by distinguishing three types of parameters: () '' 1 ki k ki k k ki k ki k axxC αδββθ −− =+ + + (13) 9 The slight inconvenience in estimating such a model stems from the starting values to be used. If all of the parameters k γ are simultaneously zero, then . The apparent marginal utilities from a naive random utility specification would therefore be obtained from the first model in equation (12) only if the starting values for the '0.5 ki ki k aFZ γ ⎡⎤ == ⎣⎦ k β parameters were all to be doubled. The default assumption (i.e. ) is therefore that the “true” marginal utilities are actually twice what they appear to be in the naive model. As the index 0 k= γ ' ki Zk γ is larger, the true marginal utility will be less than twice its apparent naive value; as the index is smaller, the true marginal utility will be more than twice its apparent naive value. In the limit, as Z' ki k γ goes to +∞, the implied propensity to attend goes to 1.0 and the associated k β corresponds to the true marginal utility. The counterfactual of interest in this model corresponds to the question of what would have been the marginal utilities if the subject had been paying full attention to all attributes. The answers are contained in the estimates of each k β from this specification. 10 Certainly, if a nonlinear model is used, and if analytical derivatives are to be employed, this formulation would be easier than the inverse log-odds transformation suggested for the case where marginal propensities are constrained to lie on the 0,1 interval. 11 Fortunately, the ratios of estimated marginal utilities are all that matter for welfare estimates, and the “true” marginal utilities can be known only up to a scale factor. The relevant counterfactual in this case again concerns what would be the size of the estimated marginal utilities if all attributes received equal attention. A logical value for this equal propensity to attend would be 1.0, which would also constitute a logical starting assumption, since if , which will be the case if the vector of parameters () ' exp 1 ki ki k aZ γ == '0 ki k Z γ = k γ is initially assumed to be a zero vector. 87 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 where the vector is a set of variables, when available, that capture the individual’s cognitive marginal costs of evaluating attribute k. If , then the propensity to attend, , equals exactly one for all attributes, the desired case. ' ki C 0 kkk αθδ === ki a We do not necessarily expect the expressions that capture the benefits of attention ( ' kkik kkik xx αβθ ) β −− + or the costs of attention to be the same across attributes, because the constructed variables ( ' ki k C δ ) ' ki k x β −− , ki xk β , and possibly the relevant vector of variables will differ across attributes. However, perhaps the incremental effects of the choice set design variables and individual characteristics that determine the net benefits of attention to attributes (i.e. the coefficients and for benefits, or ' ki C k α k θ k γ for costs) could be the same across attributes k=1,...,K, so that the corresponding coefficients can be constrained to be equal across attributes: ' 1( ) ki ki k ki k ki axxC βα ' β θ −− =+ + + δ ' (14) Where possible, one should estimate models with and without these restrictions and test whether the restrictions can be rejected. These types of restrictions are possible because all of the variables in question for and are in utility-units, not the units of the raw attributes. α θ If the cost-of-attention variables do not differ quantifiably across attributes, so that , it may actually be necessary to constrain . Otherwise, the attribute-specific parameters are likely to be difficult or impossible to identify separately from the marginal utilities, ' ki i CC=k δδ = k δ k β . If we assume that the effects of the cost variables are the same across all attributes—e specially if we adopt the version of in the third line of equation (12)— then can be readily factored into two components: ki a ki a () ' exp kkik kkik xx αβθ β −− + and , and the implied form of the indirect utility difference will be: ( ' exp ki C δ ) k () () () () 10 ' ' 1 '' 1 exp exp exp exp K i i k k ki k k ki k i ki i k K ikkkikkkikki k VV x x Cx Cxx i x β αβθβ δ δβαβθβ −− = −− = −= + + =+ ∑ ∑ ε ε + ) (15) Since is strictly positive and utility is invariant to the scale of measurement, we could divide through by to produce a heteroskedastic model: ( ' exp i C δ () ' exp i C δ () () 10 ' ' 1exp exp Ki i i k k ki k k ki k ki ki VV x x x C ε βαβθβ δ −− = −= + + ∑ (16) In the model with a strictly positive propensity to attend to attributes, the variables that capture cognitive cost differences across individuals can enter the model equivalently as factors that affect the dispersion in the conditional logit error term. Any variable that ' i C 88 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 89 increases the overall cost of attention should tend to decrease the respondent’s propensity to attend to every attribute. Lesser attention can be expected to increase the error variance in the model. 4 Empirical Example 4.1 A stated preference survey concerning morbidity/mortality risk reduction programs As a simple illustration, we use choice data from a stated preference survey concerning individuals’ preferences over health-risk reduction programs. Cameron and DeShazo (2009) use SP methods to elicit preferences for programs to reduce the risk of morbidity and mortality in a general-population sample of adults in the United States. The survey was fielded by Knowledge Networks, Inc., to their standing consumer panel, using a combination of internet and WebTV interfaces. The details of the survey and its basic analysis have been documented very extensively elsewhere, so we do not repeat the information here.12 In brief, the Cameron/DeShazo survey consists of five modules. We will outline these models only to explain the sources of some of the covariates we will use in our illustration. The first module asks respondents, among a variety of other questions, to rate their subjective risks, from low (-2) to high (+2), of contracting each of a range of major illnesses or injuries. The second module is a tutorial that explains the concept of an “illness profile.” This is a description of a sequence of future health states associated with a specified major illness or injury that the respondent may face over his or her remaining lifetime. An illness profile includes the years before the individual becomes sick (latency), illness-years while the individual is sick, recovered/post-illness-years after the individual more-or-less recovers from the illness, and lost life-years if the individual dies earlier than he would have in the absence of the illness or injury. After the tutorial about illness profiles, the individual is informed that he might be able to purchase new programs that would reduce his risk of experiencing certain illness profiles. Each illness-related risk-reduction program described in the survey consists of diagnostic blood tests, drug therapies, and life-style changes, and would be available at a specified annual cost, to be paid on a recurring basis as long as the individual is neither sick with this illness nor dead. The key module of each survey involves a set of five different three-alternative conjoint choice tasks where the individual is asked to choose one of two possible health-risk reducing programs or a status quo alternative. Each program reduces the individual’s risk of experiencing the corresponding illness profile. The illness profiles are described succinctly for each of the choice tasks—in terms of the baseline probability, age at onset, duration, and eventual outcome (recovery or death). Each corresponding risk reduction program is defined in terms of the extent to which it can be expected to reduce this risk, and its monthly and annual cost. Figure 1 provides one instance of the type of a stated choice scenario posed to respondents. 12 For more information on the survey instrument and the data, see the appendices which accompany Cameron and DeShazo (2009): Appendix A - Survey Design & Development, Appendix B - Stated Preference Quality Assurance and Quality Control Checks, Appendix C - Details of the Choice Set Design, Appendix D - The Knowledge Networks Panel and Sample Selection Corrections, Appendix E - Model, Estimation and Alternative Analyses, and Appendix F - Estimating Sample Codebook. Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 Choose the program that reduces the illness that you most want to avoid. But think carefully about whether the costs are too high for you. If both programs are too expensive, then choose Neither Program. If you choose “neither program”, remember that you could die early from a number of causes, including the ones described below. Program A for Diabetes Program B for Heart Attack Symptoms/ Treatment Get sick when 77 years old 6 weeks of hospitalization No surgery Moderate pain for 7 years Get sick when 67 years old No hospitalization No surgery Severe pain for a few hours Recovery/ Life expectancy Do not recover Die at 84 instead of 88 Do not recover Die suddenly at 67 instead of 88 Risk Reduction 10% From 10 in 1,000 to 9 in 1,000 10% From 40 in 1,000 to 36 in 1,000 Costs to you $12 per month [ = $144 per year] $17 per month [ = $204 per year] Your choice Reduce my chance of diabetes Reduce my chance of heart attack Neither Program Figure 1: An Example of a Choice Set Summary Table Module 4 contains debriefing questions to cross-check the internal consistency of responses. Module 5 is collected separately from our survey and contains detailed socio-demographic data for the individual and their household, as well as responses to a battery of health-related questions (including any illnesses the individual has already faced). Table 1 contains descriptive statistics for the empirical models to be used in this paper. It summarizes only those variables pertinent to the present illustration. These include raw data, but this information is processed before use, based on the economic theory of discounted expected utility, to yield the necessary constructed variables for our analysis to be discussed below. Finally, in some of our specifications, we allow for preferences to differ systematically with exogenous characteristics of the respondent (gender and age). We also allow preferences to differ according to the individual’s subjective risk rating for the illness/injury in question, with the average of their subjective risk ratings for all of the other major risk categories covered by our survey, and with any individual subjective adjustments of the surveys’ statements about the existence of likely benefits and the latency of the health risk. 90 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 Table 1: Descriptive Statistics for Variables Used in Estimating Specifications Variable Description Mean Std. Dev. Min Max Program attributes (14074 programs) - Raw illness/program attributes Cost Annual cost of program (paid when not sick or dead) 355.00 341.14 24 1680 AS i ΔΠ Risk change (i.e. negative, a risk reduction) -0.0034 0.0017 -0.006 -0.001 Latency Years until illness/injury begins 19.65 12.03 1 60 Sick years Duration of illness/injury (years) 6.53 7.21 0 52 Recovered years N umbe r of years in post-illness health state 1.62 4.62 0 55 Lost life-years N umber of life-years lost 10.87 10.32 0 55 - Constructed variables (income term) N et income under each alternative -0.052747 0.048772 -0.2513 0.1083 AS i ΔΠ log(pdvi+1) Term in present discounted sick-years -0.003111 0.003006 -0.01710 0 AS i ΔΠ log(pdvr+1) Term in present discounted recovered-years -0.003374 0.003189 -0.01711 0 AS i ΔΠ log(pdvl+1) Term in present discounted lost life-years -0.000746 0.001841 -0.01648 0 Sasubrsk (mean = msasubrsk) Same-illness subjective risk rating (-2 = low, 2=high) -0.2593 1.2531 -2 2 Cosubrsk (mean = mcosubrsk) Average subjective risk rating (other major health risks) -0.2537 0.8670 -2 2 (benefits never) =1 if expects never to benefit from this program 0.0759 0.2648 0 1 (min overest latency) Minimum overestimate of the latency of the health risk -7.483 11.98 -58 29 Respondent characteristics (1519 respondents) Income Annual income (dollars) 51048 33781 5000 150000 Female =1 if female 0.5135 0.5000 0 1 age (mean = mage) Age in years at time of response 50.11 15.18 25 93 91 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 4.2 Estimating specification for the naïve choice model We will use a simplified version of the theoretical model presented by Cameron and DeShazo (2009). In that paper, it is established that stated choices in this general population sample appear to be best predicted by a model that involves discounted expected utility from durations in different adverse future health states. Here, we will outline the model only briefly, to justify why the variables which explain choices are constructed in the ways that they are. The choice scenarios involve probabilistic sequences of future events, so a model based on discounted expected utility is about the simplest reasonable starting point. To understand the basic model, consider just a pair-wise choice between Program A and the status-quo alternative (N). Define the discount rate as r and let . For individual i, let be the probability of suffering a given adverse health profile (i.e. getting “sick”) if the status quo alternative is selected, and let be the (reduced) probability of suffering this adverse health profile if Program A is chosen. Thus is negative, since this is the risk reduction to be achieved by Program A. () 1t tr δ − =+ AS NS i Π AS i Π i Π AS NS ii ΔΠ = Π − The sequence of health states that makes up the illness profile to be addressed by Program A is captured by a set of mutually exclusive and exhaustive (0,1) indicator variables associated with each future time period, : for years in the latency period prior to any symptoms, 1for illness-years, for recovered, remission, or post-illness years, and 1for a year of premature mortality. Individuals are modelled as expecting to pay the annual cost of the risk reduction program only if they are neither sick nor dead. i T y ( 1A it pre illness ) 1re () A it ear lost ) ( A it illness life () A it covered The algebra of calculating present discounted expected utility differences is simplified considerably because we model health states as being uniform within specified intervals, as are income and program costs in this model. This feature allows us to discount health states first, and then take expectations. The present discounted number of years making up the remainder of the individual’s nominal life expectancy, , is given by . Other relevant discounted spells, also summed from t=1 to include: , i T1 i Tt it pdvc δ = =∑ A it pdve = i T () 1tA it pre illness δ ∑ () 1 At it it pdvi illness δ =A ∑ , () A it ecovered1 At δ it pdvr r= ∑ , and . () A it r lost1At it pdvl life yea δ =∑ Since the different health states exhaust the individual’s nominal life expectancy, i . Finally, to accommodate the assumption that each individual expects to pay program costs only during the pre-illness or recovered postillness periods, i is defined as the present discounted (healthy) time over which payments must be made. This can be interpreted as the expected discounted duration of program costs, with the expectation taken across whether or not the individual gets sick. AA AA iii i pdve pdvi pdvr pdvl pdvc+++= AA ii pdvp pdve pdvr=+ A A i To further simplify notation, let () 1 AAS AS iiii cterm pdvc pdvp ⎡ ⎤ =−Π +Π ⎣ ⎦ SAASA iii and let () 1 ANS N iiii y term pdvc pd ⎡ =−Π +Π vp pdvi ⎤ −ΔΠ ⎣⎦ . These two terms account for the 92 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 pattern of income net of program costs over time as a function of probabilistic health states. Then the expected utility-difference that drives the individual’s choice between Program A and the status quo can be defined as follows (where there will be an analogous term for the utility difference between Program B and the status quo in our three-alternative model): () () { } {}{}{} 1 23 4 AAAA iiiiii AA A A AA ii ii ii PDV E V Y c cterm Y yterm pdvi pdvr pdvl β A i βββ ε ⎡⎤ Δ=−+ ⎣⎦ +ΔΠ +ΔΠ +ΔΠ + (17) The four terms in braces can be constructed from the data, given specific assumptions about the discount rate.13 In this application, these constructed variables are the ki x (the differences in the attribute levels between each substantive alternative and the status quo). The empirical results described in Cameron and DeShazo (2009) suggest that a basic four-parameter, homogeneous-preferences model such as that in equation (17) is dominated by a specification that is not merely linear in the terms involving present discounted healthstate years. Factoring out the probability difference from the final substantive term in equation (17) gives: { } { } { } 23 4 23 4 j jjjj ii ii ii jj j j ii i i j p dvi pdvr pdvl pdvi pdvr pdvl βββ βββ ΔΠ + ΔΠ + ΔΠ ⎡⎤ =ΔΠ + + ⎣⎦ (18) where j=A,B,N, and for 0 N i pdvX =,, X irl= since the durations in adverse health states are all normalized to zero for the numeraire status quo alternative. However, this simple linear specification does not explain respondents’ observed choices as well as a model that employs shifted logarithms of the j i Xpdv terms: ()()( 23 4 log 1 log 1 log 1 jj j ii i i pdvi pdvr pdvl βββ ⎡⎤ ΔΠ + + + + + ⎣⎦ ) j (19) The basic discounted expected utility-difference specification that is presumed to drive respondent’s choices is therefore: 13 In this paper, we assume a common discount rate of five percent. In Cameron and DeShazo (2009), we explore the consequences of assuming either a three percent or a seven percent alternative discount rate. Work in progress involves the estimation of individual-specific discount rates simultaneously with these stated choices concerning health risk reduction programs, using additional data on inter-temporal choices by a separate sample of respondents from the same population. 93 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 () () { } () {} () { () {} 1 2 3 4 log 1 log 1 log 1 } j jj iiiii jj ii jj ii j i j jj ii PDV E V Y c cterm Y yterm pdvi pdvr pdvl β β β i β ε ⎡⎤ Δ=−+ ⎣⎦ +ΔΠ + +ΔΠ + +ΔΠ ++ jj ii x βε =+ (20) There is an analogous term for Program B in the three-way choice context. In the empirical estimates that follow, our “basic linear” model involves these four constructed variables in the sets of braces in equation (20), and their four estimated parameters, )( 1234 ,,, ββββ . We assume that a researcher who ignores the effects of scenario design (specifically, the mix of attribute levels presented in a choice set) would merely estimate this simple model. In Cameron and DeShazo (2009) we show how the estimated model can be used to build estimates of willingness to pay for a microrisk reduction (10-6) in the chance of suffering from a specific type of illness profile. This construct is essentially a generalization of the more-restrictive concept of the value of a statistical life (VSL) commonly employed in the mortality risk valuation literature. For this paper, however, we concentrate mainly on the estimation of the four parameters in equation (20) and the extent to which attention to these four different attributes may be biased as a result of the design of our choice sets. 4.3 Potential attention biases: measurement and control To construct measures for the similarity of alternatives based on all attributes other than the one in question, it is necessary to have measures of the “true” marginal utilities of each attribute, uncontaminated by attention biases. To identify these true marginal utilities, however, it is necessary to control for attention biases. Ideally, one would specify a conditional logit choice model where each additively separable marginal utility parameter is allowed to shift with variables which measure () ' ki k dissim x β −− and ( ki k dissim x ) β , for each attribute. However, these dissimilarity variables will each be a fairly complicated function of the same basic vector of “true” marginal utility parameters that they modify. It is straightforward (if tedious) to write down the log-likelihood for full information maximum likelihood estimation of this model (using any of the practical candidates for these dissimilarity measures). However, given the complex and repeated manner in which the basic utility parameters enter the model, one can expect the log-likelihood function to be somewhat difficult to maximize. To allow us to explore these data for evidence of unequal attention bias across attributes, however, it is possible to implement a crude correction without resorting to custom-programmed nonlinear optimization models. Estimation can be accomplished by employing an iterative algorithm that relies solely on sequential utilization of packaged conditional logit algorithms. This iterative algorithm is described in detail in Appendix A (this seems to mimic the method used by Swait and Adamowicz (2001a,b) in their work with entropy as a measure of choice set complexity). Upon convergence, the last set of parameters can be used to compute the “final” estimated values of the shift variables capturing, for each attribute, the similarity of the available alternatives based on the other attributes, and the dissimilarity of the available alternatives based on this attribute. In our 94 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 model with four marginal utility parameters, there will be eight additional (estimated) regressors to be interacted with the basic attribute variables. When the model is estimated iteratively in this fashion, using packaged conditional logit software, the parameter variance-covariance matrix in the last round, of course, does not reflect the estimated nature of the estimated otherand own-attribute standard deviations (or “leads”) in utility. Full information maximum likelihood estimation is required to estimate all of the parameters of the two models simultaneously, so that a full parameter variance-covariance matrix can be obtained.14 4.4 Empirical results In this example, we expect to find a positive marginal utility of income () 1 β , and negative marginal utilities associated with the logarithms of (shifted) present discounted sick-years, recovered-years, and lost life-years ( 234 ,, ) βββ . We expect that the greater the disparity in utilities across alternatives, based on other attributes, the less will be the individual’s apparent responsiveness to differences in the level of any given attribute. We also expect that the greater the difference in utility derived from the attribute in question, the greater will be the individual’s apparent responsiveness to differences in the level of any given attribute. 4.4.1 Models using () ' ki k sd x β −− and () ki k sd x β : Table 2 shows the results of a succession of fixed-effects conditional logit-type models where the disparities in indirect utility based on all other attributes, and based on just this attribute, are measured as the standard deviation across alternatives. Model SD1 (homogeneous preferences) is a baseline specification, with no attention-related correction terms, involving only the four utility parameters in our most basic specification. The signs on all three estimated parameters are as anticipated, and each is strongly statistically significantly different from zero. Model SD2 (heterogeneous preferences) generalizes this specification to allow for systematically varying preference parameters. The marginal utility of net income is statistically significantly higher for women. The coefficient capturing the marginal disutility of expected discounted sick-years is negative. It is more negative, the higher the individual’s subjective risk of suffering the illness or injury targeted by the risk-reduction program in question. It is less negative, the higher the individual’s average subjective risk of suffering from any of the other major categories of health risks addressed in the survey. If the individual indicates, ex post, that they expect never to benefit from the program in question, the disutility from illness in this case is drastically reduced. Finally, the greater the individual’s overestimate of the latency period before benefits begin (i.e. before the illness will cause pain or disability), the lesser the implied disutility from present discounted sicktime. If recovered-years are viewed as a return to perfect health, we would expect utility in that state should be identical to pre-illness utility, but our estimates suggest that most 14 We have explored a number of FIML specifications for the overall optimization process, using Matlab’s general function-optimizing software. As would be expected, however, it can be very difficult to achieve convergence in this context because the various utility parameters in the model enter multiplicatively. One can expect the iterated estimates used in the body of this paper to understate the amount of noise in the estimates, to a degree, because the dissimilarity variables are treated as non-stochastic when they are actually estimated quantities. 95 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 96 individuals do not view this to be the case. There appears to be negative utility associated with “recovered” years, and this disutility is greater, the older the respondent at the time when these stated choices are being made. These are major illnesses, including five types of cancers, heart disease or heart attack, respiratory disease, stroke, diabetes, Alzheimer’s disease and traffic accidents. An expectation of lingering morbidity is reasonable. Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 Table 3: Sizes of the effects of dissimilarity variables on estimated marginal utilities (Model SD5) Denominator of WTP ↓ In numerator of WTP   Dissimilarity variables normalized so that sample mean = 0 Income term (1 β ) Sick-years term (2 β ) Recovered-years term ( 3 β ) Lost life-years term ( 4 β ) MU at “mean” dissimilarity = 1.514 -7.124 -33.92 -20.23 Effects of other-attribute utility dissimilarity (percentiles): 5th 2.24 -38.04 -49.38 -47.83 25th 2.01 -28.85 -44.52 -38.91 50th 1.68 -16.37 -37.64 -27.48 75th 1.19 6.03a -27.63 -8.16 95th 0.21 50.81a -4.99 30.20a Effects of own-attribute utility dissimilarity (percentiles): 5th 1.01 20.08a -35.90 -7.53 25th 1.12 12.77a -35.90 -10.76 50th 1.33 1.08a -35.90 -15.94 75th 1.70 -18.80 -33.47 -25.07 95th 2.67 -62.93 -25.82 -48.23 a Unexpected signs on some of these fitted marginal utilities result from estimation without constraints. In a non-linear adaptation of this model, it would be possible to estimate the negative of the logarithm of each marginal utility, and to allow this log-transformed parameter to shift systematically with the two types of dissimilarity measures. This would constrain the fitted marginal utility to remain strictly negative. different framework, where attention was applied to tradeoffs between attributes, rather than to the attributes themselves. In a linear and additively separable model with four basic attributes, there are six possible ratios of marginal utilities (marginal rates of substitution) that respondents might consider. Unfortunately, it is beyond the scope of this paper to develop and implement an alternative specification where six attention parameters are associated with these six ratios of the four basic marginal utility parameters, so we leave this option for subsequent research as well. Such a specification would of course be interesting, since this referee notes that it might permit us to discriminate between fully compensatory strategies, conjunctive strategies, disjunctive strategies or even simply an adding up of attributes that exceed a threshold. This alternative approach might also be valuable if individuals actually have lexicographic preferences (e.g. always just choose the cheapest option—although in our example, this would be the status quo in every choice set). But consider a case of forced choice, where every alternative comes at a cost. If the individual chooses solely on the basis of cost, and costs are identical, he or she may pick randomly 103 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 among alternatives with these equal costs. Our model would not perform well if everyone chose in this manner. We resort to our “alternating estimator” because of the fundamental challenge of identification when the same set of utility parameters shows up in so many places in the objective function for FIML estimation. It has been suggested that we might consider breaking the link between the preferences used in the utility function and the preferences used in the attention component of the model by adopting the equal weights proposed by Dawes (1979). A system of equal weights would fix the “utility” coefficients in the calculation of the attention variables arbitrarily at unity. We feel that our alternating estimator is preferable, however, since it fixes these coefficients at their estimated values from the previous iteration of the richer heterogeneous-preferences specification for the “true” underlying utility function (which we assume would be unobserved by the typical practitioner in search of a simple linear and additively separable model for the representative consumer). We note that it is tempting to consider using time-on-task durations for each choice as a direct proxy for attention to that choice. In a laboratory setting, this might be viable. For our internet-based field survey, it would be a risky strategy. A longer duration on a choice task can just as easily mean that the respondent was distracted from the task for a few seconds up to a few hours. In a laboratory setting, one might also use eye-tracking software to measure how long the respondent’s gaze dwells upon an area of the screen occupied by each attribute. With this more-direct information about attention, it would be much easier to build explicit constraints for the attention-related terms in our model. With the current data, however, it is very difficult to come up with additional constraints to aid in the identification of the attention and utility-related parameters.16 In our empirical application, it is much more convenient to use the first variant of equation (12) for the propensity to attend to different attributes. Had it been equally tractable to employ the third variant, it would be straightforward to assess whether the attention factor could be equal across all attributes for any given respondent. This would suggest a model equivalent to one wherein the scale factor for the utility-difference function merely differs across people. This type of test is left for future research where the third variant in equation (12) can be implemented. Subjects in our study each have the opportunity to make five different choices. In this case, it may not be the standard deviation across the current choice set in utility contributions for a particular attribute which determines attention to that attribute. Instead, it may be the standard deviation in utility contributions across both the current and all previous choice sets that determine the attention devoted to an attribute. We do not pursue this possibility here. To allow the baseline marginal utility parameters to be comparable across our various specifications, we first normalize our dissimilarity variables on their mean values across all respondents. This permits us to consider the case where all dissimilarity variables might match the sample-wide mean as equivalent to the case where the shift variables we actually use are all simultaneously zero. We have not addressed the possibility that one might normalize on the within-individual means, but to allow these mean dissimilarity measures to differ across individuals. An a priori sense of the promise of such a strategy is harder to come by, since the relevant quantities are factors in interaction terms, rather than basic variables in the model. 16 Chabris et al. (2009) consider the allocation of time across choice tasks according to the “value gap” between the two options in each choice. However, they do not consider the allocation of attention among attributes, or among alternatives within a choice task. 104 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 One might argue that our use of the richer heterogeneous-preferences model to generate our eight fitted dissimilarity measures is merely an alternative strategy for bringing respondent heterogeneity into the naive homogeneous preferences model. This criticism may be supported by the fact that the same dissimilarity measures make no real difference when they are added to the heterogeneous-preferences model. But this does not take away from the intriguing finding that respondent heterogeneity—exclusively via its influence on the two types of theoretically motivated measures of alternative similarity—contributes so very much to explaining differences in apparent marginal utilities in the naive model (which is where many practical conjoint choice analyses begin and end). 17 5 Conclusions and Potential Implications In conventional random utility choice models, researchers usually assume complete and costless information. However, subjects’ cognitive resources are typically scarce. Individuals presumably must compare the expected marginal benefits and marginal costs of attention to different dimensions of a choice task, and optimize their allocation of attention. In this paper, we focus on the individual’s allocation of his or her attention across the different attributes which can be used to describe each alternative in a choice set. Inattention to differences in the levels of a particular attribute may masquerade empirically as a lower marginal utility associated with that attribute. Marginal utilities from choice models are the key ingredients in the calculation of willingness-to-pay in many applications. Distortions in these marginal utilities can lead to distortions in the sorts of willingness-to-pay estimates which are critical to an understanding of demands for the goods in question. Our illustrative empirical example represents a first partial attempt to implement an attention-corrected choice model with a sample of “field” data from a conjoint choice experiment in a large stated preference survey. When we use a four-parameter homogeneous-preferences model to build the two dissimilarity measures associated with each attribute, and use these two measures to shift each marginal utility in what is otherwise the same four-parameter homogeneous model, we find no evidence of the effects predicted by our theory. We then generalize our model to make each of our four marginal utilities a systematically varying parameter, allowing for heterogeneity in preferences. If these heterogeneous preferences are used to build the two dissimilarity measures associated with each attribute, and these measures are the used to shift each of the four marginal utilities in the same heterogeneous-preferences model, they likewise fail to produce the effect predicted by our theory. However, we subsequently assume heterogeneous preferences in the process of constructing the dissimilarity measures, so that our pairs of dissimilarity measures associated with each attribute differ across individuals because their preferences are different. Using these heterogeneous dissimilarity measures as estimates of “latent” variables that have the capability to shift the four basic marginal utilities in a homogeneouspreferences model produces highly significant results fully consistent with our theory. Choice modellers do often explore homogeneous-preferences specifications, seeking to estimate preferences for a representative consumer. Our results certainly suggest that such “representative preferences” may be biased by heterogeneity in perceived dissimilarities. Our theoretical and empirical explorations of criteria that may affect a respondent’s optimal allocation of attention to attributes also have some implications for other regularities 17 One can in principle impose sign restrictions by assuming, for example, a log-normal rather than normal distribution for each parameter (using the negative of the variable in estimation if a strictly negative coefficient is desired). For this application, however, models with such restrictions failed to converge. 105 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 which have been observed in different types of choice behaviour. Most of these issues are basically familiar to choice modellers. What our research introduces is an additional theoretical justification and attention-based rationale for why these patterns are observed. Our research also suggests the extent to which WTP estimates could be unintentionally biased by arbitrary decisions about the mix of attributes in a choice set. Our work is certainly not the last word on this subject, however, since the limited number of dimensions of variability across the choice sets in our empirical application do not permit us to test all of the possible predictions based on this type of approach to bounded rationality in choice situations. SP too different from RP choice sets.—Choice set designs used for stated preference surveys may produce uneven attention to different attributes. This might be of little consequence if the corresponding real choice contexts were assured of being similar. However, if the conditions surrounding the choice are sufficiently different in the context wherein a choice prediction is desired—so that the marginal benefits and/or marginal costs of attention to attributes are different—a model calibrated under an implicit assumption of complete attention (when this is not so) may produce misleading forecasts of future choices. This suggests that when SP data are to be used to predict likely RP choices, it may be important to design those SP choice sets to feature the types of attribute mixes that will be encountered in the real world, to ensure that the distribution of attention across attributes in the SP case will be similar to that in the future RP case. Consequentiality.—In purely hypothetical SP choice contexts, where stated choices may be viewed as inconsequential, the marginal benefits from attention to all attributes could be perceived to be very low (see Carson et al. 2003; 2004). In contrast, the marginal costs of attention to any additional attribute may be very similar to those in a real choice context. A lack of perceived consequentiality would thus be predicted to lead to a lower overall optimal level of attention being paid to the choice task. If attention to each attribute is reduced proportionally, we have argued that the only substantive effect may be observationally equivalent to an increase in the error dispersion, relative to the full attention case, with no resulting bias in the relative sizes of the estimated marginal utility parameters, and thus no distortion in any resulting estimates of the expected WTP. However, levels of attention to different attributes may not be scaled down uniformly across all attributes when less-than-complete attention is optimal. Our theory focuses on the marginal benefits part of the story, and suggests that the marginal benefits from attention to an additional attribute depend in a fairly complex fashion upon the pattern of attributes in the choice set and on the individual’s marginal utilities from each attribute. Simply the design of a choice set can steer the subject’s attention toward some attributes and away from others. If the purpose of the SP choice task is to measure social preferences for a non-market public good, such as environmental quality, one may wish to simulate the preference parameter estimates that would emerge under the counterfactual case where everyone in the sample devotes their complete attention to every attribute in each choice set. These would simulate the “full information” case that would be highly desirable in the estimation of preferences in such a context. Price listed last.—In SP experiments, the utility loss from making a wrong choice can be negligible, since the individual may believe that he or she will not have to live with the consequences of an “incorrect” choice—in particular, the knowledge that they have paid good money for something that turned out to be not exactly what they wanted or expected. If respondents do not fully take into account the fact that they would actually have to pay the cost of the preferred alternative, they may pay less attention than they should to any differences in costs (especially if these costs are listed at the bottom of the conjoint choice table). If order effects increase the relative cost of attention to the cost attribute, and attention is steered toward other attributes by listing them first, it may be unsurprising that 106 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 the propensity to attend to other attributes will be greater than the propensity to attend to cost. The marginal utility of income may be underestimated by more than the marginal utilities of the other attributes. The predicted result would be an upward bias in WTP, since the marginal utility of income forms the denominator in WTP calculations. Cheap talk scripts.—In SP surveys, since the publication of Cummings and Taylor (1999), researchers have been encouraged to employ a so-called “cheap talk” script wherein subjects are specifically reminded to consider their budget constraint carefully before stating their preferred option. This section of a survey will typically draw special attention to the cost attribute, immediately prior to the choice task. In a conjoint choice context, this effort can be expected to increase attention to the cost attribute without treating the other attributes symmetrically. Our theory suggests that this can be expected to lead to a larger-thanotherwise estimated marginal utility of income and perhaps smaller-than-otherwise estimated marginal utilities for other attributes (if scarce attention is reallocated), which will tend to “bias,” rather than “correct” the resulting WTP estimates. However, if a lack of consequentiality for the entire choice exercise has already produced lower attention to the cost attribute, the cheap talk effort may be warranted. However, our results strongly suggest that any attempt to direct the subject’s attention specifically towards one attribute or another should be examined very carefully. We know that the cost attribute is frequently downplayed in RP contexts: restaurant menus list the price of the entree last, and advertisements encourage prospective customers to “contact the dealer for price information.” Effort is often made to divert attention from other attributes as well, especially where they may convey negative marginal utilities: some less desirable attributes of goods for sale are listed in the fine print (e.g. pharmaceutical side effects), SP attributes sometimes “too orthogonal.”—To maximise estimation efficiency for marginal utilities associated with a whole range of attributes, the joint distribution of attributes in SP studies often has greater orthogonality or greater variance than might be present in the corresponding real-world choice context. As Jordan Louviere has pointed out, “Realism is not a design property” for a choice set. Our theoretical results suggest that the degree of orthogonality in attributes may have a systematic effect on the sizes of naively estimated marginal utilities. In the corresponding real choice context, subjects may face alternatives where the differences in many attributes across alternative may be much smaller than they had been in the SP estimating sample. This lesser difference changes the expected net benefits from considering the different attributes and changes the extent to which the individual is likely to take into account each of these attributes in the real-choice context. A choice model estimated on SP data would incorrectly predict choices under the different choice regimes in a subsequent RP setting. More “ceteris paribus” than in real choices.—While attribute levels may be more different in some SP studies than they are in real life, in other cases the researcher’s goal is merely to obtain a precise estimate of just one marginal utility. In this situation, the choice sets might consist of alternatives where all other attributes are held constant and only the attribute of interest is varied across alternatives. In some cases, the choice scenario may not even list other important attributes and will simply ask respondents to assume that all other features of the alternatives are identical. Our theory suggests that the greater the number of attributes held essentially constant across alternatives, the larger will be the apparent marginal utilities associated with the attributes which do vary.18 18 For example, had the identical resumes with different names been sent to the same prospective employers in the Bertrand and Mullainathan (2004) study, one might expect that race, as implied by the different names, might have been found to have an exaggerated influence on choices compared to a choice context where the resumes differed in many other dimensions as well. 107 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 Marginal cost differences.—Our theory does not explicitly derive the factors which should determine the marginal cost of attention to an additional attribute. However, intuition suggests that the marginal costs of attention to different attributes may also differ. A variety of conditions could affect the marginal cost of attention to an incremental attributes. One is the accessibility of the information about each attribute (e.g. its position in the order of attributes in a conjoint choice scenario). For a decision-maker faced by different levels of distraction or time pressure in making a choice, the fact that the marginal cost of attention is likely increasing in the number of attributes in a choice set will also be relevant. Some attributes, such as risk for example, may be more difficult to understand for some types of subjects. This suggests that there will likely be differences in the marginal propensity to attend to each attribute whenever cognitive constraints are binding. Differences in attention can lead to biases in estimated marginal utilities and thereby to distortions in estimated WTP. We have derived, from optimizing behaviour, results that seem to match closely with casual empiricism about how people make choices. Individuals are motivated to pay attention to additional attributes in a choice exercise to the extent that this behaviour will reduce their expected lost utility from making an incorrect choice. They pay more attention to any given attribute if the alternatives look more similar in terms of utility based on the other attributes under consideration. They also pay more attention to an attribute if the utility derived from that attribute differs greatly across alternatives. These are simple insights. In our empirical adaptation of this theory, we encounter some difficulty in estimation of an appropriate specification using full-information maximum likelihood methods. This is because the same utility parameters appear in so many places in the log-likelihood. Nevertheless, we have implemented the estimation in an alternating sequence of steps that appears to lead to stable converged parameter estimates. We demonstrate that the apparent marginal utilities from different attributes can vary dramatically with the mix of attribute levels presented across all alternatives, and thus so can the implied WTP. This is more evidence that, by manipulating the mix of attribute levels in a choice set, it may be possible to “steer” respondent attention (inadvertently or strategically) to either exaggerate or downplay apparent marginal utilities and hence the resulting average WTP (benefits estimates). Our findings may therefore have important implications for how researchers approach the problem of experimental design in specifying choice sets in SP research. They may also explain problems in using RP data from one type of choice context to infer likely behaviour in another context where the patterns of attributes across alternatives are too different. 6 Acknowledgements Our recognition of the need for this paper was crystallized by discussions in the session entitled “Dissecting the Random Component” at the Fifth Triennial Invitational Choice Symposium hosted by UC Berkeley at Asilomar in Pacific Grove, CA, in June 2001. Valuable feedback was obtained from presentation of the theory section in the session entitled “Recent Progress on Endogeneity in Choice Modeling” at the Sixth Triennial Invitational Choice Symposium in Estes Park, CO in June 2004. Initial empirical results supporting the model were presented in the session entitled “Behavioral Frontiers in Choice Models” at the Seventh Triennial Invitational Choice Symposium at the Wharton School in Philadelphia, PA, in June 2007. We are grateful to all of the participants at these sessions who have influenced our thinking on this topic, as well as to members of the Triangle Resource and Environmental Economics seminar sponsored by NC State, RTI, and Duke University, and the 10th Occasional Workshop on Environmental and Resource Economics 108 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 at UCSB. Jason Lindo has provided some helpful editorial comments. This research was supported in part by the National Science Foundation (SES-0551009) and by the Raymond F. Mikesell Foundation at the University of Oregon. 7 Appendices 7.1 Alternating Estimation Algorithm Step 0: Estimate the model without any attention corrections. Save these temporary estimates of the k=1,...,K marginal utility parameters. These might be scalars, (00 1 ˆˆ ,..., ) K ' ββ , or systematically varying parameters, 0 1 ˆ (' 0' ' 1ˆ ,..., ) iKKi Z Z ββ ch depending on a sub-vector of parameters 0 ˆk , ea β and a vector ki Z of individual characteristics. Step 1: Based on these initial estimates of the four marginal utilities, construct the contribution to net indirect utility associated with each attribute, relative to that for the numeraire alternative, J. This may be a single scalar marginal utility times its associated attribute level, ˆ j ki k x β ,or it may be a systematically varying marginal utility times the associated attribute level, ' ˆ ( j ki k ki ) x Z β , for k=1,...,K. Sum these contributions across all attributes to calculate 'ˆ j i x β , the net indirect utility index associated with each alternative, j=1,...,J-1. For the numeraire alternative, this net utility will be zero. a.) For each attribute, subtract from total systematic utility the contribution made by just that attribute to leave the model’s prediction about net indirect utility based only on the other attributes in the model, ˆ jki k x β −− 0 ˆ () ki k x (or in the systematically varying parameter case). Construct a measure of the dissimilarity of the alternatives on the basis of these other attributes, ' ˆ ( jki k ki xZ β −−− ) dissim β −− _( ki k m x , or . We have suggested several candidates: the size of the lead, in utility units, the standard deviation across alternatives in these other-attribute utility levels, and the skewness in these measures across alternatives. Adjust the location of these measures by using their deviations from the overall sample mean values (or any other target value to be simulated by zeroing out this dissimilarity measure): 0' ) k ki Z −−− )mean ˆ (( ki dissim x β 00 ˆˆ ) ( ki k dissim x ) 0 ˆ )_( ki k dissim xd dissi ββ β −− −− −− =− . b.) For each attribute, construct a measure of the dissimilarity of the three alternatives on the basis of just this attribute: 0 ˆ () ki k dissim x β or . Again, possible candidates include the lead of the highest utility contribution due to this attribute, over the second-highest across alternatives, or the standard deviation, or the skewness in these utilitycontributions across alternatives. Again, adjust these measures by using their deviations from the overall sample mean values (or some other target value to be simulated when the deviations are all zero), to yield d 0' ˆ (( )) ki k ki dissim x Z β 0 ˆ _( ki k dissim x ) β . Step 2: Re-estimate the model, but now allow the marginal utility from each attribute (or the intercept of the marginal utility expression, if it is modelled as a systematically varying parameter) to vary systematically with the calculated dissimilarity of the alternatives in this choice set based on net utility from other attributes, as well as the dissimilarity of the alternatives based on net utility only from this attribute. Each “observed” marginal utility parameter is now modelled as also varying systematically with 0 ˆ () ki k dissim x β −− and 109 Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 110 ) 0 ˆ ( ki k dissim x β . In this second iteration, the new vector of “true (corrected)” underlying marginal utility parameters for each attribute, 1 ˆk β , is supplemented by the estimated coefficients on each of these two dissimilarity terms, yielding 111 ˆˆ ˆ (,,) kkk β αθ d dissi for k=1,...,K. If the marginal utilities in the model are scalars, this generalization will triple the number of estimated parameters. If the marginal utilities are systematic varying parameters, the number of estimated parameters will increase by 2K. Step 3: Net out the estimated biases in systematic utility due to 0 ˆ _( ki k m x ) β −− and 0 ˆ _( ki k d dissim x ) β by setting these 2K different constructed variables to zero. This simulates the case where, for all attributes, the dissimilarity of alternatives based on all other attributes, and based on each specific attribute, is the same for all attributes in all choice sets. We then interpret the other utility parameters in the model as the “true” utility parameters (corrected for attention biases created (unintentionally?) by the mix of attributes designed into the choice set). Step 4: Repeat Step 1, now using these updated estimates of the basic utility parameters, 11 ˆˆ ,..., ) K 1' 11 ˆ ( ,..., 1' ' ˆ ) iKK ' Z 1 ( ββ , or systematically varying parameters, i Z β _(issim x β d d , as the “true” utility parameters to construct updated measures of dissimilarity, 1 ˆ ki k ) β −− and 1 ˆ) ki k x_(ddissim β . Continue to iterate through Step 1 through 3 until the length of the step-to-step permutation in the parameter vector becomes arbitrarily small. 7.2 Corrected Variance-Covariance Matrix We have also estimated Models SD1 through SD5 as conditional logit-type specifications without fixed effects, with results as shown in Table A. A fixed effects specification is less crucial in this context because the attributes of the alternatives in our choice sets were randomized, subject only to exclusions for implausibility. Results of a similar flavour emerge, with the analogue to Model SD5 again providing evidence of the types of attentiondiverting effects suggested by our theory. These non-fixed-effects models also allow us to address the problem that the standard errors at the last iteration of the steps in the estimation algorithm do not reflect the fact that the variables used for the eight dissimilarity terms are calculated based on the last round of point estimates from the heterogeneous specification (which also involves fitted dissimilarity variables from the most recent round of estimates). Ideally, one would estimate all parameters of the model simultaneously by full information maximum likelihood. However, since the basic utility parameters appear in so many different places in these models, we are not surprised to find that such likelihood function is very difficult to optimize by standard methods. Using the alternating algorithm, convergence seems to be straightforward and unambiguous. When the converged point estimates from the alternating algorithm are inserted into the full likelihood function for the same problem and numerical derivatives are calculated for the full set of parameters, there is some shrinkage of the asymptotic t-test statistics on most parameters, but everything that was statistically significant at better than the 10 percent level at the end of the alternating algorithm remains significant in terms of the full log-likelihood function. However, we note one markedly larger t-test statistic for the very last parameter in the model. The standard step-sizes for numeric derivatives may be inappropriate for this parameter. This particular test statistic needs yet to be understood. Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 111 Table A: “Standard Deviation” Variant: Uncorrected and Attention-corrected Non-Fixed-Effects Conditional Logit Models for Health-Risk Reduction Programs (with homogeneous and heterogeneous preferences) (SD1) iterative (SD2) iterative (SD3) iterative (SD4) iterative (SD5) iterative (One-step eff.) FIML Exp. sign Homogeneous preferences Heterogeneous preference Attention homogenoushomogeneous Attention heterogeneousheterogeneous Attention heterogeneoushomogeneous Attention heterogeneoushomogeneous Income term ( 1 β ) (income term) [ + ] 3.364 3.72 4.769 3.222 2.475 2.475 (8.28)*** (6.66)*** (0.38) (3.37)*** (2.78)*** ( 2.22)** ... (sd(U othr attr)-mean sd) ×[ - ] - - 58.84 3.646 -2.227 -2.227 (2.05)** (3.16)*** (2.30)** ( -1.94)* ... (sd(U this attr)-mean sd) ×[ + ] - - 470.6 -3.995 2.256 2.256 (1.75)* (1.91)* (1.64) (1.03) … female × - 3.253 - 4.952 - - (5.00)*** (4.93)*** Sick-years term ( 2 β ) AS i ΔΠ log(pdvi+1) [ - ] -28.87 -17.8 -14.3 -17.77 -18.57 -18.57 (4.76)*** (2.37)** (0.78) (1.34) (1.59) ( -1.7)* ... (sd(U othr attr)-mean sd) ×[ + ] - -29.67 -22.4 109.9 109.9 (0.65) (1.00) (5.83)*** ( 4.46)** ... (sd(U this attr)-mean sd) ×[ - ] - -71.4 11.54 -121.1 -121.1 (0.37) (0.45) (6.47)*** ( -6.89)** ... (sasubrsk-msasubrsk) × - -22.49 - -26.18 - - (3.83)*** (4.27)*** ... (cosubrsk-mcosubrsk) × - 29.68 - 33.21 - - (3.47)*** (3.80)*** ... (benefits never) × - 126.4 - 122.3 - - (3.86)*** (3.65)*** ... (min overest latency) × - 7.614 - 8.456 - - (11.85)*** (11.36)*** Recovered-years term ( 3 β ) AS i ΔΠ log(pdvr+1) [ - ] -24.29 -40.35 -76.69 -30.9 -34.08 -34.08 (2.53)** (3.92)*** (2.79)*** (1.23) (1.78)* ( -1.57) Cameron and DeShazo, Journal of Choice Modelling, 3(3), pp. 73-115 112 ... (sd(U othr attr)-mean sd) ×[ + ] - - -19.88 -14.13 23.73 23.73 (0.21) (0.48) (1.04) -0.69 ... (sd(U this attr)-mean sd) ×[ - ] - - 127.6 -38.32 -29.48 -29.48 (2.15)** (0.26) (0.31) (-0.611) ... (age-mage) × - -1.614 - -1.335 - - (2.40)** (1.37) Lost life-years term ( 4 β ) AS i ΔΠ log(pdvl+1) [ - ] -30.73 -23.03 -57.94 -16.33 -43.93 -43.93 (5.88)*** (3.50)*** (3.88)*** (1.43) (4.43)*** ( -4.73)** ... (sd(U othr attr)-mean sd) ×[ + ] - - 84.86 -2.609 104.7 104.7 (1.54) (0.10) (4.74)*** ( 3.8)*** ... (sd(U this attr)-mean sd) ×[ - ] - - 64.72 -23.58 -36.28 -36.28 (1.55) (1.20) (2.92)*** ( -36.8)*** a ... (sasubrsk-msasubrsk) × - -40.74 - -40.28 - (7.44)*** (6.97)*** ... (cosubrsk-mcosubrsk) × - 33.15 - 32.62 - - (4.19)*** (4.06)*** ... (benefits never) × - 204.2 - 209.9 - - (6.33)*** (6.35)*** ... (min overest latency) × - 7.869 - 8.019 - - (13.11)*** (12.31)*** Observations 21111 21111 21111 21111 21111 21111 Log L -7682.953 -7050.867 -7670.404 -7042.261 -7603.081 -14645.34 Iterations 40 40 40 1 Marginal WTP for incr. in log(pdvi+1) 5% -11.37 -10.32 -18.1 50% -8.58 -3.31 -17.72 95% -5.92 9.06 -9.9 Marginal WTP for incr. in log(pdvl+1) 5% -11.97 -34.65 -41.43 50% -9.11 -.93 -7.49 95% -6.64 32.46 .34 Absolute value of z statistics in parentheses; * significant at 10%; ** significant at 5%; *** significant at 1% a The reason for this unexpectedly small standard error (large t-test statistic) is unclear. These standard errors are calculated by substituting the converged values of the parameters from the sequential method into the full maximum likelihood function, followed by calculation of numeric derivatives at this optimum.