Intertemporal discrete choice
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Pennesi, Daniele Working Paper Intertemporal discrete choice Quaderni - Working Paper DSE, No. 1061 Provided in Cooperation with: University of Bologna, Department of Economics Suggested Citation: Pennesi, Daniele (2016) : Intertemporal discrete choice, Quaderni - Working Paper DSE, No. 1061, Alma Mater Studiorum - Università di Bologna, Dipartimento di Scienze Economiche (DSE), Bologna, https://doi.org/10.6092/unibo/amsacta/4715 This Version is available at: https://hdl.handle.net/10419/159899 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/3.0/
ISSN 2282-6483 Intertemporal discrete choice Daniele Pennesi Quaderni - Working Paper DSE N°1061
Intertemporal discrete choice∗ Daniele Pennesi† February 2016 Abstract The discounted logit is widely used to estimate time preferences using data from field and laboratory experiments. Despite its popularity, it exhibits the "problem of the scale": choice probabilities depend on the scale of the value function. When applied to intertemporal choice, the problem the scale implies that logit probabilities are sensitive to the temporal distance between the choice and the outcomes. This is a failure of an intuitive requirement of stationarity although future values are discounted geometrically. As a consequence, patterns of choice following from the structure of the logit may be attributed to non-stationary discounting. We solve this problem introducing the discounted Luce rule. It retains the flexibility and simplicity of the logit while it satisfies stationarity. We characterize the model in two settings: dated outcomes and consumption streams. Relaxations of stationarity give observable restrictions characterizing hyperbolic and quasi-hyperbolic discounting. Lastly, we discuss an extension of the model to recursive stochastic choices with the present bias. Keywords: Discrete Choice, Intertemporal Choice, Quasi-hyperbolic Discounting Jel Classification: D01 ∗Some results contained in the current work appeared in a paper circulated under the title "The intertemporal Luce rule". †University of Bologna, Department of Economics, Piazza Scaravilli 2, 40126, Bologna (Italy). Email: [email protected] 1
1 Introduction The discounted multinomial logit (DML) is the most common model of probabilistic choice among those used in the estimation of time preference (e.g. Chabris et al., 2008;Louie and Glimcher,2010;Tanaka et al.,2010). Despite its popularity, the DML may be inappropriate to study intertemporal choice due to the "problem of the scale": choice probabilities depend on the scale of the value function (Fosgerau and Bierlaire,2009). Consider the probability of choosing now an outcome xat time t over an alternative yat time s.1Should this probability remain unchanged when both rewards are equally delayed? A positive answer is motivated by an intuitive form of stationarity borrowed from deterministic choice (Fishburn and Rubinstein, 1982): choice probabilities should be independent of the temporal distance between the choice and the dates of the rewards. The only quantity that matters is the relative time distance between the rewards, i.e. |t−s|. Therefore, a stationary stochastic choice rule should assign the same probability to the choice of (x, t)over (y, s)and the choice of (x, t +r)over (y, s +r). Although the discounted multinomial logit discounts future values geometrically, it fails to satisfy an even weaker version of this intuitive form of stationarity2(see Section 2). Choice probabilities in the discounted multinomial logit do depend on the temporal distance between the choice and the delivery of the rewards. As a consequence, the estimation of time preferences may be distorted. For example, attributing to quasi-hyperbolic discounting patterns of choice probabilities that follows from the structure of the discounted multinomial logit, rather than violations of geometric discounting (see Section 2). We propose a new model, the discounted Luce rule, that solves the stationarity problem while maintaining the flexibility and simplicity of the discounted logit. The model is completely characterized through testable restrictions on choice probabilities. Differently from the discounted logit, it allows a stark identification of the observational consequences of hyperbolic and quasi-hyperbolic discounting, since they correspond to certain violations of stationarity. We axiomatize the model in two settings: dated rewards and consumption streams. Discrete choice over dated rewards (x, t), meaning 1According to the DML this probability is PLogit((x, t)≥(y, s)) = eδtv(x) eδtv(x)+eδsv(y). 2More precisely, in the discounted multinomial logit the ratio between the probability of choosing (x, t) over (y, s)is different from the ratio between the probability of choosing (x, t +r)over (y, s +r). 2
a reward xdelivered at time t, is interesting for two reasons: first, in this setting all the properties characterizing the model are intuitive and easy to test. Second, choice among dated rewards includes the Multiple Price List method (Coller and Williams, 1999), the workhorse of experimental and field studies devoted to the elicitation of time preferences.3In the dated reward setting, the probability of choosing (x, t)∈A according to the discounted Luce rule is given by the relative discounted value of (x, t) in A: P((x, t), A) = δtv(x) X (y,s)∈A δsv(y) with δ∈(0,1] representing the discount factor. The model is completely characterized by stochastic stationarity and stochastic impatience. The former implies that choice probabilities are not sensitive to the temporal distance between the choice and the timing of rewards. Generalizations of the previous rule accounting for quasi-hyperbolic and hyperbolic discounting are provided through a relaxation of the stationarity axiom. Quasi-hyperbolic discounting, for example, predicts that the probability of choosing an immediate reward over a delayed one decreases when both rewards are equally delayed (Section 3.1). The discounted logit may generate the same behavior while discounting future values geometrically. The more general setting of consumption streams, x= (x0, x1, . . . , xT), allows us to compare the properties characterizing the discounted Luce rule to axiomatizations in the deterministic setting of additively separable discounted utility with geometric (Koopmans,1960) or quasi-hyperbolic discounting (Hayashi,2003;Montiel Olea and Strzalecki,2014). In this case, the probability of choosing a consumption stream x from a set Aof alternatives is given by the relative present value of xin A: P(x, A) = T X t=0 δtv(xt) X y∈A T X t=0 δtv(yt) Beyond an intuitive requirement of stationarity adapted to this setting, the main novelty characterizing the previous rule is a property called separability: the relative probability of choosing a consumption stream xover yin a set Ais equal to the sum 3Among the others: Harrison et al. (2002)Andersen et al. (2008), Tanaka et al. (2010), Halevy (2015). 3
of the relative probabilities of choosing its "components", properly defined (see Section 4). Separability is the stochastic counterpart of additive separability in deterministic choice. Once stationarity is relaxed, we give observable conditions to characterize quasi-hyperbolic discounting: P(x, A) = v(x0) + β T X t=1 δtv(xt) X y∈A"v(y0) + β T X t=1 δtv(yt)# The properties are comparable to the ones used in deterministic choice (e.g. in Montiel Olea and Strzalecki,2014). The restrictions characterizing stochastic choice with quasi-hyperbolic discounting can inform applied works estimating time preferences from field data (for example durable good adoptions, Chevalier and Goolsbee,2009; Dubé et al.,2014) and works estimating recursive models of stochastic choice with quasi-hyperbolic discounting (see Section 6). These are structural econometric models that retain the recursive dynamic structure of Rust (1987), but allow quasi-hyperbolic discounting of the continuation value. Their behavioral characterization has not been studied yet. Only recently, Fudenberg and Strzalecki (2015); Matêjka et al. (2015) studied the particular case of geometric discounting and logit choice probabilities. Therefore, our results can be helpful in understanding the implicit restrictions entailed by the use of such models. In the recursive quasi-hyperbolic discounted model, the individual stochastically selects in each period, an immediate consumption xtand a continuation plan At+1. The value of an action at time t,at= (xt, At+1)is equal to: Ut(xt, At+1) = v(xt) + βδ Emax at+1∈At+1 Ut+1(at+1) + at+1 and the choice from the menu at time tis probabilistic: Pt(at, At) = Prob Ut(at) + at≥max bt∈At Ut(bt) + bt The individual correctly anticipates the shocks to her future utility and discounts the continuation value quasi-hyperbolically. Non-experimental data on job search (Paserman,2008), insecticide treated nets’ adoption (Tarozzi and Mahajan,2011), drug 4
compliance (An et al.,2014), mammography decisions (Fang and Wang,2015) and cellphone usage (Yao et al.,2012), have been used to separately identify time preferences δand the present bias factor β. The recursive Luce rule corresponds to the previous model when the error terms are i.i.d. and distributed according to a Fréchet (see Section 2.1). We show a simple application of the recursive Luce rule with quasihyperbolic discounting to the purchase of a durable good and we study the elasticity of demand with respect to a transitory or a permanent price shock. We also discuss a possible extension of the model that relaxes the IIA axiom and retains the stationarity properties. Allowing the dependence of the value function on a parameter, as in the mixed logit model, the discounted mixed Luce rule has the same stationary properties of the Luce rule and can accommodate realistic substitution patterns among elements that are excluded by the IIA. Lastly, we study the elasticity and crosselasticity of choice probability when an element of a consumption stream varies. We find that choice probabilities of the discounted Luce rule contain relevant information concerning elasticities and cross-elasticities. For example, the probability of choosing a consumption stream in a set is inversely related to the sum of the elasticities of its components. The paper is organized as follows: after a review of the relevant literature, Section 2introduces the shortcoming of using the discounted multinomial logit. In Section 3we provide an axiomatization of the discounted Luce rule when choices are over dated outcomes. We then relax stationarity to pin down the observable restrictions of hyperbolic and quasi-hyperbolic discounting. In Section 4we extend the model to consumption streams. Elasticity and cross-elasticity of choice probabilities is studied in Section 5. Section 6discusses an extension of the model to recursive stochastic choice. Section 7proposes a version of the model that relaxes the independence of irrelevant alternatives assumption. 1.1 Related literature In the static setting, the multinomial logit and the Luce choice rule are equivalent. The latter has been introduced by Luce (1959). Recent works provided various foundations for the static multinomial logit based on bounded rationality (Mattsson and 5
Weibull,2002), rational inattention (Matêjka and McKay,2015) and neuroscientific models of choice (Webb,2015). Two recent models generalize the Luce rule to account for violations of the IIA: Gul et al. (2014) proposed a model similar to the nested logit, Echenique et al. (2014) included perception priorities. The axioms we introduce in the present work can be extended to their model, since stationarity and the IIA are not related. Concerning the dynamic setting, Fudenberg and Strzalecki (2015) axiomatized a general version of the recursive stochastic choice model in which larger menus may be disliked due to choice aversion. The individual stochastically chooses at each time an action and a continuation menu. The dynamic logit, widely used in applied works4(Rust,1987;Hendel and Nevo,2006;Gowrisankaran and Rysman, 2012;Chen et al.,2013), is a particular case of their model. The aim of their paper is different from ours since, we are interested in the effect of discount on choice probabilities and we focus on "static stochastic choice" over consumption streams. They are interested in the dynamics of stochastic choice. However, the notion of stationarity they use is comparable to ours. We show that it is weaker and it cannot distinguish the discounted logit from the discounted Luce rule (see Section 4). The axiomatic characterization of the quasi-hyperbolic Luce rule of Theorem 3can be used to characterize (for example, adapting the axioms of Fudenberg and Strzalecki (2015)) a recursive model of stochastic choice, that allows for quasi-hyperbolic discounting. Such model is receiving an increasing attention in applied works (e.g. Paserman,2008;Tarozzi and Mahajan,2011;Yao et al.,2012;An et al.,2014;Fang and Wang,2015). The interaction of discounting and stochastic choice has been understudied so far. Recently, Lu and Saito (2016) introduced a model where stochastic choice follows uncertainty about the discount function. They characterize geometric and quasi-hyperbolic discounting. Concerning critiques to the multinomial logit, Fosgerau and Bierlaire (2009) proposed a random utility model with multiplicative error that solves the scale problem. A more general critique of the use of stochastic choice models in the study of risk aversion and time preference comes from Apesteguia and Ballester (2015). They show that a large class of models including the logit is not monotone with respect to parameters measuring risk aversion or impatience. In other words, an increase in the risk aversion or impatience parameter is not necessarily followed by a larger probability of selecting a 4See (Aguirregabiria and Mira,2010) for a review. 6
less risky or more impatient option. Our critique of the multinomial logit is different since it is not related to parametric restrictions. Moreover, the model we introduce in Section 7belongs to the class of random parameter models proposed by Apesteguia and Ballester (2015) as a solution to the monotonicity problem. In general, both papers highlight substantial flaws of the logit model that should be taken into account in applied works. 2 The "problem of the scale" and its consequences Consider a multinomial logit model of choice from a general set Aand two "value functions" uand u0with u0=ku for some k≥0, the "problem of the scale" (Fosgerau and Bierlaire,2009) comes from the following inequality: Pu0 Logit(a, A) = eku(a) X b∈A eku(b)6=eu(a) X b∈A eu(b)=Pu Logit(a, A) The choice probabilities are sensitive to the scale of the value function. The problem is particularly relevant when we consider the discounted logit. Assume that the elements of Aare dated outcomes and the individual has to choose today between a dated outcome (x, t), meaning xdelivered/payed at time t, and an alternative (y, s)and that u((x, t)) = δtv(x). For simplicity, if A={(x, t),(y, s)}, we write P((x, t), A) = P((x, t)≥(y, s)). The scale problem produces the following: PLogit((x, t)≥(y, s)) = eδtv(x) eδsv(y)+eδtv(x) and when both payments are delayed by r > 0periods, PLogit((x, t +r)≥(y, s +r)) = eδt+rv(x) eδs+rv(y)+eδt+rv(x) It is easy to see that PLogit((x, t)≥(y, s)) 6=PLogit((x, t +r)≥(y, s +r)) 7
(SSA). For all A∈ A and t, s, r ≥0, P((x, t), A) P((y, s), A)=P((x, t +r), Ar) P((y, s +r), Ar) The relative probability of choosing xat time tover yat time s, does not vary when both payments are equally delayed. The condition holds for the sets Arand says nothing concerning the interaction with dated outcomes that can be added or subtracted to the set. In other words, without assuming IIA, the relative probability of choosing (x, t)over (y, s)from A={(x, t),(y, s)}, can be different from the probability of choosing (x, t+r)over (y, s+r)from B={(x, t +r),(y, s +r),(z, q +r)}. So the SSA does not restrict possible interactions among outcomes. Consider now the following Stationary version of the IIA: (SIIA). For all A, B ∈ A and t, s, r ≥0, P((x, t), A) P((y, s), A)=P((x, t +r), B) P((y, s +r), B) The SIIA axiom implies the IIA axiom for r= 0 and for B=Arit implies the SSA. Next lemma shows the opposite implications: Lemma 1. The SIIA holds, if and only if, the IIA and the Stochastic Stationarity Axiom hold. Assuming the SIIA is equivalent to assume both stochastic stationarity and the IIA axiom. In this work, we will retain the IIA and we will study the role of stationarity and the consequences of its weakening. Section 7discusses a simple extension of the discounted Luce rule that relaxes the IIA. The next condition imposes a stochastic form of impatience: if two rewards have the same probability of being selected when one is payed later, the equality is broken in favour of the latter when both are delivered at the same date. (Stochastic Impatience). For all x, y ∈Zand t≥0, if P((x, t), A) = P((y, t + 1), A) then P((x, t), B)≤P((y, t), B). The next theorem characterizes the discounted Luce rule: 14
Theorem 1. The SIIA axiom and Stochastic Impatience hold, if and only if, choice probabilities are represented by the discounted Luce rule: P((x, t), A) = δtv(x) X (y,s)∈A δsv(y) for some random scale v:Z→R++ and δ∈(0,1]. The IIA axiom, implied by the SIIA, gives choice probabilities the Luce’s relative weight form. The Stochastic Stationarity part of the SIIA imposes separable and geometric discounting. Lastly, Stochastic Monotonicity allows us to interpret δas a discount factor. Due to the great amount of empirical and theoretical research in nongeometric discounting, the rest of the section focuses on violations of the SSA axiom, while maintaining IIA. Relaxing the latter and maintaining stationarity represents an interesting line for future research. As a final note, we show how to elicit time preferences with the discounted Luce rule. Consider the following ratios P((x,t+1),A) P((y,0),A)and P((y,0),B) P((x,t),B)for some t≥0and x, y ∈Z. It is immediate to see that P((x,t+1),A) P((y,0),A)·P((y,0),B) P((x,t),B)=δ. The discount factor δ can be inferred directly from choice probabilities. 3.1 Implications of non-geometric discounting Geometric discounting of future rewards is normatively plausible, but it is often challenged by the experimental evidence of diminishing impatience (for example, Thaler, 1981). Well-known alternatives are the quasi-hyperbolic discounting of Laibson (1997) and the hyperbolic discounting of Prelec (2004). Deviations from geometric discounting necessarily induce violations of the Stochastic Stationarity Axiom, for example, let consider the quasi hyperbolic discounting model of Laibson (1997), 1, βδ, βδ2, βδ3. . .. for some β∈[0,1) and δ∈(0,1]. With quasi-hyperbolic the trade-off between consumption in two consecutive periods is maximum at the present. It is plausible to imagine that, for a general stochastic choice rule, quasi-hyperbolic discounting implies the following violation of the SSA: Prob((x, 0), A) Prob((y, 1), A)>Prob((x, t), At) Prob((y, t + 1), At)(3) 15
the relative probability of choosing xnow over ytomorrow decreases when both outcomes are equally delayed. This is the stochastic counterpart of the present bias. The discounted Luce rule of Theorem 1cannot accommodate the previous inequality. Therefore, we introduce the general discounted Luce rule. Let define a discount function D(t), as a decreasing function D:N+→(0,1], with D(0) = 1 and limt→∞ D(t)=0then, the choice probabilities according to the generalized discounted Luce rule are: P((x, t), A) = D(t)v(x) X (y,s)∈A D(s)v(y) The following axiom is the required relaxation of the SSA that allows for general discounting of future consequences. (Weak SSA). For all A∈ A and all t, r ≥0, P((x, t), A) P((y, t), A)=P((x, t +r), Ar) P((y, t +r), Ar) It imposes invariance of the relative probability of choosing between two payoffs, only when they are payed at the same date. Intuitively, this ratio is not influenced by intertemporal trade-offs, since both outcomes are delivered on the same date. Then we have the following result: Proposition 1. The IIA axiom, Weak SSA and Stochastic Impatience hold, if and only if, choice probabilities are represented by a generalized discounted Luce rule., i.e. P((x, t), A) = D(t)v(x) X (y,s)∈A D(s)v(y) for some random scale v:Z→R++ and discount function D:{0, . . . , T} → (0,1]. We imposed the static IIA axiom to give probabilities a simple structure and we relax stochastic stationarity to allow non-geometric discounting. The result is a flexible rule that accommodates common discount functions, such as the hyperbolic and the quasi-hyperbolic. Violations of the SSA can be related to the degree of impatience of D(t), defined as I(t) = D(t) D(t+ 1) 16
we say that D(t)exhibits the present bias if I(0) > I(t)for all t > 0. We say that D(t)exhibits strict diminishing impatience if I(t)> I(t+1) for all t. Quasi-hyperbolic discounting, D(0) = 1,D(t) = βδt, exhibits the present bias and does not exhibit strict diminishing impatience. Hyperbolic discounting, D(t) = 1 1+kt , exhibits both. Then we have the following simple consequences: •P((x,0),A) P((y,1),A)>P((x,t),At) P((y,t+1),At),∀t > 0, if and only if, D(t)exhibits the Present Bias. •P((x,t−1),At−1) P((y,t),At−1)>P((x,t),At) P((y,t+1),At),∀t > 0, if and only if, D(t)exhibits strict Diminishing Impatience. With DI, the relative probability of choosing an earlier over a later payoff is always greater than the same probability when both are delayed by an additional period. One may be interested in distinguishing quasi-hyperbolic of Laibson (1997) from the general discount function D(t). The next axiom contains the required restrictions: (Quasi-hyperbolic SSA). For all A∈ A: 1. (Delayed SSA): P((x,t),A) P((y,s),A)=P((x,t+r),Ar) P((y,s+r),Ar), for all t, s > 0, r ≥0. 2. (Present Bias): P((x,0),A) P((y,t),A)≥P((x,r),A) P((y,t+r),Ar), for all t≥0,r > 0. 3. (Invariance): For all x, y ∈Z,P((x,0),A) P((y,0),A)=P((x,t),At) P((y,t),At). All the intuitive features of the quasi-hyperbolic discounting affect relative choice probabilities. For non-immediate outcomes, the relative choice probabilities are constant when outcomes are equally delayed, since the (delayed) SSA holds. However, the relative probability of choosing an immediate payment over a delayed one is strictly greater than the same proportion when both payments are equally delayed, this is the present bias. The last part, Present Weak SSA, imposes equality of relative probability only when the outcomes are payed at the same date (it is a weakening of the weak SSA). All the restrictions of the Quasi-hyperbolic SSA are easily observable in laboratory or fields experiments. Then we have: Proposition 2. The IIA axiom, Quasi-hyperbolic SSA and Stochastic Impatience hold, if and only if, there exists a positive ratio scale v:Z→R++ and a discount function D(t)such that: P((x, t), A) = D(t)v(x) X (y,s)∈A D(s)v(y) 17
and D(t) = βδtif t > 0and D(0) = 1, for some β, δ ∈(0,1]. 4 Consumption streams The dated outcomes setting is quite restrictive and cannot be compared with axiomatizations of additively separable discounted utility with geometric discounting (Koopmans,1960;Fishburn,1970) or quasi-hyperbolic discounting (Hayashi,2003; Montiel Olea and Strzalecki,2014). In this section we fill the gap extending to finite consumption streams the intuitions of the previous section. For T > 0, let ZT+1 =Z×Z×· · · Zbe a T+1 product of a finite set of alternatives. An element of ZT+1 represents a consumption stream x= (x0, x1, . . . , xT). A choice set is an element of A= 2ZT+1 \{∅}. Given a set A, the discounted Luce rule probability of choosing x∈Ais given by P(x, A) = T X t=0 δtv(xt) X y∈A T X t=0 δtv(yt) for some δ∈(0,1]. The probability of selecting a given consumption stream is given by its relative weight in the choice set. The weight is a discounted sum of the values of its components. For an arbitrary x∈Z, denote x(t)the consumption stream x(t)=(z, z, . . . , x, z, . . . , z), i.e. xis payed at tand zotherwise. The next condition postulates the existence of special z∈Z: (Separability). There exists z∈Zsuch that, for all x,y∈ZT+1, with x,y∈A, y,xt(t)∈Bfor all t≥0and with P(y, A)>0,P(y, B)>0implies, P(x, A) P(y, A)=PT t=0 P(xt(t), B) P(y, B) Separability implies that the relative probability of choosing xfrom a menu Awhen yis available, is equal to the sum of the probabilities of choosing its "components" xt(t) = (z, z, . . . , xt, z, . . . , z), relative to y. To gain intuition, assume T= 1,x= (x0, x1),y= (y0, y1),A={(x0, x1),(y0, y1)}and B={(x0, z),(z, x1),(y0, y1)}. Then 18
Separability implies P((x0, x1), A) P((y0, y1), A)=P((x0, z), B) + P((z, x1), B) P((y0, y1), B) Suggesting that the probability of selecting a consumption stream can be decomposed in the probability of selecting its components. Separability is the stochastic choice counterpart of additive separability in deterministic choice. Let z= (z, z, . . . , z), the first consequence of Separability is the following: Lemma 2. For all z∈Zsatisfying Separability, P(z, A) = 0 for all A∈ A with x,z∈Aand x6=z. The elements zcan be interpreted as being "nothing" and the probability of selecting them is zero whenever there is an alternative. Given the existence of special z∈Z, we can turn to the stationarity properties of the discounted Luce rule in the consumption stream setting. For each A∈ A, we define A+1 ={(z, x) : x∈A}, where the notation (z, x)indicates (z, x)=(z, x0, . . . , xT−1). Differently, for x∈ZT+1, (x, z)=(x0, x1, . . . , xT−1, z)and (z, x, z)=(z, x0, x1, . . . , xT−2, z). As for the interpretation, (z, x)is a "shift forward" of (x, z). Following the intuitions of the dated outcome setting, we consider a choice rule to be stationary if (assume the probabilities at the denominator are strictly positive): P((z, x), A) P((z, x0), A)=P((x, z), A+1) P((x0, z), A+1)(4) The relative probability of choosing (z, x)over (z, y)remains unchanged after a shift forward of both. The property resembles the definition of stationarity of deterministic choice (see Fishburn,1970, Def. 7.3). Equation (4) holds for the discounted Luce rule if v(z)=0, indeed: P((z, x), A) P((z, x0), A)=v(z) + δv(x0) + δ2v(x1) + PT t=3 δtv(xt−1) v(z) + δv(x0 0) + δ2v(x0 1) + PT t=3 δtv(x0 t−1)=P((x, z), A+1) P((x0, z), A+1) but this follows from Lemma 2. Therefore, a discounted Luce rule satisfies the previous equality. Consider now the discounted multinomial logit. The probability of choosing 19
x∈Ais: PLogit(x, A) = exp T X t=0 δtv(xt)! X y∈A exp T X t=0 δtv(yt)! and the equality in Eq. (4) does not hold for the discounted logit: PLogit((z, x), A) PLogit((z, x0), A)=exp(v(z) + PT t=1 δtv(xt−1)) exp(v(z) + PT t=1 δtv(x0 t−1)) 6=exp(PT−1 t=0 δtv(xt) + δTv(z)) exp(PT−1 t=0 δtv(x0 t) + δTv(z)) =PLogit((x, z), A+1) PLogit((x0, z), A+1) Therefore, according to the discounted logit, the relative probability of choosing a consumption stream xover yin a set A, changes when all the elements in Aare "shifted" by one period. This form of stationarity is then able to tell apart the two models. We call it Stochastic Fishburn Stationarity: (SFS). For all x,x0∈ZT+1,A∈ A with (x, z),(x0, z)∈Aand P((x0, z), A)>0, P((z, x0), A+1)>0: P((x, z), A) P((x0, z), A)=P((z, x), A+1) P((z, x0), A+1) Fudenberg and Strzalecki (2015) proposed an alternative notion, called stream stationarity, that imposes the following: P((x, y), A)≥P((x0, y), A)⇐⇒ P((y, x), B)≥P((y, x0), B) Stream stationarity is too weak to distinguish the discounted logit and the discounted Luce rule, since it is satisfied by both models. Separability and the SFS axiom are the main innovations of the section, the following axioms are standard. Since we deal with zero probability events, we use a more general axiom than the IIA.10 10GIIA is equivalent to the original Luce choice axiom (see Luce,1959), which is more general than the IIA. 20
(GIIA). For all A∈ A and x∈A, there exists u:ZT+1 →R+, such that: P(x, A) = u(x) X y∈A u(y) Lastly, we define a stochastic notion of impatience similar to the one in the delayed rewards setting: (Stochastic Impatience). For all x, y 6=z∈Zand t≥0, if P(x(t), A) = P(y(t+1), A), then P(x(t), B)≤P(y(t), B). The next theorem characterizes the discounted Luce rule: Theorem 2. The GIIA axiom, SFS, Separability and Stochastic Impatience hold, if and only if, choice probabilities are represented by a discounted Luce rule, i.e. P(x, A) = T X t=0 δtv(xt) X y∈A T X t=0 δtv(yt) for a ratio scale v:Z→R+and δ∈(0,1]. Also in this setting, we are interested in relaxing stationarity to determine the observable restrictions following quasi-hyperbolic discounting. Indeed, although v(z) = 0, quasi-hyperbolic discounting violates SFS. Let, P(x, A) = v(x0) + β T X t=1 δtv(xt) X y∈A"v(y0) + β T X t=1 δtv(yt)# 21
then SFS is violated: P((x, z), A) P(x0, z), A)= v(x0) + β T−1 X t=1 δtv(xt) + βδTv(z) v(x0 0) + β T−1 X t=1 δtv(x0 t) + βδTv(z) 6= v(z) + β T X t=1 δtv(xt−1) v(z) + β T−1 X t=1 δtv(x0 t−1) = T X t=1 δtv(xt−1) T−1 X t=1 δtv(x0 t−1) =P((z, x), A+1) P((z, x0), A+1) The following relaxation of SFS is parallel to that introduced in the delayed rewards setting and includes all the intuitive features of a present-biased Luce rule. They are comparable with the axioms characterizing quasi-hyperbolic discounting in deterministic choice (e.g. Montiel Olea and Strzalecki,2014). It contains three properties: Delayed Stochastic Fishburn Stationarity, Present Bias and Invariance. (Quasi-hyperbolic Stationarity). 1. (DSFS): For all A∈ A, with (z, x),(z, x0)∈A, if P((z, x0, z), A)>0and P((z, z, x0), A+1)>0: P((z, x, z), A) P((z, x0, z), A)=P((z, z, x), A+1) P((z, z, x0), A+1) 2. (PB): For all A∈ A, with x,(z, x)∈A, if P((z, x, z), A+1)>0and P((z, z, x0), A+1)>0: P((x, z, z), A) P((z, x, z), A+1)≥P((z, x0, z), A) P((z, z, x0), A+1) 3. (Invariance): For all x, y ∈Z\ {z},t > 0and all A∈ A, if P(y(0), A)>0 and P(y(1), A+1)>0: P(x(0), A) P(y(0), A)=P(x(1), A+1) P(y(1), A+1) 22
The DSFS imposes stationarity to non-immediate shifts. The present bias imposes a greater probability of choosing a stream of consumption when the immediate outcome has a positive value. Invariance excludes variations of the relative likelihood of choosing two outcomes when they are payed at the same date. Theorem 3. The GIIA axiom, Quasi-hyperbolic Stationarity, Separability and Stochastic Impatience hold, if and only if, choice probabilities are represented by a quasihyperbolic discounted Luce rule, i.e. P(x, A) = v(x0) + β T X t=1 δtv(xt) X y∈A"v(y0) + β T X t=1 δtv(yt)# The axioms characterizing the quasi-hyperbolic Luce rule can be used to extend the recursive axiomatization of Fudenberg and Strzalecki (2015) to a dynamic discrete choice model that accounts for quasi-hyperbolic discounting (see Section 6for a discussion). Such model has recently gained attention in applied works (Paserman,2008; Tarozzi and Mahajan,2011;Yao et al.,2012;An et al.,2014;Fang and Wang,2015) but, its behavioral restrictions are not yet understood. 5 Elasticity and cross-elasticity Elasticity measures how a variation in an observable factor, for example a component of the consumption stream xt, affects choice probabilities (Train,2009). The elasticities of the discounted Luce rule have similar properties to the elasticities of the logit, although they inherit the scale-free property of the model. Fosgerau and Bierlaire (2009) provides calculations that are valid for a Luce rule. Since we are studying the interaction of intertemporal preferences and stochastic choice, we will exploit the structure of the discounted Luce rule to study elasticity. We are interested in the elasticity of the probability of choosing xwhen one element xtof xchanges. Formally: E[P(x, A); xt],∂P(x, A) ∂xt xt P(x, A) 23
the Delayed SSA, P((x, t), A) P((y, s), A)=P((x, t +r), Ar) P((y, s +r), Ar)=P((x, t +r), B) P((y, s +r), B) is equivalent to u(x, t) u(y, s)=u(x, t +r) u(y, s +r) for some ratio scale u:X→R++. Let x=y,t= 2 and s= 1, then u(x, 2) u(x, 1) =u(x, r + 2) u(x, r + 1) hence u(x, r + 2) = u(x,2) u(x,1)u(x, r + 1), going back until r= 0 implies u(x, r + 2) = u(x,2) u(x,1) r+1u(x, 1), now, let δx=u(x,2) u(x,1), by S. Impatience δx∈(0,1], then u(x, t) = δt−1 xu(x, 1). By Delayed SSA with t, s = 1 and r= 1, P((x, 1), A) P((y, 1), A)=P((x, 2), A1) P((y, 2), A1)=⇒u(y, 2) u(y, 1) =u(x, 2) u(x, 1) for all x, y ∈Z, hence δx=δyfor all x, y ∈Z. For t= 1, r = 1, Present Bias implies u(x, 0) u(y, 1) ≥u(x, 1) u(y, 2) or u(x,1) u(x,0) ≤u(y,2) u(y,1) =δhence u(x, 0)δ≥u(x, 1), then for some βx≤1,βxδu(x, 0) = u(x, 1), then we have u(x, t) = βxδtu(x, 0) for t > 0. By Present Weak SSA and IIA, P((x,0),A) P((y,0),A)=P((x,t),B) P((y,t),B)implies u(x,0) u(y,0) =βxδtu(x,0) βyδtu(y,0) or βx βy= 1, hence βx=βy, since it is true for arbitrary x, y, the result follows, defining v(x) = u(x, 0) and D(t) = βδtfor t > 0and D(0) = 1. Proof. Of Lemma 2. Take A=B, by Separability, P(z, A) = PT t=0 P(zt(t), A)for some Acontaining z, but zt(t) = zfor all 0≤t≤T, hence, P(z, A)=(T+ 1)P(z, A) and this can be true only if P(z, A) = 0 since T > 0. In turn, P(z, A) = 0 implies u(z, z, . . . , z)=0since zis unique. Proof. Of Theorem 2. By the GIIA, there exists a random scale u:ZT+1 →R+ such that P(x, A) = u(x0,x1,...,xT) Py∈Au(y0,y1,...,yT). For an arbitrary x∈Z, let define v(x) = u(x, z, z, . . .) = u(x(0)) for some zsatisfying Separability, moreover v(z)=0for all z∈Z0, where Z0={z∈Z:zsatisfies Separability}. For an arbitrary x∈Z\Z0, let define δx=u(x(1)) u(x(0)) . By SFS and x, y ∈Z, P(x(1), A+1) P(y(1), A+1)=P(x(0), A) P(y(0), A) or equivalently, P(x(1), A+1) P(x(0), A)=P(y(1), A+1) P(y(0), A)=⇒u(x(1)) u(x(0)) =u(y(1)) u(y(0)) and this implies δx=δy=δfor all x, y ∈Zand define δz=δfor all z∈Z0. SFS implies P(x(1), A) P(x(0), A)=P(x(2), A+1) P(x(1), A+1) 30
another application of SFS implies P(x(2), A+1) P(x(1), A+1)=P(x(3),(A+1)+1) P(x(2),(A+1)+1) equivalently, u(x(1)) u(x(0)) =u(x(2)) u(x(1)) =u(x(3)) u(x(2)) By Impatience δ≤1. The first equality implies u(x(2)) = u(x(1))·u(x(1)) u(x(0)) and the second, u(x(3)) = u(x(2)) ·u(x(1)) u(x(0)), and together u(x(3)) = u(x(1)) ·u(x(1)) u(x(0))2, repeating the same argument gives u(x(t)) = u(x(1)) ·u(x(1)) u(x(0))t−1 multiplying and dividing by u(x(0)), gives u(x(t)) = u(x(0)) ·u(x(1)) u(x(0))t (5) in our notation this becomes u(x(t)) = δtv(x). To conclude, we need to prove that u(x0, x1, x2, . . .) = P∞ t=0 δtv(xt). To see this, consider P((x0, x1, . . . , xT), A)for some A, Separability implies u(x0, x1, . . . , xT) u(y0, y1, . . . , yT)=PT t=0 u(xt(t)) u(y0, y1, . . . , yT) then, by Eq. (5) u(x0, x1, . . . , xT) = T X t=0 u(xt(t)) = T X t=0 δtv(xt) Proof. Of Theorem 3. By the GIIA, there exists a random scale u:ZT+1 →Rsuch that P(x, A) = u(x0,x1,...,xT) Py∈Au(y0,y1,...,yT). For a given x∈Z, let define v(x) = u(x, z, z, . . .) for some z∈Z0. For an arbitrary x∈Z\Z0, let define δx=u(x(2)) u(x(1)) By QSFS, P(x(2), A+1) P(y(2), A+1)=P(x(1), A) P(y(1), A) or equivalently, P(x(2), A+1) P(x(1), A)=P(y(2), A+1) P(y(1), A)=⇒u(x(2)) u(x(1)) =u(y(2)) u(y(1)) and this implies δx=δy=δfor all x, y ∈Z\Z0and define δz=δfor all z∈Z0. QSFS implies P(x(2), A) P(x(1), A)=P(x(3), A+1) P(x(2), A+1)) another application of QSFS implies P(x(3), A+1) P(x(2), A+1)) =P(x(4),(A+1)+1) P(x(3),(A+1)+1) 31
equivalently, u(x(2)) u(x(1)) =u(x(3)) u(x(2)) =u(x(4)) u(x(3)) The first equality implies u(x(3)) = u(x(2))·u(x(2)) u(x(1)) and the second, u(x(4)) = u(x(3))· u(x(2)) u(x(1)) , and together u(x(4)) = u(x(2)) ·u(x(2)) u(x(1))2, repeating the same argument for an arbitrary t > 0, gives u(x(t)) = u(x(2)) ·u(x(2)) u(x(1))t−2 multiplying and dividing by u(x(1)), gives u(x(t)) = u(x(1)) ·u(x(2)) u(x(1))t−1 (6) By PB. with x0=y(0) and x=x(0): u(x(0)) u(x(1)) ≥u(y(1)) u(y(2)) hence u(x(1)) u(x(0)) ≤u(y(2)) u(y(1)) =δthen, there exists βx∈[0,1] such that u(x(0)) = βxδu(x(1)) Multiplying and dividing Eq. (6) by u(x(0)) and using the last fact gives u(x(t)) = u(x(0)) ·βxu(x(2)) u(x(1))t for all t > 0. In our notation u(x(t)) = βxδtv(x)for t > 0and u(x(0)) = v(x). By Invariance, P(x(1), A+1) P(x(0), A)=P(y(1), A+1) P(y(0), A) and, u(x(1)) u(x(0)) =u(y(1)) u(y(0)) ⇐⇒ v(x)βxδ v(x)=v(y)βyδ v(y) that implies βx=βyfor all x, y ∈Z\Z0. To conclude, we need to prove that u(x0, x1, x2, . . .) = v(x0) + βPT t=1 δtv(xt). To see this, consider P((x0, x1, . . . , xT), A) for some Athat satisfies the condition, then Separability implies u(x0, x1, . . . , xT) Py∈Au(y0, y1, . . . , yT)= T X t=0 u(xt(t)) Py∈Au(y0, y1, . . .) then, u(x0, x1, . . . , xT) = v(x0) + β T X t=1 δtv(xt) References Aguirregabiria, V. and P. Mira (2010). Dynamic discrete choice structural models: A survey. Journal of Econometrics 156(1), 38–67. 32
An, Y., Y. Hu, and J. Ni (2014). Dynamic models with unobserved state variables and heterogeneity: time inconsistency in drug compliance. Technical report. Andersen, S., G. W. Harrison, M. I. Lau, and E. E. Rutström (2008). Eliciting risk and time preferences. Econometrica 76(3), 583–618. Apesteguia, J. and M. A. Ballester (2015). Monotone stochastic choice models: The case of risk and time preferences. Economics Working Papers 1499, Department of Economics and Business, Universitat Pompeu Fabra. Chabris, C. F., D. Laibson, C. L. Morris, J. P. Schuldt, and D. Taubinsky (2008). Individual laboratory-measured discount rates predict field behavior. Journal of Risk and Uncertainty 37(2-3), 237–269. Chen, J., S. Esteban, and M. Shum (2013). When do secondary markets harm firms? American Economic Review 103(7), 2911–34. Chevalier, J. and A. Goolsbee (2009). Are durable goods consumers forward-looking? Evidence from college textbooks. The Quarterly Journal of Economics 124(4), 1853–1884. Coller, M. and M. B. Williams (1999). Eliciting individual discount rates. Experimental Economics 2(2), 107–127. Dubé, J., G. J. Hitsch, and P. Jindal (2014). The joint identification of utility and discount functions from stated choice data: An application to durable goods adoption. Quantitative Marketing and Economics 12(4), 331–377. Echenique, F., K. Saito, and G. Tserenjigmid (2014). The perception-adjusted Luce model. Technical report. Fang, H. and Y. Wang (2015). Estimating dynamic discrete choice models with hyperbolic discounting, with an application to mammography decisions. International Economic Review 56(2), 565–596. Fishburn, P. C. (1970). Utility theory for decision making. New York: Wiley. Fishburn, P. C. and A. Rubinstein (1982). Time preference. International Economic Review 23(3), 677–694. Fosgerau, M. and M. Bierlaire (2009). Discrete choice models with multiplicative error terms. Transportation Research Part B: Methodological 43(5), 494–505. Fudenberg, D. and T. Strzalecki (2015). Dynamic logit with choice aversion. Econometrica 83, 651–691. Giglio, S., M. Maggiori, and J. Stroebel (2015). Very long-run discount rates. The Quarterly Journal of Economics 130(1), 1–53. Gowrisankaran, G. and M. Rysman (2012). Dynamics of consumer demand for new durable goods. Journal of Political Economy 120(6), 1173–1219. Gul, F., P. Natenzon, and W. Pesendorfer (2014). Random choice as behavioral optimization. Econometrica 82(5), 1873–1912. Halevy, Y. (2015). Time consistency: stationarity and time invariance. Econometrica 83(1), 335–352. 33
Harrison, G. W., M. I. Lau, and M. B. Williams (2002). Estimating individual discount rates in Denmark: A field experiment. American Economic Review 92(5), 1606– 1617. Hayashi, T. (2003). Quasi-stationary cardinal utility and present bias. Journal of Economic Theory 112(2), 343 – 352. Hendel, I. and A. Nevo (2006). Measuring the implications of sales and consumer inventory behavior. Econometrica 74(6), 1637–1673. Koopmans, T. C. (1960). Stationary ordinal utility and impatience. Econometrica 28, 287. Laibson, D. (1997). Golden eggs and hyperbolic discounting. The Quarterly Journal of Economics 112(2), 443–77. Louie, K. and P. W. Glimcher (2010). Separating value from choice: delay discounting activity in the lateral intraparietal area. The Journal of Neuroscience 30(16), 5498– 5507. Lu, J. and K. Saito (2016). Random intertemporal choice. Technical report. Luce, R. D. (1959). Individual choice behavior. New York: Wiley. Marschak, J. (1959). Binary choice constraints on random utility indicators. Cowles Foundation Discussion Papers 74, Cowles Foundation for Research in Economics, Yale University. Matêjka, F. and A. McKay (2015). Rational inattention to discrete choices: A new foundation for the multinomial logit model. American Economic Review 105(1), 272–98. Matêjka, F., J. Steiner, and C. Stewart (2015). Rational inattention dynamics: Inertia and delay in decision-making. CEPR Discussion Papers 10720. Mattsson, L. and J. W. Weibull (2002). Probabilistic choice and procedurally bounded rationality. Games and Economic Behavior 41(1), 61–78. Mattsson, L., J. W. Weibull, and P. O. Lindberg (2014). Extreme values, invariance and choice probabilities. Transportation Research Part B: Methodological 59, 81–95. Montiel Olea, J. L. and T. Strzalecki (2014). Axiomatization and measurement of quasi-hyperbolic discounting. The Quarterly Journal of Economics 129, 1449–1499. Paserman, M. D. (2008). Job search and hyperbolic discounting: Structural estimation and policy evaluation. The Economic Journal 118(531), 1418–1452. Prelec, D. (2004). Decreasing impatience: A criterion for non-stationary time preference and "hyperbolic" discounting. Scandinavian Journal of Economics 106(3), 511–532. Rust, J. (1987). Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher. Econometrica 55(5), 999–1033. Simonson, I. (1989). Choice based on reasons: The case of attraction and compromise effects. Journal of Consumer Research, 158–174. 34
Simonson, I. and A. Tversky (1992). Choice in context: Tradeoff contrast and extremeness aversion. Journal of Marketing Research. Tanaka, T., C. F. Camerer, and Q. Nguyen (2010). Risk and time preferences: linking experimental and household survey data from Vietnam. American Economic Review 100(1), 557–71. Tarozzi, A. and A. Mahajan (2011). Time inconsistency, expectations and technology adoption: The case of insecticide treated nets. Economic Research Initiatives at Duke (ERID) Working Paper (105). Thaler, R. (1981). Some empirical evidence on dynamic inconsistency. Economics Letters 8(3), 201–207. Train, K. E. (2009). Discrete choice methods with simulation. Cambridge university press. Webb, R. (2015). The dynamics of stochastic choice. Available at SSRN 2226018. Yao, S., C. F. Mela, J. Chiang, and Y. Chen (2012). Determining consumers’ discount rates with field studies. Journal of Marketing Research 49(6), 822–841. 35