Moment inequalities for multinomial choice with fixed effects
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Pakes, Ariel; Porter, Jack Article Moment inequalities for multinomial choice with fixed effects Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Pakes, Ariel; Porter, Jack (2024) : Moment inequalities for multinomial choice with fixed effects, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 15, Iss. 1, pp. 1-25, https://doi.org/10.3982/QE1776 This Version is available at: https://hdl.handle.net/10419/296339 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 15 (2024), 1–25 1759-7331/20240001 Moment inequalities for multinomial choice with fixed effects Ariel Pakes Department of Economics, Harvard University and NBER Jack Porter Department of Economics, University of Wisconsin This paper proposes a new approach to identification of the semiparametric multinomial choice model with fixed effects. The framework employed is the semiparametric version of the traditional multinomial logit with the fixed-effects model (Chamberlain (1980)). This semiparametric multinomial choice model places no restrictions on either the joint distribution of the random utility disturbances across choices or their within group (or across time) correlations. We show that a novel within-group comparison leads to a set of conditional moment inequalities. Our main finding shows that the derived conditional moment inequalities yield the sharp identified set for the random utility covariate index, while avoiding the incidental parameter problem. Specializing this result to the binary choice case shows that Manski (1987)’s conditional moment inequalities still lead to sharp bounds without restrictions on covariates. Keywords. Discrete choice, panel data, fixed effects, sharp partial identification. JEL classification. C14, C23, C25. 1. Introduction This paper characterizes identification of the semiparametric multinomial choice model with fixed effects and a group (or panel) structure. A standard multinomial framework (McFadden (1974)) is employed with random utility that is additively separable between unobservables, which include a disturbance and choice-specific fixed effects, and a covariate index function. The key semiparametric assumption, replacing the multinomial logit specification (Chamberlain (1980)), is a familiar group stationary condition on the disturbances. This assumption places no restrictions on either the joint distribution of the disturbances across choices or the correlation of disturbances across time (or within group). Under this specification, a novel within-group comparison leads to a set of conditional moment inequalities, which are the basis for our result on sharp partial identification. Our main finding establishes sharp nonparametric identification of the covariate index function in the semiparametric multinomial choice model with fixed effects. Under Ariel Pakes: [email protected] Jack Porter: [email protected] We thank two anonymous referees and discussions with many seminar participants. Mark Shepard and the Massachusetts Health Connector generously shared data. ©2024 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1776
2Pakes and Porter Quantitative Economics 15 (2024) the group stationarity assumption alone, we find that our full set of derived conditional moment inequalities contains all of the model’s potential identifying information in the sense that the bounds provided by these inequalities are sharp. Sharpness is shown by a constructive proof. In particular, given a distribution of observables and a parameter value in the identified set, we demonstrate that there exists a distribution of unobservables that can be combined with the parameter value to generate the given distribution of observables.1 The semiparametric model considered here does not place parametric restrictions on the disturbance distribution. The only restriction on the disturbances is a group or time stationary assumption. Since the joint distribution of disturbances across choices is left unrestricted, the model contains no vestiges of independence of irrelevant alternatives or limits on cross price elasticities. Within group or across time disturbance correlation is also left completely unrestricted in this specification. The panel aspect of the model allows for an additive choice-specific fixed effect in the random utility specification. The fixed effects are allowed to be arbitrarily correlated with the observed covariates. We focus on the case with only two time periods (or group observations). The derived conditional moment inequalities are based on only within variation and in this sense is the discrete choice analog to the familiar within transformation in linear models. As a result, the incidental parameter problem is fully circumvented in our identification results. The most closely related work is Shi, Shum, and Song (2018), which obtains point identification with a linear covariate index using cyclic monotonicity in the semiparametric multinomial setup. Since our conditional moment inequalities provide sharp bounds, it is not surprising that we are able to show that the Shi, Shum, and Song (2018) conditional moment inequalities are implied by our conditional moment inequalities. It follows that, under the additional conditions on covariates given in Shi, Shum, and Song (2018), our conditional moment inequalities also yield point identification in the linear covariate index case. When we specialize our setup to the binary choice case with a linear covariate index, we find that our conditional moment inequalities match the weak version of Manski (1987)’s conditional moment inequalities. It follows that these conditional moment inequalities yield sharp bounds even when point identification fails due to either insufficient variation in the covariates or nonlinearities in the covariate index functions. Further, we establish that Manski (1987)’s maximum score criterion can be derived as an aggregation of the conditional moments that make up our inequalities and extended to allow for nonlinear covariate indices. We prove that the identified set determined by our conditional moment inequalities is exactly the set of parameters that maximize the maximum score criterion function. This new result shows that the maximum score criterion can be used for (sharp) identification (and hence its sample counterpart can be used for estimation) even when point identification fails in the binary choice panel model. 1The distribution of unobservables can be expressed as a nonnegative solution to a system of linear equations, so existence is constructive in the sense that a solution can then be determined in a finite number of matrix manipulation steps.
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 3 Multinomial discrete choice models are extensively used in almost all fields that empirically analyze the determinants of agents’ choices. Applications have typically employed parametric forms of the multinomial model. With panel data problems in mind, Chamberlain (1980) uses an assumption of logistic disturbances to provide a novel conditional likelihood method of identification and estimation. An alternative application is in the demand literature where markets are the grouping device, the within group observations are consumers, and the choice-specific fixed effects represent product level unobservables (e.g., Berry, Levinsohn, and Pakes (1995)). Markets are also used as a grouping device when analyzing firm decision making (e.g., entry decisions) with the market-specific fixed effect representing unobserved determinants of the market’s profitability (e.g., Pakes (2014)). We apply the findings of this paper to consider whether the price sensitivity of demand for health insurance depends on income in a subsidized health insurance market for low-income consumers. The data come from the Commonwealth Care program in Massachusetts, which allowed consumers to choose among competing private plans. We compare the price sensitivity of consumers with incomes one to two times the Federal Poverty Level with those whose income is two to three times the Federal Poverty Level. Implementing the conditional moment inequalities derived here via Andrews and Soares (2010)andAndrews and Shi (2013), we obtain a confidence set for the ratio of the price coefficients in the two income groups and reject the hypothesis of no income dependence at the 5% level. Manski (1975) introduced a semiparametric, maximum score approach to point identification and estimation for multinomial choice without choice-specific fixed effects. Assuming independent and identical distributions of the unobservable components of the different choices, Manski uses differences in the observable, parametric component of random utility across choices for identification. Using Manski’s identification approach, Fox (2007) shows that exchangeability of the unobservable component across choices is sufficient for identification, and Yan (2013) obtains the limiting distribution for a smoothed version of the multinomial maximum score estimator. Lee (1995) provides an alternative semiparametric approach to multinomial choice for models without choice-specific fixed effects using an assumption of an i.i.d. distribution of disturbances across agents. Rather than imposing conditions on the joint distribution of the disturbances across choices, our approach requires that the joint distribution of the choice-specific unobservables does not differ across observations in a group, but leaves the distribution of disturbances across choices unrestricted. The different assumptions are likely to be useful in different applications. Kahn, Ouyang, and Tamer (2019) develop an approach to identification that can be used in both static and dynamic semiparametric multinomial choice models. Chesher, Rosen, and Smolinski (2013)andChesher and Rosen (2017) obtain sharp identification for nonseparable instrumental variable models that include discrete choice. Using a nonparametric multinomial choice model with endogeneity for the California Health Insurance Exchange, Tebaldi, Torgovitsky, and Yang (2018)identify and estimate bounds on counterfactuals. Gao and Li (2018) develops estimation and
4Pakes and Porter Quantitative Economics 15 (2024) identification in a nonseparable version of the panel multinomial choice model without restricting the joint distribution of disturbances. This work also continues a substantial literature that has focused on extending nonlinear econometric models to allow for fixed effects while relaxing parametric distributional assumptions on disturbances. Manski (1987) applied his maximum score approach to the binary choice model with fixed effects. Honore (1992)furtherdeveloped Powell (1986)’s trimmed least squares approach to estimate the censored regression model with fixed effects. Abrevaya (1999) developed a new approach to estimation to allow for fixed effects in the transformation model, and further extended Han (1987)’s generalized regression model to include fixed effects in Abrevaya (2000). Ahn, Ichimura, Powell, and Ruud (2017) develop an approach to identification and estimation of semiparametric index models that can also be used in various fixed effects cases. The multinomial choice setup considered in the current work presents an additional complexity relative to the models in this previous literature. In particular, the multinomial choice model depends on multiple index functions of the covariates, where each index function corresponds to a choice-specific random utility. The main insight of our identification strategy is that a comparison of the multiple index functions for any two within-group observations has observable implications on the relative likelihood of certain choice outcomes. The paper is structured as follows. Section 2sets up a semiparametric version of the standard random utility model for multinomial choice with fixed effects. We then introduce our main stochastic disturbance assumption and derive a set of conditional moment inequalities. In Section 3, we show that the conditional moment inequalities provide sharp bounds on the parameters (or functions) of interest. In Section 4, we address point identification and binary choice. Section 5implements the conditional moment inequalities in an empirical exercise using data on health insurance choices through the Commonwealth Program exchange in Massachusetts. Section 6concludes. Proofs are in the Online Supplemental Material (Pakes and Porter (2024)). 2. Conditional moment inequalities for multinomial choice 2.1 Setup The data will be assumed to have a group/panel structure, where i=1, ,nindexes the groups and t=1, ,Tindexes observations within a group. There are a number of familiar multinomial choice applications with this group structure. In panel data applications in Labor and Public Finance, itypically indexes individuals, and tindexes time periods, though alternative groupings can also be relevant (an example from the study of hospital choice has iindexing the Cartesian product of illness category and hospital and tindexing patients, see Ho and Pakes (2014)). In Industrial Organization and Marketing applications, iwould typically index markets and twould index either the different consumers in those markets (in demand analysis) or the firms that compete in them (in the analysis of a firm’s choice of controls). Observation (i,t)faces a number of choices. Each choice dhas an associated random utility, Ud,i,t, and the observed choice, yi,t, maximizes the random utility over
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 5 choices. Suppose that d∈{0, ,D}, so that the number of choices is D+1. We consider the case of unordered response, where the numbering associated with each choice is arbitrary,2and 2 ≤D+1<∞. Given covariates xd,i,tfor each choice dassociated with observation (i,t),therandom utility for choices d=0, ,Dtakes the form Ud,i,t=gd(xd,i,t,θ0)+λd,i+εd,i,t,(1) where the term λd,idenotes choice-specific fixed effects, which account for unobserved characteristics of choice dthat do not vary across t. No restrictions are placed on the correlation between covariates xd,i,tand choice-specific fixed effects λd,i,sothesefixed effect terms generate a potential incidental parameter problem. The term εd,i,trepresents any remaining unobserved, idiosyncratic determinants of the random utility. The covariates enter random utility through the covariate index function gd(·,θ0),whereθ0 is used to index the functions gdand is unknown to the researcher. The most commonly assumed form for the index function is linear, for example, x d,i,tθ0. However, the parameter space for θ0is unrestricted and need not even be finite-dimensional. That is, θ0could index functions in an arbitrary function space, and the index functions are allowedtovarybychoice. 3The additive separability between the covariate index and the unobserved terms λd,i+εd,i,tis critical to the results that follow. However, the additive separability between the fixed effect λd,iand disturbance εd,i,tcould be relaxed. That is, λd,i+εd,i,tcould be replaced by a term of the form fd(λd,i,εd,i,t),wherefdis an unknown nonlinear function for choice d. In fact, under the assumptions below, the fixed effect could be absorbed into the disturbance without loss of generality.4Normaliza- tions to the model, such as pinning down the scale of coefficients in the linear covariate index case, can be incorporated as restrictions on the space of parameters, covariates, and the dimension of the conditional distribution of unobservables. The observed choice, yi,t, for agent (i,t)maximizes the random utility Ud,i,tover choices d. When a single choice uniquely maximizes random utility, then that choice is the observed choice for (i,t). We also allow for situations where there is a nonzero probability that two or more choices maximize utility. This situation could occur if the distribution of εi,thas mass points, as could be allowed under the flexible nonparametric assumptions on disturbances here, and is useful to extend our results to applications that involve set-valued regressors.5To fully specify the choice decision, we adopt a simple rule for resolving ties among maximizing random utility choices. If choices d1 2Inequalities for models with ordered responses are considered in Pakes, Porter, Ho, and Ishii (2015). 3The index functions could also be allowed to depend on time without any change in the results that follow. 4This point will appear again below when we see that the constructed distribution in the proof of sharpness has fixed effects set to zero. 5In the set-valued regressors case, the researcher does not know the specific value of some regressors, but does observe a set that contains their values. Pakes and Porter (2014) use the tools developed here to analyze this case. Two familiar examples are when the regressor is: (i) income (or wealth) and all the econometrician knows is that the income of each observation lies in a particular interval; and (ii) the distance from home to a service (or retail) outlet when the home location is only observed as a zip code (with known geographic boundaries).
6Pakes and Porter Quantitative Economics 15 (2024) and d2both maximize random utility (Ud1,i,t=Ud2,i,t=maxdUd,i,t)andd1<d 2,then assume that the choice with the largest choice index, in this case d2,istheobserved choice for (i,t). And, in general, if there are multiple utility maximizing choices, then the observed outcome is assumed to be the largest choice number among the utility maximizing choices. More generally, any rule for resolving utility maximizing ties that is a nonstochastic function of argmaxdUd,i,twill lead to the same form of conditional moment inequalities derived below6and so knowledge of such a rule is not needed for their implementation. The setup thus far is a random utility formulation of multinomial choice except that a choice-specific group fixed effect is included and general covariate indices are allowed. It will be useful to establish notation for the mapping defined by this setup from the covariates, parameter θ0, fixed effects, and disturbances to the observed outcome yi,t.Letxi,t=(x 0,i,t,,x D,i,t),λi=(λ0,i,,λD,i),εi,t=(ε0,i,t,,εD,i,t),andUi,t= (U0,i,t,,UD,i,t). When it is helpful to be explicit about the dependence of random utility on its components, we will use the notation Ud(xi,t,θ0,λi,εi,t)to denote Ud,i,t= gd(xd,i,t,θ0)+λd,i+εd,i,t. The observed outcome yi,tcan also be written as a function of these same components: yi,t=y(xi,t,λi,εi,t,θ0)=max argmaxdUd(xi,t,θ0,λi,εi,t), where yis the mapping that represents the random utility formulation for multinomial choice given above. Our identification results will correspond to the case where Tis fixed at T=2. We will denote the two time periods or observations within each group by sand t,rather than 1 and 2 to avoid confusion, especially in the variable subscripts, with the choices d,whicharenumbered0,,D. The key stochastic assumption for this framework is within-group/time homogeneity of the disturbances. This assumption is a form of strict exogeneity and is a common condition imposed in panel data models (Chernozhukov, Fernández-Val, Hahn, and Newey (2013)).7 Assumption 1. (a) (xi,s,xi,t,λi,εi,s,εi,t)is independently and identically distributed for i=1, ,n; (b) Given the conditioning set (xi,s,xi,t,λi),the conditional distributions of εi,sand εi,tare the same: εi,s|xi,s,xi,t,λi∼εi,t|xi,s,xi,t,λi. The second part of the assumption mirrors the stochastic assumption made for panel data binary choice models in Manski (1987) and for discrete choice in Shi, Shum, and Song (2018). No parametric distributional restrictions are placed on the distribution of εi,t.Notethatεi,tis, in general, a vector of individual choice disturbances, in contrast to the binary choice case. Importantly, for a given time t, the marginal distribution of 6See also footnote 8for allowable tie-breaking rules. 7Mean independence and zero covariance forms of strict exogeneity also appear commonly in the literature, especially in linear panel data model cases.
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 7 these choice disturbances is allowed to vary arbitrarily across choices (d), and there are no restrictions on the joint behavior of these disturbances across choices. As a result, neither independence of irrelevant alternatives, nor any other limitation on the substitutability of different choices induced by the covariance structure of disturbances (such as the limited substitutability property discussed in Berry and Pakes (2007)) is a source of concern. This assumption also allows the disturbances for the different choices to be freely correlated across time. Assumption 1nests both the familiar panel data model with individual choice-specific fixed effects and i.i.d. disturbances, a special case of which is Chamberlain’s (1980) conditional logit model, and many differentiated product demand models for micro data (e.g., Berry, Levinsohn, and Pakes (2004)). Assumption 1does restrict the relationship between the disturbances and the covariates. For instance, heteroskedasticity would need to take a specific form where the heteroskedasticity in εi,tisthesameasεi,seven when xi,t= xi,s. For example, if the heteroskedasticity in both εi,tand εi,sdepended on xi,t+xi,s, then Assumption 1would not be violated. Of course, the typical assumption of independence of disturbances and covariates across different sand twould suffice to satisfy Assumption 1. Assumption 1(b) means that “within” variation could be useful for identification. By restricting the conditional joint distribution of the disturbances across the random utility choices to be the same for observations in group i, Assumption 1enables us to learn about relative response probabilities by comparing the observable components of random utilities across tfor that group i. This within-group comparison will not depend on the joint distribution of disturbances across choices in any way. To simplify notation, below we eliminate the group iindex with the understanding that all variables below are associated with the same group unless otherwise indicated. 2.2 Illustrative moment inequality Given the random utility framework above along with Assumption 1,wecanderivea set of moment inequality conditions that can be taken to data for inference on the parameter θ0. We begin with a single conditional moment inequality that makes both the assumptions and logic underlying our conditional moment inequality analysis transparent. Following this derivation, we show how an extension of this logic leads to a collection of conditional moment inequalities. Our moment inequalities are based on a within comparison of choice probabilities for individual/group iat times sand t. We can express the conditional probability of observing choice dat time tthrough the corresponding region of the disturbance space, Ed,t=εt:y(xt,λ,εt,θ0)=d.(2) Given this definition, Pr(yt=d|xs,xt,λ)=Pr(εt∈Ed,t|xs,xt,λ).Toconsiderhowvariation in the covariates across time affects choice probabilities, it is useful to note the explicit dependence of the region Ed,ton the covariates: Ed,t=εt:εd,t≥max c<d gc(xc,t,θ0)−gd(xd,t,θ0)+(λc−λd)+εc,t ∩εt:εd,t>max c>d gc(xc,t,θ0)−gd(xd,t,θ0)+(λc−λd)+εc,t,(3)
8Pakes and Porter Quantitative Economics 15 (2024) where the sets with weak and strict inequalities follow from our rule for resolving utility maximizing ties. We compare the time tregions E0,t,,ED,tto the analogous regions at time s,E0,s,,ED,s. From this comparison, we will be able to show that for one of the D+1 choices, the region at time scontains the corresponding region at time t.Moreover, the choice with this property is determined completely by the covariate indices. Assumption 1then implies that the corresponding choice probability at time swill be at least as large as the choice probability at time t. To find a choice with this special property, we order the covariate index differences across time by choice. In particular, find the choice with the largest change in covariate index: d∗=argmaxcgc(xc,s,θ0)−gc(xc,t,θ0).(4) Ifthereismorethanonechoiceintheargmax set, then set d∗to any element of this set. Note that gd∗(xd∗,s,θ0)−gd∗(xd∗,t,θ0)≥gc(xc,s,θ0)−gc(xc,t,θ0),∀c =⇒ gc(xc,t,θ0)−gd∗(xd∗,t,θ0)+(λc−λd∗) ≥gc(xc,s,θ0)−gd∗(xd∗,s,θ0)+(λc−λd∗),∀c. (5) The latter covariate index differences on either side of the inequality are the same differences that define Ed∗,tand Ed∗,sin (3). And the inequality (5) ensures that Ed∗,t⊂Ed∗,s. Hence, Prys=d∗|xs,xt,λ=Pr(εs∈Ed∗,s|xs,xt,λ) =Pr(εt∈Ed∗,s|xs,xt,λ) ≥Pr(εt∈Ed∗,t|xs,xt,λ) =Pryt=d∗|xs,xt,λ.(6) The first and last equalities follow from the definition of the disturbance regions in (2). The second equality follows from Assumption 1, and the inequality follows from the set inclusion derived above. Since the inequality holds regardless of the values of the fixed effects (λ), the fixed effects can be integrated out of the inequality in (6) yielding a corresponding conditional moment inequality below. We also extend the argument behind this inequality to generate additional conditional choice probability comparisons and their related conditional moment inequalities, which can then be used for identification of the parameter θ0. To illustrate the key intuition behind this inequality, consider the case with three choices, a linear covariate index, and d∗=2 implying that E2,t⊂E2,s. To show the regions Ed,sand Ed,ton two-dimensional graphs, these regions can be reexpressed in terms of (ε1,s−ε0,s,ε2,s−ε0,s)and (ε1,t−ε0,t,ε2,t−ε0,t). In Figure 1(a), the time s
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 15 is still satisfied. The conclusion of Theorem 2follows. Practically, this enables the researcher to use the conditional moment inequalities in (10) to consider the identifying power of a particular empirical covariate design and investigate the potential identifying power of alternatives. 4. Additional remarks Next, we discuss two topics related to our sharp identification result: point identification and the special case of binary choice. 4.1 Point identification In Section 3, we showed that the proposed conditional moment inequalities produce a sharp identified set. With a linear covariate index function, point identification is established by imposing further conditions that ensure the identified set reduces to a singleton. Assumption 1places no restrictions on the covariates. The key additional conditions for point identification ensure sufficient variation in the covariates, in particular an assumption of unboundedness (see Chamberlain (2010)). Shi, Shum, and Song (2018) derived conditional moment inequalities implied by cyclic monotonicity for the multinomial choice model with a linear covariate index function under conditions including Assumption 1and absolute continuity of the error distribution with respect to Lebesgue measure. Under assumptions on the covariates, they show that their conditional moment inequalities are sufficient for point identification. It is straightforward to compare the conditional moment inequalities in (10)witha linear covariate index function to the corresponding Shi, Shum, and Song (2018)cyclic monotonicity conditional moment inequalities. We adopt the normalization for choice zero in Shi, Shum, and Song (2018), x0,i,t=0, λ0,i=ε0,i,t=0soU0,i,t=0 and similarly at time s.FromShi, Shum, and Song (2018) Lemma 3.1, the length 2-cycle conditional moment inequality can be expressed as 0≤ D d=1Pr(yi,s=d|xi,s,xi,t)−Pr(yi,t=d|xi,s,xi,t)x dθ0. (13) Now denote a (weak) ordering of covariate index differences as follows: x (D)∗θ0≥x (D−1)∗θ0≥···≥x (0)∗θ0. (14) And suppose that choice zero has the j+1th smallest covariate index difference, that is, (j)∗=0, so that x (j)∗θ0=0. Then rewriting the sum in (13), D d=1Pr(yi,s=d|xi,s,xi,t)−Pr(yi,t=d|xi,s,xi,t)x dθ0 = D d=j+1Pryi,s=(d)∗|xi,s,xi,t−Pryi,t=(d)∗|xi,s,xi,tx (d)∗θ0
16 Pakes and Porter Quantitative Economics 15 (2024) + j−1 d=0Pryi,s=(d)∗|xi,s,xi,t−Pryi,t=(d)∗|xi,s,xi,tx (d)∗θ0 = D d=j+1D d=dPryi,s=d∗|xi,s,xi,t−Pryi,t=d∗|xi,s,xi,t ·x (d)∗θ0−x (d−1)∗θ0 + j−1 d=0D d=dPryi,s=d∗|xi,s,xi,t−Pryi,t=d∗|xi,s,xi,t ·x (d)∗θ0−x (d+1)∗θ0 = D d=j+1Pryi,s∈(d)∗,,(D)∗|xi,s,xi,t−Pryi,t∈(d)∗,,(D)∗|xi,s,xi,t ·x (d)∗θ0−x (d−1)∗θ0 + j−1 d=0Pryi,s∈(0)∗,,(d)∗|xi,s,xi,t−Pryi,t∈(0)∗,,(d)∗|xi,s,xi,t ·x (d)∗θ0−x (d+1)∗θ0 = D d=j+1Pryi,s∈(d)∗,,(D)∗|xi,s,xi,t−Pryi,t∈(d)∗,,(D)∗|xi,s,xi,t ·x (d)∗θ0−x (d−1)∗θ0 + j−1 d=0Pryi,s∈(d+1)∗,,(D)∗|xi,s,xi,t −Pryi,t∈(d+1)∗,,(D)∗|xi,s,xi,t ·x (d+1)∗θ0−x (d)∗θ0. (15) The terms (x (d)∗θ0−x (d−1)∗θ0)and (x (d+1)∗θ0−x (d)∗θ0)in (15) are nonnegative by the ordering defined in (14). The relationship between the conditional moment inequalities in (10) and the cyclic monotonicity conditional moment inequalities in (13) follows immediately. The inequalities in (10) imply that the probability difference terms in square brackets in (15) are nonnegative, which further implies that (13)holds.We can then conclude that under Assumption 1and the additional conditions given in Shi, Shum, and Song (2018), the conditional moment inequalities in (10) yield point identification.9 9Actually, we find that point identification can be achieved using a subset of the inequalities in (10)that correspond to partitions of the choice set of a fixed size. Fix δ∈{1, ,D}. Then, for point identification, it suffices to consider the subset of conditional moment inequalities in (10)with|D|=δor (D+1)−δ.
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 17 The implication illustrated also highlights a point made in the remarks following Theorem 2. The conditional moment inequalities of this paper are actually conditionally sharp. Any given covariate value yields one Shi, Shum, and Song (2018) length 2-cycle conditional moment inequality. Since this inequality is a linear inequality, it rules out a half-space in the parameter space. The same covariate value yields 2D+1−2 conditional moment inequalites from (10) and each of these conditional moment inequalities rules out a cone in the parameter space. Equation (15) shows formally that the union of ruledout cones must contain the ruled out half-space. We also find that, in general, the inequalities (13)donotimply(10). Inequality (13) generates the half-space whose boundary is given in (15) (which is a weighted combination of covariate values where the weights are given by the choice probability differences), while the inequalities (10) generate 2D+1−2 (cone) regions (each of which is determined by the covariate values and does not depend on the choice probabilities). So, in general the boundaries from (15)and(10) will not coincide, that is, the ruled out half-space given by (15) will be strictly contained in the ruled out region given by (10)for each conditioning value of the covariate. 4.2 Binary choice The sharpness result in Section 3can, of course, be specialized to the case of binary choice. In this case, Theorem 2provides a (to our knowledge) new supplement to the point identification finding in Manski (1987). In particular, even when point identification fails, the weak version of Manski (1987)’s conditional moment inequalities provide sharp bounds on the binary choice random utility covariate coefficient, θ,underAssumption 1.10 Moreover, we show here that Manski’s maximum score criterion used for point identification can also be used to obtain sharp partial identification when point identification does not hold. Additionally, while Manski (1987) considers the linear index case, the result shown here allows for parametric or nonparametric covariate index functions as in (1). In the point identified binary choice case, Manski (1987) proposes an alternative method of estimation, commonly referred to as maximum score estimation. We see below the close connection between the maximum score criterion and the conditional moment inequalities defining the identified set. The conditional moment inequalities are: EmD(ys,yt,xs,xt,θ)|xs,xt≥0, ∀D∈D, (16) and in the binary choice case, D(xs,xt,θ)=⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ {1},forg(xs,xt,θ)>0 {0},forg(xs,xt,θ)<0 {0},{1},forg(xs,xt,θ)=0, (17) This collection of conditional moment inequalities is nonnested with the cyclic monotonicity inequalities in (13). 10In the binary choice case, Assumption 1(a) is exactly Manski’s Assumption 3, and Assumption 1(b) is exactly Manski’s Assumption 1(a). So, the sharpness result relaxes Manski’s Assumptions 1(b) and 2.
18 Pakes and Porter Quantitative Economics 15 (2024) where g(xs,xt,θ)=g1(x1,s,θ)−g1(x1,t,θ)−g0(x0,s,θ)−g0(x0,t,θ). The maximum score criterion function is H(θ)=Esgng(xs,xt,θ)(ys−yt). To connect these expressions, define a function which “aggregates” the conditional moments across the sets D∈D(xs,xt,θ): H(xs,xt,θ)= D∈D EmD(ys,yt,xs,xt,θ)|xs,xt = D∈D(xs,xt,θ) E1{ys∈D}−1{yt∈D}|xs,xt. (18) Then H(θ)=EH(xs,xt,θ). (19) That is, the maximum score criterion is an aggregation of the unconditional version of the moments from the conditional moment inequalities. Clearly, H(θ0)≥0and0⊂{θ∈|H(θ)≥0}, but in general this set inclusion would be strict and {θ∈|H(θ)≥0}would not be sharp. Instead of checking nonnegativity of this criterion, maximum score seeks to maximize it. According to Manski (1987)in the binary choice case with a linear covariate index function, under conditions implying point identification, the maximum score criterion H(θ)is uniquely maximized at θ0. The following proposition shows that, in the binary choice case, the maximum score criterion is useful even when point identification is not achieved. In particular, under Assumption 1alone, 0could be either a set or a point, and the maximum score criterion H(θ)exactly identifies this set (or point) 0. Proposition 3. Suppose D+1=2(binary choice)and Assumption 1holds.Then 0=argmax θ∈H(θ)=θ∈:H(θ)=H(θ0). Under the conditions of Theorem 2,0is itself sharp, showing that the maximum score criterion will identify the sharp bounds for the covariate index function in the binary choice model. 5. Empirical example We implement our conditional moment inequalities in an empirical example that analyzes health insurance choices in the Commonwealth Care (or “CommCare”) program in Massachusetts, enacted as part of the state health reforms in 2006. The program pro-
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 19 vided subsidized health insurance to low-income citizens via an insurance exchange that let consumers choose among competing private plans. We focus on a classic question in demand analysis; does the response of demand to price changes depend on the income of consumers? CommCare was for citizens whose earnings were less than 300% of the Federal Poverty Level (or the FPL; this was $10,830 in 2010, and increased by the CPI-U annually thereafter). We examine whether the response to price movements differed between two groups of individuals: those with incomes between one and two times the FPL and those whose income was between two and three times the FPL (those with incomes less than the FPL were fully subsidized, and hence not included in the analysis). Differences in the price coefficient between these two groups has distributional implications for the welfare generated by this and other programs directed at low income households. Our model has consumer iat time tchoosing among plans dto maximize Ud,i,t=−pd,i,tβ0−pd,i,t1Ii,t∈[2FPL,3FPL]γ0+λd,i+εd,i,t, (20) where individual i’s price coefficient is β0if that individual’s income Ii,tis contained in [FPL,2FPL]and β0+γ0if Ii,t∈[2FPL,3FPL].Theλd,icapture individual (additive) product preferences, and εd,i,tcaptures the remaining unobserved variation in random utility. Our focus is on γ0, the difference between the price sensitivity of individuals when they are in the higher versus the lower income group. Since the parameters are only identified up to scale, we can only learn about γ0/β0. Assuming downward sloping demand (β0>0), we can normalize β0=1sothatγ0represents the desired ratio.11 We consider regions where four insurers participate in the market during our data period, with each insurer (by rule) offering a single plan. Program rules required each enrollee to make a separate choice; there was no family coverage, and kids were covered in the separate Medicaid program. We analyze plan choices in an annual open enrollment month each year. Individuals are also allowed to choose plans when they change their income group and we treat changes occurring at these times as separate choices for estimation purposes.12 For more detail on the data and the CommCare program, see 11To provide an example connecting the hypotheses on γ0with a demand effect, consider a case where prices of all choices in period sare the same and in period tthe price of choice ddecreases while all other prices stay the same or increase. Let 1=Pr(yt=d|pd,s,pd,t,p−d,s,p−d,t,Is,It∈[FPL,2FPL],λ) −Pr(ys=d|pd,s,pd,t,p−d,s,p−d,t,Is,It∈[FPL,2FPL],λ). Assuming downward sloping demand, 1≥ 0. We compare 1to the change in demand for the same choice under the same price dynamics for a higher income individual, 2=Pr(yt=d|pd,s,pd,t,p−d,s,p−d,t,Is,It∈[2FPL,3FPL],λ)−Pr(ys= d|pd,s,pd,t,p−d,s,p−d,t,Is,It∈[2FPL,3FPL],λ). Assume that the disturbance distribution is conditionally independent of incomes. Then, under the null γ0=0, 1=2. Under the alternative that γ0<0, 1> 2. That is, the demand is less sensitive to this price change for higher income individuals. And, both of these conclusions would still hold after integrating the fixed effects out over a given distribution to yield “average” demand effects. 12Coverage is heavily regulated, with all cost sharing and covered medical services completely standardized across insurers. The only flexible plan attributes are provider networks. These were largely stable during our sample period with one major exception. Network Health (one of our plans) drops Partners Health-
20 Pakes and Porter Quantitative Economics 15 (2024) Shepard (2020), Finkelstein, Hendren, and Shepard (2019), McIntyre, Shepard, and Wagner (2021). We want to capture choices that are not induced by changes in the individual’s choice environment, just by prices, and we require the choice set (plan availability) to be the same four plans in the two periods we compare. We therefore remove comparisons for individuals who changed regions (there are five in the data), or who faced different plan offerings in the comparison periods. We divide the remaining data into cells that reflect the discrete values of consumers’ observed characteristics. The characteristics of a cell are defined by the Cartesian product of; (a) pair of years, (b) region, and (c) the income groups in each of the two periods being compared. So, the λd,irepresent differences in tastes among consumers with the same characteristics. Our data contains all cells as defined above that have more than 20 members. There were large changes in relative prices between 2010 and 2012. During this period, the provider BMC, which had the largest share in 2010 with over a third of the market, increased its average price from below $50 per member per month to over $90. By the end of 2012, it was clear that the price increases cost them almost half of their subscribers, and in 2013 they reduced their prices to an average price of just over $40.13 We focus on the differential responses to these price changes and use the (s,t)combinations of (2010, 2012) and (2012, 2013). This generates 13,169 pairs of choices and 14 moments corresponding to the four choices.14 To use the conditional moment inequalities for inference, there are many recently developed methods (Andrews and Shi (2013), Armstrong (2015), Armstrong and Chan (2016), Chernozhukov, Lee, and Rosen (2013), Chetverikov (2018), and Lee, Song, and Whang (2013)). Given the discreteness of the CommCare data, there are limited choices for instruments to translate the conditional moments into unconditional moments. The price variation across year-pairs motivates our use of indicators for year pairs as instruments, which yields 28 unconditional moments. The confidence set is then estimated using the generalized moment selection procedure in Andrews and Soares (2010)following the recommended tuning parameters choices in Andrews and Shi (2013). Implementing a squared negative part criterion function and bootstrap critical values yields a 95% confidence interval for γ0,[−0.79, −0.31], indicating that when individuals transit to a higher income group their price sensitivity falls significantly. Care (the state’s largest medical system) from its hospital network at the start of 2012. To account for this, we treat Network as two different plans, one before and one after 2012, and apply the rules above with that understanding. 13There were no major changes in BMC’s network or other quality attributes at this time. There was, however, a change in the rules governing the exchange in 2012, which set up an auction-like environment to determine the plans that were available to the fully subsidized individuals and likely induced price experimentation. 14The data combines the unbalanced panel of observations from both year pairs, 2010–2012 and 2012– 2013. Alternatively, we could have used only individuals observed across both year pairs and then generated fourteen separate moment conditions for each year pair. Unfortunately, this approach reduced our data set to less than 10% of the data set we use. As a result, we combine all the data from both year pairs and used year pair as an instrument. The variance estimator does not account for possible correlation of observations across year pairs.
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 21 Figure 2. Criterion function and critical values. The solid line is the negative part criterion function, and the dashed line is the corresponding 95% bootstrap critical value function. More detail is provided in Figure 2. The blue line graphs the sample test statistic. The red line graphs the 95% bootstrap critical values. Using the max criterion function (Armstrong (2014)) and using various instruments that aggregate less (yielding more unconditional moments) leads to similarly shaped test statistic graphs (though exact magnitudes depend on the number of unconditional moments). We also considered the Shi, Shum, and Song (2018) length 2-cycle conditional moment inequalities. Using the year pair instruments, as above, yields two linear inequalities. The lower bound information in these moment inequalities is the same as obtained previously. The two linear inequalities are necessarily one-sided inequalities and neither provides upper bound information.15 6. Conclusion We have provided a new approach to identification for multinomial choice models. Our focus has been on the multinomial choice model, which allows for choice-specific fixed effects with a group (or panel) structure and a nonparametric distribution of disturbances only restricted to satisfy a stationarity assumption. We show that this specification generates conditional moment inequalities, which can be used for identification of the covariate index function and avoids the incidental parameter problem using only two time periods. These conditional moment inequalities provide sharp bounds without restrictions on the covariates. When T>2, each pair of time periods generates a set of conditional moment inequalities as described above. A sharpness result for this case is left as an open problem for future research. 15We note that using different instruments we are able to find some upper bound information in the cyclic monotonicity conditional moment inequalities.
22 Pakes and Porter Quantitative Economics 15 (2024) Our empirical example illustrates how these techniques can be applied to examine differential responses by individuals with different characteristics to a given determinant of choices. This should lead to a better understanding of distributional implications of different policies. Often the focus of empirical studies is not on θ0per se but rather on different functionals that could depend on θ0(e.g., Tebaldi, Torgovitsky, and Yang (2018)). In the discrete choice panel data setting, Chernozhukov et al. (2013) suggest particular functionals of interest such as the conditional quantile or average structural effects. In such cases, our conditional moment inequalities provide a new source of identifying information. Without restricting the disturbance distribution across choices, our conditional moment inequalities are relatively easy to compute and provide sharp (and sometimes point) identifying information on θ0. This additional “within” information can be used to improve upon the estimation of the various effects by narrowing the range of parameter values to be considered together with the possible disturbance distributions (including fixed effects) that are consistent with the “between” variation specified in the Chernozhukov et al. (2013) paper. The issue of what is consistent with the “between” variation opens up the question, which we have left for future research, of what information is available on the fixed effects per se. In addition to helping us analyze responses to changes in characteristics, there are cases where knowledge of the fixed effects are of inherent interest and should be analyzable. For example, in Ho and Pakes (2014)’s investigation of the impact of capitation on allocation of patients to hospitals, the fixed effects represent the perceived qualities of (the 194) different hospitals for each of (the 106) alternative illness categories. They examine whether the perceptions of the providers from different insurance networks coincide, and are able to rank hospital by their perceived quality for the major illness categories. References Abrevaya, Jason (1999), “Leapfrog estimation of a fixed-effects model with unknown transformation of the dependent variable.” Journal of Econometrics, 93, 203–228. [4] Abrevaya, Jason (2000), “Rank estimation of a generalized fixed-effects regression model.” Journal of Econometrics, 95, 1–23. [4] Ahn, Hyungtaik, Hidehiko Ichimura, James Powell, and Paul Ruud (2017), “Simple estimators for invertible index models.” Journal of Business Economics and Statistics, 36, 1–10. [4] Andrews, Donald and Xiaoxia Shi (2013), “Inference based on conditional moment inequalities.” Econometrica, 81, 609–666. [3,20] Andrews, Donald and Gustavo Soares (2010), “Inference for parameters defined by moment inequalities using generalized moment selection.” Econometrica, 78, 119–157. [3,20] Armstrong, Timothy (2014), “Weighted ks statistics for inference on conditional moment inequalities.” Journal of Econometrics, 181, 92–116. [21]
Quantitative Economics 15 (2024) Moment inequalities for multinomial choice 23 Armstrong, Timothy (2015), “Asymptotically exact inference in conditional moment inequality models.” Journal of Econometrics, 186, 51–65. [20] Armstrong, Timothy and Hock Peng Chan (2016), “Multiscale adaptive inference on conditional moment inequalities.” Journal of Econometrics, 194, 24–43. [20] Berry, Steven, James Levinsohn, and Ariel Pakes (1995), “Automobile prices in market equilibrium.” Econometrica, 63, 841–890. [3] Berry, Steven, James Levinsohn, and Ariel Pakes (2004), “Differentiated products demand systems from a combination of micro and macro data: The new car market.” Journal of Political Economy, 112, 68–105. [7] Berry, Steven and Ariel Pakes (2007), “The pure characteristics demand model.” International Economic Review, 48, 1193–1225. [7] Chamberlain, Gary (1980), “Analysis of covariance with qualitative data.” The Review of Economic Studies, 47, 225–238. [1,3,7] Chamberlain, Gary (2010), “Binary response models for panel data: Identification and information.” Econometrica, 78, 159–168. [15] Chernozhukov, Victor, Iván Fernández-Val, Jinyong Hahn, and Whitney Newey (2013), “Average and quantile effects in nonseparable panel models.” Econometrica, 81, 535– 580. [6,22] Chernozhukov, Victor, Sokbae Lee, and Adam Rosen (2013), “Intersection bounds: Estimation and inference.” Econometrica, 81, 667–737. [20] Chesher, Andrew, Adam Rosen, and Konrad Smolinski (2013), “An instrumental variable model of multiple discrete choice.” Quantitative Economics, 4, 157–196. [3] Chesher, Andrew and Adam Rosen (2017), “Generalized instrumental variable models.” Econometrica, 85, 959–989. [3] Chetverikov, Denis (2018), “Adaptive test of conditional moment inequalities.” Econometric Theory, 34, 186–227. [20] Finkelstein, Amy, Nathaniel Hendren, and Mark Shepard (2019), “Subsidizing health insurance for low-income adults: Evidence from Massachusetts.” American Economic Review, 109, 1530–1567. [20] Fox, Jeremy (2007), “Semiparametric estimation of multinomial discrete-choice models using a subset of choices.” RAND Journal of Economics, 38, 1002–1019. [3] Gao, Wayne and Ming Li (2018), “Robust semiparametric estimation in panel multinomial choice models.” Yale, Working Paper. [3] Han, Aaron (1987), “Nonparametric analysis of a generalized regression model.” Journal of Econometrics, 35, 303–316. [4] Ho, Kate and Ariel Pakes (2014), “Hospital choice, hospital prices and financial incentives to physicians.” American Economic Review, 104, 3841–3884. [4,22]
24 Pakes and Porter Quantitative Economics 15 (2024) Honore, Bo (1992), “Trimmed lad and least squares estimation of truncated and censored regression models with fixed effects.” Econometrica, 60, 533–565. [4] Kahn, Shakeeb, Fu Ouyang, and Elie Tamer (2019), “Inference on semiparametric multinomial response models.” Working Paper, Harvard University. [3] Lee, Lung-fei (1995), “Semiparametric maximum likelihood estimation of polychotomous and sequential choice models.” Journal of Econometrics, 65, 381–428. [3] Lee, Sokbae, Kyyngchul Song, and Yoon-Jae Whang (2013), “Testing functional inequalities.” Journal of Econometrics, 172, 14–32. [20] Manski, Charles (1975), “Maximum score estimation of the stochastic utility model of choice.” Journal of Econometrics, 3, 205–228. [3] Manski, Charles (1987), “Semiparametric analysis of random effects linear models from binary panel data.” Econometrica, 55, 357–362. [1,2,4,6,17,18] McFadden, Daniel (1974), “Conditional logit analysis of qualitative choice behavior.” In Frontiers in Econometrics (P. Zarembka, ed.), 105–142, Academic Press, New York. [1] McIntyre, Adrianna, Mark Shepard, and Myles Wagner (2021), “Can automatic retention improve health insurance market outcomes?” Technical report, Harvard University Working Paper. [20] Pakes, Ariel (2014), “Behavioral and descriptive forms of choice models.” International Economic Review, 55, 603–624. [3] Pakes, Ariel and Jack Porter (2014), “Moment inequalities for multinomial choice with fixed effects.” Working Paper, University of Wisconsin. [5] Pakes, Ariel and Jack Porter (2024), “Supplement to ‘Moment inequalities for multinomial choice with fixed effects’.” Quantitative Economics Supplemental Material, 15, https://doi.org/10.3982/QE1776.[4] Pakes, Ariel, Jack Porter, Kate Ho, and Joy Ishii (2015), “Moment inequalities and their application.” Econometrica, 80, 315–334. [5] Powell, James (1986), “Symmetrically trimmed least squares estimation for Tobit models.” Econometrica, 54, 1435–1460. [4] Shepard, Mark (2020), “Hospital network competition and adverse selection: Evidence from the Massachusetts health insurance exchange.” Technical report, Harvard University Working Paper. [20] Shi, Xiaoxia, Matthew Shum, and Wei Song (2018), “Estimating semiparametric panel multinomial choice models using cyclic monotonicity.” Econometrica, 86, 737–761. [2, 6,15,16,17,21] Tebaldi, Pietro, Alexander Torgovitsky, and Hanbin Yang (2018), “Nonparametric estimates of demand in the California health insurance exchange.” Working Paper, University of Chicago. [3,22]