Specification testing of discrete choice models: A note on the use of a nonparametric test
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Fosgerau, Mogens Article Specification testing of discrete choice models: A note on the use of a nonparametric test Journal of Choice Modelling Provided in Cooperation with: Journal of Choice Modelling Suggested Citation: Fosgerau, Mogens (2008) : Specification testing of discrete choice models: A note on the use of a nonparametric test, Journal of Choice Modelling, ISSN 1755-5345, University of Leeds, Institute for Transport Studies, Leeds, Vol. 1, Iss. 1, pp. 26-39 This Version is available at: https://hdl.handle.net/10419/66833 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc/2.0/uk/
Journal of Choice Modelling, 1(1), pp. 26-39 www.jocm.org.uk Specification testing of discrete choice models: a note on the use of a nonparametric test Mogens Fosgerau∗ Institute for Transport, Technical University of Denmark, Knuth-Winterfeldts All´e, Bygning 116 Vest, 2800 Kgs. Lyngby, Denmark Received 21 August 2007, revised version received 4 December 2007, accepted 7 January 2008 Abstract Model misspecification is a serious issue since misspecification generally renders statistical inference invalid. However, specification testing of discrete choice models is rarely applied. This paper describes a nonparametric test procedure which uses a combination of smoothed residual plots and a test statistic able to detect general misspecification. Nonparametric methods require large datasets when the number of independent variables is more than a few. A way to circumvent this problem is indicated, increasing the usefulness of the approach also with limited datasets. Keywords: discrete choice, specification test, nonparametric, functional form 1 Introduction It is standard practice in regression models to perform model control using the residuals of the estimated model. Residuals are plotted to verify whether they are in fact white noise unrelated to the independent variables. Residuals are less easily defined in discrete choice models and similar model control is rarely performed for such models. In fact, a variety of specification tests are available in the literature for some discrete choice models, but are not widely used. Lechner (1991) presents some specification tests for the binary logit model. Gourieroux et al. (1987a) and Gourieroux et al. (1987b) present tests based on generalised residuals for a range of models including the multinomial logit model. McFadden (1987) presents regression-based specification tests for the multinomial logit model similar in nature to the test presented here. The seminal paper by McFadden and Train ∗T: +45 4525 6521, F: +45 4593 6533, [email protected] Licensed under a Creative Commons Attribution-Non-Commercial 2.0 UK: England & Wales License http://creativecommons.org/licenses/by-nc/2.0/uk/ ISSN 1755-5345
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 (2000) provides specification tests of MNL and mixed logit against alternatives with more mixing. Finally, software exists that allows the comparison of observed choices to predictions when data are grouped on a categorical variable. This approach may be viewed as a kind of residual test.1 The point of this paper is to describe how the nonparametric test of functional form in Zheng (1996) may be applied to discrete choice models of general form. This means that the test applies even to complicated models such as the mixed generalised extreme value model. The Zheng test is based on nonparametric kernel regression of the parametric model residuals against the independent variables. With residuals defined as the difference between choice 0-1 indicators and predicted probabilities, the Zheng test applied to a discrete choice model is based on the comparison of predicted and observed choices. Pagan and Ullah (1999) review a range of nonparametric tests of functional form that are more or less similar to the Zheng test. The procedure presented in this paper of applying a test to a function of the independent variables is not restricted to the Zheng test but may be applied with other tests as well. On the use of nonparametrics in a discrete choice context and in transport applications see, e.g. Fosgerau (2006), Fosgerau (2007) and Fosgerau and Bierlaire (2007). Pagan and Ullah (1999), Yatchew (2003) and H¨ardle (1990) give general introductions to nonparametric and semiparametric methods.2 A general problem in using nonparametric techniques is that the demands on data increase exponentially in the dimension of the space of independent variables, this is the so-called curse of dimensionality. This issue is addressed in this paper by showing how such tests may be applied to a subset of variables or more generally to functions of variables such as the index representing the indirect utility of a choice alternative. It is thus possible to apply nonparametric tests to a model, while regressing only on a low number of variables or even just one. This reduces the demand on data and is of practical importance when the model to be tested has several independent variables. Furthermore, the test may be applied to variables that are not included in the model. Thus the test may serve as a test of omitted variables. The paper is organised as follows. Section 2 presents an exposition of the Zheng test. Section 3 shows how such tests may be applied to low-dimensional functions of the independent variables. Section 4 presents an example of application of the test to a discrete choice model while section 5 concludes. 2 The Zheng test Consider a model E(y|x), predicting the expectation of a variable y∈Rconditional on an m-dimensional vector of observables x∈Rm, which are also taken 1E.g. Alogit produces so-called apply tables. 2Tests based on regression of residuals against independent variables are also discussed in Ellison and Ellison (2000). 27
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 as random having density p(x). The dependent variable ymay be continuous, but it may also be binary or a binary dummy indicator of an alternative in a multinomial model. In such cases E(y|x) = P(y= 1|x). In models with many alternatives it may also be useful to let ybe a dummy indicator for the choice of a group of alternatives. This situation is also accommodated. In situations with unlabelled alternatives it could be useful to reorder the alternatives according to some characteristic such that yindicates, e.g. choice of the fastest alternative. The expectation of yconditional on xis a function of x. We let g(x) denote the true but unknown conditional expectation. The researcher specifies a parametric model f(x:θ) where θ∈Θ. We wish to test whether there is a parameter θ0∈Θ such that f(x:θ0) = g(x). Formally, we formulate our null hypothesis as H0:P[f(x:θ0) = g(x)] = 1.(1) The null hypothesis thus states that f(·:θ0) coincides with gfor almost all x. The corresponding alternative hypothesis is that, for all θ∈Θ, fdiffers from g on a set of probability greater than zero, or formally, H1:∀θ∈ΘP[f(x:θ) = g(x)] <1.(2) The alternative hypothesis is the negation of the null hypothesis. It thus encompasses all possible departures from the null. Zheng (1996) defines a statistic Tn, where nis sample size, and shows that, under the null hypothesis, Tnconverges in distribution to a standard normal as sample size increases, whereas it converges in probability to infinity under the alternative. Therefore the test is consistent against all departures from the parametric model f. The idea of the test is the following. Define residuals by ε=y−f(x: θ0). Then under the null hypothesis, E(ε|x) = 0 almost surely and hence also E[εE(ε|x)p(x)] = 0. However, under the alternative hypothesis, E[εE(ε|x)p(x)] = E[E(ε|x)2p(x)] =E[(g(x)−f(x:θ0))2p(x)] >0.(3) where the first equality follows from the law of iterated expectations. Consider now an i.i.d. sample (yi, xi). The test statistic is formed from a sample analogue of E[εE(ε|x)p(x)] = 0, constructed using kernel regression and kernel density estimation (see for example Pagan and Ullah 1999). In order to apply these methods we need a non-negative, bounded, continuous and symmetric kernel Kwith RK(u)du = 1 and a bandwidth hdepending on the sample size n. The application in section 4uses a standard normal density for K. Then we construct first an estimate of the density of x∈Rmas ˆp(xi) = 1 n−1X j≤n,j6=i 1 hmKxi−xj h(4) 28
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 and next an estimate of the expected residual as ˆ E(εi|xi) = 1 n−1Pj≤n,j6=i 1 hmKxi−xj hεj ˆp(xi).(5) Letting ˆ θbe consistently estimated, define ei=yi−f(xi:ˆ θ) and define next a sample analogue of E[εE(ε|x)p(x)] by 1 n(n−1) X i≤nX j≤n,j6=i 1 hmKxi−xj heiej(6) Zheng then defines a standardised version of the test statistic by Tn=Pi≤nPj≤n,j6=iKxi−xj heiej hPi≤nPj≤n,j6=i2K2xi−xj he2 ie2 ji1 2 (7) and shows that under some regularity conditions, if h→0 and nhm→ ∞, then under the the null hypothesis, Tn→dN(0,1), while under the alternative hypothesis, Tn nhm/2→pc > 0.(8) Since Tnis a one-sided test statistic, we reject at level αif Tn> Zα, where Zαis the upper α-percentile of a standard normal variable. For example, we reject H0 at the 5 percent level if Tn>1.645 (Li and Racine, 2007). 2.1 Some comments Although the Zheng statistic looks complicated, there is a straight-forward intuition behind. The denominator in the expression for Tntakes care of the standardisation so we concentrate on the numerator. This is just a weighted sum of products of residuals. The statistic is built up as a sum of contributions from each observation i. The weighting ensures that only observations near the i’th receive non-negligible weight. The weighted product eiejis positive if residuals ejnear eihave the same sign as ei. When the true and the estimated models are smooth functions of x, this will happen if 06=g(xi)−f(xi:ˆ θ) = E(yi−f(xi:ˆ θ)|xi) = E(e|xi).(9) The weighting ensures that local discrepancies between gand fare detected even when positive and negative discrepancies would net out globally. The estimator in eq. (5) is a nonparametric kernel regression of the residuals εagainst x. It is useful to use this regression to produce smoothed plots of 29
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 the residuals against the independent variables. Confidence bands around the regression can be computed using that (nh)1/2[ˆ E(εi|xi)−E(εi|xi)] ∼N0, σ2(xi)p−1(xi)ZK2(u)du,(10) where σ2is the variance of εconditional on x.3For the purpose of testing the specification in the parametric model fwe may estimate this variance using the null hypothesis that E(y|x) = f(x|θ) a.e. In the case of a uni-dimensional x, we may estimate σ(x) by ˆσ2(x) = ˆ f(x)(1 −ˆ f(x)) ˆp(x)nh ZK2(u)du, (11) where ˆ fis a kernel regression of the parametric model probabilities against xand ˆpis the estimated density of x. For a normal density kernel, RK2(u)du =1 2√π. The idea that the bandwidth hshould depend on sample size may be unfamiliar. If the bandwidth is too large then averaging over too large neighbourhoods will smooth away relevant differences and hence introduce bias. If the bandwidth is too small then estimates will be noisy, they will be too much influenced by random fluctuations in data and hence variance will be high. Choosing the bandwidth appropriately as a function of sample size balances bias and variance and ensures that the bandwidth tends to zero at an appropriately slow rate, such that the asymptotical results obtain. These considerations suggest that using a large bandwidth will tend to reduce the Zheng statistic as positive differences between fand gin some regions of the data will then be averaged with negative differences in other regions. For a given sample size n, the bandwidth may be chosen by a rule such as, e.g. h=n−1 2m. This choice is easy and agrees with the conditions for the Zheng test. It is also possible to select a bandwidth using cross-validation in the regression of the residuals against x. This is however time consuming and the gain from doing it is not clear.4In applied research, it may be sufficient or even preferable to select a bandwidth by just inspecting the resulting regression and the confidence bands visually, this is so-called eye-balling, suggested by Pagan and Ullah when the dimension of xis low enough to make this feasible. 3 Reducing dimensionality A concern with the application of tests like the Zheng test, just as with any nonparametric technique, is the curse of dimensionality. The size of the dataset 3See Pagan and Ullah (1999). 4See the discussion in Li and Racine (2007). Generally all that is required for the test to be consistent is that bandwidths tend to zero as sample size tends to infinity but slowly enough that nmultiplied by the product of bandwidths tends to infinity. This is, however, not very helpful in finite samples. 30
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 necessary to achieve a given degree of precision increases exponentially in the dimension of x. Rather than use the test directly it is therefore useful instead to consider the test defined over a function of the data. Let t(·) : Rm→Rmtbe a measurable function such that t(x) has a density. Note that mt≤m, such that tis not in general injective. The inverse function of tis then a set-valued function defined by t−1(t0) = {x:t(x) = t0}. The idea is now to work with the mt-dimensional stochastic variable trather than with the m-dimensional stochastic variable x. So we apply the Zheng test to E(y|t) rather than to E(y|x). When mt< m there is a reduction of dimensionality and the test will be easier to apply. Note that E(y|t0) = E(y|x∈t−1(t0)) = E(E(y|x)|x∈t−1(t0)) (12) such that g(t) = E(y|t) and f(t:θ) = E(f(x:θ)|t) are the corresponding true and modelled conditional means, now defined over trather than over x. It is straight-forward that our null hypothesis H0implies the version reformulated in terms of functions t. If the model is true, then the model is also true as a function of any function t. H0⇒H∗ 0:∃θ0∀t P [f(t:θ0) = g(t)] = 1.(13) It is also clear that the converse hypothesis H1implies the converse of H∗ 0. That is, if the model is false, then there exists a function tsuch that the model is false as a function of t. H1⇒H∗ 1:∀θ∈Θ∃t P [f(t:θ) = g(t)] <1.(14) To verify the latter, simply choose a function tas t(x) = g(x)−f(x:θ). The reverse implications also hold: H∗ 0implies the converse of H∗ 1, which implies the converse of H1, which in turn implies H0. So H0is equivalent to H∗ 0and H1is equivalent to H∗ 1. 3.1 Some comments It is now possible to proceed by selecting various functions tof the data and to conduct the Zheng test (or another similar test) on the model conditional on these t’s. We cannot be sure to reject a false model for any given t, as the discrepancy between the parametric model and the true model may not lie in the direction of t. But if we reject the model for any t, then we are under the alternative hypothesis H∗ 1and hence also under H1, which means that the model has been rejected. Otherwise the model is accepted. The test based on functions twill be less powerful since we may not be able to choose appropriate t’s. On the other hand we gain the advantage that the dimension of tcan be chosen to be low, which greatly facilitates the use of nonparametric kernel methods. Note that it is not necessary that every variable in xactually appears in the model f. If a variable xois omitted from the parametric model we may define t(x) = x0to form a nonparametric omitted variable test. 31
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 In a discrete choice model where the systematic utility of an alternative is defined by βx with βfixed, we may test the model by defining t(x) = βx and take yas the binary indicator for choice of the same or another alternative. When the dimension of tis 1 or 2 it is feasible to select the bandwidth visually by plotting the nonparametric regression of the residuals against t(x) and varying the bandwidth until an appropriate number of features arise. It may be informative to inspect the data visually in this way and the Zheng test may then be used to check if detected deviations of the expected residual from zero are in fact significant. It is hard to very specific about how this should be done. But usually one has some sense of how complicated a potential relationship might be. A low bandwidth tends to lead to an estimated relationship with many peaks. If there are more peaks than one thinks is likely in the true relationship, then it is appropriate to increase the bandwidth. On the other hand, choosing a too large bandwidth will smooth away many differences. This may introduce bias so it might be better to use a somewhat smaller bandwidth. 4 Examples 4.1 Mode choice This section provides an illustrative application of the test to a mode choice model. The model is not intended as a serious model. It merely serves the purpose of showing how the Zheng test may be used to detect significant differences between model and data. I use a mode choice data set from Sweden comprising 1799 observations of trips to Stockholm using either air, bus, car or train. For each observation the data give travel time and cost for each alternative as well as some background variables. I estimated an MNL with alternative specific constants, mode-specific parameters for travel time and a common cost parameter. Write this model with utilities Ui=Vi+εi, where εiare iid. extreme value and Vi=βxi. Then I computed the predicted mode choice probabilities for the observations in the sample. Residuals were computed for each mode by subtracting the predicted probability from a dummy indicator for the observed choice. All the tests are carried out using uni-dimensional kernel regressions and density estimates were carried out using a normal density kernel and a bandwidth of n−1/2, where nis sample size and data (t(x)) are scaled to the unit interval. The choice of kernel is generally not important. Judging from the graphs in the following, the bandwidth seems appropriate. For computing confidence intervals around the kernel regression I used equation (10).5I trimmed away 1 percent of the sample, excluding the up5The MNL was estimated in Biogeme (Bierlaire, 2003, 2005), which was also used to compute the predicted choice probabilities. The Zheng test and the nonparametric regressions were performed in Ox (Doornik, 2001). The Ox code for the test is given in the appendix. 32
M. Fosgerau, Journal of Choice Modelling, 1(1), pp. 26-39 2.75 3.00 3.25 3.50 3.75 4.00 4.25 4.50 4.75 5.00 5.25 5.50 5.75 6.00 6.25 6.50 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Density_realincome 09:51:26 17-Jan-2008 Figure 1: Density of log of real income per and lower 0.5 percent after sorting on the independent variable, in order to exclude regions with very thin data and hence uncertain estimates. 4.1.1 Omitted variable test An income variable is available in the data but has not been included in the model. Suppose we want to test whether income could improve the prediction of choices. In this simple model one could of course just try estimating some specifications including an income variable. For more complicated models it could however be prohibitively time-consuming to estimate many different versions of the same model. In such cases an omitted variable test could be useful. Figure 1shows first a kernel density estimate of the log of real income. The Zheng statistics for the four alternatives are 3.3 (Air), 0.7 (Bus), 1.1 (Car) and 3.7 (Train), which means that the residuals depend systematically on income for the Air and Train alternatives. Figures 2and 3show the corresponding plots for the air and bus modes. It is clear that the residuals of choosing air increases with income. Thus there is a dependency of choices on income that is not captured by the model. A similarly clear relationship is not visible for the bus residuals. The dependency on income is, of course, unsurprising. The point here is that the test is well able to detect it. Income may be included in the model in various ways and the test does not indicate how it should be done. The smoothed residual plots suggest a specification that is linear in log income. It is 33