The provision of wage incentives: A structural estimation using contracts variation
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
D'Haultfœuille, Xavier; Février, Philippe Article The provision of wage incentives: A structural estimation using contracts variation Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: D'Haultfœuille, Xavier; Février, Philippe (2020) : The provision of wage incentives: A structural estimation using contracts variation, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 11, Iss. 1, pp. 349-397, https://doi.org/10.3982/QE597 This Version is available at: https://hdl.handle.net/10419/217190 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 11 (2020), 349–397 1759-7331/20200349 The provision of wage incentives: A structural estimation using contracts variation Xavier D’Haultfœuille CREST Philippe Février CREST We address empirically the issues of the optimality of simple linear compensation contracts and the importance of asymmetries between firms and workers. For that purpose, we consider contracts between the French National Institute of Statistics and Economics (Insee) and the interviewers it hired to conduct its surveys in 2001, 2002, and 2003. To derive our results, we exploit an exogenous change in the contract structure in 2003, the piece rate increasing from 202to 229euros. We argue that such a change is crucial for a structural analysis. It allows us, in particular, to identify and recover nonparametrically some information on the cost function of the interviewers and on the distribution of their types. This information is used to select correctly our parametric restrictions. Our results indicate that the loss of using such simple contracts instead of the optimal ones is no more than 16%, which might explain why linear contracts are so popular. We also find moderate costs of asymmetric information in our data, the loss being around 22% of what Insee could achieve under complete information. Keywords. Incentives, asymmetric information, optimal contracts, nonparametric identification. JEL classification. C14, D82, D86. 1. Introduction Over the past three decades, extensive attention has been devoted to asymmetries of information and their consequences in economics. These asymmetries play, in particular, a fundamental role in the economics of the firms (see Prendergast (1999) for a survey). Firms have to provide the right incentives to their workers, and design appropriate compensation plans, even when restricting to simple contracts such as piece rate, commissions at quota, or lump-sum bonuses. Indeed, a growing empirical literature shows that overall, incentives substantially increase workers’ productivity (see, e.g., Lazear (2000) Xavier D’Haultfœuille: [email protected] Philippe Février: [email protected] We would like to thank Gaurab Aryal, Steve Berry, Raicho Bojilov, Thierry Magnac, Arnaud Maurel, Aviv Nevo, Martin Pesendorfer, Rob Porter, Jean-Marc Robin, Jean-Charles Rochet, Mathieu Rosenbaum, Bernard Salanié, Michael Visser, and the participants of various seminars and conferences for helpful discussions and comments. We finally acknowledge Daniel Verger and the Unité Méthodes Statistiques of Insee for providing us with the data. ©2020 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE597
350 D’Haultfœuille and Février Quantitative Economics 11 (2020) or Paarsch and Shearer (2000)), and that the form of the payment scheme matters (Ferrall and Shearer (1999), Copeland and Monnet (2009), Chung, Steenburgh, and Sudhir (2014)). Our paper adds to this empirical personnel literature by quantifying the loss of using simple linear compensation contracts instead of nonlinear, optimal ones, and the importance of asymmetries between firms and workers. We use for this purpose contract data between the French National Institute of Economics and Statistics (Insee) and its interviewers. Insee is a public institute that conducts each year between twelve and twenty household surveys on different topics such as labor force, consumption or health. It hires interviewers to contact the households and conduct the corresponding interviews. We have data on three successive surveys on household living conditions (“enquête Permanente sur les Conditions de Vie des Ménages,” PCV hereafter) that took place in October 2001, 2002, and 2003. For each survey and all interviewers, we observe their average response rates, defined as the ratio of the number of respondents to the number of households each interviewer has to interview. These response rates vary with the effort the interviewers make to contact the households and to persuade them to accept the interview. Response rates also differ from one interviewer to another because of the heterogeneity in interviewers’ cost of effort, and differences between the geographical areas attributed to them. These unobserved effort and heterogeneity are the reasons why Insee faces an asymmetric information problem. To give incentives to its interviewers, Insee then uses a simple compensation scheme. Interviewers receive a basic wage (around 47euros in the three surveys), which does not depend on whether the interview is achieved or not, plus a bonus for each interview they conduct. The key point of the paper is to exploit the fact that the bonus changed in 2003, increasing from 202euros in 2001 and 2002 to 229euros in 2003. Moreover, we have reasons to believe that this increase was not due to a change in the cost of interviewers. To investigate the efficiency of simple linear compensation contracts and the importance of asymmetries between Insee and its interviewers, we rely on a structural principal-agent model that incorporates both adverse selection and moral hazard. We show that the cost function and the distribution of the interviewers’ types are partially identified nonparametrically using the exogenous change in contracts. An important feature of this result is that the information on the functions of interest are recovered using the interviewers’ program solely. This is convenient because it is very likely that Insee does not implement the optimal contracts, but only optimizes over linear ones. More generally, aiming at testing the optimality of the principal precludes any identification method relying precisely on this optimality. Importantly, also, our identification result is robust to the presence of selection effects, namely whether or not the new compensation scheme has attracted better interviewers. If the identification argument developed for the moral hazard part is specific, our result on the adverse selection part could actually apply to many adverse selection models, including regulatory contracts, and nonlinear and price discrimination models.1All of these models share a common underlying structure for which our procedure 1For an incomplete list of empirical papers in these fields, see Ivaldi and Martimort (1994), Wolak (1994), Gagnepain and Ivaldi (2002), Miravete (2002), Leslie (2004), Miravete and Roller (2005), Lavergne
Quantitative Economics 11 (2020) The provision of wage incentives 351 is well adapted and can be useful to study their nonparametric identification. Though the models somewhat differ, our identification result is therefore connected with those of Perrigne and Vuong (2011b), Aryal, Perrigne, and Vuong (2016), Luo, Perrigne, and Vuong (2018), and Aryal and Gabrielli (2018) on regulation, insurance models, unidimensional and multidimensional nonlinear pricing, respectively. An important difference with these papers is that we neither rely on the knowledge of the principal’s objective function, nor on the optimality of observed contracts. On the other hand, the identification of the cost function and the distribution of the interviewers’ types relies on exogenous variation in contracts, and is only partial with one exogenous change. Our identification argument is also related to the identification of first-price auctions models with risk-adverse bidders, using exogenous variations in the number of bidders (Guerre, Perrigne, and Vuong (2009)). Interestingly also, we show that our problem boils down to the identification of nonparametric transformation models or, equivalently in duration models, generalized accelerated failure time model, with discrete regressors. This question has been studied by Abbring and Ridder (2015), but under some large support conditions and regularity conditions at the boundary of this support. We show that without such conditions, the model is still partially identified. Beyond identification, we also develop a nonparametric estimation procedure using our identification method. We estimate nonparametrically bounds on the cost function and the distribution of interviewers’ type. In a second step, we introduce parametric specifications in line with the nonparametric estimates of the interviewers’ cost function and distribution of types. As the model is not point identified nonparametrically, such restrictions are necessary to estimate the policy effects we are interested in. However, contrary to most papers in the personnel literature, which adopt directly a parametric framework, our specifications are driven by the nonparametric analysis. Studying Insee and its interviewers, our method allows us, first, to conclude that the loss of using a simple contract instead of an optimal one is rather small, around 16%. Even if the theoretical literature concludes that optimal contracts are in general nonlinear (see Laffont and Martimort (2002), for a survey),2simple compensation schemes such as piece rates and bonuses are usually thought of as the best compromise between efficiency and ease of implementation (Raju and Srinivasan (1996)). Our result supports this claim and may explain why simple contracts are so popular and widely used by firms. This idea is also in line with the theoretical findings of Wilson (1993, Section 6.4), Rogerson (2003), and Chu and Sappington (2007), who show that simple tariffs can secure more than 70% of the maximal surplus. Firms can adopt simple compensation systems and still give the right incentives to workers. Little empirical work has however tried to estimate the loss associated with the use of simple compensation scheme and the empirical personnel literature mentioned previously usually abstracts from these issues. An exception is Miravete (2007), who reports a loss of only 3%.Ferrall and Shearer (1999), on the other hand, concluded that simple nonlinear compensation plans lead to substantial inefficiencies. and Thomas (2005), Crawford and Shum (2007), Perrigne and Vuong (2011a), Miravete (2007), Gagnepain, Ivaldi, and Martimort (2013), Lim and Yurukoglu (2018), and Kang and Silveira (2018). 2An exception is the result of Holmstrom and Milgrom (1987).
352 D’Haultfœuille and Février Quantitative Economics 11 (2020) Our method also allows us to recover what Insee’s surplus would have been under complete information. Independently of the issue of contracts’ optimality, asymmetries create inefficiencies because of the informational rent captured by the agents. Measuring this rent is therefore important for the firm. This question is central in the insurance literature (see Chiappori and Salanié (2003), for a survey), or in the auction literature (see Perrigne and Vuong (1999), for a survey). On the contrary, few empirical works have focused on quantifying the magnitude of such asymmetries between firms and workers in the personnel literature. We find moderate cost of asymmetric information, the estimated expected surplus under incomplete information being 78% of the full information surplus. This loss (22%) is in particular smaller than the one reported by Ferrall and Shearer (1999) who found an efficiency loss of 33%. Overall, in our data, the surplus under asymmetric information and with a simple linear compensation plan is 66% of what it could be under complete information. The main part of this loss (65%)isdueto incomplete information whereas the last 35% are associated with the simple payment scheme. The paper is organized as follows. Section 2presents institutional details and the data at our disposal. In Section 3, we focus on the interviewers’ behavior. We develop a simple theoretical model and show that it is partially identified thanks to the exogenous change in the contract. We then propose estimators for the corresponding bounds and show their consistency. Finally, we estimate these bounds on the data. Section 4focuses on the policy analysis. We show how the information on interviewers can be used to recover counterfactual parameters. We then study the optimality of the linear contracts used by Insee and the importance of asymmetries in this context. Section 5concludes. 2. Institutional details and data description The French National Institute of Economics and Statistics (Insee) conducts each year between twelve and twenty household surveys on different topics such as labor force, consumption, or health. For that purpose, Insee used to draw, until 2009 and approximately every 10 years, a large sample of housings from the exhaustive census database. This sample consisted of geographical areas called primary units. All survey samples were then drawn from these primary units. To conduct the interviews, Insee hired interviewers who live close to the primary units, in order to limit their traveling costs. Interviewers’ work is similar for almost all surveys. First, Insee gives them a list of sampled households to interview in their designated area, as well as some characteristics of the housings and households, as described in the census database. Interviewers then have to locate precisely the housings of their sample in order, for instance, to identify unoccupied or destroyed housings. After that, they try to contact the households. This stage is the main part of their job and usually takes several days. Usually, interviewers have to go to the housings several times and leave phone messages before coming in contact with the household. Finally, once contacted, interviewers have to convince the households to accept the survey. In theory, it is usually mandatory to participate to a survey by Insee. In practice, during the period we consider hereafter, more than 90% of
Quantitative Economics 11 (2020) The provision of wage incentives 353 households accepted to participate, once they had been contacted.3In a typical household survey, it takes around one hour to go through all the questions. In compensation, interviewers were paid in a similar way for all household surveys until 2013. They received a basic wage for each household they have to interview, plus a bonus for each interview they achieved. They were also reimbursed for all their expenses, such as the travel costs or the meals they have to take during their work. We have data on three successive surveys on household living conditions (“enquête Permanente sur les Conditions de Vie des Ménages”, PCV hereafter), which took place in October 2001, 2002, and 2003.4Each survey comprises a fixed part, which is identical for each edition (representing more than half of the questions), and a complementary part, which changes every year. In 2001, 2002, and 2003, the focus of the survey was put respectively on the use of new technologies, participation in associations and education practices in the family. For each survey, our dataset consists of the list of all housings in the survey sample, excluding secondary, unoccupied and destroyed housings. For each housing, we observe some of its characteristics in the 1999 census, namely the number of rooms, the household size, and the age of the reference person. We also observe the identification number of the interviewer in charge of interviewing the corresponding household, and a dummy indicating whether the interview was conducted or not. Table 1summarizes the main information about the three surveys, on the whole sample of households. There were between 379 and 478 interviewers in each survey. On average, each interviewer was assigned around 16 households in 2001 and 2002, and 28 in 2003. The 2001 and 2002 surveys display very similar patterns. In particular, their average response rates, defined as the ratio of the number of respondents to the number of housings, are not significantly different at the 5% level (785and 777%, resp.). Their distribution functions are also very close (see Figure 1), with a p-value of the two-sided Kolmogorov–Smirnov test equal to 087. On the other hand, the average response rate is significantly higher in 2003 (807%), and the distribution function of the 2003 survey stochastically dominates the one of 2001–20025(see Figure 1), with a p-value of the onesided Kolmogorov–Smirnov test equal to 0003. We also note that the distribution functions displayed in Figure 1exhibit several jumps, especially at 05,067,and1.These Table 1. Descriptive statistics on the full sample. Number of Number of Average Year Interviewers Households Response Rate 2001 379 173785% 2002 478 154777% 2003 453 280807% 3Insee never fines households that do not participate, but interviewers can use the argument that the survey is mandatory to convince households to participate. 4We also have some limited information on interviewers that we use at the end of our analysis; see Appendix Afor details on these data. 5The average response rate on 2001–2002 is defined as the ratio between the total number of interviews and the total number of households, where the 2001 and 2002 data are pooled.
354 D’Haultfœuille and Février Quantitative Economics 11 (2020) Figure 1. Distribution functions of the response rates on all interviewers, for all households. jumps are due to the fact that the response rates are ratios of two integers, and the number of households to interview is rather small.6 There are two main differences between the 2003 and the other two surveys. The first one is related to its sampling design, and the second to its payment scheme. As previously mentioned, the PCV surveys are drawn from primary units. This was the case for the three surveys we consider. However, the sample was approximately twice as large in 2003 as in 2001 and 2002. Besides, because the 2003 survey focused on families, housings in which a family lived at the time of the census were overrepresented in 2003. As a result of this overrepresentation, housings in which a family lived at the time of the census represent 545% of the housings in 2003, as opposed to 444% and 483% in 2001 and 2002. Because families are on average easier to contact than, for instance, single persons, this difference may partly explain why response rates were higher in 2003. To control for this sampling effect and make comparisons possible for the three surveys, we restrict hereafter our attention to such housings occupied by families. These were the only differences in the survey designs of the three surveys. In particular, the corresponding subsamples of families were drawn similarly. Table 2shows that, as expected, the average response rates for families are higher than in the general population (resp., 790%,798%,and831% versus 785%,777%, Table 2. Descriptive statistics on the subsample of families. Payment per Household Average Income Number of Number of Average Year Interviewers Families Response Rate Basic Bonus Basic Bonus Total 2001 377 835 790% 47203393 1350 1743 2002 471 685 798% 47202322 1119 1441 2003 453 1524 831% 46229701 2897 3598 6Because of this small numbers of households, it is logical, from a pure statistical point of view, to observe more jumps at 05or 067 as more integers can be divided by 2or 3.
Quantitative Economics 11 (2020) The provision of wage incentives 355 and 807%). When comparing the three surveys on families only, we find however the same pattern as in Table 1. The difference between the 2001 and 2002 surveys is not significant (790% and 798%, resp.), whereas interviewers achieve significantly higher response rates in 2003 (831%). There is also a second difference in the three surveys, namely their payment schemes. Whereas the basic wage is nearly constant the 3years, at a low level (47euros in 2001, 46euros in 2002 and 2003),7the bonus for achieving an interview with a family was 229euros in 2003, compared to 203and 202euros in 2001 and 2002. We use this modification afterwards to identify the principal-agent model that we consider in the following section. 3. The interviewer’smodel We first model the interviewers’ decision, in particular to recover their utility function. We use this utility function in the next section to quantify the loss due to linear contracts and asymmetric information. 3.1 The interviewers’ program We suppose that interviewers decide on the effort they spend to try to contact each household. Instead of modeling effort, we model directly the probability of contact that each interviewer fixes for each household. These households are heterogeneous and may be easy or difficult to contact, depending on their characteristics. Single persons living in urban areas are difficult to contact, for instance, because they spend relatively little time at home, and digital locks make a direct contact more difficult to establish. Interviewers do not face such barriers in the countryside, and families are on average more at home. Once we restrict our attention to an interviewer’s area and to the housings in which a family was living in 1999, however, households appear to be almost homogeneous ex ante. To support this claim, we regress the response rates of interviewers on the mean of the 1999 census characteristics (household size, number of rooms, and age of the reference person), controlling for interviewers and years fixed effects. While household size has a positive and significant effect when considering the whole sample, this effect disappears when restricting to the sample of families. None of the other census variables are significantly different from zero. As each interviewer works in a small and specific geographic area, this result does not really come as a surprise. In each restricted area, housings in which a family was living are, ex ante, quite similar and homogeneous for the interviewers. Because families are homogeneous in terms of contact ease, we suppose that interviewers treat them similarly and take the same decision for all of them. An interviewer thus decides, for each household, with which probability yhe wants to survey it, and produces his effort accordingly. As detailed below, this probability is not equal to the actual response rate because of randomness in interviewers’ work. The expectation of the cost to reach a probability yis supposed to depend on the survey x,thenumbern 7All figures are in 2002 euros.
356 D’Haultfœuille and Février Quantitative Economics 11 (2020) of households to interview but also the area and the interviewer herself. We summarize by θthe heterogeneity term in cost related to the interviewers and their areas. For simplicity, we refer subsequently to interviewers’ type, but one should keep in mind this dual aspect of θ. At the end, we denote by C(nxyθ) the expected cost of reaching a probability yin survey x, for an interviewer of type θwith ninterviews to conduct. To give the interviewers incentives to achieve high response rates, Insee provides them with a bonus if they realize the interview. Let δ(x) and w(x) denote respectively the bonus and basic wage chosen by Insee for survey x. In this case, the interviewer receives w(x)+δ(x) when the interview is achieved and w(x) otherwise. Hence, if the interviewer with nhouseholds to interview implements a probability yof conducting the survey for each household in his sample, he obtains on average a total wage of n(δ(x)y +w(x)). We suppose hereafter that interviewers are risk-neutral and have a quasi-linear utility function. In this case, an interviewer of type θchooses a probability y(nxθ) satisfying y(nxθ) ∈argmax ynw(x) +δ(x)y−C(nxyθ) (3.1) We denote by Y=y(NXθ) the actual probability chosen by the interviewer.8Y is not observed by the principal Insee (nor the econometrician), which is the source of moral hazard here.9Instead, it only observes the number of interviews Rthe interviewer eventually does in survey X. Our first assumption relates this observable output Rwith Y. Assumption 1 (Independence in households reactions). R|NXY ∼Binomial(N Y). We thus suppose that each household reacts independently from each other. Independence between households seems very likely here, as the households to interview are not neighbors in general, contrary to what happens in labor force surveys for instance. Next, we impose a separability and regularity conditions on the cost of interviewers. To cope with potential selection effects, we introduce here S, the dummy of being a “stayer.” Specifically, S=1if the interviewer participates to the 2003 and either the 2001 or 2002 surveys, S=0otherwise. Besides, for any random variables Uand V,weletFU (resp., FU|V) denote the cumulative distribution function of U(resp., of Uconditional on V). Finally, with a slight abuse of notation, we denote by [a b]closed intervals of the real line, even if bis possibly infinite. Assumption 2 (Cost separability and continuous distribution of types). C(nxyθ) = θC(nxy),where C(nx·)is twice continuously differentiable with ∂2C/∂y2>0for all y∈(01).Moreover,Fθ|XS has support [θθ]with 0≤θ≤θ≤+∞and it is continuously differentiable with density fθ|XS. 8As usually, capital Latin letters correspond to random variables, while their lowercase counterpart are realizations of these variables. 9Even if we assume that interviewers are risk-neutral, moral hazard affects the design of contracts, because it means that Insee can only design contracts based on Nand R, rather than on Nand Y.Wecome back to this issue in Section 4.2.
Quantitative Economics 11 (2020) The provision of wage incentives 363 and for any k≥1,H(k+1)(y) =H(k)[H(y)].ThenCis identified in two steps. First, one can show that αis identified by α=lnδ(2)/δ(1) lnlim y→0H(y)/y Second, C(y) can be shown to be point identified by C(y) =c0⎡ ⎢ ⎢ ⎢ ⎣ lim k→∞δ(1) δ(2)k H(k)(y) lim k→∞δ(1) δ(2)k H(k)(y0) ⎤ ⎥ ⎥ ⎥ ⎦ α This formula shows that any small variation in Haround 0is amplified when taking the limit, rendering estimation of Cbased on this equation very difficult. Related to this issue, no paper has addressed so far the estimation of nonparametric transformation or GAFT models with only discrete regressors. Theorem 3.2 is also related to Guerre, Perrigne, and Vuong (2009), who show that exogenous changes are necessary but also sufficient to point identify first-price auction models with risk averse bidders. The reason why they obtain point identification rather than partial identification as here is that in their framework, the bidders’ strategies cross at the lowest valuation, and this crossing point can be used for identification. In our framework but with other types of contracts, the functions FY|X=1S=1and FY|X=2S=1could cross inside the support of Y, leading also to point identification (see D’Haultfœuille and Février (2010)).15 Finally, our result imply that standard parametric models on Cand Fθ|X=xS=s are identified with an exogenous change. For instance, the parameters of a lognormal or Weibull distribution are identified thanks to the knowledge of Fθ|X=xS=son the sequence (θk)k+1{x=2}∈K. Actually, because we retrieve an infinite sequence of points on Cand Fθ|XS, such standard parametric models are overidentified. The sequences (C(yk))k∈Kand (Fθ|X=xS=s(θk))k+1{x=2}∈Kmay thus serve as a guidance for choosing appropriate parametric restrictions, as will be the case in Section 4.3 below. 3.3 Nonparametric estimation of the cost function and the distribution of types We now turn to the nonparametric estimation of Cand Fθ|XS.Inparticular,westudy the behavior of these estimators when the number of interviewers tend to infinity in the sense that L≡min(xs)∈{12}×{01}#{i:Xi=x Si=s}→∞. We impose hereafter the following standard assumption of independent sampling. 15 Another solution to recover point identification would be to use the principal’s program together with restrictions on its objective function. In our framework yet, this program has no identification power on C and Fθ|XS|, but simply allows us to recover its objective function. See Section 4.2 below for more details.
364 D’Haultfœuille and Février Quantitative Economics 11 (2020) Assumption 6 (Independent sampling). For any (x s) ∈{12}×{01},the sample (θiRiNi)i:Xi=xSi=sis made of i.i.d.vectors. Our nonparametric estimation method follows closely the identification strategy and may be decomposed into two steps. We first estimate the conditional distribution FY|XS of the unobserved probabilities. We then estimate bounds on the primitive functions Cand Fθ|XS, using the result of Theorem 3.2. For the first step, we use a sieve maximum likelihood estimator (see, e.g., Chen (2007), for a survey on sieve estimation). We choose to approximate the densities16 fY|XS by functions of the sieve space FL=f:0≤f≤MlnKL1 0 f(x)dx=1and f∈PKL where PJdenotes the space of polynomials of order at most J,Mis a constant and (KL)L∈Nis an increasing sequence tending to infinity. We thus approximate the conditional density fY|XS by squares of polynomials that integrate to one. Squares of polynomials are convenient because they ensure that the estimated density is positive, are easy to integrate, and lead to a simple likelihood.17 To see this, let us consider f(·; a)∈FL defined by f(x;a)=KL k=0 akxk2 ≡ 2KL k=0 bk(a)xk where a=(a0aKL)and bk(a)=min(kKL) =max(0k−KL)aak−. The likelihood of an observation corresponding to f(·; a)is, by independence between Yand Nconditional on X=xS =s, Pr(R =r|N=nX =x S =s) =EPr(R =r|N=n Y X =xS =s)|X=xS =s =r nEy(xθ)r1−y(xθ)n−r|X=xS =s =r n1 0 2KL k=0 bk(a)yr+k(1−y)n−rdy =r n2KL k=0 bk(a)B(r +k+1n−r+1) (3.4) 16Assumption 2and the equality Fθ|X=xS=s(δ(x)/C(y)) =1−FY|X=xS=s(y) ensure that the density of Yconditional on XS does exist. 17We also restrict ourselves to bounded polynomials. This ensures that FLis compact and simplifies the consistency proof.
Quantitative Economics 11 (2020) The provision of wage incentives 365 where B(··)denotes the beta function. We let fY|XS denote the maximum likelihood estimator (over FL)offY|X=xS=s.18 We then estimate FY|XS and F−1 Y|XS by FY|XS(y) = y 0 fY|XS(u) du and F−1 Y|XS(u) = F−1 Y|XS(x). We now turn to the estimation of Cand Fθ|XS.First,weestimateHand Yx(x ∈ {12})by respectively H(y) = F−1 Y|X=2S=1◦ FY|X=1S=1(y) and Yx=[ F−1 Y|X=xS=1(τL) F−1 Y|X=xS=1(1−τL)], for a sequence τLtending to 0. Second, we define Kand ( yk)k∈ K as before. 0∈ Kand if k≥0is such that yk∈ Y1,thenk+1∈ Kand yk+1= H( yk). Similarly, if k≤0and yk∈ Y2, then k−1∈ Kand yk−1= H−1( yk). Note that θkdoes not need to be estimated. Then we consider plug-in estimators for the bounds on Cand Fθ|XS: C(y) =sup0δ(2)/δ(1)kc0:k∈ K yk≤y C(y) =inf+∞δ(2)/δ(1)kc0:k∈ K yk≥y Fθ|X=xS=s(θ) =sup01− FY|X=xS=s( yk):k+1{x=2}∈ Kθk≤θ Fθ|X=xS=s(θ) =inf11− FY|X=xS=s( yk):k+1{x=2}∈ Kθk≥θ Theorem 3.3 below establishes the consistency of these bounds under the following regularity conditions. Assumption 7. For all k∈K,yk/∈{infY2supY1}.For all (xs) ∈{12}×{01}, limθ→θθ2fθ|X=xS=s(θ) =0and C is bounded on (01/2).Either θ=0and limy→1C(y) C(y)2 exists and is finite,or θ > 0and fθ|X=xS=s(θ)=0.For all u>0,E(uN|X=xS =s) < ∞. The condition on the sequence hold automatically when Supp(θ) =R+.Otherwise, we just impose that the smallest (resp., largest) value of ykis not equal to inf Y2(resp., supY1), which is a very mild restriction on the choice of y0. The conditions on fθ|X=xS=s and Censure that fY|X=xS=sis continuous on [01],whetherornot1/θ and θare finite. The last condition imposes light tails for the conditional distribution of N. Theorem 3.3. Suppose that Assumptions 1–7hold,KL→∞and K2 LlnKL/L →0.Then FY|X=xS=sis uniformly consistent on [01].Moreover,for any sequence τL→0such that P(supy∈[01]| FY|X=xS=s(y) −FY|X=xS=s(y)|<τ L)→1, Fθ|XS(θ) and Fθ|XS(θ) are consistent for all θ>0. C(y) and C(y) are consistent on every y∈(01)\{ykk∈K\{0}}.Finally,for all k∈K, yk C( yk)= C( yk)P −→ ykC(yk) Theorem 3.3 has four parts. The first establishes the uniform consistency of the nonparametric estimator of FY|X=xS=s. The second shows that the estimated bounds on 18 Intuitively, this estimator weights more interviewers with a large N, because for them, the distribution R/N is closer to the one of Y: by the central limit theorem, we have approximately R/N =Y+ Y(1−Y)/Nε,withε|NY ∼N(01). Because of the error term Y(1−Y)/Nε,itismoredifficulttodiscriminate between two parametric distributions on Ywhen Nis small.
366 D’Haultfœuille and Février Quantitative Economics 11 (2020) Fθ|XS are consistent. The third shows the convergence of Cand Coutside the sequence (yk)k∈K. Even if consistency fails in general on this sequence, the last part of the theorem shows point consistency in R2of the estimated sequence ( yk C( yk)). As a consequence Cand Fθ|XS are well estimated on the sequences where they are point identified, while sharp bounds are consistently recovered anywhere else. Consistency of the bounds requires to choose τLappropriately, which is difficult because the rate of convergence of FY|X=xS=sis unknown. But interestingly, if we consider τLfixed, independent of L, one can show that we still estimate consistently the bounds on Cand Fθ|XS, but only on subsets of (01)and R+, respectively. Elsewhere, we get outer bounds on these functions because, basically, KKwith probability approaching one. Therefore, letting τLtend to zero at the appropriate rate is required only to estimate optimal bounds everywhere. 3.4 Results We estimate in a first step FY|X=xS=sby the sieve MLE proposed above. As usual, there is a trade-off between bias and variance in the choice of KL. Empirically, the estimates do not seem to be too smooth or too erratic for KLbetween 3and 6. Results are quite similar in this range, and we choose KL=3for the stayers and KL=2for the movers. The corresponding estimates are displayed in Figure 3. As predicted by the theory, the distribution function of y(2θ)for stayers dominates stochastically the one of y(1θ)on most part of (01)(see the left graph). We also observe that for movers, the estimated distribution of y(2θ) dominates stochastically the one of y(1θ) (see the right graph). This arises because of incentive effects but also possibly because of selection effects. We discuss in the next section the existence of selection effects in our context. We now estimate nonparametrically the sharp bounds on Fθ|X=2and C.Fθ|X=2is interesting as it corresponds to the distribution of interviewers’ types on the 2003 survey. We obtain similar patterns for Fθ|X=1.Wefirstchooseastartingvaluey0close to the median of FY|X=1S=1, namely y0=08, in order to get more precise estimates for central Figure 3. Sieve MLE estimates of FY|XS.Notes:161 observations for X=1S =0,79 for X=2S=0,and374 for S=1.
Quantitative Economics 11 (2020) The provision of wage incentives 367 Figure 4. Estimated bounds on Fθ|X=2and C.Notes:the95% confidence intervals are computed by bootstrap. values of Fθ|X=2and C.19 For that y0, we impose the normalization C(y0)=δ(1),which is equivalent to imposing θ(1y0)=1. We then choose τL=005, which leads to estimating 14 points in K. Of course, other choices of τLenlarge or shrink K. For instance, with τL=001 and τL=010, we get respectively 22 and 11 points. But the bounds on the functions are not altered substantially; they only change for large values where standard errors are large anyway. Figure 4displays the estimates of the bounds on Fθ|X=2and C,andtheir95% confidence interval obtained by bootstrap. The bounds on both functions are close and we are able to correctly retrieve their shape. The highly convex form of the cost function shows in particular that incentives are relatively large for small values of the production but significantly lower for higher ones. Finally, the width of the confidence intervals on the bounds of Fθ|X=2(resp., C) increases with |θ−1|(resp., |y−08|), reflecting the fact that, as expected, the estimation error increases with |k|. 3.5 Tests of the model and robustness checks The results above rely on a few assumptions that we now check. The first part of Assumption 3implies that y(nxθ) does not depend on n,sothatYis only a function of (Xθ). Then, by the second part of Assumption 3, N⊥⊥ Y|XS (3.5) This condition is testable. To see this, note that by definition of Y,E(R|NYXS|)= NY.Then E(R/N|NXS) =E(Y|NXS)=E[Y|XS](3.6) 19We have checked that other values of y0do not modify the choice of the parametric families that is made using our nonparametric estimates.
368 D’Haultfœuille and Février Quantitative Economics 11 (2020) In other words, R/N is mean independent of Nconditional on XS.Wetesttherestriction (3.6) by considering the model R/N =ζ(XS) +g(N) +ε where E(ε|NS) =0,ζis a year times participation status fixed effect and we distinguish in Xbetween the 2001 and 2002 survey to be robust to the exogeneity condition (Assumption 5). Equation (3.6) implies that the function gshould be equal to zero. We perform a test of this restriction using linear, quadratic, and flexible parametric specifications for g. We consider for this latter specification the piecewise linear function g(n) =g1n+g2(n −10)++g3(n −20)+(with x+=max(0x)), which can detect more complex dependence between R/N and Nunder the alternative.20 Results are presented in Table 3.Weacceptthenullhypothesisthatg=0at standard levels in any of the three specification. Moreover, the point estimates are very small. The results of the linear specification for instance imply that a very substantial increase of 10 households to interview is associated to a small increase of one percentage point in the average response rate. Next, our results crucially hinge upon Assumption 5. We present hereafter three suggestive tests of X⊥⊥ θ|S=1. The idea behind the first is that households are more or less difficult to contact depending on their characteristics. If the distribution of their characteristics changed in 2003, this could induce a shift in the distribution of θ, thus violating Assumption 5. We thus check whether the average characteristics of the housings attributed to stayers differ systematically between 2001–2002 and 2003. We can only use housings’ characteristics that are available for both respondents and nonrespondents. Table 3. Test of Assumption 3based on regressions of response rates on functions gof subsample sizes. Piecewise Variable Linear gQuadratic gLinear g Subsample size 0001 (0001)00028 (00028)00045 (00032) Subsample size squared – −00001 (00001) – (Subsample size −10)+––−00058 (00043) (Subsample size −20)+––00026 (00031) Participation status ×Yes Yes Yes year included R20014 0014 0016 p-value of the test g=0029 054 052 Note:984 observations. We control for participation status interacted with the year. Standard errors are clustered by interviewers to take into account the dependence arising because of θi. 20We consider in Appendix B.1 another test using not only the first moment of R, but its whole distribution. However, this test also relies on Assumption 1, and thus does not solely test Assumption 3.Again,we fail to reject the null at standard level with this alternative test.
Quantitative Economics 11 (2020) The provision of wage incentives 369 Table 4. Stability test of housing characteristics (p-values). Stayers Variable All Stayers With exp >10 Proportion of collective housings 021 020 Average number of rooms 099 093 Proportion of housings in rural areas 011 012 Proportion of housings in small towns 028 049 Proportion of housings in large towns 083 093 Number of observations 748 280 Note:p-values of the Kolmogorov–Smirnov test that the distribution of proportions of, for example, individual housing, remains constant between 2001–2002 and 2003. We use hereafter variables that are known to be correlated with nonresponse, namely the dummy of being in a collective housing, the number of rooms and the dummies of being in rural areas, in small towns (of less than 100,000 inhabitants), and in large towns (of more than 100,000 inhabitants). We then perform a Kolmogorov–Smirnov test that the distributions of the averages over stayers of these variables did not change between 2001–2002 and 2003. The results are displayed in the first column of Table 4.Notestis significant at the 10% level, which clearly supports our assumption. Our second test aims at testing whether areas could have been attributed to interviewers differently in 2003, as a function of the areas’ and the interviewers’ characteristics. A change in the “match” between interviewers and areas could change the distribution of θ, even though both marginal distributions (of interviewers and areas) have remained the same. To test for this issue, we investigate whether experienced interviewers had a different distribution of housing characteristics in 2003 than in 2001–2002. We rely on the same Kolmogorov–Smirnov tests as above, but now focusing only on interviewers’ with more than 10 years of experience (results are similar with other thresholds). The results, displayed in the second column of Table 4, support again our claim that the assignment of areas to interviewers was not different in 2003. A final concern on the assumption that X⊥⊥ θ|S=1is related to experience: stayers have accumulated experience between 2001 or 2002 and 2003, and could therefore be more efficient in 2003. This learning-by-doing effect is unlikely to be of first order here since interviewers have already 9years of experience on average. Yet, we can test for this possibility under some assumptions. Let Edenote the experience of an interviewer and suppose that ∂C/∂y(y;E) =β(E))(y/(1−y))ζ. Combined with Assumption 1,thisimplies that the dummy yixk of whether interviewer imanaged or not to interview household kin survey x∈{012}satisfies yixk =1γ1{x=2}+ β(Eix)+ θi+εixk ≥0(3.7) where γ=lnδ(2)/ζ, β(Eix)=−ln(β(Eix)), θi=−1 ζln(θi)and the (εij k )ijk are independent and follow a logit distribution. We recall that γmeasures the incentive effect of 2003 with respect to 2001–2002, and is therefore key in our analysis. Hence, we want
370 D’Haultfœuille and Février Quantitative Economics 11 (2020) Table 5. Effects of interviewers’ experience versus incentive effects. p-Value of No Specification of β(E) Estimate of γEstimate of bEffect of Experience β(E) =1024 (006)–– β(E) =E−b026 (007)−009 (012) 044 β(E) =1+(exp(−b) −1)1{E≥3}025 (006)−008 (015) 060 β(E) =1+(exp(−b) −1)1{E≥5}024 (006)005 (023)084 Note: Estimates of b1in the fixed-effect logit model (3.7), with various parametrizations of β(E). Standard errors under parentheses. 9851 observations. to check whether the estimate of γis sensitive to the introduction of the effect of experience. We estimate for that purpose Model (3.7) using various parametric specifications for β(Eix). Given that Eix =Ei+x, we cannot identify flexible functions β, since β(Eix)would become collinear with 1{x=2}. For instance, we cannot identify (b1b2)if we let β(E) =exp(b1E+b2E2). We consider hereafter three specifications: β(E) =E−b, β(E) =1+(exp(−b) −1)1{E≥3}and β(E) =1+(exp(−b) −1)1{E≥5}. The results are displayed in Table 5. We obtain two conclusions. First, the coefficient of γis hardly affected by the introduction of experience. Second, experience has no significant effect in the three specifications we consider on β. Hence, this test suggests that the effect of experience would threaten our conclusion. Finally, when combining Assumptions 1,3, and the polynomial restriction on fY|XS behind the sieve MLE, we obtain a relatively parsimonious parametric model for the distribution of Rconditional on N. Specifically, the probabilities Pr(R =r|N=n X = xS =s) only depend on KL+1parameters (see equation (3.4)). Hence, if seen as a parametric model (where KLis fixed as Ltends to infinity), the model is largely overidentified, as there are many more moment conditions corresponding to all the equalities implied by equation (3.4), than parameters. To assess whether this model fits well the data, and thus whether the polynomial restriction of fY|XS is reasonable (under the maintained Assumptions 1and 3), we consider a GMM overidentification test, for each of the four subpopulations {X=xS =s},(xs) ∈{01}2. Given the moderate subsample sizes and the large number of moment conditions, we can expect the quantiles of the asymptotic distribution of this test to underestimate the true quantiles under the null. To conduct more reliable inference, we thus use the bootstrap instead.21 We compute bootstrap critical values by first drawing with replacement the (Ni)i:Xi=xSi=sand then drawing R|Ni=nusing (3.4), with areplaced by the sieve MLE estimator. The p-values of the tests are displayed in Table 6. For the four subpopulations {X=xS =s},(x s) ∈{01}2, we fail to reject the null hypothesis at all usual levels. This suggests that under the maintained Assumptions 1and 3, the polynomial restriction we rely on in the sieve MLE is reasonable. 21When using the asymptotic distribution instead of the bootstrap, we indeed obtain much higher pvalues.
Quantitative Economics 11 (2020) The provision of wage incentives 371 Table 6. Overidentification tests of Assumption 3and the polynomial restriction on fY|XS. Stayers 2001–2002 Stayers 2003 Movers 2001–2002 Movers 2003 p-value 070 023 074 048 Number of obs. 348 374 161 79 Note: GMM overidentification test of (3.4), with KL=3for the stayers and KL=2for the movers. The optimal weighting matrix is estimated using the sieve MLE estimator of aand the critical values are estimated by bootstrap. 4. Policy analysis In this section, we compare the current contracts with the optimal, nonlinear ones, and with settings without asymmetries of information. The information that we have recovered so far on the interviewers’ type and their cost function is used for that purpose. But before performing this analysis, we have to check that there is no selection effects in our context. If these potential effects were not an issue for identifying the interviewers’ utility function, they would complicate substantially the policy analysis because basically, different contracts would select different types of interviewers. 4.1 Testing the absence of selection effects To evaluate the average response rate that would prevail under contracts that differ from the actual ones, we have to take into account selection effects, namely that more attractive contracts may attract betters interviewers, for instance. We provide here statistical evidence that this is likely not the case in our context. More precisely, we test whether the distribution of movers of the 2001–2002 surveys are identical to the one of the movers of the 2003 survey (while the distribution of stayers could differ from them). Formally, this amounts to test H0:θ⊥⊥ X|S=0. To perform such a test, remark that under H0,we have, for s∈{01}, FY|X=1S=s(y) =Fθ|X=1S=sθ(1y) =Fθ|X=2S=sθ2y2θ(1y) =FY|X=2S=sy2θ(1y) This shows that under H0,F−1 Y|X=2S=s◦FY|X=1S=sdoes not depend on s.Asaresult, FY|X=2S=0=FY|X=1S=0◦H−1(4.1) In other words, if we transform the distribution of the probabilities chosen by the 2002 movers using the quantile-quantile transform of the stayers, we should obtain the distribution of the probabilities chosen by the 2003 movers. Equation (4.1) holds on the domain of definition Y2of H−1, but by letting H−1(y) =0for y<infY2and H−1(y) =1 for y>supY2,(4.1) actually holds on the whole interval [01]. Equation (4.1) suggests the use of the test statistic T=supy∈[01]| (y)|,where is the nonparametric estimator of =FY|X=2S=0−FY|X=1S=0◦H−1. The logic behind this
372 D’Haultfœuille and Février Quantitative Economics 11 (2020) test statistic is that under the null hypothesis H0,=0. The main challenge here is to derive the distribution of Tunder H0. We estimate this distribution by the distribution of T∗, the test statistic of bootstrap samples drawn under H0.TodrawunderH0,we consider estimators of the distributions of Rconditional on (NXS) that satisfy H0 and are consistent under this hypothesis: 1. For the stayers and the 2001–2002 movers, we first draw Nfrom its empirical distribution, and independently of N,Yaccording to the sieve MLE estimator FY|X=xS=s. We then draw R|NY ∼Binomial(N Y). 2. For the 2003 movers, we draw Nfrom its empirical distribution, and independently of N,Yaccording to FY|X=1S=0◦ H−1.WethendrawR|NY ∼Binomial(N Y). Our estimators of FY|X=xS=sare consistent by Theorem 3.3. Moreover, the bootstrap distribution corresponding to FY|X=2S=0satisfies the null hypothesis by construction. Thus, the bootstrap distribution we consider is consistent under the null hypothesis. We obtain T≃0038 and a p-value of 088, and thus cannot reject the absence of selection effects. This result may seem surprising, given the importance of selection effects obtained by, for example, Lazear (2000). This difference may stem from the pattern in workers’ turnover. Whereas new workers were hired by the car glass company in Lazear’s application, Insee always relies on the same pool of interviewers. Thus, selection effects could only occur through a reallocation of interviewers among this pool. The result of our test suggests that such reallocations are not related to interviewers’ productivity. 4.2 Insee’s program and counterfactual contracts 4.2.1 Insee’s program Turning to Insee’s program, we suppose that Insee values each interview in survey xas λ(x).λ(x) represents the “price” of the information contained in a household’s answers. The dependence in xreflectthefactthatsurveysmaydiffer in the “social value” of the information that can be recovered from it. The 2003 survey on education may have been considered by Insee more important than the other ones, as there was much debate at that time in France on the relationship between families, education, and the emergence of inequalities (see, for instance, the report of the Haut Conseil de l’Education in 2007 on this topic). More formally, more publications from Insee and other institutions were based on this survey and the questionnaire was slightly longer in 2003. We suppose that Insee maximizes its objective function by choosing among linear contracts only. The rationale for this assumption is that Insee uses linear contracts for all its household surveys, not only the PCV ones. This feature seems too peculiar to assume that Insee maximizes its objective function among all contracts. Note however that the linear contract chosen by Insee could well be optimal among the larger set of nonlinear contracts defined below. One of our aims is to evaluate the loss by Insee due to restricting to linear contracts, keeping in mind that there could actually be no loss. On a related note, Insee also violates the Informativeness Principle, which states that all factors correlated with performance should be included in the contracts (Prendergast
Quantitative Economics 11 (2020) The provision of wage incentives 379 Figure 7. Evolution of tn=E[tr2/n2|n2=n y]with n. the interviewers (the less efficient ones) would have benefitted from a change from the linear contract used in 2003 to the optimal, nonlinear contract. Second, we find no significant cost of moral hazard here. If Insee was able to use contracts based on the probability of response rather than on the realized number of respondents, it would only increase its surplus by 00037% only. To understand this, recall that the cost of moral hazard is small when absent any moral hazard, the optimal polynomial contracts tn(y) ≡1 n n k=0n kyn(xθt)k1−yn(xθt)n−kt∗ nk where t∗ n=(t∗ n0t∗ nn), can approximate well the unconstrained optimal contracts t∞(y) ≡tWM n(xy)/n.24 In Figure 7, we plot the functions tnfor n=123,andn=+∞. While there is an important gap between n=1and n=2, t2provides already a good approximation of t∞, while the fit is almost perfect for n=3. This explains the overall negligible loss, as the number of households for which n≤2only represent 014% of the whole sample of households. Third, we find moderate cost of incomplete information, the optimal surplus under asymmetric information being 78% of the optimal one under full information. This loss of 22% is in particular smaller than the one reported by Ferrall and Shearer (33%). Moreover, the surplus under asymmetric information and with the linear contract is 66% of what it could be under complete information. The main part of this loss (65%)isdueto incomplete information whereas 35% is associated with the simple tarification. The rather mild degree of asymmetric information between Insee and its interviewers may explain why Insee chooses not to use some information at its disposal. To confirm this intuition, we investigate what Insee would obtained if it relied on interviewers’ 24One can show using (4.7) and (4.8)thattWM n(xy)/n does not depend on n.
380 D’Haultfœuille and Février Quantitative Economics 11 (2020) Table 9. Estimation of the parametric model with interviewers’ covariates. Unrestricted Restricted Variable Model Model Intercept −0005 (0093) −0102 (0065) Experience ≤50333 (0103)0333 (0097) Experience between 5and 15 0194 (0077)0186 (0071) Rural or small urban area −0229 (0074) −0212 (0066) Female 0001 (0066)– Married −0053 (0051) – Stayer −0089 (0059) – σ0421 (0107)0398 (0096) β1222 (0304)1156 (027) R2=1−σ2/V (lnθ) 0158 0145 Note:600 and 601 observations are used in the unrestricted and restricted model. A positive parameter indicates larger θ,andthus,onaverage,lowerresponserates.Thestandard errors are computed using the inverse of the estimated hessian, and the delta method. characteristics. We first estimate how the characteristics Wof an interviewer relate to θ, by positing lnθ=Wγ+ν where, in line with our lognormal specification, we suppose ν|W∼N(0σ2).Thecharacteristics include experience, gender, marriage status, the interviewer’s status, and the type of area (large urban areas versus others).25 We reestimate the model keeping our preferred specification of the cost function. The results are displayed in Table 9.Notsurprisingly, we find that interviewers with larger experience and living in smaller urban areas perform better on average. Gender and marital status do not seem to be correlated with interviewers’ types. Once controlling for experience and the type of area, we also see that stayers do not perform better on average. This can be seen as a confirmation of our previous result that participation’s decision is not endogenous. Overall, the part of the variance of interviewers’ type that is explained by their observable characteristics is quite small, around 15%. Note also that the estimators of βis very similar to the one we obtained before. These results suggest that experience and the type of areas are the major determinants of interviewers’s type. Using the same model restricted to these covariates (see the second column of Table 9), we estimate what would be the optimal bonus to provide to the six types of interviewers defined by the interactions of these two variables. Insee would propose bonuses ranging from 202for interviewers with more than 15 years 25We do not include the dummy of having another professional activity because of missing data. The interviewer’s area is considered as a large urban areas if most of the housings he has to interview are in towns with more than 100,000 inhabitants.
Quantitative Economics 11 (2020) The provision of wage incentives 381 of experience in rural or small urban areas to 255for interviewers working for Insee for less than 5years in large urban areas. Overall, however, the gain in terms of surplus would remain nearly constant, with a negligible gain of only 015%. This very small gain can be explained by two things. First, the characteristics we use only explain 63% of the variance of the interviewers’ types. The adverse selection problem remains therefore relatively important. Second, we still consider linear contracts here, and they are not optimal. At the end, the cost of discriminating between interviewers is thus likely to exceed these expected gains. In addition to implementation costs mentioned by Ferrall and Shearer (1999), Insee faces social costs due to quite strong unions opposed to such discriminations. 5. Conclusion This work contributes to the empirical personnel literature by showing, in a context of moderate asymmetric information, that interviewers react to incentives and that the simple contracts proposed by Insee are nearly optimal. Beyond these empirical results, we also propose a new approach that extensively uses the exogenous change in 2003 in the compensation scheme, the piece rate increasing from 202to 229euros. This change allows us, in particular, to identify and recover nonparametrically some information on the cost function of the interviewers and on the distribution of their types. This information is used to select correctly the parametric restrictions that we need to impose to derive our results. More generally, we believe that such an exogenous change, associated with a nonparametric estimation in a first step, is essential to estimate and test the optimality of contracts or the presence of asymmetries. Appendix A: Details on interviewers data Besides the data on the surveys, we also have some limited data on interviewers who participate to these surveys and, more generally, on Insee’s households interviewers at the beginning 2001. The most striking fact emerging from Table 10 is the large average experience of interviewers: 85years for the whole set of interviewers and around 10 years for PCV interviewers. Moreover, out of the 12 surveys conducted by Insee in 2001 and for which we have information about interviewers, a typical interviewer conducts more than 5surveys a year in his designated area. This is not surprising, given that Insee basically relies on the same pool of interviewers for all its surveys, even if the precise set of interviewers may vary from one survey to another. By doing so, Insee avoids sunk costs stemming from the recruitment of new interviewers. This sunk cost includes the recruitment procedure itself, as well as a 3-day training period received by interviewers before they conduct their first survey. A second reason is that experience matters for this job. It is well documented that interviewers may influence households and bias their responses (see, e.g., Mensh and Kandel (1988)orO’Muircheartaigh and Campanelli (1998)). It seems, however, that experienced interviewers are less prone to this socalled interviewer’s effect (see, e.g., Cleary, Mechanic, and Weiss (1981), Singer, Frankel, and Glassman (1983), or Campanelli, Martin, and Rothgeb (1991)). Finally, most surveys
382 D’Haultfœuille and Février Quantitative Economics 11 (2020) Table 10. Descriptive statistics on Insee household interviewers. PCV Interviewers All Interviewers Variable in 2001 2001 2002 2003 Experience at Insee (in years) 855 (697)1047 (652)97 (694)956 (724) Yearly income 4054 (3075)6146 (3139)2976 (1702)5169 (2300) Number of surveys done during the year 545 (373)859 (289)821 (385)797 (282) Female 084 (037)085 (035)085 (035)084 (036) Married 066 (047)069 (046)069 (046)068 (047) Age 479 (907)497 (786)492 (800)495 (849) Other professional activity 041 (049)039 (049)040 (049)042 (049) Number of obs. 939 379 469 453 Note: For each column, we indicate the mean and standard deviation (in parenthesis) of the variables. Some observations are missing for the dummy of other professional activity. The income is computed using most but not all of the household surveys. are repeated over time. As interviewers receive a specific training corresponding to each survey, relying on the same pool of interviewers from one edition to another also allows Insee to avoid the duplication of these training costs. Table 10 also shows that the typical interviewer is a middle-aged woman who is out of the labor market. Conversations with them reveal that their job at Insee is usually not the main source of income for the household. It is a flexible job that allows them to complement the revenue of the family. Even if there is a large variability among interviewers and across years, the annual income of 4545 euros earned on average by household interviewers in 2001 corresponds to the minimum wage for a third time job. Appendix B: The effects of the interviewers’sample size n B.1 Alternative test of Assumption 3 We consider another test of (3.5), based not only on the conditional expectation of R,but on its whole distribution. We rely for that purpose also on the binomial model posited in Assumption 1.26 These two conditions imply a known link between Pr(R =r|N=n X = xS =s) and the first nmoments of y(xθ). Specifically, integrating equation (C.1)below over Sleads to Px n=Qnmx n(B.1) where Px n=(Pr(R =1|X=xN =n)Pr(R =n|X=x N =n)),Qnis a nonsingular matrix whose terms are given in the proof of Theorem 3.1 below and mx n= (E(y(x θ)1|X=x)E(y(xθ)n|X=x)).Equation(B.1) should hold for all n∈ 26The previous test only uses the first moment of R. So it is not relying on the independence in households reactions, but just on the fact any household within an area had the same probability of being interviewed.
Quantitative Economics 11 (2020) The provision of wage incentives 383 Supp(N|X=x). Because mx nand mx nhave n∧ncommon terms, many overidentifying restrictions are available. We thus consider a test close to usual overidentification tests for minimum distance estimation. A difference, though is that we also incorporate the constraint that mx nshould be a vector of moments, which implies several restrictions such as variance positivity. We refer to D’Haultfœuille and Rathelot (2017) for details on how to incorporate these constraints.27 If some of these constraints are binding, the test statistic does not have an asymptotic chi-squared distribution. To estimate the critical value, we therefore draw bootstrap samples under the null distribution. D’Haultfœuille and Rathelot (2017) established the validity of a very similar bootstrap test (see their TheoremC.1).Attheend,weobtainp-values of 093 and 071 for the two surveys, supporting again the validity of Assumption 3. B.2 Partial identification of marginal costs without Assumption 3 We show here that we can actually weaken the condition that C(nxy)/C(xy) =nand still obtain bounds on the effect of non the cost function. Specifically, let us assume that C(nxy) =f(n)C(xy) for some function f(·)satisfying without loss of generality f(n0)=1for some n0∈Supp(N|X=1)∩Supp(N|X=2). Combining this with our exclusion restriction C(xy)=C(y),weobtain y(nxθ) =C−1nδ(x) f(n)θ Then, for all x∈{12}and n∈Supp(N|X=x), E(R|N=n X =x) n=EC−1nδ(x) f(n)θ Moreover, C−1is strictly increasing. This means that for any (n1n2)∈Supp(N|X=1)× Supp(N|X=2), sgnE(R|N=n1X =x) n1 −E(R|N=n2X =x) n2=sgnn1δ(1) f(n1)−n2δ(2) f(n2) This allows us to obtain bounds on f(·), and possibly point identify f(·)if this function is parametrized. B.3 The effect of a bounded support on N The point identification in Theorem 3.1 relies on a large support assumption on N(see Assumption 4). If sup Supp(N|X=x) =Nx<+∞, the proof of Theorem 3.1 reveals that we identify only the first NXmoments of Y|XS. This implies that the distribution of 27Also, we actually perform our test conditional on N<35. We faced numerical issues otherwise, due in particular to constrained optimization in large dimensional spaces. This restriction only removes less than 2% of the data in the two cases.
384 D’Haultfœuille and Février Quantitative Economics 11 (2020) Figure 8. D’Haultfœuille and Rathelot (DH-R)’s and sieve ML estimates of FY|X=xS=1. Y|XS is not point identified in general. We can still obtain bounds on FY|XS(y) by minimizing or maximizing F(y)over all cumulative distributions Fwith prescribed first NX moments, taking also into account that FY|X=2S=1(y) ≤FY|X=1S=1(y) for all y. This problem is difficult and, to our knowledge, has not been addressed in the literature. On the other hand, if we do not include the inequality constraints, we can use Theorem 2.1 of D’Haultfœuille and Rathelot (2017), who show that the optimization can be conducted without loss of generality on discrete distributions with at most NX+1 support points. Of course, the corresponding bounds FY|XS(y) and FY|XS(y) are not sharp, as they do not incorporate the inequality constraints above. We computed the estimators of FY|XS(y) and FY|XS(y) suggested by D’Haultfœuille and Rathelot (2017). The results are displayed in Figure 8, which also presents for comparison the sieve estimates considered in Section 3.4. We actually obtain a point estimate on FY|XS. This could be expected, given that here NX≥50.D’Haultfœuille and Rathelot (2017) showed in their setting that the upper and lower estimated bounds typically collapse for usual sample sizes (up to 10,000,say)whenNX≥6.Also,theestimator corresponds to that of a finitely supported distribution, which is also expected given that optimization is run over distributions with at most NX+1support points. Other than this feature, the estimator looks quite similar to the sieve estimator. Appendix C: Proofs Proof of Theorem 3.1. First, by Assumption 3and the first-order condition, y(xθ) does not depend on n.Wedenoteitbyy(xθ). Then, for all nin the support of N|X= xS =sand all 1≤r≤n,wehave Pr(R =r|N=nX =x S =s) =EPrR=r|N=ny(x θ)|N=nX =x S =s
Quantitative Economics 11 (2020) The provision of wage incentives 385 =En ry(xθ)r1−y(xθ)n−r|N=nX =x S =s =En ry(xθ)r1−y(xθ)n−r|X=x S =s = n−r i=0n rn−r i(−1)n−r−iEy(xθ)n−i|X=xS =s = n i=rn ii r(−1)i−rEy(xθ)i|X=xS =s = n i=1n ii r(−1)i−rEy(xθ)i|X=xS =s where the first equality follows from the law of iterated expectation, the second from Assumption 1, the third stems from independence between θand Nconditional on X=xS =s(Assumption 3), the fourth from the decomposition of (1−y(xθ))n−r,the fifth is obtained by setting i=n−iand remarking that n rn−r i=n ii r,andthelast by noting that the j−1first terms in the sum are zero. Hence, letting Pxs n=(Pr(R = 1|N=nX =x S =s)Pr(R =n|N=n X =xS =s)),mxs n=(E(y(x θ)1|X=x S = s)E(y(xθ)n)|X=x S =s)and Qnbe the n×nmatrix of typical (i r) element n ii r(−1)i−r,weget Pxs n=Qnmxs n(C.1) Moreover, Qnis invertible as an upper triangular matrix with nonzero diagonal elements. Thus, mxs nis identified by Q−1 nPxs n. Conditional on X=xS =s,thenfirst moments of y(xθ) are identified from the distribution of Rconditional on N=n X = xS =s. Because sup{n:Pr(N =n|X=xS =s) > 0}=+∞, all moments of y(xθ) (conditional on X=x S =s) are identified. This, together with y(xθ) bounded, ensures that the distribution of y(xθ) conditional on X=xS =sis identified (see, e.g., Gut (2005)). Proof of Theorem 3.2. It follows from the discussion before Theorem 3.2 that C, Fθ|X=1S=sand Fθ|X=2S=sare point identified on (yk)k∈K,(θk)k∈Kand (θk)k+1∈K,respectively. Elsewhere, θk(y) =sup k∈K:yk≥y θ(1yk)≤θ(1y)≤inf k∈K:yk≤yθ(1yk)=θk(y) Thus, C(y) =δ(1)/θ(1y)≥δ(1)/θk(y) =C(y) and similarly, C(y) ≤C(y). Besides, yk(θ) =sup k∈K:θk≥θ yk≤y(1θ)≤inf k∈K:θk≤θyk=yk(θ)
386 D’Haultfœuille and Février Quantitative Economics 11 (2020) Hence, Fθ|X=1S=s(θ) =1−FY|X=1S=sy(1θ)≥1−FY|X=1S=s(yk(θ))=Fθ|X=1S=s(θ) and similarly for the upper bound. The bounds on Fθ|X=2S=s(θ) follow by remarking that y(2θ)=y1δ(1)θ/δ(2)∈[yk(θ)+1yk(θ)+1] Wenowshowthatforally0∈(01)\{yk:k∈K},andθ0∈R+\{0θ(1yk):k∈K},the bounds on C(y0)and Fθ|XS(θ0)aresharp.WefocusonC(y0)as the proof is similar for C(y0),Fθ|XS(θ0)and Fθ|XS(θ0). More precisely, we want to construct a function C such that C(y0)is arbitrarily close to C(y0)and all the restrictions given by the data and the model hold. We consider separately two cases, whether or not there exists k∈K such that yk<y0<y k+1. In the first case, fix εsuch that 0<ε<δ(1)[1/θk+1−1/θk]. We first define Con [ykyk+1). To do so, we consider any strictly increasing, continuously differentiable function Csuch that C(yk)=δ(1)/θk, C(y0)=C(y0)−ε,limy↑yk+1 C(y) =δ(1)/θk+1,and lim y↑yk+1 C(y) =δ(2) C(yk) δ(1)H(yk+1)(C.2) Such a function exists because δ(1)/θk< C(y0)−ε<δ(1)/θk+1and H(y)=C−1[δ(2)× C(y)/δ(1)]is differentiable with positive derivative at any y>0. We then extend Con (01)using (3.3). For instance, assume that k+2∈K.Thenwe define Con [yk+1yk+2)by C(y) =δ(2) δ(1) CH−1(y) Moreover, because His continuously differentiable, Cis continuously differentiable on (yk+1yk+2). It also admits a right derivative at yk+1given by lim y↓yk+1 C(y) =δ(2) C(yk) δ(1)H(yk) and equation (C.2) ensures that Cis differentiable at yk+1. By induction, using either H or H−1, we can then extend Con YK≡!k:k∈Kk+1∈K[ykyk+1).IfYK=(01),wehave defined this way a continuously differentiable function on the whole interval (01).This function is also strictly increasing as both Hor H−1are strictly increasing. If YK= (01), we still have to extend Con intervals of the form [0y)or [y1). Consider the first case (the second is similar). We simply consider any strictly increasing, continuously differentiable function Csuch that C(0)=0,limy↑y C(y) = C(y),andlimy↑y C(y) = C(y). Again, this defines a continuously differentiable function on the whole interval (01). To complete the proof for this case, we have to show that with such a function C, we can rationalize the model and the data. For that purpose, let θ(xy) be defined
Quantitative Economics 11 (2020) The provision of wage incentives 387 by θ(xy) =δ(x)/ C(y) for all y∈Yx. By construction, θ(x·)is strictly decreasing. Let y(x·)denote its inverse, and let Fθ|X=xS(θ) =1−FY|X=xS y(xθ) By construction, Cand Fθ|XS rationalize the data and the first-order condition (3.2). The second-order condition also holds since Cis strictly increasing. Thus, these functions rationalize Model (3.1) as well. Because εcould be arbitrarily close to 0, this shows that C(y) is sharp. We now consider the case where there is no k∈Ksuch that yk<y 0<y k+1.Equivalently, y0/∈YK, and either y0∈[0y)or y0∈(y1). In both case, we simply let C=C on YK. Suppose that y0∈[0y)(the case y0∈(y1)is similar). Note that y∈(yk)k∈K, which implies that C(y0)=C(y).Fixεsuch that 0<ε<C (y). Then define Con [0y)as any strictly increasing, continuously differentiable function such that C(0)=0, C(y0)=C(y0)−ε,limy↑y C(y) =C(y)and lim y↑y C(y) =C(y) Inthecasewheresup YK<1, define similarly Con [y1)by C(y)=C(y),limy↑1 C(y) = +∞ and limy↓y C(y) =C(y). By construction, C is then strictly increasing and continuously differentiable on (01). The rest of the proof is identical as above. Nonidentification with one menu of contracts. With only one menu of contracts, the variables Xand Sare irrelevant, so we drop them here. Let us consider a strictly increasing and differentiable function C, different from the true one C. Define then θ(y) by θ(y) =δ/ C(y). θis strictly decreasing and admits an inverse function y. Then define Fθby Fθ(θ) =1−FY y(θ) By construction, Cand Fθare consistent with the firstand second-order conditions and the identified distribution FY.Asaresult,Cand Fθare not identified. Proof of Theorem 3.3. The proof proceeds in four steps. We first prove that FY|X=xS=sis uniformly consistent. We then prove that His uniformly consistent on each compact set included in (01).Third,weprovethatforallk∈K, ykis consistent. Finally, we show that the estimated bounds of Cand Fθ|XS are consistent. 1. Uniform consistency of FY|X=xS=s. For any function gon [01]let g=supx∈[01]|g(x)|. We actually prove the stronger result that for all (x s) ∈{12}×{01}, fY|X=xS=s−fY|X=xS=sP −→ 0(C.3)
388 D’Haultfœuille and Février Quantitative Economics 11 (2020) For all yin the interior of Yxs,fY|X=xS=s(y) =∂θ/∂y(x y)fθ|X=xS=s(θ(xy)).Hence, fY|X=xS=sis continuous in the interior of Yxs. Moreover, differentiating the first-order condition, we obtain ∂θ ∂y (xy) =−θ(xy)C(y) C(y) =−δ(x)C(y) C2(y) =−θ(xy)2C(y) δ(x) (C.4) By Assumption 7,limy↓infYxs θ(xy)2fθ|X=xS=s(θ(xy)) =0. Because C is bounded, this implies that limy↓inf Yxs fY|X=xS=s(y) =0.Hence,fY|X=xS=sis continuous or can be extended by continuity on [0sup Yxs]. Similarly, limy↑supYxs C(y)/C2exists. Hence, limy↑supYxs fY|X=xS=s(y) exists as well. If θ>0,fθ|X=xS=s(θ)=0and so limy↑sup Yxs fY|X=xS=s(y) =0.Ifθ=0,supYxs =1. Hence, in both cases we can extend by continuity fY|X=xS=son [01]. Let Fdenote the space of continuous density functions on [01].Forf∈F,n∈N and r∈{0n},let (fr n) =ln1 0 yr(1−y)n−rf(y)dy let Qxs(f) =E((fRN)|X=xS =s),and Qxs(f) = i:Xi=xSi=s (fRiNi) By definition of fY|X=xS=s, fY|X=xS=s=arg maxf∈FL Qxs(f) is a sieve M-estimator. We use Theorem 3.1 of Chen (2007) and its associated Remark 3.2 to prove (C.3). To this end, we check the following conditions: a. Qxs is uniquely maximized at fY|X=xS=sand Qxs(fY|X=xS=s)>−∞. b. For all L,FL⊂FL+1and for all f∈F,thereexistsfL∈FLsuch that fL−f→0. c. Qxs is continuous for ·. d. FLis compact. e. E[supf∈FL|(fR N)||X=x S =s]<∞. f. There exists U(··)such that E(U(RN)|X=x S =s) < ∞and for all (fg) ∈F2 L, |(fR N) −(g R N)|≤f−gU(RN). g. The minimal number of δ-balls that cover FL, denoted Nb(δ FL·), satisfies lnNb(δ FL·)=o(L). a. First, for all g∈F, Eexp(g R N) exp(fY|X=xS=sRN)|N=n X =xS =s = n r=0 Pr(R =r|N=nX =x S =s) r n1 0 yr(1−y)n−rg(y)dy Pr(R =r|N=nX =x S =s)
Quantitative Economics 11 (2020) The provision of wage incentives 395 Chu, C. S., P. Leslie, and A. Sorensen (2011), “Bundle-size pricing as an approximation to mixed bundling.” The American Economic Review, 101, 263–303. [378] Chu, L. Y. and D. E. M. Sappington (2007), “Simple cost-sharing contracts.” The American Economic Review, 97, 419–428. [351,378] Chung, D., T. J. Steenburgh, and K. Sudhir (2014), “Do bonuses enhance sales productivity? A dynamic structural analysis of bonus-based compensation plans.” Marketing Science, 33, 165–187. [350] Cleary, P. D., D. Mechanic, and N. Weiss (1981), “The effect of interviewer characteristics on responses to a mental health interview.” Journal of Health and Social Behavior, 22, 183–193. [381] Copeland, A. and C. Monnet (2009), “The welfare effect of incentive schemes.” Review of Economic Studies, 76, 93–113. [350] Crawford, G. S. and M. Shum (2007), “Monopoly quality degradation and regulation in cable television.” Journal of Law and Economics, 50, 181–219. [351] D’Haultfœuille, X. and P. Février (2010), “Identification of a class of adverse selection models with contracts variation.” CREST Working Paper, https://ideas.repec.org/p/crs/ wpaper/2011-27.html.[357,363] D’Haultfœuille, X. and R. Rathelot (2017), “Measuring segregation on small units: A partial identification analysis.” Quantitative Economics, 8 (1), 39–73. [358,383,384] Ferrall, C. and B. Shearer (1999), “Incentives and transaction costs within the firm: Estimating an agency model using payroll records.” Review of Economic Studies, 99, 309–338. [350,351,352,357,377,381] Gagnepain, P. and M. Ivaldi (2002), “Incentive regulatory policies: The case of public transit systems in France.” The RAND Journal of Economics, 33, 605–629. [350] Gagnepain, P., M. Ivaldi, and D. Martimort (2013), “The cost of contract renegotiation: Evidence from the local public sector.” American Economic Review, 103, 2352–2383. [351] Guerre, E., I. Perrigne, and Q. Vuong (2009), “Nonparametric identification of risk aversion in first-price auctions under exclusion restrictions.” Econometrica, 77, 1193–1227. [351,363] Gut, A. (2005), Probability: A Graduate Course. Springer-Verlag, New York, NY. [358,385] Holmstrom, L. and P. Milgrom (1987), “Aggregation and linearity in the provision of intertemporal incentives.” Econometrica, 55, 303–328. [351] Ivaldi, M. and D. Martimort (1994), “Competition under nonlinear pricing.” Annales d’Economie et de Statistique, 34, 71–114. [350] Kang, K. and B. Silveira (2018), “Understanding disparities in punishment: Regulator preferences and expertise.” Working Paper. [351]
396 D’Haultfœuille and Février Quantitative Economics 11 (2020) Laffont, J. J. and D. Martimort (2002), The Theory of Incentives: The Principal-Agent Model. Princeton University Press. [351] Laffont, J. J. and J. Tirole (1993), A Theory of Incentives in Procurement and Regulation. MIT Press. [357] Lavergne, P. and A. Thomas (2005), “Semiparametric estimation and testing in a model of environmental regulation with adverse selection.” Empirical Economics, 30, 171–192. [350,351,357] Lazear, E. (2000), “Performance pay and productivity.” American Economic Review, 90, 1346–1361. [349,372] Le Lan, R. (2009), “Enquêtes ménages de l’insee: vers la fin de la baisse des taux de réponse?” Courrier des Statistiques, 128, 33–41. [359] Leslie, P. (2004), “Price discrimination in broadway theater.” The RAND Journal of Economics, 35, 520–541. [350] Lim, C. and A. Yurukoglu (2018), “Structural analysis of nonlinear pricing.” Journal of Political Economy, 126, 2523–2568. [351] Luo, Y., I. Perrigne, and Q. Vuong (2018), “Structural analysis of nonlinear pricing.” Journal of Political Economy, 126, 2523–2568. [351] Mensh, B. S. and D. B. Kandel (1988), “Underreporting of substance use in a national longitudinal youth cohort individual and interviewer effects.” The Public Opinion Quarterly, 52, 100–124. [381] Miravete, E. J. (2002), “Estimating demand for local telephone service with asymmetric information and optional calling plans.” Review of Economic Studies, 69, 943–971. [350] Miravete, E. J. (2007), “The limited gains from complex tariffs.” Working Paper Series 3971, Victoria University of Wellington, The New Zealand Institute for the Study of Competition and Regulation, https://ideas.repec.org/p/vuw/vuwcsr/3971.html.[351,378] Miravete, E. J. and L.-H. Roller (2005), “Estimating markups under nonlinear pricing competition.” Journal of the European Economic Association, 2, 526–535. [350] Neeman, Z. (2003), “The effectiveness of English auctions.” Games and Economic Behavior, 43, 214–238. [378] O’Muircheartaigh, C. and P. Campanelli (1998), “The relative impact of interviewer effects and sample design effects on survey precision.” Journal of the Royal Statistical Society: Series A (Statistics in Society), 161, 63–77. [381] Paarsch, H. and B. Shearer (2000), “Piece rates, fixed wages and incentive effects: Statistical evidence from payroll records.” International Economic Review, 41, 59–92. [350] Perrigne, I. and Q. Vuong (1999), “Structural econometrics of first-price auctions: A survey of methods.” Canadian Journal of Agricultural Economics, 47, 203–223. [352] Perrigne, I. and Q. Vuong (2011a), “Nonlinear pricing in yellow pages.” Report. [351]
Quantitative Economics 11 (2020) The provision of wage incentives 397 Perrigne, I. and Q. Vuong (2011b), “Nonparametric identification of a contract model with adverse selection and moral hazard.” Econometrica, 79, 1499–1539. [351] Prendergast, C. (1999), “The provision of incentives in firms.” Journal of Economic Literature, 37, 7–63. [349,372,373] Raju, J. S. and V. Srinivasan (1996), “Quota-based compensation plans for multiterritory heterogeneous salesforces.” Management Science, 6, 1454–1462. [351] Rogerson, W. P. (2003), “Simple menus of contracts in cost-based procurement and regulation simple menus of contracts in cost-based procurement and regulation.” The American Economic Review, 93, 919–926. [351,378] Singer, E., M. R. Frankel, and M. B. Glassman (1983), “The effect of interviewer characteristics and expectations on response.” Public Opinion Quarterly, 47, 68–83. [381] van der Vaart, A. W. (1998), Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics. [392] van der Vaart, A. W. and J. Wellner (1996), Weak Converence and Empirical Process. Springer. [392] Wilson, R. (1993), Nonlinear Pricing. Oxford University Press. [351,357,378] Wolak, F. (1994), “An econometric analysis of the asymmetric information, regulatorutility interaction.” Annales d’Economie et de Statistique, 34, 13–69. [350,357] Wood, G. R. (1999), “Binomial mixtures: Geometric estimation of the mixing distribution.” Annals of Statistics, 27, 1706–1721. [358] Co-editor Rosa L. Matzkin handled this manuscript. Manuscript received 24 July, 2015; final version accepted 25 May, 2019; available online 1 July, 2019.