scieee AI-readable full text Open interactive document viewer

Inference on breakdown frontiers

Masten, Matthew A.,Poirier, Alexandre

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Masten, Matthew A.; Poirier, Alexandre Article Inference on breakdown frontiers Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Masten, Matthew A.; Poirier, Alexandre (2020) : Inference on breakdown frontiers, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 11, Iss. 1, pp. 41-111, https://doi.org/10.3982/QE1288 This Version is available at: https://hdl.handle.net/10419/217183 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Quantitative Economics 11 (2020), 41–111 1759-7331/20200041 Inference on breakdown frontiers Matthew A. Masten Department of Economics, Duke University Alexandre Poirier Department of Economics, Georgetown University Given a set of baseline assumptions, a breakdown frontier is the boundary between the set of assumptions which lead to a specific conclusion and those which do not. In a potential outcomes model with a binary treatment, we consider two conclusions: First, that ATE is at least a specific value (e.g., nonnegative) and second that the proportion of units who benefit from treatment is at least a specific value (e.g., at least 50%). For these conclusions, we derive the breakdown frontier for two kinds of assumptions: one which indexes relaxations of the baseline random assignment of treatment assumption, and one which indexes relaxations of the baseline rank invariance assumption. These classes of assumptions nest both the point identifying assumptions of random assignment and rank invariance and the opposite end of no constraints on treatment selection or the dependence structure between potential outcomes. This frontier provides a quantitative measure of the robustness of conclusions to relaxations of the baseline point identifying assumptions. We derive √N-consistent sample analog estimators for these frontiers. We then provide two asymptotically valid bootstrap procedures for constructing lower uniform confidence bands for the breakdown frontier. As a measure of robustness, estimated breakdown frontiers and their corresponding confidence bands can be presented alongside traditional point estimates and confidence intervals obtained under point identifying assumptions. We illustrate this approach in an empirical application to the effect of child soldiering on wages. We find that sufficiently weak conclusions are robust to simultaneous failures of rank invariance and random assignment, while some stronger conclusions are fairly robust to failures of rank invariance but not necessarily to relaxations of random assignment. Matthew A. Masten: [email protected] Alexandre Poirier: [email protected] This paper was presented at the 2017 IAAE Conference, the 2017 CEME Conference, the 2017 U Chicago Interactions Conference, the 2017 Midwest Econometrics Group, the 2017 SEA Meetings, the 2018 Winter Meeting of the Econometric Society, the 2018 NU Econometrics Alumni Conference, Duke, UC Irvine, U Penn, UC Berkeley, Yale, UGA, Auburn, UC Boulder, Georgetown, UNC-Chapel Hill, Rice, Boston University, UCL, University of Cambridge, and the LSE. We thank audiences at those conferences and seminars, the referees, as well as Federico Bugni, Ivan Canay, Joachim Freyberger, Guido Imbens, Ying-Ying Lee, Chuck Manski, Anna Mikusheva, Jim Powell, Pedro Sant’Anna, and Andres Santos for helpful conversations and comments. We thank Margaux Luflade for excellent research assistance. This paper uses data from phase 1 of SWAY, the Survey of War Affected Youth in northern Uganda. We thank the SWAY principal researchers Jeannie Annan and Chris Blattman for making their data publicly available and for answering our questions. ©2020 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1288 42 Masten and Poirier Quantitative Economics 11 (2020) Keywords. Nonparametric identification, partial identification, sensitivity analysis, selection on unobservables, rank invariance, treatment effects, directional differentiability. JEL classification. C14, C18, C21, C25, C51. 1. Introduction Traditional empirical analysis combines the observed data with a set of assumptions to draw conclusions about a parameter of interest. Breakdown frontier analysis reverses this ordering. It begins with a fixed conclusion and a set of baseline assumptions and asks, “What are the weakest assumptions needed to draw that conclusion, given the observed data?” For example, consider the impact of a binary treatment on some outcome variable. The traditional approach might assume random assignment, point identify the average treatment effect (ATE), and then report the obtained value. The breakdown frontier approach instead begins with a conclusion about ATE, like “ATE is positive,” and reports the weakest assumption—relative to random assignment—on the relationship between treatment assignment and potential outcomes needed to obtain this conclusion, when such an assumption exists. When more than one kind of assumption is considered, this approach leads to a curve, representing the weakest combinations of assumptions which lead to the desired conclusion. This curve is the breakdown frontier. At the population level, the difference between the traditional approach and the breakdown frontier approach is a matter of perspective: an answer to one question is an answer to the other. This relationship has long been present in the literature initiated by Manski on partial identification (e.g., see Manski (2007)orSection3ofManski (2013)). In finite samples, however, which approach one chooses has important implications for how one does statistical inference. Specifically, the traditional approach estimates the parameter or its identified set. Here we instead estimate the breakdown frontier. The traditional approach then performs inference on the parameter or its identified set. Here we instead perform inference on the breakdown frontier. Thus the breakdown frontier approach puts the weakest assumptions necessary to draw a conclusion at the center of attention. Consequently, by construction, this approach avoids the nontight bounds critique of partial identification methods (e.g., see Section 7.2 of Ho and Rosen (2017)). One distinction is that the traditional approach may require inference on a partially identified parameter. The breakdown frontier approach, however, only requires inference on a point identified object. The breakdown frontier we study generalizes the concept of an “identification breakdown point” introduced by Horowitz and Manski (1995), a one-dimensional breakdown frontier.1Their breakdown point was further studied and generalized by Stoye (2005, 2010). Our emphasis on inference on the breakdown frontier follows Kline and Santos 1The identification breakdown point is distinct from the breakdown point introduced earlier by Hampel (1968,1971) in the robust statistics literature that began with Huber (1964); also see Donoho and Huber (1983). Horowitz and Manski (1995) gave a detailed comparison of the two concepts. Throughout this paper we use the term “breakdown” in the same sense as Horowitz and Manski’s identification breakdown point. Quantitative Economics 11 (2020) Inference on breakdown frontiers 43 (2013), who proposed doing inference on a breakdown point. Finally, our focus on multidimensional frontiers builds on the graphical sensitivity analysis of Imbens (2003)and the multidimensional sensitivity analysis of Manski and Pepper (2018). We discuss these papers and others in detail in Appendix A. The breakdown frontier approach The breakdown frontier approach requires six main steps: (a) specify a parameter of interest, (b) specify a set of baseline assumptions, (c) define a class of assumptions indexed by a sensitivity parameter which deliver a nested sequence of identified sets, with the baseline assumptions obtained at one extreme and the no assumptions bounds obtained at the other, (d) characterize identified sets for the parameter of interest as a function of the sensitivity parameter, (e) use those identified sets to define the breakdown frontier for a conclusion of interest, and (f) develop estimation and inference procedures for that frontier based on its characterization. In principle this analysis can be done for a general class of models, for example, by using the general identification analysis in Chesher and Rosen (2017)orTorgovitsky (2019) for step (d), and then applying general tools for nonparametric estimation and inference for step (f). While such a general analysis is an important next step for future work, in this paper we focus on just one important and widely used model: the potential outcomes model with a binary treatment. By focusing on a single concrete model, we can clearly illustrate how to do the six main steps (a)–(f) required for a breakdown frontier analysis in any model. While the mathematical analysis will differ from model to model, the general ideas and approach do not. In the rest of this section and this paper, we use the binary treatment potential outcomes model to explain and illustrate the breakdown frontier approach. Our main parameter of interest is the proportion of units who benefit from treatment. Under random assignment of treatment and rank invariance, this parameter is point identified. One may be concerned, however, that these two assumptions are too strong. We relax rank invariance by supposing that there are two types of units in the population: one type for which rank invariance holds and another type for which it may not. The proportion tof the second type measures the relaxation of rank invariance. We relax random assignment using a propensity score distance c≥0as in our previous work, Masten and Poirier (2018a). We give more details on both of these relaxations in Section 2.Wederive the identified set for P(Y1>Y 0)as a function of (ct). For a specific conclusion, such as P(Y1>Y 0)≥05, this identification result defines a breakdown frontier. Figure 1illustrates this breakdown frontier. The horizontal axis measures c, the relaxation of the random assignment assumption. The vertical axis measures t, the relaxation of rank invariance. The origin represents the baseline point identifying assumptions of random assignment and rank invariance. Points along the vertical axis represent random assignment paired with various relaxations of rank invariance. Points along the horizontal axis represent rank invariance paired with various relaxations of random as- 44 Masten and Poirier Quantitative Economics 11 (2020) Figure 1. An example breakdown frontier, partitioning the space of assumptions into the set for which our conclusion of interest holds (the robust region) and the set for which our evidence is inconclusive. signment. Points in the interior of the box represent relaxations of both assumptions. The points in the lower left region are pairs of assumptions (ct) such that the data allow us to draw our desired conclusion: P(Y1>Y 0)≥05. We call this set the robust region. Specifically, no matter what value of (c t) we choose in this region, the identified set for P(Y1>Y 0)always lies completely above 05. The points in the upper right region are pairs of assumptions that do not allow us to draw this conclusion. For these pairs (c t), the identified set for P(Y1>Y 0)contains elements smaller than 05. The boundary between these two regions is precisely the breakdown frontier. The area under the breakdown frontier—the robust region—is a quantitative measure of robustness. The breakdown frontier is similar to the classical Pareto frontier from microeconomic theory: It shows the trade-offs between different assumptions in drawing a specific conclusion. For example, consider Figure 1. If we are at the top left point, where the breakdown frontier intersects the vertical axis, then any relaxation of random assignment requires strengthening the rank invariance assumption in order to still be able to draw our desired conclusion. The breakdown frontier specifies the precise marginal rate of substitution between the two assumptions. Figure 2illustrates how the breakdown frontier changes as our conclusion of interest changes. Specifically, consider the conclusion that P(Y1>Y 0)≥p for five different values for p. The figure shows the corresponding breakdown frontiers. As pincreases toward one, we are making a stronger claim about the true parameter, and hence the set of assumptions for which the conclusion holds shrinks. For strong enough claims, the claim may be refuted even with the strongest assumptions possible. Conversely, as pdecreases toward zero, we are making progressively weaker claims about the true parameter, and hence the set of assumptions for which the conclusion holds grows larger. Quantitative Economics 11 (2020) Inference on breakdown frontiers 45 Figure 2. Example breakdown frontiers for the claim P(Y1>Y 0)≥p, for five different values of p. Under the strongest assumptions of (ct) =(00), the parameter P(Y1>Y 0)is point identified. Let p00denote its value. The value p00is often strictly less than 1.Inthiscase, any p∈(p001]yields a degenerate breakdown frontier: This conclusion is refuted under the point identifying assumptions. Even if p00<p, the conclusion P(Y1>Y 0)≥p may still be correct. This follows since, for strictly positive values of cand t,theidentified sets for P(Y1>Y 0)do contain values larger than p00. But they also contain values smaller than p00. Hence there do not exist any assumptions for which we can draw the desired conclusion. We provide additional interpretation of population breakdown frontiers in Section 2, Section 4, and Appendix E in the Online Supplemental Material (Masten and Poirier (2020)). We discuss estimation and inference in Section 3and Appendix D in the Online Supplemental Material. Finally, we illustrate how our results can be used as a sensitivity analysis within a larger empirical study in Section 4. In that section we also provide a detailed and largely self-contained guide to implementing our approach. 2. Model and identification In this section we study identification of the standard potential outcomes model with a binary treatment. We focus on two key assumptions: random assignment and rank invariance. We discuss how to relax these two assumptions and derive identified sets for various parameters under these relaxations. While there are many different ways to relax these assumptions, our goal is to illustrate the breakdown frontier methodology, and hence we focus on just one kind of relaxation for each assumption. Setup Let Y1and Y0denote the unobserved potential outcomes. Let X∈{01}be an observed binary treatment. Let W∈supp(W ) be a vector of observed covariates. This vector may contain discrete covariates, continuous covariates, or both. We observe the scalar outcome variable Y=XY1+(1−X)Y0(1) Let px|w=P(X =x|W=w) denote the observed propensity score. We maintain the following assumption on the joint distribution of (Y1Y0XW)throughout. 46 Masten and Poirier Quantitative Economics 11 (2020) Assumption A1. For each x x∈{01}and w∈supp(W ): 1. Yx|X=x,W=whas a strictly increasing and continuous distribution function on its support,supp(Yx|X=xW =w). 2. supp(Yx|X=xW =w) =supp(Yx|W=w) =[yx(w) yx(w)]where −∞ ≤ yx(w) < yx(w) ≤∞. 3. px|w>0. Via Assumption A1.1, we restrict attention to continuously distributed potential outcomes. Assumption A1.2states that the support of Yx|X=x,W=wdoes not depend on x, and is a possibly infinite closed interval. This assumption implies that the endpoints yx(w) and yx(w) are point identified. We maintain Assumption A1.2for simplicity, but it can be relaxed using similar derivations as in Masten and Poirier (2016). Assumption A1.3is an overlap assumption. Define the conditional rank random variables U0=FY0|W(Y0|W) and U1= FY1|W(Y1|W). Since FY1|W(·|w) and FY0|W(·|w) are strictly increasing (by Assumption A1.1), U0|Wand U1|Ware uniformly distributed on [01]. The value of unit i’s conditional rank random variable Uxtells us where unit ilies in the distribution of Yx|W. Identifying assumptions It is well known that the joint conditional distribution of potential outcomes (Y1Y0)|W is point identified under two assumptions: 1. Conditional random assignment of treatment: X⊥⊥Y1|Wand X⊥⊥Y0|W. 2. Conditional rank invariance: U1=U0almost surely. Note that the joint conditional independence assumption X⊥⊥(Y1Y0)|Wprovides no additional identifying power beyond the marginal conditional independence assumption stated above. Any functional of FY1Y0|Wis point identified under these random assignment and rank invariance assumptions. The goal of our identification analysis is to study what can be said about such functionals when one or both of these point identifying assumptions fails. To do this, we define two classes of assumptions: one which indexes the relaxation of random assignment of treatment, and one which indexes the relaxation of rank invariance. These classes of assumptions nest both the point identifying assumptions of random assignment and rank invariance and the opposite end of no constraints on treatment selection or the dependence structure between potential outcomes. A key feature of the relaxations we use is that they are “orthogonal” in the sense that we can relax each of the two assumptions separately: The amount by which we relax one assumption does not constrain the amount by which we can relax the other assumption. This feature is important since a key goal of our analysis is to quantify the trade-off between relaxations of these two assumptions. We begin with our measure of distance from conditional independence. Quantitative Economics 11 (2020) Inference on breakdown frontiers 47 Definition 1. Let cbe a scalar between 0 and 1. Say Xis conditionally c-dependent with Yxgiven Wif sup y∈supp(Yx|W=w)P(X =x|Yx=yW =w) −P(X =x|W=w)≤c(2) for x∈{01}and w∈supp(W ). For c=0, conditional c-dependence implies X⊥⊥Y1|Wand X⊥⊥Y0|W.Forc>0, however, it allows for some deviations from conditional independence. Specifically, it allows the unobserved treatment assignment probability P(X =1|Yx=yW =w) to be at most cprobability units away from the observed propensity score p1|w.Wediscuss one way to interpret the magnitude of con page 68. We give further discussion in our previous paper, Masten and Poirier (2018a). Our second class of assumptions constrains the dependence structure between Y1 and Y0, conditional on W. By Sklar’s theorem (Sklar (1959)), write FY1Y0|W(y1y0|w) =CFY1|W(y1|w)FY0|W(y0|w) |w where C(··|w) is a unique conditional copula function. See Nelsen (2006)foran overview of copulas and Fan and Patton (2014) for a survey of their use in econometrics. Restrictions on Cconstrain the dependence between potential outcomes. For example, if C(u1u0|w) =min{u1u0} then U1=U0almost surely. Thus conditional rank invariance holds. In this case the potential outcomes Y1and Y0are sometimes called conditionally comonotonic and min{··} is called the comonotonicity copula. At the opposite extreme, when Cis an arbitrary copula, the dependence between Y1|Wand Y0|Wis constrained only by the Fréchet–Hoeffding bounds, which state that max{u1+u0−10}≤C(u1u0|w) ≤min{u1u0} We next define a class of assumptions which includes both conditional rank invariance and no assumptions on the dependence structure as special cases. Definition 2. The potential outcomes (Y1Y0)satisfy (1−t)-percent conditional rank invariance given Wif for all w∈supp(W ) their conditional copula Csatisfies C(u1u0|w) =(1−t)min{u1u0}+tH(u1u0|w) (3) where His some conditional copula. This assumption says that within each covariate cell the population is a mixture of two parts: In one part, rank invariance holds. This part contains 100(1−t)%oftheoverall population in that cell. In the second part, rank invariance may fail in an arbitrary, unknown way. Hence, for this part, the dependence structure is unconstrained beyond 48 Masten and Poirier Quantitative Economics 11 (2020) the Fréchet–Hoeffding bounds. This part contains 100 ·t% of the overall population in that cell. Thus for t=0the usual conditional rank invariance assumption holds, while for t=1no assumptions are made about the dependence structure. For t∈(01),we obtain a kind of partial conditional rank invariance. Note that by exercise 2.3 on page 14 of Nelsen (2006), a mixture of copulas like that in equation (3)isalsoacopula. To see this mixture interpretation formally, let Tfollow a Bernoulli distribution with P(T =1|W=w) =t,whereT⊥⊥ Y1|Wand T⊥⊥ Y0|W,butTis not independent of (Y1Y0)|Wjointly. This implies that Thas an effect on the dependence structure of (Y1Y0)|Wbut not on their conditional marginal distributions. Suppose that individuals for whom Ti=1have an arbitrary dependence structure, while those with Ti=0have conditionally rank invariant potential outcomes. Then by the law of total probability, FY1Y0|W(y1y0|w) =(1−t)FY1Y0|WT(y1y0|w0)+tFY1Y0|WT(y1y0|w1) =(1−t)CFY1|WT(y1|w0)FY0|WT(y0|w0)|w0 +tCFY1|WT(y1|w1) FY0|WT(y0|w1)|w1 =(1−t)minFY1|W(y1|w)FY0|W(y0|w)+tHFY1|W(y1|w)FY0|W(y0|w) |w Our approach to relaxing rank invariance is an example of a more general approach. In this approach we take a weak assumption and a stronger assumption and use them to define a continuous class of assumptions by considering the population as a mixture of two subpopulations. The weak assumption holds in one subpopulation while the stronger assumption holds in the other subpopulation. The mixing proportion tcontinuously spans the two distinct assumptions we began with. This approach was used earlier by Horowitz and Manski (1995) in their analysis of the contaminated sampling model. While this general approach may not always be the most natural way to relax an assumption, it is often available and hence can be used to facilitate breakdown frontier analyses. Throughout the rest of this section we impose both conditional c-dependence and (1−t)-percent conditional rank invariance given W. Assumption A2. The following hold: 1. Xis conditionally c-dependent with the potential outcomes Yxgiven W,where for all w∈supp(W ) we have c<min{p1|wp0|w}. 2. (Y1Y0)satisfies (1−t)-percent conditional rank invariance given W,where t∈ [01]. For brevity we focus on the case c<min{p1|wp0|w}throughout this paper. This allows us to explain the key ideas while keeping the notation and derivations relatively simple. All of our results, however, can be relaxed to the general case where c∈[01]. Quantitative Economics 11 (2020) Inference on breakdown frontiers 55  px|w= 1 N N  i=1 1(Xi=xWi=w) 1 N N  i=1 1(Wi=w) (12) and  qw=1 N N  i=1 1(Wi=w) (13) denote the sample analog estimators of these three quantities, which converge uniformly to a Gaussian process at a √N-rate; see Lemma C1 in Appendix C. Next consider the bounds (4)and(5) on the marginal distributions of potential outcomes. These population bounds are a functional φ1evaluated at (FY|XW (·| ··)p(·|·)q(·))where p(·|·)denotes the probability px|was a function of (x w) ∈{01}× supp(W ),andq(·)denotes qwas a function of w∈supp(W ). We estimate these bounds by a plug-in estimator φ1( FY|XW (·|··)  p(·|·) q(·)). If this functional is differentiable in an appropriate sense, √N-convergence in distribution of its arguments will carry over to the functional by the delta method. The type of differentiability we require is Hadamard directional differentiability, first defined by Shapiro (1990)andDümbgen (1993), and further studied in Fang and Santos (2019). We use the functional delta method for Hadamard directionally differentiable mappings (e.g., Theorem 2.1 in Fang and Santos (2019)) to show convergence in distribution of our estimators. Such convergence is usually to a non-Gaussian limiting process. We do not use this distribution to do inference since obtaining analytical asymptotic confidence bands would be challenging. Instead, we use a bootstrap procedure to obtain asymptotically valid uniform confidence bands for our breakdown frontier and associated estimators. Returning to our population bounds (4)and(5), we estimate these by  Fc Yx|W(y |w) =min FY|XW (y |x w) px|w  px|w−c FY|XW (y |x w) px|w+c  px|w+c  Fc Yx|W(y |w) =max FY|XW (y |x w) px|w  px|w+c FY|X(y |xw) px|w−c  px|w−c (14) Note that these estimators may not perform well when cis close to px|w. In our analysis we assume cis bounded away from px|w. We show that these estimators converge in distribution to a nonstandard distribution in Appendix B.1, Lemma 1. In addition to Assumption A1, we make the following regularity assumptions. Assumption A5. For each x∈{01}and w∈supp(W ): 1. −∞<yx(w) < yx(w) < +∞. 2. FY|XW (y |x w) is continuously differentiable everywhere.Its density fY|XW (y | xw) is uniformly continuous in y,uniformly bounded from above,and uniformly bounded away from zero on supp(Y |X=xW =w). 56 Masten and Poirier Quantitative Economics 11 (2020) Assumption A5.1combined with our earlier Assumption A1.2constrain the potential outcomes to have compact support. This compact support assumption is not used to analyze our cdf bounds estimators (14), but we use it later to obtain estimates of the corresponding quantile function bounds uniformly over their arguments u∈(01), which we then use to estimate the bounds on P(QY1|W(U |w) −QY0|W(U |w) ≤z).This is a well-known issue when estimating quantile processes; for example, see van der Vaart (2000, Lemma 21.4(ii)). Assumption A5.2requires the density of Y|XW to be bounded away from zero uniformly. This ensures that conditional quantiles of Y|XW are uniquely defined. It also implies that the limiting distribution of the estimated quantile bounds are well behaved. Uniform continuity of the density implies that the derivatives of the conditional quantile function with respect to τare uniformly continuous. For our main result in this section (Theorem 2) along with some of our preliminary results (Lemmas 1,3,and4in Appendix B.1), we establish convergence uniformly over c∈Cfor some finite grid C={c1c2cJ}⊂[0 C]where C∈(0min{p1|wp0|w})for all w∈supp(W ). We discuss the choice of these grid points in Appendix B.3.Weconstrain this grid to be below min{p1|wp0|w}solely for simplicity, as all our results can be extended to grids C⊂[01]by combining our present bound estimates with estimates based on the c≥min{p1|wp0|w}case given in Masten and Poirier (2018a). Weak convergence of the breakdown frontier estimators does not hold uniformly over an interval of csince the associated functional is not Hadamard directionally differentiable when its codomain is a set of functions on that interval. To resolve this issue, we propose two ways of conducting inference on the breakdown frontier uniformly over intervals of c. The first is to use the fixed grid and monotonicity of the breakdown frontier to construct a uniform band. The second is to smooth the population breakdown frontier such that it is Hadamard differentiable when viewed as a function of c. We use the first approach in this section and in the empirical illustration, and discuss the second approach in Appendix G in the Online Supplemental Material. Next, we consider the conditional quantile bounds (6)and(7), which we estimate by  Qc Yx|W(τ) = QY|XW τ+c  px|w min{τ 1−τ}xw  Qc Yx|W(τ) = QY|XW τ−c  px|w min{τ 1−τ}xw (15) In Appendix B.1, Lemma 2, we establish their uniform convergence in distribution to a Gaussian process. By applying the functional delta method, we can show asymptotic normality of smooth functionals of these conditional quantile bounds. A first set of functionals are the CQTE bounds of equation (8), which are a linear combination of the quantile bounds. Let  CQTE(τc |w) = Qc Y1|W(τ |w) − Qc Y0|W(τ |w) (16) and  CQTE(τc |w) = Qc Y1|W(τ |w) − Qc Y0|W(τ |w) (17) Quantitative Economics 11 (2020) Inference on breakdown frontiers 57 Both of these estimators converge in distribution to Gaussian processes. A second set of functionals are the CATE bounds from equation (9). These bounds are continuous linear functionals of the CQTE bounds. Therefore the joint asymptotic distribution of these bounds can be established by the continuous mapping theorem. Let  CATE(c |w) =1 0  CQTE(u c |w)du and  CATE(c |w) =1 0  CQTE(u c |w)du Then, by the linearity of the integral operator, these estimated CATE bounds converge to their population counterpart at a √N-rate. We give their asymptotic distribution in Appendix B.1. We estimate the unconditional ATE bounds by integrating over the empirical distribution of the covariates W:Let  ATE(c) =1 N N  i=1  CATE(c |Wi)and  ATE(c) =1 N N  i=1  CATE(c |Wi) (18) We show that these estimated ATE bounds converge weakly to a Gaussian element in Appendix B.1. Next consider estimation of the breakdown point for the claim that ATE ≥μwhere μ∈R. To focus on the nondegenerate case, suppose the population value of ATE obtained under full independence is greater than μ,ATE (0)>μ. This implies c∗>0.Let  c∗=infc∈[01]: ATE(c) ≤μ(19) be the estimated breakdown point. This is the estimated smallest relaxation of independence such that we cannot conclude that the ATE is strictly greater than μ.Bytheproperties of the quantile bounds as a function of c, the function ATE(c) is nonincreasing and differentiable in c. We now formally present a result about the asymptotic distribution of  c∗. Proposition 1. Suppose Assumptions A1,A3,A4,and A5 hold.Assume c∗∈(0 C].Then √N(  c∗−c∗)Zbp,a Gaussian random variable. The assumption that c∗∈(0 C]can again be relaxed to the general case where c∗∈(01]but we maintain the stronger assumption for brevity. We characterize Zbp,the asymptotic distribution of the estimated breakdown point, in the proof of Proposition 1. Under conditional rank invariance, we can also establish asymptotic normality of bounds for P(QY1|W(U |w) −QY0|W(U |w) ≤z) for a fixed z∈R. These bounds are given by P(c |w)P(c |w)≡CDTE(z c0|w)CDTE(z c0|w) 58 Masten and Poirier Quantitative Economics 11 (2020) We keep zimplicit in the notation for these bounds. Estimates for these quantities are provided by  P(c |w) =1 0 1 Qc Y1|W(u |w) − Qc Y0|W(u |w) ≤zdu  P(c |w) =1 0 1 Qc Y1|W(u |w) − Qc Y0|W(u |w) ≤zdu (20) Their convergence to nonstandard distributions can be established using the Hadamard directional differentiability of the mapping from the differences in quantile bounds to the bounds (P(c |w)P(c |w)). We do this in Appendix B.1, Lemma 3. There we also discuss an additional technical assumption on the smoothness of the conditional quantile functions, Assumption A6. If conditional random assignment holds (c=0) in addition to conditional rank invariance (t=0), then the CDTE is point identified and Lemma 3in Appendix B.1 gives the asymptotic distribution of the sample analog CDTE estimator (in this case the upper and lower bound functions are equal). This can be considered an estimator of the CDTE in one of the models of Matzkin (2003). For cand tvalues greater than or equal to zero, we estimate the CDTE bounds by  CDTE(zct |w) =(1−t) P(c |w) +tmaxsup y∈Yz(w) Fc Y1|W(y |w) − Fc Y0|W(y −z|w)0  CDTE(zct|w) =(1−t) P(c |w) +t1+mininf y∈Yz(w) Fc Y1|W(y |w) − Fc Y0|W(y −z|w)0 We estimate the unconditional DTE bounds by integrating the estimated CDTE bounds over the empirical distribution of the covariates: Let  DTE(zct)=1 N N  i=1  CDTE(zct|Wi)  DTE(zct)=1 N N  i=1  CDTE(zct|Wi) Lemma 3in Appendix B.1 shows that the terms P(c |w) and P(c |w) are estimated at a √N-rate by the Hadamard directional differentiability of the mapping linking empirical cdfs and these terms. The second components of the CDTE bounds are a Hadamard directionally differentiable functional as well, leading to the √Njoint convergence of the DTE bounds to a tight, random element uniformly in cand t. Lemma 4in Appendix B.1 shows this formally. Having established the convergence in distribution of the DTE, we can now show that the breakdown frontier also converges in distribution uniformly over its arguments. Quantitative Economics 11 (2020) Inference on breakdown frontiers 59 Denote the estimated breakdown frontier for the conclusion that P(Y1>Y 0)≥pby  BF(c p)=minmax bf(c p)01(21) where  bf(c p)= num(c p)  denom(c) (22) with  num(c p)=1−p−1 N N  i=1 P(c |Wi)  denom(c) =1+1 N N  i=1mininf y∈Y0(Wi) Fc Y1|W(y |Wi)− Fc Y0|W(y |Wi)0− P(c |Wi) We next show that the estimated breakdown frontier converges in distribution. Theorem 2. Suppose Assumptions A1,A3,A4,A5,and A6 hold.Let P⊂[01]be a finite grid of points.Then √N BF(c p)−BF(cp)Zbf(c p) a tight random element of ∞(C×P). This result essentially follows from the convergence of the preliminary estimators established in Lemma C1 in Appendix Cand by showing that the breakdown frontier is a composition of a number of Hadamard differentiable and Hadamard directionally differentiable mappings, implying convergence in distribution of the estimated breakdown frontier. Breakdown frontiers for more complex conclusions can typically be constructed from breakdown frontiers for simpler conclusions. For example, consider the breakdown frontier for the joint conclusion that P(Y1>Y 0)≥pand ATE ≥μ. Then the breakdown frontier for this joint conclusion is the minimum of the two individual frontier functions. Alternatively, consider the conclusion that P(Y1>Y 0)≥por ATE ≥μ, or both, hold. Then the breakdown frontier for this joint conclusion is the maximum of the two individual frontier functions. Since the minimum and maximum operators are Hadamard directionally differentiable, the samaple analog estimators of these joint breakdown frontiers will also converge in distribution. Since the limiting process is non-Gaussian, inference on the breakdown frontier is not based on standard errors as with Gaussian limiting theory. Our processes’ distribution is characterized fully by the expressions in Appendices B.1 and C, but obtaining analytical estimates of quantiles of functionals of these processes would be challenging. In the next subsection we give details on the bootstrap procedure we use to construct confidence bands for the breakdown frontier. 60 Masten and Poirier Quantitative Economics 11 (2020) Bootstrap inference As mentioned earlier, we use a bootstrap procedure to do inference on the breakdown frontier rather than directly using its limiting process. In this subsection we discuss how to use the bootstrap to approximate this limiting process. In the next subsection we discuss its application to constructing uniform confidence bands. First we define some general notation. Let Zi=(YiXiWi)and ZN={Z1ZN}. Let θ0denote some parameter of interest and let  θbe an estimator of θ0based on the data ZN.LetA∗ Ndenote √N( θ∗− θ) where  θ∗is a draw from the nonparametric bootstrap distribution of  θ. Suppose Ais the tight limiting process of √N( θ−θ0).Denote bootstrap consistency by A∗ N P Awhere P denotes weak convergence in probability, conditional on the data ZN. Weak convergence in probability conditional on ZNis defined as sup h∈BL1EhA∗ N|ZN−Eh(A)=op(1) where BL1denotes the set of Lipschitz functions into Rwith Lipschitz constant no greater than 1. We focus on the following specific choices of θ0and  θ: θ0=⎛ ⎜ ⎝ FY|XW (·|··) p(·|·) q(·) ⎞ ⎟ ⎠and  θ=⎛ ⎜ ⎝ FY|XW (·|··)  p(·|·)  q(·) ⎞ ⎟ ⎠ For these choices, let Z∗ N=√N( θ∗− θ).LetZ1denote the limiting distribution of √N( θ−θ0); see Lemma C1 in Appendix C. Theorem 3.6.1 of van der Vaart and Wellner (1996) implies that Z∗ N P Z1. Our parameters of interest are all functionals φof θ0. For Hadamard differentiable functionals φ, the nonparametric bootstrap is consistent. For example, see Theorem 3.1 of Fang and Santos (2019). They further show that φis Hadamard differentiable if and only if √Nφ( θ∗)−φ( θ)P φ θ0(Z1) where φ θ0denotes the Hadamard derivative at θ0. This implies that the nonparametric bootstrap can be used to do inference on the QTE and ATE bounds since they are Hadamard differentiable functionals of θ0. A second implication is that the nonparametric bootstrap is not consistent for the DTE or for the breakdown frontier for claims about the DTE since they are Hadamard directionally differentiable mappings of θ0,but they are not ordinary Hadamard differentiable. In such cases, Fang and Santos (2019) showed that a different bootstrap procedure is consistent. Specifically, let  φ θ0be a consistent estimator of φ θ0. Under some regularity conditions, their results imply that  φ θ0Z∗ NP φ θ0(Z1) Quantitative Economics 11 (2020) Inference on breakdown frontiers 61 Analytical consistent estimates of φ θ0are often difficult to obtain, so Dümbgen (1993) and Hong and Li (2018) proposed using a numerical derivative estimate of φ θ0.Their estimate of the limiting distribution of √N(φ( θ) −φ(θ0)) is given by the distribution of  φ θ0√N( θ∗− θ)=φ θ+εN√N( θ∗− θ)−φ( θ) εN (23) across the bootstrap estimates  θ∗. Under the rate constraints εN→0and √NεN→∞, and some measurability conditions stated in their Appendix, Hong and Li (2018)showed  φ θ0√N( θ∗− θ)P φ θ0(Z1) where the left-hand side is defined in equation (23). This bootstrap procedure requires evaluating φat two values, which is computationally simple. It also requires selecting the tuning parameter εN,whichwediscussin Appendix B.4. Note that the standard, or naive, bootstrap is a special case of this numerical delta method bootstrap where εN=N−1/2. Uniform confidence bands for the breakdown frontier In this subsection, we combine all of our asymptotic results thus far to construct uniform confidence bands for the breakdown frontier. As in Section 2,weusethefunction BF(·p)to characterize this frontier. We specifically construct one-sided lower uniform confidence bands. That is, we will construct a lower band function  LB(c) such that lim N→∞ P LB(c) ≤BF(cp)for all c∈[01]=1−α We use a one-sided lower uniform confidence band because this gives us an inner confidence set for the robust region. Specifically, define the set RRL=(c t) ∈[01]2:t≤ LB(c) Then validity of the confidence band  LB implies lim N→∞ PRRL⊆RR(0p)=1−α Thus the area underneath our confidence band, RRL, is interpreted as follows: Across repeated samples, approximately 100(1−α)% of the time, every pair (ct) ∈RRLleads to a population level identified set for the parameter P(Y1>Y 0)which lies weakly above p. Put differently, approximately 100(1−α)% of the time, every pair (ct) ∈RRLstill lets us draw the conclusion we want at the population level. Hence the size of this set RRLis a finite sample measure of robustness of our conclusion to failure of the point identifying assumptions. We discuss an alternative testing-based interpretation in Appendix D in the Online Supplemental Material. One might be interested in constructing one-sided upper confidence bands if the goal was to do inference on the set of assumptions for which we cannot come to the 62 Masten and Poirier Quantitative Economics 11 (2020) conclusion of interest. This might be useful in situations where two opposing sides are debating a conclusion. But since our focus is on trying to determine when we can come to the desired conclusion, rather than looking for when we cannot, we only describe the one-sided lower confidence band case. When studying inference on scalar breakdown points, Kline and Santos (2013)constructed one-sided lower confidence intervals. Unlike for breakdown frontiers, uniformity over different points in the assumption space is not a concern for inference on breakdown points. See Appendix D in the Online Supplemental Material for more discussion. We consider bands of the form  LB(c) = BF(c p)− k(c) (24) for some function  k(·)≥0. This band is an asymptotically valid lower uniform confidence band of level 1−αif lim N→∞ P BF(c p)− k(c) ≤BF(cp)for all c∈[01]=1−α or, equivalently, if lim N→∞ Psup c∈[01] √N BF(c p)−BF(c p)− k(c)≤0=1−α In our theoretical analysis, we consider  k(c) = z1−ασ(c) for a scalar z1−αand a function σ. We focus on known σfor simplicity. We start by deriving a uniform band over a grid C, then extend it over an interval using monotonicity of the breakdown frontier. As discussed earlier, we only derive uniformity of the band over c∈[0C]rather than over c∈[01], but this is also for brevity and can be relaxed. The choice of σaffects the shape of the confidence band, and there are many possible choices of the function σwhich yield valid level 1−αuniform confidence bands. See Freyberger and Rai (2018)foradetailed analysis. A simple choice of σis the constant function: σ(c) =1, which delivers an equal width uniform band. Alternatively, as we do below, one could choose σ(c) to construct a minimum width confidence band (equivalently, maximum area of RRL). Proposition 2. Suppose Assumptions A1,A3,A4,A5,and A6 hold.Define φ:∞(R× {01}×supp(W )) ×∞({01}×supp(W )) ×∞(supp(W )) →∞(C)such that  BF(c p)=φ( θ)(c) Then φis Hadamard directionally differentiable.Suppose that εN→0and √NεN→∞. Let  θ∗denote a draw from the nonparametric bootstrap distribution of  θ.Then  φ θ0√N( θ∗− θ)P φ θ0(Z1)≡Zbf(25) For a given function σ(·)such that infc∈Cσ(c)>0,define  z1−α=infz∈R:Psup c∈C φ θ0√N( θ∗− θ)(c p) σ(c) ≤zZN≥1−α(26) Quantitative Economics 11 (2020) Inference on breakdown frontiers 63 Finally,suppose also that the cdf of sup c∈Cφ θ0(Z1)(c p) σ(c) =sup c∈C Zbf(c p) σ(c) is continuous and strictly increasing at its 1−αquantile,denoted z1−α.Then  z1−α= z1−α+op(1). This proposition is a variation of Corollary 3.2 in Fang and Santos (2015). Note that this result is pointwise in the underlying dgp; we are unsure if it can be extended to hold uniformly over the dgp and leave this question to future work. Proposition 2implies that the lower 1−αband  LB(c) = BF(c p)− z1−ασ(c)is valid uniformly on the grid C.To extend the uniformity to all of [0C], we propose the following lower confidence band: LB(c) = ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩  LB(c1)if c∈[0c1]     LB(cj)if c∈(cj−1cj],forj=2J    0if c∈(cJC] This band is a step function which interpolates between grid points using the least monotone interpolation. The following result shows its validity. Corollary 1. Let the assumptions of Proposition 2hold.Then LB(c) is a uniform lower 1−αband for BF(cp)over c∈[0 C]. Corollary 1shows that, for any fixed J≥1, the interpolated lower confidence band preserves the exact 1−αcoverage on the grid points. This follows by monotonicity of the breakdown frontier; see Lemma C6 in Appendix C. That said, this interpolated band might not be taut, in the sense that there may exist other lower bands with 1−αcoverage that are weakly larger than LB(c) for all cand strictly larger at some values of c.See Freyberger and Rai (2018) for further discussion of taut confidence bands. Proposition 2can be extended to estimated functions σ, although we leave the details for future work. We use an estimated σin our application, as described next. When both z1−αand σare estimated, we work directly with  k(c) = z1−ασ(c). We choose  k(c) to minimize an approximation to the area between the confidence band and the estimated function; equivalently, to maximize the area of RRL. Specifically, we let  k(c1) k(cJ) solve min k(c1)k(cJ)≥0 J  j=2 k(cj)(cj−cj−1)(27) subject to Psup c∈{c1cJ} √N BF(c p)−BF(c p)−k(c)≤0=1−α (28) 64 Masten and Poirier Quantitative Economics 11 (2020) where we approximate the left-hand side probability via the numerical delta method bootstrap. The criterion function here is just a right Riemann sum over the grid points. This optimization is not computationally costly: It is only performed once per value of p and εN. Moreover, in our empirical illustration it takes an average of 15 seconds per run on a mid-2013 MacBook Air. 4. Empirical illustration:The effects of child soldiering In this section we use our results to examine the impact of assumptions in determining the effects of child soldiering on wages. We first briefly discuss the background and motivation and then we present our analysis. Background We use data from phase 1 of SWAY, the Survey of War Affected Youth in northern Uganda, conducted by principal researchers Jeannie Annan and Chris Blattman (see Annan, Blattman, and Horton (2006)). As Blattman and Annan (2010) discuss on page 882, a primary goal of this survey was to understand the effects of a twenty year war in Uganda, where “an unpopular rebel group has forcibly recruited tens of thousands of youth.” In their paper, they use this data to examine the impacts of abduction on educational, labor market, psychosocial, and health outcomes. In our illustration, we focus solely on the impact of abduction on wages. Blattman and Annan noted that self-selection into the military is a common problem in the literature studying the effects of military service on outcomes. They argue that forced recruitment in Uganda led to random assignment of military service in their data. They first provide qualitative evidence for this, based on interviews with former rebels who led raiding parties. After murdering and mutilating civilians, the rebels had no public support, making abduction the only means of recruitment. Youths were generally taken during nighttime raids on rural households. According to the former rebel leaders, “targets were generally unplanned and arbitrary; they raided whatever homesteads they encountered, regardless of wealth or other traits.” This qualitative evidence is supported by their survey data, where Blattman and Annan show that most pretreatment covariates are balanced across the abducted and nonabducted groups (see their Table 2). Only two covariates are not balanced: year of birth and prewar household size. They say this is unsurprising because “a youth’s probability of ever being abducted depended on how many years of the conflict he fell within the [rebel group’s] target age range. Moreover, abduction levels varied over the course of the war, so youth of some ages were more vulnerable to abduction than others. The significance of household size, meanwhile, is driven by households greater than 25 in number. We believe that rebel raiders, who traveled in small bands, were less likely to raid large, difficult-to-control households.” (page 887) Hence they use a selection-on-observables identification strategy, conditioning on these two variables. Quantitative Economics 11 (2020) Inference on breakdown frontiers 71 the empirical cdf using the bootstrap data. We use empirical cdfs here since they are substantially faster to compute than smoothed cdfs. For each xand wvalue, we then monotonize this perturbed cdf and truncate it to lie between 0and 1,toensurethat it is a valid cdf. This is not necessary asymptotically, but it is helpful in finite samples. Moreover, this operation is Hadamard directionally differentiable and hence does not affect the asymptotic theory. (b) Next we compute the breakdown frontier at all grid points cand all pusing these perturbed values as the first step plug-in estimators. This gives us φ( θ+ε√N( θ∗− θ)) from equation (23). At this step we compute the integral defining the term  P(c |w) differently than above. The root-based approach is not always reliable in these bootstrap perturbed samples since the term inside the indicator function can oscillate dramatically. So here we instead take a grid of 500 equally spaced values between 0and 1,evaluate the integrand at all values and take the average. We also use a slightly smoothed indicator function for the integrand. (c) Continuing from the previous step, subtract the estimated breakdown frontier using the original data and divide by εto finish computing equation (23). 3. Fixing an εand a c, the distribution of the term from the previous step across the bootstrap samples approximates the sampling distribution of √N BF(c p)−BF(c p) This continues to hold if we look at the vector of these terms for all cin the grid. Thus, for any given vector of constants k(c), we can approximate the left-hand side of equation (28) by looking at the distribution of the supremum over cof equation (23)fromthe previous step minus √Nk(c). Specifically, note that the inequality inside the probability in equation (28) can be written as a product of indicator functions for each grid point. Since we will optimize subject to this constraint, it is helpful to smooth it. So we approximate each indicator function by a smoothed indicator. We use the logitstic kernel and for each cwe pick the bandwidth by Hansen’s (2004) method, using the distribution of the term from equation (23) across the bootstrap samples as data. For a given bootstrap dataset, we compute the product of these indicators. We then average that product over all bootstrap datasets to get an estimate of the probability on the left-hand side of equation (28). 4. Solve equation (27) subject to equation (28). Since we smoothed the constraint function this can be done using any nonlinear optimizer. At this step we must pick our desired coverage probability. We compute 95% confidence bands. Denote the solution vector by  k(c). 5. Finally, we compute the lower confidence band using equation (24). For certain values of ε, this band can be nonmonotonic; in these cases we monotonize the band. This procedure produces seven different confidence bands—one for each ε.Topickour preferred band, we use the double bootstrap procedure described in Appendix B.4.We implement this as follows: 72 Masten and Poirier Quantitative Economics 11 (2020) 1. Draw and store 500 bootstrap samples of the data. For this we use the smoothed bootstrap. Specifically, for each draw i=1N, (a) Draw a value of the covariate vector from its empirical distribution. (b) Draw a treatment value from the Bernoulli distribution with success parameter equal to the estimated propensity score evaluated at the covariate draw from the previous step. (c) Draw from the Unif[01]distribution. Evaluate the smoothed quantile function  F−1 Y|XW (·|x w) at this draw, where wis from the first step and xis from the second step. This defines the outcome variable value for this draw. We compute the smoothed quantile by linear interpolation of the smoothed cdf, as described earlier. 2. For each ε, (a) For each of the 500 bootstrap samples, treat it as if it was the true data and compute a confidence band as we described above using this choice of ε. Check whether this confidence band lies under the breakdown frontier estimated from the original data. Note that this estimate corresponds to the true breakdown frontier in the “bootstrap world” where we treat the smoothed empirical cdf  FY|XW as the true distribution of outcomes given treatment and covariates. (b) Compute the estimated coverage probability for this choice of εby averaging the indicators of coverage for each of the 500 samples. 3. Pick the εwith estimated coverage probability closest to 95%. Part 2(a) can be done in parallel on a computing cluster. Breakdown frontier for P(Y1>Y 0):Results Figure 3shows the results of the procedure we just described. As in our earlier plots, the horizontal axis plots c, relaxations of full conditional independence, while the vertical Figure 3. Estimated breakdown frontiers (solid lines) and confidence bands (dotted/dashed lines) for the claim P(Y1>Y 0)≥p.Lefttoright:p=01,025,05. The light dotted lines are confidence bands for all seven values of εNconsidered. The darker dashed line is the band selected by the double bootstrap procedure described in Appendix B.4. See text for additional implementation details. Quantitative Economics 11 (2020) Inference on breakdown frontiers 73 axis plots t, relaxations of conditional rank invariance. As mentioned earlier, the natural upper bound for cis about 073. Since all of the breakdown frontiers intersect the horizontal axis at much smaller values, we have cut off the part of the overall assumption space with c≥02. Remember that, for the following analysis, it is valid to examine various (ct) combinations since we use uniform confidence bands. First consider the left plot, p=01. Since this is the weakest conclusion of the five we consider, the estimated breakdown frontier and the corresponding robust region are the largest among the three plots. If we impose full conditional independence, then our estimated frontier suggests that we can completely relax conditional rank invariance and still conclude that at least 10% of people benefit from not being forced into military service. Even accounting for sampling uncertainty, we can still draw this conclusion. Moreover, looking at all choices of εN—not just our selected one—the lowest the vertical intercept ever gets is about 61%. Next, suppose we relax full conditional independence. Recall that the maximal relaxation between the observed propensity score and the “leave out variable k” propensity scores gave cage =00625 and chhSize =00403.Bothofthese numbers are substantially smaller than the horizontal intercept of our selected confidence band. Hence, if we impose full conditional rank invariance, our conclusion that P(Y1>Y 0)≥01is robust to relaxations of full conditional independence. Suppose instead that we think selection on unobservables is at most the largest cvalue, about 006. Then for c’s in the range [0006], and accounting for sampling uncertainty, we can still conclude P(Y1>Y 0)≥01so long as at least 30% of the population satisfies rank invariance. Thus we can relax full independence within this range without paying too high a cost in terms of requiring stronger rank invariance assumptions. If we are willing to restrict selection on unobservables to be smaller than the largest cvalue, then we can allow for larger relaxations of conditional rank invariance. To quantify this trade-off, we can compute the difference between the values of the estimated breakdown frontier at two different points. As a starting point we recommend computing  BF(c(K)p)− BF(c(K−1)p) where Kdenotes the number of observed regressors, c(K) denotes the largest value of c,andc(K−1)denotes the second largest value. One could also divide this difference by c(K) −c(K−1)to get a secant line. In our empirical application, c(K) =cage and c(K−1)= chhSize.Thus  BF(cage01)− BF(chhSize01)=−6% The corresponding secant line has slope about −3. Thus if we assume selection on unobservables is at most as large as the second largest amount of variation in “leave out variable k” propensity scores, we can allow for an additional 6% of the population to violate conditional rank invariance. Put differently, around these values of c, allowing latent conditional propensity scores to vary an extra 1percentage point requires us to impose that an additional 3% of the population must satisfy conditional rank invariance. This rate of substitution generally increases as cgets larger. Our ability to quantify 74 Masten and Poirier Quantitative Economics 11 (2020) this kind of trade-off between assumptions is a primary goal of our breakdown frontier analysis. Overall, our results from this top left plot suggest that the conclusion that at least 10% of people benefit from not being forced into military service is robust to relaxations of full conditional independence up to twice the size we see between the observed and leave out variable kpropensity scores, depending on how much conditional rank invariance failure we allow. For relaxations of full conditional independence up to the largest value of c, we can allow up to 70% of the population to deviate from conditional rank invariance, accounting for sampling uncertainty. Next consider the middle plot, p=025. Since this is a stronger conclusion than the previous one, all the frontiers are shifted toward the origin. Consequently, by construction, this conclusion is not as robust as the previous one. Our qualitative conclusions, however, are similar to those obtained for p=01. If we impose full conditional independence, we can allow conditional rank invariance to fail for about 70% of the population. Conversely, if we impose full conditional rank invariance, we can allow the latent conditional propensity scores to vary by about 10 percentage points—well beyond the largest observed variation c.Forp=025,wehave  BF(cage025)− BF(chhSize025)=−17% Hencetheslopearoundourobservedmaximalc’s is much larger for p=025 as compared to p=01. An important caveat to our conclusions for both p=01and p=025 is that there is substantial variation in confidence bands as εNchanges. This point underscores the need for future work on the choice of εN. Next consider the right plot, p=05. Here we consider the conclusion that at least half of people benefit from not being forced into military service. If we impose full conditional independence, and accounting for sampling uncertainty, then we can allow conditional rank invariance to fail for about 25% of the population. This is quite large, but it relies on full conditional independence holding exactly. If we also relax conditional independence to c=003, then we need conditional rank invariance to hold for everyone if we still want to conclude that at least 50% of people benefit from not being forced into military service. 003 is smaller than both cage and chhSize.Hencewemightnotbe comfortable with such small values of c. This suggests the data do not definitively support the conclusion P(Y1>Y 0)≥05, even though our point estimate under the baseline assumptions is 067. Empirical conclusions In this section we used our breakdown frontier methods to study the robustness of conclusions about ATE and P(Y1>Y 0)to failures of conditional independence and conditional rank invariance. We first considered the conclusion that the average treatment effect of not being abducted on log wages is nonnegative. Our point estimates suggest that this conclusion is robust to deviations in unobserved latent propensity scores of up to the same value as cage, which is also about two-thirds as large as chhSize;thisrobustness does not hold up when accounting for sampling uncertainty, however. We then Quantitative Economics 11 (2020) Inference on breakdown frontiers 75 considered the conclusion that at least p%of people earn higher wages when they are not abducted. This conclusion is robust to large simultaneous relaxations of conditional rank invariance and conditional independence for p=10%.Forp=25%, this conclusion continues to be robust to reasonable relaxations, although after accounting for the variation in confidence bands over εN, this conclusion appears to be more sensitive to conditional independence than to conditional rank invariance. This robustness to rank invariance matches the findings of Heckman, Smith, and Clements (1997), who imposed full independence and studied deviations from rank invariance. In their table 5B they found that, in their empirical application, one could generally conclude that P(Y1>Y 0) was at least 50%, regardless of the assumption on rank invariance. In our empirical application our results are not quite as robust to rank invariance failures, which could be because we use a different measure of relaxation of rank invariance, and also because of differences in the empirical applications. 5. Conclusion Summary In this paper we advocated the breakdown frontier approach to sensitivity analysis. Given a set of baseline assumptions, this approach defines the population breakdown frontier as the weakest set of assumptions such that a specific conclusion of interest holds. Sample analog estimates and lower uniform confidence bands allow researchers to do inference on this frontier. The area under the confidence band is a quantitative, finite sample measure of the robustness of a conclusion to relaxations of point identifying assumptions. To examine this robustness, empirical researchers can present these estimated breakdown frontiers and their accompanying confidence bands along with traditional point estimates and confidence intervals obtained under point identifying assumptions. We illustrated this general approach in the context of a treatment effects model, where the robustness of conclusions about ATE and P(Y1>Y 0)to relaxations of random assignment and rank invariance are examined. We applied these results in an empirical study of the effect of child soldiering on wages. We found that weak conclusions about P(Y1>Y 0)are fairly robust to failures of both rank invariance and random assignment, but stronger conclusions are more sensitive to relaxations of random assignment. Breakdown frontier analysis for other models and other relaxations As discussed in Section 1, breakdown frontier analysis can in principle be done in most models. In that section we outlined the six main steps required for any breakdown frontier analysis. In this paper we illustrated this general approach by studying a single important and widely used model: the potential outcomes model with a binary treatment. In future work it would be helpful to perform breakdown frontier analyses in other models. In particular, it may be possible to do breakdown frontier analyses in a large class of models by using the general identification analysis in Chesher and Rosen (2017)orTorgovitsky (2019). 76 Masten and Poirier Quantitative Economics 11 (2020) A key conceptual step in any breakdown frontier analysis is deciding how to define the indexed classes of assumptions such that the magnitude of the relaxation can be reasonably interpreted. This is not easy, and will generally depend on the model, the specific kind of assumption being relaxed, and the empirical context. Moreover, this choice may affect our findings: A conclusion can be robust with respect to one measure of relaxation but not another. Thus one goal of future research is to explore this space of assumption relaxations, to understand their substantive interpretations, and to chart their implications for the robustness of empirical findings. In Masten and Poirier (2016) we have already compared three different measures of relaxation of the random assignment assumption, including the one used here. We further studied quantile independence, a common relaxation of random assignment, in Masten and Poirier (2018b). In the present paper, we also used a general method for spanning two discrete assumptions by defining a (1−t)-percent relaxation, as we did with rank invariance. But much work still remains to be done. Appendix A: Related literature We begin with the identification literature on breakdown points; as mentioned earlier, here we use “breakdown” in the same sense as Horowitz and Manski’s (1995) identification breakdown point. This breakdown point idea goes back to the one of the earliest sensitivity analyses, performed by Cornfield, Haenszel, Hammond, Lilienfeld, Shimkin, and Wynder (1959). They essentially asked how much correlation between a binary treatment and an unobserved binary confounder must be present to fully explain an observed correlation between treatment and a binary outcome, in the absence of any causal effects of treatment. This level of correlation between treatment and the confounder is a kind of breakdown point for the conclusion that some causal effects of treatment are nonzero. Their approach was substantially generalized by Rosenbaum and Rubin (1983), which is discussed in detail in Chapter 22 of Imbens and Rubin (2015). Neither Cornfield et al. (1959)norRosenbaum and Rubin (1983) formally defined breakdown points. Horowitz and Manski (1995) gave the first formal definition and analysis of breakdown points. They studied a “contaminated sampling” model, where one observes a mixture of draws from the distribution of interest and draws from some other distribution. An upper bound λon the unknown mixing probability indexes identified sets for functionals of the distribution of interest. They focus on a single conclusion: That this functional is not equal to its logical bounds. They then define the breakdown point λ∗ as the largest λsuch that this conclusion holds. Put differently, λ∗is the largest mixing probability we can allow while still obtaining a nontrivial identified set for our parameter of interest. They also relate this “identification breakdown point” to the earlier breakdown point concepts studied in the robust statistics literature (e.g., Hampel, Ronchetti, Rousseeuw, and Stahel (1986, pp. 96–98) and Huber and Ronchetti (2009,Section1.4and Chapter 11)). More generally, much work by Manski distinguishes between informative and noninformative bounds (which the literature also sometimes calls tight and nontight Quantitative Economics 11 (2020) Inference on breakdown frontiers 77 bounds; see Section 7.2 of Ho and Rosen (2017)). The breakdown point is the boundary between the informative and noninformative cases. For example, see his analysis of bounds on quantiles with missing outcome data on page 40 of Manski (2007). There the identification breakdown point for the τth quantile occurs when max{τ 1−τ}is the proportion of missing data. Similar discussions are given throughout the book. Stoye (2005,2010) generalizes the formal identification breakdown point concept by noting that breakdown points can be defined for any claim about the parameter of interest. He then studies a specific class of relaxations of the missing-at-random assumption in a model of missing data. Kline and Santos (2013) studied a different class of relaxations of the missing-at-random assumption and also define a breakdown point based on that class. While all of these papers study a scalar breakdown point, Imbens (2003)studieda model of treatment effects where deviations from conditional random assignment are parameterized by two numbers r=(r1r2). His parameter of interest θ(r) is point identified given a fixed value of r. Imbens’ Figures 1–4 essentially plot estimated level sets of this function θ(r), in a transformed domain. While suggestive, these level sets do not generally have a breakdown frontier interpretation. This follows since nonmonotonicities in the function θ(r) lead to level sets which do not always partition the space of sensitivity parameters into two connected sets in the same way that our breakdown frontier does. Manski and Pepper (2018) also studied a model where relaxations of baseline assumptions are parameterized by a vector of numbers r. Unlike Imbens, however, they derive identified sets indexed by r. These sets are weakly increasing (in the set inclusion order) in each component of r, and hence the nonmonotonicity issue does not arise. For a two-dimensional relaxation, their Table 2 presents identified sets as a function of a grid of r=(r1r2)values. The boundary between the italicized identified sets in that table and the nonhighlighted sets is essentially a discrete approximation to the breakdown frontier in their model, for the claim that the parameter of interest is positive. Similarly, the boundary between the bold identified sets in that table and the nonhighlighted sets is essentially a discrete approximation to the breakdown frontier in their model, for the claim that the parameter of interest is negative. Neither Horowitz and Manski (1995)norStoye (2005,2010) discussed estimation or inference of breakdown points. Imbens (2003) estimated his level sets in an empirical application, but does not discuss inference. Manski and Pepper (2018)alsodonotdiscuss estimation of or inference on breakdown frontiers, although inference in their setting is conceptually complicated—see their discussion on pages 234–235. Kline and Santos (2013), on the other hand, is the first and only paper we are aware of that explicitly suggests doing inference on a breakdown point. We build on their work by proposing to do inference on the multidimensional breakdown frontier. This allows us to study the trade-off between different assumptions in drawing conclusions. They do study something they call a “breakdown curve,” but this is a collection of scalar breakdown points for many different claims of interest, analogous to the collection of frontiers presented in Figure 2. Inference on a frontier rather than a point also raises additional issues they did not discuss; see our Appendix D in the Online Supplemental Material for more details. 78 Masten and Poirier Quantitative Economics 11 (2020) Moreover, we study a model of treatment effects while they look at a model of missing data, hence our identification analysis is different. Building on Horowitz and Manski (1995), Kreider, Pepper, Gundersen, and Jolliffe (2012) combine a continuous relaxation sensitivity analysis for assumptions regarding measurement error with various discrete relaxations of assumptions regarding treatment selection. This allows them to study the interaction between these two kinds of assumptions in drawing conclusions. For inference, they present confidence intervals for partially identified parameters for a variety of values of the relaxations, rather than doing inference on breakdown frontiers. See Gundersen, Kreider, and Pepper (2012)and Kreider, Pepper, and Roy (2016) for further examples of identification analysis combining discrete and continuous relaxations. Our breakdown frontier is a known functional of the distribution of outcomes given treatment and covariates and the observed propensity scores. This functional is not Hadamard differentiable, however, which prevents us from applying the standard functional delta method to obtain its asymptotic distribution. Instead, we show that it is Hadamard directionally differentiable, which allows us to apply the results of Fang and Santos (2019). We then use the numerical bootstrap of Dümbgen (1993)andHong and Li (2018) to construct our confidence bands. For other applications of Hadamard directional differentiability, see Kaido (2016), Hansen (2017), and Lee and Bhattacharya (2019). Our identification analysis builds on two strands of literature. First is the literature on relaxing statistical independence assumptions. There is a large literature on this, including important work by Rosenbaum and Rubin (1983), Robins, Rotnitzky, and Scharfstein (2000), and Rosenbaum (1995,2002). We apply results from our paper Masten and Poirier (2018a), which discusses that literature in more detail. In that paper we did not study estimation or inference. Second is the literature on identification of the distribution of treatment effects Y1−Y0, especially without the rank invariance assumption. In their introduction, Fan, Guerre, and Zhu (2017) provided a comprehensive discussion of this literature; also see Abbring and Heckman (2007, Section 2). Here we focus on the papers most related to our sensitivity analysis. Heckman, Smith, and Clements (1997) performed a sensitivity analysis to the rank invariance assumption by fixing the value of Kendall’s τfor the joint distribution of potential outcomes, and then varying τfrom −1to 1; see Tables 5A and 5B. Their analysis is motivated by a search for breakdown points, as evident in their Section 4 title, “How far can we depart from perfect dependence and still produce plausible estimates of program impacts?” Nonetheless, they do not formally define identified sets for parameters given their assumptions on Kendall’s τ, and they do not formally define a breakdown point. Moreover, they do not suggest estimating or doing inference on breakdown points. Gechter (2016) performed a sensitivity analysis to the rank invariance assumption by fixing a lower bound on the value of Spearman’s ρ. Under this assumption he derives the identified set for a certain average treatment effect. He then studies estimation and inference on this set for a fixed value of the sensitivity parameter. Fan and Park (2009) provided formal identification results for the joint cdf of potential outcomes and the distribution of treatment effects under the known Kendall’s τassumption. They also discuss how to extend those results Quantitative Economics 11 (2020) Inference on breakdown frontiers 79 to known Spearman’s ρin their Remark 1. They provide estimation and inference methods for their bounds, but do not study breakdown points. Finally, none of these papers study the specific relaxation of rank invariance we consider (as defined in Section 2). In this section, we have focused narrowly on the papers most closely related to ours. We situate our work more broadly in the literature on inference in sensitivity analyses in Appendix D in the Online Supplemental Material. In that section we also briefly discuss Bayesian inference, although we use frequentist inference in this paper. Appendix B: Additional inference results and discussion B.1 Asymptotic results for the bound functionals In this section we provide asymptotic results for the various bound functionals discussed in this paper. These are preliminary results used in our breakdown analysis done in Section 3. They can also be used on their own as inputs to traditional inference on partially identified parameters. First we establish convergence in distribution of the cdf bound estimators (14). Here and throughout the paper we use the following notation: For an arbitrary set Aand a Banach space B,∞(AB)denotes the set of all maps z:A→Bwith finite sup-norm z=supa∈Az(a)B, equipped with this norm. For example, see van der Vaart and Wellner (1996, p. 381). Lemma 1. Suppose Assumptions A1,A3,and A4 hold.Let Y⊂Rbe a finite grid of points. Then √N% Fc Yx|W(y |w) −Fc Yx|W(y |w)  Fc Yx|W(y |w) −Fc Yx|W(y |w)&Z2(yxwc) a tight random element of ∞(Y×{01}×supp(W ) ×CR2). Z2is not Gaussian itself, but it is a continuous transformation of Gaussian processes. For given (xcw), the limit will be Gaussian at all values of yexcept for y∈QY|XW px|w−c 2px|wxwQ Y|XW px|w+c 2px|wxw Next we consider the conditional quantile bound estimators. Recall that C∈ (0min{p1|wp0|w})for all w∈supp(W ). Lemma 2. Suppose Assumptions A1,A3,A4,and A5 hold.Then √N% Qc Yx|W(τ |w) −Qc Yx|W(τ |w)  Qc Yx|W(τ |w) −Qc Yx|W(τ |w)&Z3(τxwc) a mean-zero Gaussian process in ∞((01)×{01}×supp(W ) ×[0C]R2)with continuous paths. 80 Masten and Poirier Quantitative Economics 11 (2020) This result is uniform in con an interval, in x∈{01},w∈supp(W ), and in τ∈(01). This result directly implies convergence over c∈Cas well. Unlike the distribution of the cdf bounds estimators, this process is Gaussian. This follows by Hadamard differentiability of the mapping between θ0≡(FY|XW (·|··) p(·|·)q(·))and the conditional quantile bounds. Let the superscript Z(j) denotes the jth component of the vector Z. Since the CQTE bounds are functionals of these conditional quantile bounds, Lemma 2implies the following convergence result: √N% CQTE(τc |w) −CQTE(τ c |w)  CQTE(τc |w) −CQTE(τ c |w)&%Z(1) 3(τ 1wc)−Z(2) 3(τ 0wc) Z(2) 3(τ 1wc)−Z(1) 3(τ 0wc)& Similarly, Lemma 2also implies √N% CATE(c |w) −CATE(c |w)  CATE(c |w) −CATE(c |w)&⎛ ⎜ ⎜ ⎝1 0Z(1) 3(u1wc)−Z(2) 3(u0wc)du 1 0Z(2) 3(u1wc)−Z(1) 3(u0wc)du ⎞ ⎟ ⎟ ⎠ a mean-zero Gaussian process in ∞(supp(W ) ×[0C]R2)with continuous paths. Next consider the ATE bounds. The following decomposition implies that the estimated ATE upper bound converges weakly to a Gaussian element: √N ATE(c) −ATE(c) =√N K  k=1 qwk CATE(c |wk)−CATE(c |wk)+ K  k=1 CATE(c |wk)√N( qwk−qwk)  K  k=1 qwk1 0Z(1) 3(u1wkc)−Z(2) 3(u0wkc)du + K  k=1 CATE(c |wk)Z(3) 1(00wk) A similar result holds for the estimated ATE lower bound. The mapping used in equation (20) is called the pre-rearrangement operator.Chernozhukov, Fernández-Val, and Galichon (2010) showed that this operator was Hadamard differentiable when the quantile functions are continuously differentiable for all u∈ (01). In our case, the underlying quantile functions are continuously differentiable on (01/2)∪(1/21), and continuous but not differentiable at u=1/2. At this value, the left and right derivatives exist and are finite, but are generally different from one another. We extend the result of Chernozhukov, Fernández-Val, and Galichon (2010)tothecase where the quantile function has a point of nondifferentiability by showing Hadamard directional differentiability of this mapping. To do so, we make additional assumptions on the behavior of these quantile functions. Quantitative Economics 11 (2020) Inference on breakdown frontiers 87 Lemma C1. Suppose Assumption A3 and Assumption A4 hold.Then √N⎛ ⎜ ⎝ FY|XW (y |x w) −FY|XW (y |x w)  px|w−px|w  qw−qw ⎞ ⎟ ⎠Z1(yxw) a mean-zero Gaussian process in ∞(R×{01}×supp(W ) R3)with continuous paths and covariance kernel equal to Σ1(yxw˜ y ˜ x ˜ w) =EZ1(yxw)Z1(˜ y ˜ x ˜ w) =diag⎛ ⎜ ⎜ ⎜ ⎜ ⎝ FY|XW min{y ˜ y}|xw−FY|XW (y |xw)FY|XW (˜ y|xw) px|wqw 1(x =˜ xw =˜ w) px|w qw 1(x =˜ xw =˜ w) −px|wp˜ x|w qw 1(w =˜ w) qw1(w =˜ w) −qwq˜ w ⎞ ⎟ ⎟ ⎟ ⎟ ⎠  Proof of Lemma C1. By a second-order Taylor expansion,  FY|XW (y |x w) −FY|XW (y |x w) = 1 N N  i=1 1(Yi≤y)1(Xi=xWi=w) 1 N N  i=1 1(Xi=xWi=w) − P(Y ≤yX =xW =w) P(X =xW =w) = 1 N N  i=1 1(Yi≤y)1(Xi=xWi=w) −P(Y ≤yX =xW =w) P(X =xW =w) −FY|XW (y |x w) P(X =xW =w)%1 N N  i=1 1(Xi=xWi=w) −P(X =xW =w)& +Op'%1 N N  i=1 1(Yi≤y)1(Xi=xWi=w) −FY|XW (y |x w)P(X =xW =w)& ·%1 N N  i=1 1(Xi=xWi=w) −P(X =xW =w)&( +Op'%1 N N  i=1 1(Xi=xWi=w) −P(X =xW =w)&2( By standard bracketing entropy results (e.g., example 19.6 on p. 271 of van der Vaart (2000)) the function classes {1(Y ≤y)1(X =x)1(W =w) :y∈Rx∈{01}w∈supp(W )} 88 Masten and Poirier Quantitative Economics 11 (2020) and {1(X =x)1(W =w) :x∈{01}w∈supp(W )}are both P-Donsker. Hence the residual is of order Op(N−1)uniformly in (yxw) ∈R×{01}×supp(W ). Combining this with Slutsky’s theorem, we get the uniform over y,x,andwasymptotically linear representation  FY|XW (y |x w) −FY|XW (y |x w) =1 N N  i=1 1(Xi=xWi=w)1(Yi≤y)−FY|XW (y |x w) P(X =xW =w) +opN−1/2 By the same bracketing entropy arguments, the class 1(X =xW =w)1(Y ≤y)−FY|XW (y |x w) P(X =xW =w) :y∈Rx∈{01}w∈supp(W ) is P-Donsker, and hence √N( FY|XW (·|··)−FY|XW (·|··)) converges in distribution to a mean-zero Gaussian process with continuous paths. A similar argument yields the asymptotically linear representations  px|w−px|w=1 N N  i=1 1(Wi=w)1(Xi=x) −px|w qw+opN−1/2 and  qw−qw=1 N N  i=11(Wi=w) −qw The covariance kernel Σ1can be calculated as follows: Σ1(yxw˜ y ˜ x ˜ w)11 =E1(Xi=xWi=w)1(Xi=˜ xWi=˜ w)1(Yi≤y)−FY|XW (y |xw)1(Yi≤˜ y)−FY|XW (˜ y|˜ x ˜ w) P(X =xW =w)P(X =˜ xW =˜ w)  =FY|XW min{y ˜ y}|xw−FY|XW (y |xw)FY|XW (˜ y|xw) px|wqw 1(x =˜ xw =˜ w) Σ1(yxw˜ y ˜ x ˜ w)12 =E1(Xi=xWi=˜ w)1(Xi=˜ x) −p˜ x|˜ w1(Yi≤y)−FY|XW (y |xw) px|wqwq˜ w=0 Σ1(yxw˜ y ˜ x ˜ w)13 =E1(Wi=˜ w) −q˜ w1(Xi=xWi=w)1(Yi≤y)−FY|XW (y |x w) P(X =xW =w) =0 Σ1(yxw˜ y ˜ x ˜ w)21=Σ1(˜ y ˜ x ˜ wyxw)12=0 Quantitative Economics 11 (2020) Inference on breakdown frontiers 89 Σ1(yxw˜ y ˜ x ˜ w)22 =E1(Wi=w)1(Wi=˜ w)1(Xi=x) −px|w1(Xi=˜ x) −p˜ x|˜ w qwq˜ w =px|w qw 1(x =˜ xw =˜ w) −px|wp˜ x|w qw 1(w =˜ w) Σ1(yxw˜ y ˜ x ˜ w)23=E1(Wi=˜ w) −q˜ w1(Xi=x) −px|w qw=0 Σ1(yxw˜ y ˜ x ˜ w)31=Σ1(˜ y ˜ x ˜ wyxw)13=0 Σ1(yxw˜ y ˜ x ˜ w)32=Σ1(˜ y ˜ x ˜ wyxw)23=0 Σ1(yxw˜ y ˜ x ˜ w)33=E1(Wi=w) −qw1(Wi=˜ w) −q˜ w =qw1(w =˜ w) −qwq˜ w Lemma C2 (Chain rule for Hadamard directionally differentiable functions). Let D,E, and Fbe Banach spaces with norms ·D,·E,and ·F.Let Dφ⊆Dand Eψ⊆E.Let φ:Dφ→Eψand ψ:Eψ→Fbe functions.Let θ∈Dφand φbe Hadamard directionally differentiable at θtangentially to D0⊆D.Let ψbe Hadamard directionally differentiable at φ(θ) tangentially to the range φ θ(D0)⊆Eψ.Then ψ◦φ:Dφ→Fis Hadamard directionally differentiable at θtangentially to D0with Hadamard directional derivative equal to ψ φ(θ) ◦φ θ. This result is a version of proposition 3.6 in Shapiro (1990), who omits the proof. We give the proof here because this result is key to our paper. Proof of Lemma C2.Let{hn}n≥1be in Dand hn→h∈D0. By Hadamard directional differentiability of φtangentially to D0, )))) φ(θ +tnhn)−φ(θ) tn−φ θ(h)))))E=o(1) as n→∞for any tn0.Thatis, gn≡φ(θ +tnhn)−φ(θ) tn E →φ θ(h) =g where φ θ∈φ θ(D0). Therefore, by Hadamard directional differentiability of ψ,wehave ψφ(θ +tnhn)−ψφ(θ) tn=ψφ(θ) +tngn−ψφ(θ) tn F →ψ φ(θ)(g) =ψ φ(θ)φ θ(h) By Hadamard directional differentiability of φat θand ψat φ(θ),φ θand ψ φ(θ) are continuous mappings. Hence their composition ψ φ(θ) ◦φ θis continuous. This combined with our derivations above imply that ψ◦φis Hadamard directionally differentiable tangentially to D0at θ. 90 Masten and Poirier Quantitative Economics 11 (2020) Proof of Lemma 1.Letθ0=(FY|XW (·|··)p(·|·)q(·))and  θ=( FY|XW (·|··)  p(·|·)  q(·)).Forfixedyand c, define the mapping φ1:∞R×{01}×supp(W )×∞{01}×supp(W )×∞supp(W ) →∞{01}×supp(W )R2 by φ1(θ)(x w) =⎛ ⎜ ⎜ ⎜ ⎝ minθ(1)(yxw)θ(2)(xw) θ(2)(xw) −cθ(1)(yxw)θ(2)(x w) +c θ(2)(xw) +c maxθ(1)(yxw)θ(2)(xw) θ(2)(xw) +cθ(1)(yxw)θ(2)(x w) −c θ(2)(xw) −c⎞ ⎟ ⎟ ⎟ ⎠ where θ(j) is the jth component of θ. Note that %Fc Yx|W(y |w) Fc Yx|W(y |w)&=φ1(θ0)(xw) % Fc Yx|W(y |w)  Fc Yx|W(y |w)&=φ1( θ)(xw) The maps (a1a2)→min{a1a2}and (a1a2)→max{a1a2}are Hadamard directionally differentiable with Hadamard directional derivatives at (a1a2)equal to h→⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ h(1)if a1<a 2 minh(1)h(2)if a1=a2 h(2)if a1>a 2 and h→⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ h(2)if a1<a 2 maxh(1)h(2)if a1=a2 h(1)if a1>a 2 respectively, where h∈R2; for example, see equation (18) in Fang and Santos (2015). The mapping φ1is comprised of compositions of these min and max operators, along with four other functions. We can show that these four mappings are ordinary Hadamard differentiable. Here, we compute these Hadamard derivatives with respect to θ: δ1(θ)(x w) =θ(1)(yxw)θ(2)(xw) θ(2)(xw) +chas Hadamard derivative equal to δ 1θ(h)(xw) =θ(1)(yxw)h(2)(x w) +h(1)(yxw)θ(2)(xw) θ(2)(xw) +c −θ(1)(yxw)θ(2)(xw)h(2)(x w) θ(2)(xw) +c2 Quantitative Economics 11 (2020) Inference on breakdown frontiers 91 δ2(θ)(x w) =θ(1)(yxw)θ(2)(xw) −c θ(2)(xw) −chas Hadamard derivative equal to δ 2θ(h)(xw) =θ(1)(yxw)h(2)(x w) +h(1)(yxw)θ(2)(xw) θ(2)(xw) −c −θ(1)(yxw)θ(2)(xw) −ch(2)(x w) θ(2)(xw) −c2 δ3(θ)(x w) =θ(1)(yxw)θ(2)(xw) θ(2)(xw) −chas Hadamard derivative equal to δ 3θ(h)(xw) =θ(1)(yxw)h(2)(x w) +h(1)(yxw)θ(2)(xw) θ(2)(xw) −c −θ(1)(yxw)θ(2)(xw)h(2)(x w) θ(2)(xw) −c2 δ4(θ)(x w) =θ(1)(yxw)θ(2)(xw) +c θ(2)(xw) +chas Hadamard derivative equal to δ 4θ(h)(xw) =θ(1)(yxw)h(2)(x w) +h(1)(yxw)θ(2)(xw) θ(2)(xw) +c −θ(1)(yxw)θ(2)(xw) +ch(2)(x w) θ(2)(xw) +c2 All these derivatives are well defined at θ0because θ(2) 0(xw) =px|w> C ≥c.Withthis notation, we can write the functional φ1as φ1(θ) =%minδ3(θ) δ4(θ) maxδ1(θ) δ2(θ)& By the chain rule (Lemma C2), the map φ1is Hadamard directionally differentiable at θ0 with Hadamard directional derivative evaluated at θ0equal to φ 1θ0(h) = ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 1δ3(θ0)<δ 4(θ0)·δ 4θ0(h) +1δ3(θ0)=δ4(θ0)·minδ 3θ0(h)δ 4θ0(h) +1δ3(θ0)>δ 4(θ0)·δ 3θ0(h) 1δ1(θ0)<δ 2(θ0)·δ 1θ0(h) +1δ1(θ0)=δ2(θ0)·maxδ 1θ0(h)δ 2θ0(h) +1δ1(θ0)>δ 2(θ0)·δ 2θ0(h) ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠  By Lemma C1,√N( θ(yxw) −θ0(yxw))Z1(yxw). Hence we can use the delta method for Hadamard directionally differentiable functions (see Theorem 2.1 in Fang 92 Masten and Poirier Quantitative Economics 11 (2020) and Santos (2019)) to find that √Nφ1( θ) −φ1(θ0)(xw) φ 1θ0(Z1)(xw) ≡˜ Z2(xw) This result holds uniformly over any finite grid of values for y∈Rand c∈Cby considering the Hadamard directional differentiability of a vector of these mappings indexed at different values of yand c, which yields the process Z2(yxwc). Proof of Lemma 2.LetS={(yxw)∈R2:y∈[yx(w) yx(w)]x∈{01}w∈supp(W )}. Let D(S)⊂∞(S)denote the set of functions that are càdlàg in the first argument for each x∈{01}and w∈supp(W ). Define the mapping ˜ φ2:D(S)×∞{01}×supp(W )×∞supp(W )→∞(01)×{01}×supp(W )R2 by ˜ φ2(θ)(τxw)=%θ(1)−1(τxw) θ(2)(xw) & By Assumptions A1,A3,A5, and Lemma 21.4(ii) in van der Vaart (2000)thismapping is Hadamard differentiable at θ0tangentially to C(S)×∞({01}×supp(W )) × ∞(supp(W )),whereC(S)⊂∞(S)is the set functions that are continuous in the first argument for each x∈{01}and w∈supp(W ). Its Hadamard derivative at θ0=(FY|XW (·| ··)p(·|·)q(·))is ˜ φ 2θ0(h)(τxw)→−h(1)QY|XW (τ |xw)xw fY|XW QY|XW (τ |x w) |xwh(2)(x w) By the functional delta method and Theorem 7.3.3 part (iii) of Bickel and Doksum (2015), √N˜ φ2( θ) −˜ φ2(θ0)(τxw)˜ Z3(τxw) where ˜ Z3is a mean-zero Gaussian process in ∞((01)×{01}×supp(W )R2)with uniformly continuous paths. Now define the mapping φ2:∞(01)×{01}×supp(W )×∞{01}×supp(W ) →∞(01)×{01}×supp(W ) ×[0C]R2 by φ2(ψ)(τxwc)=⎛ ⎜ ⎜ ⎝ ψ(1)τ+c ψ(2)(xw) min{τ 1−τ}xw ψ(1)τ−c ψ(2)(xw) min{τ 1−τ}xw⎞ ⎟ ⎟ ⎠ Quantitative Economics 11 (2020) Inference on breakdown frontiers 93 Then %Qc Yx|W(τ |w) Qc Yx|W(τ |w)&=φ2˜ φ2(θ0)(τxwc) % Qc Yx|W(τ |w)  Qc Yx|W(τ |w)&=φ2˜ φ2( θ)(τxwc) We will show that φ2is Hadamard differentiable tangentially to the space CU((01)× {01}×supp(W ))×∞({01}×supp(W )),whereCU(A)denotes the set of uniformly continuous functions on A. The Hadamard derivative of the first component of φ2evaluated at ψ0≡˜ φ2(θ0)is φ(1) 2ψ0(h)(τxwc) =h(1)τ+c ψ(2) 0(xw) min{τ 1−τ}xw −ψ(1) 0τ+c ψ(2) 0(xw) min{τ 1−τ}xwcmin{τ 1−τ} ψ(2) 0(xw)2h(2)(x w) To see this, a Taylor expansion gives φ(1) 2(ψ0+tnhn)−φ(1) 2(ψ0) tn(τxwc) =h(1) nτ+c ψ(2) 0(xw) +tnh(2) n(xw) min{τ 1−τ}xw −ψ(1) 0τ+c ψ(2) 0(xw) +an(x w) min{τ 1−τ}xw ×cmin{τ 1−τ} ψ(2) 0(xw) +an(x w)2h(2) n(xw) using the fact that ψ(1) 0(τxw)=QY|X(τ |xw) is continuously differentiable in τby Assumption A5.2, and noting that term an(x w) satisfies |an(x w)|≤|tnh(2) n(xw)|=O(tn). Next, sup τxwch(1) nτ+cmin{τ 1−τ} ψ(2) 0(xw) +tnh(2) n(xw) xw −h(1)τ+cmin{τ 1−τ} ψ(2) 0(xw) xw ≤sup τxwch(1) nτ+cmin{τ 1−τ} ψ(2) 0(xw) +tnh(2) n(xw) xw 94 Masten and Poirier Quantitative Economics 11 (2020) −h(1)τ+cmin{τ 1−τ} ψ(2) 0(xw) +tnh(2) n(xw) xw +sup τxwch(1)τ+cmin{τ 1−τ} ψ(2) 0(xw) +tnh(2) n(xw) xw −h(1)τ+cmin{τ 1−τ} ψ(2) 0(xw) xw ≤))h(1) n−h(1)))∞+o(1) =o(1) where all three suprema are taken over τ∈(01),x∈{01},w∈supp(W ),c∈[0C].The last inequality follows from uniform continuity of h(1). The last line follows from uniform convergence of hnto h. Similarly, we have that sup τxwcψ(1) 0τ+c ψ(2) 0(xw) +an(x w) min{τ 1−τ}xw ×cmin{τ 1−τ} ψ(2) 0(xw) +an(x w)2h(2) n(xw) −ψ(1) 0τ+c ψ(2) 0(xw) min{τ 1−τ}xw ×cmin{τ 1−τ} ψ(2) 0(xw)2h(2)(x w)=o(1) by uniform continuity of ψ(1) 0(implied by Assumption A5.2)andbyan(xw) =o(1). Again, the sup is over τ∈(01),x∈{01},w∈supp(W ),c∈[0C]. Therefore φ(1) 2is Hadamard differentiable tangentially to the space of uniformly continuous functions. A similar argument can be made for φ(2) 2. By composition, φ2◦˜ φ2is Hadamard differentiable tangentially to C(S). By the functional delta method and the fact that ˜ Z3(yxw)has uniformly continuous paths, we have that √Nφ2˜ φ2( θ)−φ2˜ φ2(θ0)(τxwc)φ 2ψ0◦˜ φ 2θ0(Z1)(τxwc) ≡Z3(τxwc) a mean-zero Gaussian process with continuous paths in τ∈(01)and c∈[0C]. Proof of Proposition 1. Consider the lower CQTE bound of equation (8)asafunction of c. Its first component is the lower bound of the conditional quantile of Y1|W=w. Quantitative Economics 11 (2020) Inference on breakdown frontiers 95 By Assumption A5.2, the derivative of that conditional quantile with respect to cequals ∂ ∂c QY|XW τ−c p1|w min{τ 1−τ}1w =−min{τ 1−τ} p1|wfY|XW QY|XW τ−c p1|w min{τ 1−τ}1w1w The second component of the lower CQTE bound is the upper bound of the conditional quantile of Y0|W=w. The derivative of that conditional quantile with respect to cequals ∂ ∂c QY|XW τ+c p0|w min{τ 1−τ}0w =min{τ 1−τ} p0|wfY|XW QY|XW τ+c p0|w min{τ 1−τ}0w0w Moreover, these derivatives are bounded away from zero and infinity uniformly over c∈ (0C]. This implies that the derivative of the CQTE is negative and uniformly bounded away from zero. Next recall that CATE(c |w) =1 0 CQTE(τc |w)dτ Its derivative with respect to cexists by the dominated convergence theorem (by Assumptions A1 and A5). Moreover, it is bounded away from zero for all c∈(0C]. By taking another expectation over the marginal distribution of W,∂ATE(c)/∂c exists (by Assumption A4), is negative, and is bounded away from zero for all c∈(0C]. c∗is defined implicitly by ATE(c∗)=μ.WehaveshownthatthefunctionATE(c) satisfies the assumptions of Lemma 21.3 on page 306 of van der Vaart (2000). Thus the mapping ATE(·)→c∗is Hadamard differentiable tangentially to the set of càdlàg functions on (0C]with derivative −hc∗ ∂ ∂c ATEc∗ By the discussion following Lemma 2,√N( ATE(c) −ATE(c)) converges in distribution to a random element of ∞([0C])with continuous paths. Let ˜ c∗=infc∈[0 C]: ATE(c) ≤μ We can then apply the functional delta method to see that √N(˜ c∗−c∗)converges in distribution to a Gaussian variable we denote by Zbp. 96 Masten and Poirier Quantitative Economics 11 (2020) Since c∗∈(0 C]and by monotonicity of ATE(·),wehaveATE(C) ≤μ.By√N- convergence of the ATE bounds, P ATE(C) > μ=P−√NATE(C)−μ<√N ATE(C)−ATE(C) →0 Therefore, the set {c∈[0C]: ATE(c) ≤μ}is nonempty with probability approaching one. This implies that ˜ c∗∈[0 C]with probability approaching one and therefore P(˜ c∗=  c∗)also approaches one as N→∞. Using these results, we obtain √N c∗−c∗=√N c∗−˜ c∗+√N˜ c∗−c∗ =op(1)+√N˜ c∗−c∗ Zbp The following result extends Proposition 2(i) of Chernozhukov, Fernández-Val, and Galichon (2010) to allow for input functions which are directionally differentiable, but not fully differentiable, at one point. It can be extended to allow for multiple points of directional differentiability, but we omit this since we do not need it for our application. Lemma C3. Let θ0(ucw)=(θ(1) 0(ucw)θ(2) 0(ucw))where for j∈{12}we have that θ(j) 0(ucw)is bounded above and below,and differentiable everywhere except at u=u∗, where it is directionally differentiable.Further,assume that the two components satisfy Assumption A6.Then,for fixed z∈R,the mapping φ3:∞((01)×supp(W ) ×CR2)→ ∞(supp(W ) ×CR2)defined by φ3(θ)(wc) =⎛ ⎜ ⎜ ⎝1 0 1θ(2)(ucw)≤zdu 1 0 1θ(1)(ucw)≤zdu ⎞ ⎟ ⎟ ⎠ is Hadamard directionally differentiable tangentially to C((01)×supp(W ) ×CR2)with Hadamard directional derivative given by equations (30)and (31)below. Proof of Lemma C3. For clarity we suppress the dependence on win the expressions below. Uniformity of convergence over w∈supp(W ) follows from the discreteness of supp(W ) (Assumption A4). Our proof follows that of Proposition 2(i) in Chernozhukov, Fernández-Val, and Galichon (2010). Let U1(c) =u∈(01):θ(1) 0(uc) =z denote the set of roots to the equation θ(1) 0(uc) =zfor fixed zand c. By Assumption A6.1, this set contains a finite number of elements. We denote these by U1(c) =u(1) k(c) for k=12K(1)(c) < ∞ Quantitative Economics 11 (2020) Inference on breakdown frontiers 103 + K  k=1 CDTE(zct|wk)√N( qwk−qwk)  K  k=1 ˜ Z5(ctwk)qwk+ K  k=1 CDTE(zct|wk)Z(3) 1(00wk) A similar derivation holds for the upper bound estimator. Proof of Theorem 2. By Lemmas 3and C1, the numerator of equation (22)converges uniformly over c∈C. By Lemmas 3and 4, the denominator also converges uniformly over c∈C. By the delta method, √N bf(c p)−bf(cp) =√N%1−p− K  k=1 P(c |wk) qwk 1+ K  k=1mininf y∈Y0(wk) Fc Y1|W(y |wk)− Fc Y0|W(y |wk)0− P(c |wk) qwk − 1−p− K  k=1 P(c |wk)qwk 1+ K  k=1mininf y∈Y0(wk)Fc Y1|W(y |wk)−Fc Y0|W(y |wk)0−P(c |wk)qwk &  − K  k=1 Z(1) 4(wkc)qwk− K  k=1 P(c |wk)Z(3) 1(00wk) 1+ K  k=1mininf y∈Y0(wk)Fc Y1|W(y |wk)−Fc Y0|W(y |wk)0−P(c |wk)qwk − 1−p− K  k=1 P(c |wk)qwk %1+ K  k=1mininf y∈Y0(wk)Fc Y1|W(y |wk)−Fc Y0|W(y |wk)0−P(c |wk)qwk&2˜ Z(c) where √N%1−p− K  k=1 P(c |wk) qwk−'1−p− K  k=1 P(c |wk)qwk(& − K  k=1 Z(1) 4(wkc)qwk− K  k=1 P(c |wk)Z(3) 1(00wk) 104 Masten and Poirier Quantitative Economics 11 (2020) and √N%K  k=1mininf y∈Y0(wk) Fc Y1|W(y |wk)− Fc Y0|W(y |wk)0− P(c |wk) qwk − K  k=1mininf y∈Y0(wk)Fc Y1|W(y |wk)−Fc Y0|W(y |wk)0−P(c |wk)qwk&˜ Z(c) Here ˜ Z(c) is a random element of ∞(C)by Lemmas 3and 4. Therefore, √N bf(c p)−bf(cp) converges to a random element in ∞(C×P). As discussed in the proof of Lemma 1, the maximum and minimum operators in equation (21) are Hadamard directionally differentiable. By Lemma C2, their composition is Hadamard directionally differentiable. Therefore, by the delta method for Hadamard directionally differentiable functions, √N( BF(cp)−BF(c p)) converges in process as in the statement of the theorem. Lemma C5. Let h:A→Rwhere A⊆R.Let F(h)=supx∈Ah(x).Let ·∞denote the supnorm h∞=supx∈A|h(x)|.Then Fis Lipschitz continuous with respect to the sup-norm ·∞and has Lipschitz constant equal to one. Proof of Lemma C5. For functions hand h, sup x∈A h(x) −sup ˜ x∈A h(˜ x) =sup x∈Ah(x) −sup ˜ x∈A h(˜ x) ≤sup x∈Ah(x) −h(x) ≤sup x∈Ah(x) −h(x) By a symmetric argument, sup x∈A h(x) −sup ˜ x∈A h( ˜ x) ≤sup x∈Ah(x) −h(x) =sup x∈Ah(x) −h(x) Therefore |F(h)−F(h)|≤h−h∞. Proof of Proposition 2. Hadamard directional differentiability of φfollows from the chain rule (Lemma C2) and from the proof of Theorem 2, since the breakdown frontier is a Hadamard directionally differentiable mapping of P(·)=E[P(·|W)]and DTE(z ··), which are themselves Hadamard directionally differentiable mappings of θ0. Lemma C1 combined with Theorem 3.6.1 of van der Vaart and Wellner (1996)implies consistency of the nonparametric bootstrap for our underlying parameters: Z∗ N= Quantitative Economics 11 (2020) Inference on breakdown frontiers 105 √N( θ∗− θ) P Z1.Bythisresult,εN→0,√NεN→∞, and Theorem 3.1 in Hong and Li (2018), equation (25) holds. By 1/σ(c) being uniformly bounded, we have that  φ θ0√N( θ∗− θ)(c p) σ(c) P  Zbf(c p) σ(c)  By Lemma C5,thesup operator is Lipschitz with Lipschitz constant equal to 1. Therefore, by Proposition 10.7 on page 189 of Kosorok (2008), we can apply a continuous mapping theorem to get sup c∈C φ θ0√N( θ∗− θ)(c p) σ(c) P sup c∈C Zbf(c p) σ(c)  The rest of the proof follows from Corollary 3.2 of Fang and Santos (2015). Lemma C6. Let C>0.Let C={c1cJ}⊆[0C]be a finite grid of points.Let f:[0C]→ R+be a nonincreasing function.Let  LB(·)be an asymptotically exact uniform lower 1−α confidence band for fon the grid points: lim N→∞ P LB(cj)≤f(cj)for j=1J=1−α Define LB :[0C]→R+by LB(c) = ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩  LB(c1)if c∈[0c1]     LB(cj)if c∈(cj−1cj],for j=2J    0if c∈(cJC] Then LB(·)is an asymptotically exact uniform lower 1−αconfidence band on [0C]: lim N→∞ P LB(c) ≤f(c)for all c∈[0C]=1−α Proof of Lemma C6. Define the events A= LB(cj)≤f(c)for all c∈(cj−1cj],forj=1J and B= LB(cj)≤f(cj)for j=1J Aimmediately implies B. Since fis nonincreasing, Bimplies A.Thus P LB(c) ≤f(c)for all c∈[0C]=P LB(cj)≤f(c)for all c∈(cj−1cj]for j=1J =P LB(cj)≤f(cj)for j=1J 106 Masten and Poirier Quantitative Economics 11 (2020) The first line follows by definition of LB and since fis nonnegative. Taking limits as N→∞gives lim N→∞ P LB(c) ≤f(c)for all c∈[0 C]=lim N→∞ P LB(cj)≤f(cj)for j=1J =1−α where the last equality follows from the validity of the band  LB(·)on C. Proof of Corollary 1. This follows immediately from Proposition 2and Lemma C6. References Abbring, J. H. and J. J. Heckman (2007), “Econometric evaluation of social programs, part III: Distributional treatment effects, dynamic treatment effects, dynamic discrete choice, and general equilibrium policy evaluation.” Handbook of Econometrics, 6, 5145– 5303. [78] Altonji, J. G., T. E. Elder, and C. R. Taber (2005), “Selection on observed and unobserved variables: Assessing the effectiveness of Catholic schools.” Journal of Political Economy, 113, 151–184. [69] Altonji, J. G., T. E. Elder, and C. R. Taber (2008), “Using selection on observed variables to assess bias from unobservables when evaluating Swan–Ganz catheterization.” American Economic Review P&P, 98, 345–350. [69] Angrist, J. D. (1990), “Lifetime earnings and the Vietnam era draft lottery: Evidence from social security administrative records.” The American Economic Review, 313–336. [65] Annan, J., C. Blattman, and R. Horton (2006), “The State of Youth and Youth Protection in Northern Uganda.” Final Report for UNICEF Uganda. [64] Bedoya, G., L. Bittarello, J. Davis, and N. Mittag (2017), “Distributional impact analysis: Toolkit and illustrations of impacts beyond the average treatment effect.” Policy Research Working Paper No. 8139, World Bank. [65] Bickel, P. J. and K. A. Doksum (2015), Mathematical Statistics: Basic Ideas and Selected Topics,Vol.2.CRCPress.[92] Blattman, C. and J. Annan (2010), “The consequences of child soldiering.” The Review of Economics and Statistics, 92, 882–898. [64,65,66,67] Brent, R. P. (1973), Algorithms for Minimization Without Derivatives. Prentice-Hall. [68] Canay, I. A. and A. M. Shaikh (2017), “Practical and theoretical advances in inference for partially identified models.” In Advances in Economics and Econometrics: Eleventh World Congress, Vol. 2 (B. Honoré, A. Pakes, M. Piazzesi, and L. Samuelson, eds.), 271– 306, Cambridge University Press. [54] Quantitative Economics 11 (2020) Inference on breakdown frontiers 107 Cao, R., A. Cuevas, and W. G. Manteiga (1994), “A comparative study of several smoothing methods in density estimation.” Computational Statistics & Data Analysis, 17, 153–176. [83] Card, D. and A. R. Cardoso (2012), “Can compulsory military service raise civilian wages? Evidence from the peacetime draft in Portugal.” American Economic Journal: Applied Economics, 4, 57–93. [65] Chernozhukov, V., I. Fernández-Val, and A. Galichon (2010), “Quantile and probability curves without crossing.” Econometrica, 78, 1093–1125. [70,80,81,96,97,99] Chesher, A. and A. Rosen (2017), “Generalized instrumental variable models.” Econometrica, 85, 959–989. [43,75] Cornfield, J., W. Haenszel, E. C. Hammond, A. M. Lilienfeld, M. B. Shimkin, and E. L. Wynder (1959), “Smoking and lung cancer: Recent evidence and a discussion of some questions.” Journal of the National Cancer Institute, 22, 173–203. [76] De Angelis, D. and G. A. Young (1992), “Smoothing the bootstrap.” International Statistical Review, 45–56. [84] Donoho, D. L. and P. J. Huber (1983), “The notion of breakdown point.” A Festschrift for Erich L. Lehmann, 157–184. [42] Dümbgen, L. (1993), “On nondifferentiable functions and the bootstrap.” Probability Theory and Related Fields, 95, 125–140. [55,61,78,83] Fan, Y., E. Guerre, and D. Zhu (2017), “Partial identification of functionals of the joint distribution of “potential outcomes”.” Journal of Econometrics, 197, 42–59. [51,78] Fan, Y. and S. S. Park (2009), “Partial identification of the distribution of treatment effects and its confidence sets.” In Nonparametric Econometric Methods: Advances in Econometrics, Vol. 25, 3–70, Emerald Group Publishing Limited. [78] Fan, Y. and S. S. Park (2010), “Sharp bounds on the distribution of treatment effects and their statistical inference.” Econometric Theory, 26, 931–951. [51] Fan, Y. and A. J. Patton (2014), “Copulas in econometrics.” Annual Review of Economics, 6, 179–200. [47] Fang, Z. and A. Santos (2015), “Inference on directionally differentiable functions.” Working paper. [63,90,100,105] Fang, Z. and A. Santos (2019), “Inference on directionally differentiable functions.” The Review of Economic Studies, 86, 377–412. [55,60,78,84,91,92,100,101] Freyberger, J. and Y. Rai (2018), “Uniform confidence bands: Characterization and optimality.” Journal of Econometrics, 204, 119–130. [62,63] Gechter, M. (2016), “Generalizing the results from social experiments: Theory and evidence from Mexico and India.” Working paper. [78] 108 Masten and Poirier Quantitative Economics 11 (2020) Gundersen, C., B. Kreider, and J. Pepper (2012), “The impact of the National School Lunch Program on child health: A nonparametric bounds analysis.” Journal of Econometrics, 166, 79–91. [78] Hampel, F. R. (1968), Contributions to the Theory of Robust Estimation. PhD Dissertation, University of California, Berkeley. [42] Hampel, F. R. (1971), “A general qualitative definition of robustness.” The Annals of Mathematical Statistics, 1887–1896. [42] Hampel, F. R., E. M. Ronchetti, P. J. Rousseeuw, and W. A. Stahel (1986), Robust Statistics: The Approach Based on Influence Functions. John Wiley & Sons. [76] Hansen, B. E. (2004), “Bandwidth selection for nonparametric distribution estimation.” Working paper. [67,71,84] Hansen, B. E. (2017), “Regression kink with an unknown threshold.” Journal of Business & Economic Statistics, 35, 228–240. [78] Heckman, J. J., J. Smith, and N. Clements (1997), “Making the most out of programme evaluations and social experiments: Accounting for heterogeneity in programme impacts.” The Review of Economic Studies, 64, 487–535. [65,75,78] Ho, K. and A. M. Rosen (2017), “Partial identification in applied research: Benefits and challenges.” In Advances in Economics and Econometrics: Eleventh World Congress (B. Honoré, A. Pakes, M. Piazzesi, and L. Samuelson, eds.), Econometric Society Monographs, Vol. 2, 307–359, Cambridge University Press. [42,77] Hong, H. and J. Li (2018), “The numerical delta method.” Journal of Econometrics, 206, 379–394. [61,78,83,84,105] Horowitz, J. L. and S. Lee (2012), “Uniform confidence bands for functions estimated nonparametrically with instrumental variables.” Journal of Econometrics, 168, 175–188. [83] Horowitz, J. L. and S. Lee (2017), “Nonparametric estimation and inference under shape restrictions.” Journal of Econometrics, 201, 108–126. [83] Horowitz, J. L. and C. F. Manski (1995), “Identification and robustness with contaminated and corrupted data.” Econometrica, 63, 281–302. [42,48,76,77,78] Huber, P. J. (1964), “Robust estimation of a location parameter.” The Annals of Mathematical Statistics, 35, 73–101. [42] Huber, P. J. and E. M. Ronchetti (2009), Robust Statistics. John Wiley & Sons. [76] Imbens, G. W. (2003), “Sensitivity to exogeneity assumptions in program evaluation.” American Economic Review P&P, 93, 126–132. [43,65,69,77] Imbens, G. W. and D. B. Rubin (2015), Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press. [76] Quantitative Economics 11 (2020) Inference on breakdown frontiers 109 Kaido, H. (2016), “A dual approach to inference for partially identified econometric models.” Journal of Econometrics, 192, 269–290. [78] Kline, P. and A. Santos (2013), “Sensitivity to missing data assumptions: Theory and an evaluation of the U.S. wage structure.” Quantitative Economics, 4, 231–267. [42,43,62, 77] Kosorok, M. R. (2008), Introduction to Empirical Processes and Semiparametric Inference. Springer Science & Business Media. [105] Kreider, B., J. V. Pepper, C. Gundersen, and D. Jolliffe (2012), “Identifying the effects of SNAP (food stamps) on child health outcomes when participation is endogenous and misreported.” Journal of the American Statistical Association, 107, 958–975. [78] Kreider, B., J. V. Pepper, and M. Roy (2016), “Identifying the effects of WIC on food insecurity among infants and children.” Southern Economic Journal, 82, 1106–1122. [78] Lee, Y.-Y. and D. Bhattacharya (2019), “Applied welfare analysis for discrete choice with interval-data on income.” Journal of Econometrics, 211, 361–387. [78] Léger, C. and J. P. Romano (1990), “Bootstrap choice of tuning parameters.” Annals of the Institute of Statistical Mathematics, 42, 709–735. [83] Li, Q. and J. S. Racine (2008), “Nonparametric estimation of conditional CDF and quantile functions with mixed categorical and continuous data.” Journal of Business & Economic Statistics, 26, 423–434. [66] Makarov, G. (1982), “Estimates for the distribution function of a sum of two random variables when the marginal distributions are fixed.” Theory of Probability & Its Applications, 26, 803–806. [51,85] Manski, C. F. (2007), Identification for Prediction and Decision. Harvard University Press. [42,77] Manski, C. F. (2013), “Response to the review of ‘Public policy in an uncertain world’.” The Economic Journal, 123, F412–F415. [42] Manski, C. F. and J. V. Pepper (2018), “How do right-to-carry laws affect crime rates? Coping with ambiguity using bounded-variation assumptions.” The Review of Economics and Statistics, 100, 232–244. [43,77] Marron, J. S. (1992), “Bootstrap bandwidth selection.” In Exploring the Limits of Bootstrap, 249–262, John Wiley. [83] Masten, M. A, and A. Poirier (2020), “Supplement to ‘Inference on breakdown frontiers’.” Quantitative Economics Supplemental Material, 11, https://doi.org/10.3982/QE1288. [45] Masten, M. A. and A. Poirier (2016), “Partial independence in nonseparable models.” cemmap Working Paper, CWP26/16. [46,76] Masten, M. A. and A. Poirier (2018a), “Identification of treatment effects under conditional partial independence.” Econometrica, 86, 317–351. [43,47,49,56,67,68,78,86] 110 Masten and Poirier Quantitative Economics 11 (2020) Masten, M. A. and A. Poirier (2018b), “Interpreting quantile independence.” Working paper. [76] Matzkin, R. L. (2003), “Nonparametric estimation of nonadditive random functions.” Econometrica, 71, 1339–1375. [58] Mullahy, J. (2018), “Individual results may vary: Inequality-probability bounds for some health-outcome treatment effects.” Journal of Health Economics, 61, 151–162. [65] Nelsen, R. B. (2006), An Introduction to Copulas, second edition. Springer. [47,48] Oster, E. (2019), “Unobservable selection and coefficient stability: Theory and evidence.” Journal of Business & Economic Statistics, 37, 187–204. [69] Polansky, A. M. and W. Schucany (1997), “Kernel smoothing to improve bootstrap confidence intervals.” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 59, 821–838. [84] Robins, J. M. (1999), “Association, causation, and marginal structural models.” Synthese, 121, 151–179. [69] Robins, J. M., A. Rotnitzky, and D. O. Scharfstein (2000), “Sensitivity analysis for selection bias and unmeasured confounding in missing data and causal inference models.” In Statistical Models in Epidemiology, the Environment, and Clinical Trials, 1–94, Springer. [78] Rosenbaum, P. R. (1995), Observational Studies. Springer. [78] Rosenbaum, P. R. (2002), Observational Studies, second edition. Springer. [78] Rosenbaum, P. R. and D. B. Rubin (1983), “Assessing sensitivity to an unobserved binary covariate in an observational study with binary outcome.” JournaloftheRoyalStatistical Society, Series B, 212–218. [76,78] Rotnitzky, A., J. M. Robins, and D. O. Scharfstein (1998), “Semiparametric regression for repeated outcomes with nonignorable nonresponse.” Journal of the American Statistical Association, 93, 1321–1339. [69] Shapiro, A. (1990), “On concepts of directional differentiability.” Journal of Optimization Theory and Applications, 66, 477–487. [55,89] Sklar, M. (1959), “Fonctions de répartition à n dimensions et leurs marges.” Publications de l’Institut de statistique de l’Université de Paris, 8, 229–231. [47] Stoye, J. (2005), Essays on Partial Identification and Statistical Decisions. PhD Dissertation, Northwestern University. [42,52,77] Stoye, J. (2010), “Partial identification of spread parameters.” Quantitative Economics,1, 323–357. [42,77] Taylor, C. C. (1989), “Bootstrap choice of the smoothing parameter in kernel density estimation.” Biometrika, 76, 705–712. [83] Quantitative Economics 11 (2020) Inference on breakdown frontiers 111 Torgovitsky, A. (2019), “Partial identification by extending subdistributions.” Quantitative Economics, 10, 105–144. [43,75] van der Vaart, A. and J. Wellner (1996), Weak Convergence and Empirical Processes: With Applications to Statistics. Springer Science & Business Media. [60,79,104] van der Vaart, A. W. (2000), Asymptotic Statistics. Cambridge University Press. [56,87,92, 95] Co-editor Christopher Taber handled this manuscript. Manuscript received 2 February, 2019; final version accepted 8 August, 2019; available online 28 August, 2019.