scieee AI-readable full text Open interactive document viewer

Carleson measures, trees, extrapolation, and T(b) theorems

Auscher, P.; Hofmann, Steve; Muscalu, C.; Tao, Terence; Thiele, C.

Abstract

The theory of Carleson measures, stopping time arguments, and atomic decompositions has been well-established in harmonic analysis. More recent is the theory of phase space analysis from the point of view of wave packets on tiles, tree selection algorithms, and tree size estimates. The purpose of this paper is to demonstrate that the two theories are in fact closely related, by taking existing results and reproving them in a unified setting. In particular we give a dyadic version of extrapolation for Carleson measures, as well as a two-sided local dyadic T(b) theorem which generalizes earlier T(b) theorems of David, Journé, Semmes, and Christ.

Full text

Publ. Mat. 46 (2002), 257–325 CARLESON MEASURES, TREES, EXTRAPOLATION, AND T(b) THEOREMS P. Auscher, S. Hofmann, C. Muscalu, T. Tao and C. Thiele Abstract The theory of Carleson measures, stopping time arguments, and atomic decompositions has been well-established in harmonic analysis. More recent is the theory of phase space analysis from the point of view of wave packets on tiles, tree selection algorithms, and tree size estimates. The purpose of this paper is to demonstrate that the two theories are in fact closely related, by taking existing results and reproving them in a unified setting. In particular we give a dyadic version of extrapolation for Carleson measures, as well as a two-sided local dyadic T(b) theorem which generalizes earlier T(b) theorems of David, Journ´e, Semmes, and Christ. Contents 1. Introduction 258 2. Notation 261 2.1. Tiles and trees 261 2.2. Size and Carleson measures 266 2.3. Wavelets, phase space projections, and BMO 266 2.4. Mean 268 3. The “L∞theory” 269 3.1. Chopping big trees into little trees, and extrapolation of Carleson measures 273 3.2. An alternate argument 281 4. The “Lptheory” 284 4.1. Example: Atomic decomposition of dyadic Hp287 5. The Carleson embedding theorem and paraproducts 290 5.1. Weak-type estimates 294 6. Calder´on-Zygmund operators, and T(b) theorems 298 6.1. Accretivity and one-sided T(b) theorems 301 6.2. Adapted Haar bases, and two-sided T(b) theorems 304 References 320 2000 Mathematics Subject Classification. 42B20, 42B25. Key words. Haar wavelets, BMO, accretivity, T(b) theorems, extrapolation, Calder´on-Zygmund theory. 258 Auscher, Hofmann, Muscalu, Tao, Thiele 1. Introduction The purpose of this article is to demonstrate the close connection between two sets of techniques in harmonic analysis: the theory of Carleson measures and related objects, and the theory of trees and related objects. A Carleson measure is a positive measure µon the upper half space such that µ(I×(0,(I))) |I|for every cube I⊆Rnwith side length (I). There is also an analogous notion for domains more general than the half-space, as well as a discrete version: if µis a mapping from dyadic cubes into the non-negative reals, then µsatisfies a (discrete) Carleson measure condition I⊆JµI|J|for every dyadic cube J, where the sum runs over all dyadic sub-cubes of J. Carleson measures are intimately connected with many aspects of harmonic analysis, including non-tangential behavior of functions in the half-space (or in a domain) (see e.g. [51]), Hptheory and BMO [27], boundedness of singular integrals, square functions and maximal functions (e.g. [12], [22], [35], [50], [37], [38]), geometric measure theory (e.g. [24], [35], [36]), and PDE (e.g. [3], [28], [32], [42]). Moreover, via their connection with the theory of trees, Carleson measures have played a significant role in recent work on Bilinear Singular Integrals [39], [40], [45], [46], [53], and (rather appropriately!) Carleson’s theorem on a.e. convergence of Fourier Series [25], [41]. In these latter connections it is more convenient to work in the phase plane than in the Carleson half-space, and we have deliberately chosen our notation to reflect this fact. This article is mainly expository. Apart from one main new result (a local T(b) theorem), we shall mostly take existing results (atomic decompositions, paraproduct estimates, Carleson embedding) and reprove them in a framework which unifies both the Carleson measure theory and the theory of trees and tiles. (As such there is some overlap with the recent lecture notes in [49].) Since this is an expository article, we shall simplify matters and only work in one dimension R. Also, we shall mostly work in the dyadic setting instead of the continuous one, to avoid issues such as rapidly decreasing tails or use of the Vitali covering lemma. Thus, our results will be phrased using dyadic intervals and the Haar basis instead of arbitrary intervals and Gaussians (or similar smooth kernels). However most of our results have continuous analogues (see e.g. [29] for a comparison between dyadic and continuous harmonic analysis). We also will truncate all our spaces to be finite-dimensional to avoid technicalities. Trees, Extrapolation, T(b)259 The paper is organized as follows. After setting up the notation of dyadic Carleson measures and BMO, Haar wavelets, and tiles and trees, we will give a quick review of the standard “L∞” theory of BMO (i.e. measuring the ways in which BMO is close to L∞), but from the perspective of trees and tiles. As part of this L∞theory, we give a trees-based proof of the (dyadic analogue of the) extrapolation lemma for Carleson measures developed recently in [3], [32], [42]. We also give an alternate proof of the extrapolation lemma due to John B. Garnett. We then show how BMO is also useful in “Lp” contexts, mainly through a BMO version of the Calder´on-Zygmund decomposition. This type of lemma is used often in the recent work on Carleson’s theorem and the bilinear Hilbert transform, and is implicit in earlier work on Carleson measures and similar objects; we illustrate this by using the BMO Chebyshev inequality to re-prove the standard atomic decomposition of Hp. Next, we prove the Carleson embedding theorem and give its usual applications to paraproduct estimates and the T(1) theorem. We also give a short proof of the boundedness of paraproducts below L1;the proof is more direct than earlier proofs in that one does not go explicitly through the T(1) theorem. Finally, we consider Calder´on-Zygmund operators. We prove a twosided local T(b) theorem which generalizes the existing local and global T(b) theorems [23], [4], [11], [50]; for instance, we can prove the standard global T(b) theorem assuming that bis only in BMO rather than L∞. The T(b) Theorem, in its various guises, has its roots in a question posed by Yves Meyer, who asked whether the T(1) Theorem of David and Journ´e[22] (see also Section 6 below) remains true if the constant function 1 is replaced by some function b∈L∞with Re b≥δ(such bare said to be “accretive”). The question was motivated by its applicability to the L2boundedness of the Cauchy integral operator on a Lipschitz graph. Indeed, if Γ denotes the graph, in the plane, of a real-valued Lipschitz function A, then by Cauchy’s theorem, we have that in the sense of BMO (that is, modulo constants), 0 = p.v. Γ 1 z−wdw, for z∈Γ. But in graph co-ordinates, this amounts to saying that (again in the sense of BMO), 0 = p.v. ∞ −∞ 1+iA(y) x−y+i(A(x)−A(y)) dy ≡T(b)(x), 260 Auscher, Hofmann, Muscalu, Tao, Thiele where bis the accretive function 1 + iA, and T is the singular integral operator naturally associated to the antisymmetric Calder´on-Zygmund kernel K(x, y)=(x−y+i(A(x)−A(y)))−1. The L2boundedness of T, and hence also that of the Cauchy integral operator CΓf(x)≡p.v. Γ f(w) z−wdw, thus follows from an analogue of the T(1) Theorem in which the condition T(1), T∗(1) ∈BMO is replaced by the condition T(b)=0=T ∗(b), for some accretive function b. Just such a result was proved by McIntosh and Meyer [43], who consequently obtained an alternative proof of their earlier joint result with Coifman [13] concerning the L2boundedness of the operator CΓ. The “T(b) Theorem” of [43] was generalized by David, Journ´e and Semmes [23] to allow T(b),T∗(b)∈BMO (indeed, they allowed other generalizations as well, for example that there could be two different accretive functions b1,b2such that T(b1),T∗(b2)∈BMO, and moreover that the pointwise accretivity condition could be relaxed to a condition holding on various sorts of averages —see, e.g., the notion of “pseudoaccretivity” defined in Section 6.1 below). This led to a proof of the T(b) Theorem by constructing Haar wavelets adapted to the function b[20] (we shall base our proof on a variation of these adapted Haar wavelets). A very simple proof of a “one-sided version” of the T(b) Theorem was obtained by Semmes [50], who observed that in the special case T(b)∈ BMO, T∗(1) = 0, one can readily show that T(1) ∈BMO, thus reducing matters to the T(1) Theorem. It is worth noting that a suitable adaptation of Semmes’s argument is applicable to the solution of the square root problem of Kato. Indeed, one of the present authors (Auscher), along with Tchamitchian [5], formulated a version of the T(b) Theorem whose proof was based upon the argument of [50], and which was subsequently used to solve the Kato problem in higher dimensions [1], [31], [2]. We further note that there are local versions of the T(b) Theorem, due to M. Christ [11] (cf. Theorem 6.8 below), which also have interesting applications, namely to questions of analytic capacity; see for instance [54] for further discussion. This work was conducted at University of Missouri, University of California at Los Angeles (UCLA), and the Centre for Mathematics and its Applications (CMA) at the Australian National University (ANU). The Trees, Extrapolation, T(b)261 authors are particularly grateful to CMA for their warm hospitality during the visit of three of us (P. Auscher, C. Thiele, T. Tao). S. Hofmann is supported by NSF grant DMS 0088920. C. Muscalu is supported by NSF grant DMS 0100796. T. Tao is a Clay Prize Fellow and is supported by a grant from the Packard Foundation. C. Thiele is supported by a Sloan fellowship and NSF grants DMS 9985572 and DMS 9970469. The authors are indebted to Stephanie Molnar, John B. Garnett, Joan Verdera and the referee for many helpful corrections and comments. 2. Notation We use ABto denote the estimate A≤CB for some absolute constant Cwhich may vary from line to line. If Eis a set, we use |E|to denote the Lebesgue measure of E.We will always be ignoring sets of measure zero, thus we only consider two sets E,Fto be intersecting if |E∩F|>0. Although our functions may be complex valued, we shall use the real inner product f,g:= f(x)g(x)dx throughout. 2.1. Tiles and trees. We shall be working with dyadic intervals throughout the paper. The number of dyadic intervals is infinite, but to simplify the arguments we shall restrict ourselves to a finite set on the half-line; in applications, this restriction can always be removed by a standard translation and limiting argument. Specifically, we fix a large integer M>0; none of our estimates will depend on M. We define dyadic interval to be an interval1of the form I=[j2k,(j+ 1)2k], where j,kare integers such that −M≤k≤Mand I⊆[0,2M]. Let Idenote the set of all dyadic intervals; observe that Iis finite. All sums and unions involving Ior Jwill be assumed to be over Iunless otherwise specified. If fis a function on R, we define [f]I:= 1 |I|Ifto denote the mean of fon I. We use 2Ito denote the parent2of I, and Il,Irto denote the left and right children of I(these are undefined if |I|=2 Mor |I|=2 −Mrespectively). We refer to the intervals Iland Iras siblings. 1We will be careless about whether our intervals are closed, half-open, or open because of our convention of ignoring sets of measure zero. 2On the non-dyadic theory 2Iis often used to denote the interval with the same center as Ibut twice the length; this can be thought of as a non-dyadic version of the parent of I. However, in this paper we use 2Ito exclusively refer to the dyadic parent of I, i.e. the unique dyadic interval of twice the length which contains I. 262 Auscher, Hofmann, Muscalu, Tao, Thiele Since our dyadic intervals have been restricted to a finite set, all norms will automatically be finite and all stopping time processes will automatically terminate. This allows us to avoid some minor technicalities in our arguments, although it also means that we occasionally have to treat the smallest scale |I|=2 −Mor the largest scale |I|=2 Ma little differently from all the other scales. A major advantage of the dyadic setting is the nesting property:if I,Jare dyadic intervals which intersect each other, then either I⊆Jor J⊆I. In particular, for any collection of dyadic intervals, the maximal intervals in this collection will always be disjoint. The theory of Carleson measures is usually set in the upper half-space R2 +:= {(x, t):x∈R,t∈R+}. Actually, because of our truncation parameter Mwe will work in the compact subset (R2 +)M:= {(x, t):x∈[0,2M],t∈[2−M,2M−1]}. The variable xrepresents spatial position, while the tvariable represents time, wavelength, or spatial scale. For every dyadic interval I∈I, we let l(I)=|I|denote the side-length of I, and define the Carleson box Q(I)⊂(R2 +)Mby Q(I):=I×[2−M,l(I)] and the Whitney box Q+(I)⊂Q(I)by Q+(I):=I×l(I) 2,l(I). We remark that we have the partition Q(I)=J:J⊆IQ+(J). Meanwhile, the theory of trees and tiles is usually set in phase space R2:= {(x, ξ):x∈R,ξ∈R}. Because of our truncation, and because we are in the dyadic setting, we will instead work in the region3 (R2)M:= {(x, ξ):x∈[0,2M],ξ∈[0,2M]}. The variable xrepresents spatial position, while ξrepresents frequency. AHeisenberg tile or simply tile is a rectangle in R2of the form P:= IP× ωP, where IPand ωPare dyadic intervals such that |P|=|IP||ωP|=1. If Pand Qare tiles, we say that P≤Qif Pintersects Qand IP⊆IQ. This is a partial order on tiles. 3In truth, we are working not with the Euclidean field R, but with the Walsh field R+≡(Z2)Z. See e.g. [52]. Trees, Extrapolation, T(b)263 t 10 x ξ x 10 0 2 0 1 1 Figure 1. The geometry of the Carleson half-plane (partitioned into Whitney boxes) and phase space (partitioned into non-lacunary tiles). The heuristic t=1/ξ provides a one-to-one correspondence between the two partitions. If Iis a dyadic interval, we define the lacunary tile P+(I)by P+(I):=I×1 l(I),2 l(I) and the non-lacunary tile P0(I)by P0(I):=I×0,1 l(I). Let P+denote the set of all lacunary tiles, and P0the set of all nonlacunary tiles. We define P+(I)≤P+(J) if and only if P0(I)≤P0(J), or equivalently if I⊆J.Thus≤is a partial ordering on P+. Of course, there are many tiles which are not in either of these two sets, and many results in this paper can be extended to general tiles. However, for simplicity we shall mostly restrict ourselves to the lacunary and nonlacunary tiles. We write [f]Pas shorthand for the averages [f]IP. If P+(I) is a lacunary tile, we define the parent 2P+(I)ofP+(I)by 2P+(I):=P+(2I). Similarly define 2P0(I):=P0(2I). 264 Auscher, Hofmann, Muscalu, Tao, Thiele Alacunary4tree (henceforth abbreviated as tree) is a collection T⊆ P+of lacunary tiles with a top tile PT∈T, such that P≤PTfor all P∈T. We use ITas short-hand for IPT.IfP∈P+, we define the complete tree Tree(P) to be the tree Tree(P):={Q∈P+:Q≤P} with top P. We sometimes write Tree(I) for Tree(P+(I)). Note that every tree Tlies inside a complete tree Tree(PT). If Tis a tree inside a collection Pof tiles, we say that Tis complete with respect to Pif T= Tree(PT)∩P. Let α>0 and Tbe a tree. We define an α-packing of Tto be a set P⊂Tof tiles such that  P∈P |IP|≤α|IT|. We say that Pis a uniform α-packing5of Tif  P∈P:IP⊂J |IP|≤α|J| for all dyadic intervals J. If α<1/2 and Pis an α-packing of T, observe that the parent tiles 2P:= {2P:P∈P}form a 2α-packing of T. Similarly if Pis a uniform α-packing of T, then 2Pis a uniform 2α-packing of T. We say that a collection of lacunary tiles Pis convex if for every pair of tiles P1≤P2in P, the set {P∈P+:P1≤P≤P2}is also contained in P. We will usually be dealing with convex trees in this paper. The correspondence between the upper half-space and phase space is given by the heuristic formula6 t=1/ξ;(1) in other words, frequency is the reciprocal of wavelength. This correspondence identifies Whitney boxes Q+(I) with lacunary tiles P+(I), and identifies a Carleson box Q(I) with the complete tree Tree(I). 4Non-lacunary trees T⊆P0are also useful in the study of the bilinear Hilbert transform and Carleson’s operator; more precisely, when treating the bilinear Hilbert transform B(f,g),hone uses a triple of trees associated to f,g,hrespectively, with two of the trees lacunary and the third non-lacunary (but possibly with a non-zero frequency origin). Similarly when treating the Carleson operator CN(x)f,χEone uses a pair of trees associated to fand χErespectively, with one lacunary and one nonlacunary. See [40], [41]. However we will not use non-lacunary trees explicitly in this paper, although they appear implicitly in Lemma 4.2 and in the paraproduct theory. 5This is roughly equivalent to P∈PχIPhaving a BMO norm bounded by α. 6To be completely precise, one would have to adjust this formula when ξ≤21−M, but as this is only a heuristic anyway we will not bother to do this. Trees, Extrapolation, T(b)265 (Incomplete trees Tare identified with the portion of a Carleson box above a “dyadic Lipschitz graph”, cf. [3].) Note how this correspondence clearly gives a privileged position to the frequency origin ξ=0. The thesis of this paper is that the theory of Carleson measures can be equated with the theory of lacunary tiles. The theory of general tiles —which is needed for applications such as Carleson’s theorem and the bilinear Hilbert transform, in which the frequency origin plays no distinguished role— can then be thought of as a generalization of Carleson measure theory7. Tiles and trees Carleson measures Phase space R2Upper half-plane R2 + Lacunary tile P+(I) Whitney box Q+(I) Non-lacunary tile P0(I) “Tower” I×[l(I),∞) Complete tree Tree(I) Carleson box Q(I) Convex tree Carleson box above a Lipschitz graph Size µsize(T)Normalized mass µ(Q(I))/|I| Bounded maximal size Carleson measure (or BMO function) ξ1/t Table 1. A partial dictionary between tree terminology, and Carleson measure terminology. In our paper the two viewpoints are essentially equivalent, however the phase space viewpoint is better adapted to handle more general situations where one needs to modulate in frequency. Conversely, Carleson measures are better adapted to complex analysis applications. 7In our paper, we will only need tiles which are centered at or near the frequency origin, in which case it does not particularly matter whether we use the Carleson half-plane or the phase plane. However, we have chosen to use phase space notation (using frequency ξinstead of wavelength t) as this is more compatible with the more general theory of multilinear operators such as the bilinear Hilbert transform (or the Carleson maximal operator), which are invariant under translations of the frequency variable. Note that the modulation operation f→ e2πiξ0xfcan be represented easily in the phase plane as a translation by ξ0in the ξvariable, but is not so elegantly representable in the Carleson half-plane. Nevertheless, we will not need to modulate in frequency in this paper, so the Carleson viewpoint and the phase space viewpoint are essentially equivalent here. 272 Auscher, Hofmann, Muscalu, Tao, Thiele it suffices to find a collection Tof trees in Twhose tops are a 1 2-packing of Tsuch that |Wf|2size(T\T∈TT)1.(8) First observe from (4) that if J⊂Iand x∈Jthen ΠTree(I)f(x)−[ΠTree(I)f]J=Π Tree(J)f(x) so from hypothesis we have J |ΠTree(I)f−[ΠTree(I)f]J|p|J|.(9) Now let Cpbe a large constant to be chosen later. Let Qdenote the tiles Q∈Tree(I) such that |[ΠTree(I)f]Q|≥Cpand that Qis maximal with respect to ≤.IfCpis sufficiently large, we see from (9) IQ |ΠTree(I)f|pCp p|IQ| and hence that Qis a 1 4-packing of T. In particular the collection 2Q= {2Q:Q∈Q}of parents of tiles in Qis a 1 2-packing of T. We set T:= {Tree(2Q):2Q∈2Q}, and define F:= ΠT\T∈TTf=Π Tree(I)f− Q∈Q ΠTree(2Q)f. Then we can rewrite the left-hand side of (8) as 1 |I|F2 2. By construction we see that Fis supported on I. Since [ΠTree(2Q)f]2Q= 0, we see that Fis constant on each I2Qand that FL∞(2Q)=|[F]2Q|=|[ΠTree(I)f]2Q|Cp. If x∈Iis not in any of the I2Q, then |F(x)|=|ΠTree(I)f(x)|Cp. Thus we have F∞Cp. Combining this with the previous we obtain (8) as desired. We now give the well-known converse to the above lemma: Lemma 3.4 (John-Nirenberg inequality).Let Ibe a dyadic interval, and let f∈S0(I)be real-valued. Then we have fp(1 + p)|I|1/pfBMO for all 0<p<∞and |{x∈I:f(x)>2nfBMO}| ≤ 2−n+1|I|for all n∈Z+. Proof: It suffices to prove the latter inequality, as the former easily follows. We prove the claim by induction on n. The claim is clear for n=1. Now suppose that n>1 and the claim has already been proven for n−1. Trees, Extrapolation, T(b)273 Fix I,f. Let Pdenote those tiles Pin Tree(I) such that [f]P> 2fBMO, and such that Pis maximal with respect to ≤. For each P we have IP |f|2≥|IP||[f]P|2≥4|IP|f2 BMO. On the other hand, from (6), (7) we have I |f|2≤|I||Wf|2size(Tree(I)) ≤|I|f2 BMO. Thus Pis a 1 4-packing of Tree(I), so that the collection 2P={2P:P∈ P}of parents of tiles in Pform a 1 2-packing of Tree(I). By construction we have [f]2P≤2fBMO for all P∈P, and f(x)≤ 2fBMO for all x∈ P∈PI2P.Thus {x∈I:f(x)>2nfBMO}⊆  2P∈2P {x∈I2P:f−[f]2P(x)>2(n−1)fBMO}. The claim then follows from the inductive hypothesis. 3.1. Chopping big trees into little trees, and extrapolation of Carleson measures. Let a:P+→R+be a function. Suppose we have a convex lacunary tree T0with a large size, let’s say asize∗(T0)≤C0.(10) Let 0 <δ≤C0be a small number. An obvious question to ask is whether one can decompose the large tree T0into small trees, each of which has size less than or equal to δ. This is clearly impossible, as the example of a singleton tree T0with large size demonstrates. However, one can do the next best thing: Theorem 3.5. With the above assumptions, we have the disjoint partition T0= T∈T T∪P(11) where the trees Tin Tare convex and satisfy asize∗(T)≤δ(12) while the tiles P∈Pobey the estimate a(P)≤C0|IP|.(13) Furthermore, the tiles Pand the tree tops {PT:T∈T}are both uniform C(C0,δ)-packings of T0. 274 Auscher, Hofmann, Muscalu, Tao, Thiele Note that Lemma 2.1 gives an easy converse to the above theorem: if T0can be partitioned by (11) with the above properties then asize∗(T0) is bounded (but by a much large constant than C0). Thus, if one is willing to ignore losses in constants, the above theorem gives a complete characterization of trees of large size in terms of trees of small size. As we shall see in this section, this theorem can be applied to give extrapolation lemma for Carleson measures, and seems likely to be useful in other contexts also. A continuous parameter version of Theorem 3.5 is at least implicit in [3], where, as here, it is used to prove the “Extrapolation Lemma for Carleson Measures” (see Corollary 3.9 below). The latter, in its continuous parameter form, was then used to establish the “restricted version” of the Kato square root conjecture, for L∞perturbations of real, symmetric, elliptic coefficient matrices. The essential idea of the extrapolation method had previously been introduced by J. L. Lewis in his work with M. A. M. Murray [42] on the heat equation in noncylindrical domains, and refined further by Lewis and one of the present authors [32] in their work on parabolic and elliptic equations. Similar ideas had also appeared previously in the work of David and Semmes on uniform rectifiability: indeed Theorem 4.5 is very closely related to the “Corona Decomposition” of [24]. Roughly speaking, in applications of the extrapolation method, the idea is first to show that some “scale-invariant estimate on cubes” (like a Carleson measure estimate, a BMO estimate, or a reverse Holder or A∞ estimate for a weight) holds when some controlling Carleson measure is suitably small in a certain sense, which will be made precise in the sequel. The term “extrapolation” refers to the removal of the smallness restriction. In that sense it is analogous to G. David’s technique for bootstrapping the Lipschitz constant (see, e.g., [21]), although it is not clear whether there exists an explicit connection between the two methods. In [42], [32], for example, the controlling Carleson measure was a condition on either the boundary of the domain, or on the coefficients of the elliptic or parabolic operator, and one proved reverse Holder iequalities for the associated elliptic-harmonic or parabolic measures. In particular, in [32], the authors give an alternative proof, via extrapolation, of the main theorem of R. A. Fefferman, C. E. Kenig and J. Pipher [28], in which the controlling Carleson measure is a condition on the disagreement between the coefficients of two elliptic (or parabolic) operators, in the case that reverse Holder estimates are known to hold for the ellipticharmonic measures associated to the first operator, and one wishes to prove such estimates for the second. Trees, Extrapolation, T(b)275 In [3], the authors exploit the fact that proving Kato’s square root estimate is equivalent (by “T(1)” type reasoning) to proving that a certain positive measure in the upper half space is Carleson. Here, the controlling Carleson measure was the one associated to the original, self-adjoint operator, and the extrapolation technique was used to prove that the analogous measure, related to the square root estimate for the perturbed operator, was also Carleson. It is in this setting that Corollary 3.9, or rather its continuous parameter analogue, is directly applicable. The proof we give here follows the approach in [3]. At the end of this section we give an alternate proof, due to John B. Garnett, which gives better dependence on constants. Before we give the rigorous proof, we first informally describe the idea of the argument. Suppose the original tree T0has size asize(T0)=c. Then 0 ≤c≤C0by (10). To create a tree of maximal size less than δ, we start with T0and remove from it some sub-trees of size between c+δ/2 and c−δ/2, which we select by a straightforward stopping time argument. It then remains to control the sub-trees that were removed. By shrinking the trees slightly (putting the error into P) we can assume that the trees have size either greater than c+δ/2 or less than c−δ/2 (so that the tree that remains must have size at most δ). We call the first type of tree “heavy” and the second type “light”. Because the original tree had size c, it cannot be the case that IT0is covered by heavy subtrees, and so a positive proportion of IT0must be covered by light trees or by nothing. We then pass to the light sub-trees and iterate this process, finding a positive proportion IT0occupied by increasingly lighter subtrees. After about O(C0/δ) steps, we must terminate, finding a positive proportion of IT0which are not covered by any further sub-trees. We then pass to the remaining portion of IT0and all the heavy trees which have until now been neglected, and iterate once again; since we have replaced IT0with a strictly smaller fraction of IT0, this procedure will converge geometrically to obtain the desired estimates. We now prove Theorem 3.5. We shall drop (13) since it follows from (10). In the spirit of Lemma 3.1, it will suffice to prove the apparently weaker Theorem 3.6. With the above assumptions, we can find a (possibly empty) collection Titerate of disjoint convex trees in T0whose tops have disjoint spatial intervals and form a (1 −η)-packing of T0for some η=η(C0,δ)>0, such that we have the disjoint partition T0= T∈Titerate T∪ T∈T T∪P(14) 276 Auscher, Hofmann, Muscalu, Tao, Thiele where the trees T∈Tobey (12), and Pand the tree tops of Tare both uniform C(C0,δ)-packings of T0. Indeed, if Theorem 3.6 holds, then we can construct the collection in Theorem 3.5 by starting with the partition (14), and then taking each of the trees in Titerate and breaking them up by a further application of Theorem 3.6. We continue on in this way until the original tree T0 is completely broken up into trees Tobeying (12) and tiles Pobeying (13). The fact that Pand the tree tops of Tare C(C0,δ)-packings of T0then follows from Theorem 3.6 and the fact that the geometric series n(1 −η)nconverges. A similar argument can then be used to improve “C(C0,δ)-packing” to “uniform C(C0,δ)-packing”. We omit the details. Proof of Theorem 3.6: Define the quantity cby c:= asize(T0),(15) thus 0 ≤c≤C0. We shall prove the theorem by induction on c. Specifically, we fix 0 ≤c≤C0and assume that the theorem has already been proven in the case asize(T0)≤c−δ/2. Note that we only have to apply this induction a finite number of times (about O(C0/δ)) so we will be allowed to let the constants get worse with each induction step. The main lemma used in the proof of the theorem will be Lemma 3.7. We can partition T0= T∈Tsmall T∪Pbuffer ∪ T∈Theavy T∪ T∈Tlight T(16) where Tsmall is a collection of convex trees which all obey (12) and whose tree tops are a uniform 4-packing of T0,Pbuffer is uniform 3-packing of T0, and Theavy,Tlight are collections of disjoint convex sub-trees of T0 which are complete with respect to T0, and are such that we have the tree counting estimates  T∈Theavy |IT|+ T∈Tlight |IT|≤|IT0|(17)  T∈Theavy |IT|≤ c c+δ/2|IT0|(18) and the size bounds asize(T)≤c−δ/2for all T∈Tlight.(19) Trees, Extrapolation, T(b)277 b l h s 2M s bbs sshhbh h bbb bb l hhhhhh hhll Figure 3. A convex tree T0and its decomposition from Lemma 3.7. The circled hand ltiles are the tops of maximal sub-trees of T0for which the size of afluctuates by at least δ/2 from c; the uncircled hand ltiles are the remaining tiles in those maximal sub-trees. Theavy thus consists of the (three) htrees while Tlight consists of the (two) ltrees. Pbuffer consists of those remaining tiles (labeled b) which lie just below a heavy or light tile, or are at the very top of the phase plane. The remaining tiles (labeled s) form the (three) small trees Tsmall. Proof: Define Tfluctuate to be those sub-trees Tof T0which are complete with respect to T0, such that |asize(T)−c|≥δ/2, and such that Tis maximal with respect to set inclusion and the above two properties. Note that such trees are automatically convex. By construction, none of the trees in Tfluctuate contain the top tile PT0. We may subdivide11 Tfluctuate =Theavy ∪Tlight 11With reference to Figure 3, Tfluctuate consists of the hand ltrees, T1consists of the sand btiles, and T2consists of just the stiles. 278 Auscher, Hofmann, Muscalu, Tao, Thiele where Theavy consists of those trees T∈Tfluctuate with asize(T)≥c+δ/2(20) and Tlight consists of those trees T∈Tfluctuate with asize(T)≤c−δ/2. The trees in Tfluctuate are disjoint, convex, and have disjoint spatial supports, so (17) holds. On the other hand if one multiplies (20) by |IT| and sums over all T∈Theavy one obtains |IT0|c= P∈T0 a(P)≥ P∈T∈Theavy T a(P)≥(c+δ/2)  T∈Theavy |IT|. Dividing by c+δ/2 we obtain (18). Let T1denote the convex tree T1:= T0\T∈Tfluctuate Twith top PT0. Informally, T1represents the portion of T0below the fluctuating tiles. The tree T1contains PT0and is hence non-empty. Let Pbuffer denote the tiles12 Pbuffer := {P∈T1:P=2Qfor some Q∈ T1}∪{P∈T1:|IP|=2 −M}. In other words, Pbuffer consists of those tiles in T1which touch the upper boundary of T1(which in particular may include the tiles of minimal width |IP|=2 −M). Since the tiles Qin the definition of Pbuffer have disjoint spatial supports and |IP|=2|IQ|we see that {P∈T1:P= 2Qfor some Q∈ T1}is a uniform 2-packing of T1. Since {P∈T1: |IP|=2 −M}is clearly a uniform 1-packing of T1, we thus see that Pbuffer is a uniform 3-packing of T1. Let T2denote the (possibly empty) tree T2:= T1\Pbuffer with top PT0. This tree is not necessarily convex, however we shall invoke the following lemma to split it into convex trees. Lemma 3.8. Let Tbe a convex tree, and let P⊂Tbe a uniform α-packing of Tfor some α>0. Then T\Pcan be partitioned into T\P= T∈T T where Tis a collection of convex trees Twhose tops {PT:T∈T}form a uniform (α+1)-packing of T. 12Here we are taking advantage of our decision to work in a finite model, where the tiles have a minimal width 2−M. One can replicate this argument in the infinite setting but one has to treat the portion of T1which “goes all the way to infinity” separately. See [3]. Trees, Extrapolation, T(b)279 Proof: Let Qdenote those dyadic intervals Q⊂ITsuch that Q∈T\P and 2Q∈ T\P. For any Q∈Q, we see that either Q=PTor 2Q∈P. Since Pis a uniform α-packing, this implies that Qis a uniform (α+1)- packing. For each Q∈Q, define the convex tree TQwith top Qby TQ:= {P∈T\P:P≤Q, and there does not exist Q∈Qsuch that P< Q<Q}. If we then set T:= {TQ:Q∈Q}we see that the lemma follows. By Lemma 3.8 we may write T2=T∈Tsmall Twhere the trees in Tsmall are distinct and the tree tops of Tsmall are a uniform 4-packing of T. We now verify that each tree T∈Tsmall obeys (12). It suffices to show that  P∈Tree(I)∩T a(P)≤δ|I|(21) for all I⊆IT0. Fix I. The idea is to write Tree(I)∩Tas the difference of trees, each of which has size c+O(δ). We may assume that P+(I)∈Tsince the claim is trivial otherwise. We observe that Tree(I)∩T= (Tree(I)∩T0)\ J∈J (Tree(J)∩T0) where Jconsists of those intervals J⊆Isuch that P+(J)∈ T, and which are maximal with respect to this property. The tile P+(I)isinTand hence in T1. By construction of T1,we thus have  P∈Tree(I)∩T0 a(P)=|I|asize(Tree(I)∩T0)≤|I|(c+δ/2) (since otherwise Tree(I)∩T0would belong to Theavy, a contradiction). Similarly, for every J∈J, the tile P+(J) is contained in T1(otherwise P+(2J) would be both in Tand in Pbuffer, a contradiction), so  P∈Tree(J)∩T0 a(P)≥|J|(c−δ/2). By the construction of J, the intervals Jin Jpartition I,thus  J∈J P∈Tree(J)∩T0 a(P)≥|I|(c−δ/2). Subtracting this from the previous we obtain (21) as desired. 280 Auscher, Hofmann, Muscalu, Tao, Thiele We apply the above lemma and place Pbuffer into P,Tsmall into T, and Theavy into Titerate. For the remaining trees Tlight we use the induction hypothesis, which splits each of the trees in Tlight into Titerate,T, and P. All the desired conclusions of Theorem 3.6 are easily verified except perhaps for the claim that the tops of Titerate form a (1 −η)-packing of T, or in other words  T∈Titerate |IT|≤(1 −η)|IT0|. To prove this inequality, note that the trees in Theavy contribute T∈Theavy |IT|to the left-hand side, while from the induction hypothesis the trees in Tlight contribute at most (1−η)T∈Tlight |IT|for some η>0. There are no other contributions. The claim then follows from (17), (18) (reducing the value of ηas necessary). The constants C(C0,δ) given by this argument are about (C0/δ)CC0/δ. This bound can be improved substantially; see below. The following corollary allows one to use one Carleson measure µto prove the Carleson measure property of a related measure µ.Itisthe dyadic version of an extrapolation lemma in [3], which in turn is based on ideas in [42], [32]. Corollary 3.9 (Extrapolation of Carleson measures).Let µ:P+→R+ have bounded maximal size and let δ>0. Let µbe a non-negative measure on R2 +obeying the “weak Carleson condition” µ(P)≤C1|IP|for all P∈P+ and such that µsize(T)≤C2for all convex trees Tsuch that µsize∗(T)≤ δ. Then µalso has bounded maximal size: µsize∗(P+)≤C(µsize∗(P+),δ)(C1+C2). Proof: Let T0be any convex tree. We need to show that µsize(T0)≤C(µsize∗(P+),δ)(C1+C2). By Theorem 3.5, we can partition T0=T∈TT∪Pwhere µsize∗(T)≤δ for all T∈T, and  T∈T |IT|+ P∈P |IP|≤C(µsize∗(P+),δ)|IT0|. Trees, Extrapolation, T(b)281 From this and assumptions on µwe see that  P∈T0 µ(P)=  T∈T P∈T µ(P)+  P∈P µ(P) ≤C(µsize∗(P+),δ)C2|IT0|+C(µsize∗(P+),δ)C1|IT0| and the claim follows. As mentioned earlier, this lemma has applications to the Kato problem. In [3], this lemma was used to establish a restricted version of Kato’s conjecture, for perturbations of real, symmetric coefficient matrices. In that case, µwas a Carleson measure which controlled the original operator, and µwas the analogous measure controlling the perturbed operator. The point was to establish that µwas also a Carleson measure. We remark that the fact that the final bound on µwas linear in C2was crucial to this application. It is possible to eliminate the weak Carleson condition by allowing the tree measured by µto be a little larger than the tree measured by µ, but we will not pursue this type of generalization here. 3.2. An alternate argument. In this section we give an alternate proof of Theorem 3.5, due to John B. Garnett (personal communication). The idea of this argument is similar to some arguments in [8]. Fix T0,a. We first observe that it suffices to prove the theorem under the additional “weak Carleson” assumption a(P)≤δ 2|IP|for all P∈T0.(22) To see this, suppose that we are in the general case when (22) need not hold. We set Pto be the set of tiles where (22) fails: P:= P∈T0:a(P)≥δ 2|IP|.(23) From (10) we see that Pis a uniform 2C0/δ-packing of T0(cf. Lemma 2.1). By Lemma 3.8 we thus see that we can split T0\Pinto a collection of disjoint convex subtrees of T0, whose tops form a uniform 2C0/δ + 1-packing of T0. On each such subtree (22) holds. Thus if we apply Theorem 3.5 to each sub-tree and then combine all the decompositions, we obtain the desired decomposition for the original tree T0 (with the constants C(C0,δ) worsened by a factor of 2C0/δ + 1). Henceforth we assume (22). Under this assumption we will not need P any more, and will set it equal to the empty set. 288 Auscher, Hofmann, Muscalu, Tao, Thiele If Tree(I) is a complete tree, we define a (dyadic) Hpatom on Tree(I) to be a function a∈S0(I) such that a2≤|I|1/2−1/p. Equivalently, a∈S0is an Hpatom on Tree(I) if and only if the wavelet transform Wa of ais supported on Tree(I) and |Wa|2size(Tree(I)) ≤|I|−2/p; this is because of (6). In this section we show Theorem 4.3 (Equivalent definitions of Hp).Let f∈S0and 0<p≤ 1. Then the following statements are equivalent. (i) Sfp1. (ii) ˜ Mfp1. (iii) There exists a collection Iof dyadic intervals, and to each I∈ Ithere exists a non-negative number cIand an Hpatom aIon Tree(I)such that f=IcIaIand Icp I1. Proof: We first show that (iii) implies (i) and (ii). From the quasitriangle inequality f+gp p≤fp p+gp p we see that it suffices to verify this on atoms, i.e. to show that SaIp, ˜ MaIp1 whenever aIis a Hpatom on Tree(I). Fix I,a. By construction ˜ MaIand SaIare supported on I,soby H¨older it suffices to show SaI2,˜ MaI2|I|1/2−1/p. But this follows from the L2normalization of aIand the fact that S,˜ Mare bounded on L2. It remains to show that either one of (i) or (ii) are enough to imply (iii). Let fbe any element of S0,thusf=P∈P+Wf(P)φP. Set a:= |Wf|2. We apply Lemma 4.1 repeatedly, starting with a sufficiently large nand setting Pn:= P+, and then decrementing n indefinitely. Eventually one obtains a partition P+= n∈Z T∈Tn T∪P−∞ where the Tnare as in Lemma 4.1, and |Wf|2size∗(P−∞ )=0. Thus Wf vanishes on P−∞, and only a finite number of Tnare non-empty. We thus have f=n∈ZT∈TnΠTf. If we then set cIT:= 2n/2|IT|1/p and aIT:= ΠTf/cITthen we have f= n∈Z T∈Tn cITaIT. Trees, Extrapolation, T(b)289 By (32) (with p= 2) we see that each aITis an Hpatom. To show (iii) it thus remains to show that  n T∈Tn cp IT= n 2np/2 T∈Tn |IT|1. First suppose that (i) holds. For each nand each T∈Tn, we see that IT |SΠTf|q|IT|SΠTfq BMO =|IT|ΠTfq BMO ∼2nq/2|IT| for all 2 <q<∞by the John-Nirenberg inequality (Lemma 3.4), (41), and (31). Also, we have IT |SΠTf|2=IT |ΠTf|2∼2n|IT|. By H¨older we thus have IT |Sf|r≥IT |SΠTf|r2nr/2|IT|, for any 0 <r<p, with the implicit constant depending on r. This clearly implies x∈IT:|Sf(x)|2n/2 |Sf|r2nr/2|IT|. Summing over all T∈Tnand using the disjointness of the ITwe obtain |Sf(x)|2n/2 |Sf|r2nr/2 T∈IT |IT|. Multiplying by 2n(p−r)/2and summing over nwe obtain  n:|Sf(x)|2n/2 2n(p−r)/2|Sf|r n 2np/2 T∈IT |IT|. Since the left-hand side is comparable to Sfp p, the claim follows. Now suppose instead that (ii) holds. By (32), (42) we have IT|˜ Mf|r 2nr/2|IT|for all 0 <r<p. Now we argue as with Sf to obtain (iii) from (ii). From the above proof we see that the atoms in fact obey the BMO bound aITBMO |I|−1/p. One can improve this BMO control to L∞ control by repeating the argument in the John-Nirenberg inequality (Lemma 3.4). Namely, one locates the maximal sub-intervals where the averages of aITare large and separates off those trees, leaving behind a 290 Auscher, Hofmann, Muscalu, Tao, Thiele bounded atom. One then repeats the process until only bounded atoms remain, in the spirit of Lemma 3.1. We omit the details. Let fbe in BMO. From (7), (6), we have fBMO = sup T 1 |IT|1/2ΠTf2. Clearly it suffices to take suprema over complete trees T, thus by (4) fBMO = sup I 1 |I|1/2I |f−[f]I|21/2 . By duality we thus have fBMO = sup I sup a∈S0(I):a2=1 |I|−1/2|f,a|(43) or equivalently fBMO = sup{|f,a| :ais a H1atom}. Thus, as is well known, BMO is the dual of H1. 5. The Carleson embedding theorem and paraproducts We now give a slight variant of the above method, in which one selects trees using the averages [f]Iinstead of the sizes. This type of argument is of course very old, and the arguments here are by no means new. On the other hand, this type of tree selection method is a special case of the “mean selection” algorithm used (together with a size selection algorithm) in the proof of Carleson’s theorem in [41]. We begin with Lemma 5.1 (Carleson embedding theorem).Let Pbe a collection of lacunary tiles, a:P→R+be a function, and 1<p<∞. Then we have  P∈P a(P)|[f]IP|pasize∗(P)fp p for all locally integrable functions f, with the implicit constants depending on p. Trees, Extrapolation, T(b)291 Proof: We apply Lemma 4.2 repeatedly, starting with a sufficiently large nand decrementing nrepeatedly. This gives us a partition P= n∈ZT∈TnT∪P−∞ where the Tnare as in Lemma 4.2, and fmean∗(P−∞)= 0. The contribution of P−∞ is zero, so it suffices to control  n T∈Tn P∈T a(P)|[f]IP|p. If P∈T∈Tn, then |[f]IP|≤fmean(P)≤fmean∗(T)2n. From this and (2), (3), (39) we may estimate the previous by  n T∈Tn P∈T a(P)2np  n T∈Tn asize∗(P)|IT|2np  n asize∗(P)2np2−n|f|2n |f| ∼asize∗(P)|f|p as desired. We can apply this theorem to various linear and bilinear operators. To do this we shall need some notation. For any sequence (aP)P∈P+of real numbers, we define the wavelet multiplier W−1aPWfrom S0to S0by W−1aPWf :=  P∈P+ aPWf(P)φP. Wavelet multipliers are the discrete analogue of pseudo-differential operators, with aPbeing the discrete analogue of a symbol a(x, ξ). Observe that if aPis bounded, then W−1aPWis bounded on L2and also bounded on BMO. One can of course extend the domain W−1aPWfrom S0to S, although some of the algebra properties are lost in doing so (since Wis not injective on S). 292 Auscher, Hofmann, Muscalu, Tao, Thiele Let f,gbe elements of S. We define the “high-low”, “low-high”, and “high-high” paraproducts14 πhl(f,g):=  P∈P+ Wf(P)[g]PφP πlh(f,g):=  P∈P+ [f]PWg(P)φP πhh(f,g):=  P∈P+ Wf(P)Wg(P)χIP |IP|. These paraproducts have the symmetries πhh(f,g)h=πhl(g,h)f=πlh(h, f)g =πhh(g,f)h=πhl(f,h)g=πlh(h, g)f = P∈P+ Wf(P)Wg(P)[h]P (44) and can be expressed in terms of the Littlewood-Paley square function S: πhl(f,g)=S∗(gSf); πlh(f,g)=S∗(fSg); πhh(f,g)=Sf ·Sg. When f,ghave mean zero (i.e. f,g ∈S0), then the paraproducts decompose the pointwise product operator: fg =πhl(f,g)+πlh(f,g)+πhh(f,g).(45) To see this, it suffices by bilinearity to reduce to the case when f=φP and g=φQfor some P,Q ∈P+.IfIPand IQare disjoint then both sides are zero. Thus there are only three cases: P> Q,P< Q, and P=Q. In these three cases the reader may easily verify that fg is equal to the high-low, low-high, or high-high paraproduct of fand g respectively, and that the other two paraproducts vanish. We observe that the high-low and low-high paraproducts can be written as wavelet multipliers: πhl(f,g)=W−1[g]PWf;πlh(f,g)=W−1[f]PWg.(46) 14The continuous counterparts would be something like (Qtf)(Ptg)dt t, (Ptf)(Qtg)dt t, and (Qtf)(Qtg)dt t, where Qtis as before and Ptis a suitable approximation to the identity at width 1/t, e.g. Pt:= et2∆. The precise definition of a paraproduct is not standardized, for instance πhh is not considered a paraproduct in some texts. Trees, Extrapolation, T(b)293 The high-high paraproduct cannot be written in this way, but we have the useful relationship πhh(W−1aPWf,g)=πhh(f,W−1aPWg).(47) One can also write paraproducts using both lacunary and nonlacunary tiles, for instance int πhh(fg)h= Idyadic |I|−1/2f,φP+(I)g,φP+(I)h, φP0(I).(48) The bilinear Hilbert transform turns out to have a similar expansion, but with the sum ranging over a larger collection of triples of tiles than the ones for paraproducts (specifically, the tiles need not be lacunary or non-lacunary, and range over a three-parameter family rather than a two-parameter one). See e.g. [39], [40], [52], [45], [53], [46]. From the Carleson embedding theorem we have paraproduct estimates: Corollary 5.2 (L2×BMO →L2paraproduct estimates).We have πhl(f,g)2f2g∞ and πhh(f,g)2,πlh(f,g)2f2gBMO for all f,g ∈S. In other words, paraproducts map L2×L∞to L2(just as the pointwise product does), and the L∞factor can be relaxed to BMO as long as one only considers high frequencies of a BMO function. Note that the low frequency portion of a BMO function is somewhat ill-defined since a BMO function might only be determined up to a constant. Proof: The first bound follows from (46) since [g]Pis bounded by g∞. To prove the second bound, it suffices by (44) to consider πlh. By orthogonality we have πlh(f,g)2= P∈P+ |Wg(P)|2|[f]IP|21/2 . The claim now follows from Carleson embedding (Lemma 5.1). Lemma 5.3 (BMO ×BMO →BMO paraproduct estimate).We have πhh(f,g)BMO fBMOgBMO for all f,g ∈S. 294 Auscher, Hofmann, Muscalu, Tao, Thiele For the other paraproducts πhl,πlh one must place the “low” factor in L∞rather than BMO, as in Lemma 5.2; this is again an easy consequence of (46). Proof: By Lemma 3.3 with p=1 it suffices to show |ΠTree(I)πhh(f,g)| |I|for all dyadic intervals I. Fix I, and expand the left-hand side as  P∈P+ Wf(P)Wg(P)ΠTree(I)χIP |IP| . The summand vanishes unless P∈Tree(I). Thus we can write the above as  ΠTree(I) P∈Tree(I) Wf(P)Wg(P)χIP |IP| . Since ΠTree(I)is bounded on L1, we can bound this by  P∈Tree(I) Wf(P)Wg(P)χIP |IP| . Putting the absolute values inside and performing the integration, we can bound this by  P∈Tree(I) |Wf(P)||Wg(P)|. The claim then follows from Cauchy-Schwarz and (7). 5.1. Weak-type estimates. We now show how to use the above machinery to prove Lp,∞paraproduct estimates, where Lp,∞is the weak Lp (quasi-)norm fLp,∞:= sup λ>0 λ|{x:|f(x)|≥λ}|1/p. We need the following basic characterization of weak Lpfor 0 <p<∞: Lemma 5.4. Let 0<p<∞and A>0. Then the following statements are equivalent up to constants: (i) fp,∞A. (ii) For every set Ewith 0<|E|<∞, there exists a subset E⊂E with |E|∼|E|and |f,χE| A|E|1/p. Here pis defined by 1/p+1/p =1(note that pcan be negative!). Trees, Extrapolation, T(b)295 Proof: To see that (i) implies (ii), set E:= E\{x:|f(x)|≥CA|E|−1/p}. If Cis a sufficiently large constant, then (i) implies |E|∼|E|, and the claim follows. To see that (ii) implies (i), let λ>0 be arbitrary and set E:= {x: Re(f(x)) >λ}. Then by (ii) we have λ|E|∼λ|E|A|E|1/p, and (i) easily follows (replacing Re by −Re, Im, −Im as necessary). When p>1 we can always set E=E, and the above lemma then reflects the duality between Lp,∞and Lp,1. However for p≤1the freedom to set Eto be smaller than Eis necessary (since fneed not be locally integrable). A typical application of Lemma 5.4 is Proposition 5.5 (Lp×Lq→Lr,∞paraproduct estimates).We have πhl(f,g)r,∞fpgq whenever 1<p,q<∞and 1/p +1/q =1/r. Similarly for πlh,πhh. Note that rcan be less than 1. One can strengthen the weak Lr to strong Lrby multilinear interpolation (see e.g. [6], [45], [33]). The continuous version of these dyadic paraproduct estimates can be found in, e.g. [14]–[19]; the version for r<1 was first proven in [30] (with some special cases in [9], [14]). It is possible to obtain the continuous estimates from the dyadic ones via averaging arguments, but we shall not do so here. Proof: We first consider πhl. We may normalize fp=gq=1; we may assume that fand gare dyadic test functions. Let Ebe a measurable set with 0 <|E|<∞. We need to find a set E⊂Ewith |E|∼|E|such that |πhl(f,g),χ E| |E|1/r.(49) By rescaling (using the hypothesis 1/p+1/q =1/r) we may take |E|∼1. We choose Eas E:= E\{x:M|f|p(x)+M|g|q(x)≥C}. If Cis large enough, then |E|∼|E|by the Hardy-Littlewood maximal inequality (see e.g. [51]). We wish to show (49) with |E|∼1. By (44) it suffices to show that |P∈PWf(P)[g]IPWχE(P)|1 for all convex collections Pof tiles. 296 Auscher, Hofmann, Muscalu, Tao, Thiele We may remove all tiles in Pfor which IP∩E=∅, since WχE vanishes on these tiles. For any remaining tile Pwe then have IP |f|p|IP|(50) by construction of E. Similarly, we have gmean∗(P)= sup P∈P [|g|]IP[|g|q]1/q IP1. Thus we reduce to showing  P∈P |Wf(P)||WχE(P)|1.(51) From (50) and the Lpboundedness of the Littlewood-Paley square function (see e.g. [51]) we have IP |SΠTree(IP)f|pIP |ΠTree(IP)f|pIP |f|p|IP|, for all P∈P,thus IP  P∈Tree(IP) |Wf(P)|2χIP(x) |IP|  p/2 |IP|. Applying Chebyshev’s inequality and Corollary 3.2 we thus see that |Wf|2size∗(P)1. Also, we have |WχE|2size∗(P)≤χE2 BMO ≤χE2 ∞≤1. Thus we can find an n=O(1) such that |Wf|2size∗(Pn)≤22nand |WχE|2size∗(Pn)≤22np/s, where have set Pn:= P, and s>1 is an exponent close to 1 to be chosen later. By a finite number of applications of Lemma 4.1 with a:= |Wf|2or a:= |WχE|2we may partition Pn=T∈TnT∪Pn−1where Pn−1is a convex collection of tiles such that |Wf|2size∗(Pn−1)≤22(n−1) and |WχE|2size∗(Pn−1)≤22(n−1)p/s and Tnis a collection of convex trees with disjoint spatial supports such that either IT |ΠTf|p2np|IT| Trees, Extrapolation, T(b)297 or IT |ΠTχE|s2np|IT| for each T∈Tn. From (42) we thus have (˜ Mf)p+(˜ MχE)s2np  T∈Tn |IT|; from our assumptions on f,χEand the Hardy-Littlewood maximal inequality (see e.g. [51]) we thus have T∈Tn|IT|2−np.Wenow return to (51), and estimate the contribution of the trees in Tnby  T∈Tn P∈T |Wf(P)||WχE(P)|. We apply Cauchy-Schwarz followed by the size control on Wf and WχE we may bound this by  T∈Tn |IT||Wf|21/2 size(T)|WχE|21/2 size(T)  T∈Tn |IT|2n2np/s 2−np2n2np/s. We now turn to the contribution of the tiles in Pn−1. We may iterate the above procedure, decomposing Pn−1into Tn−1and Pn−2, and continue in this fashion until we are left with a collection of tiles P−∞ with size zero, which we can discard. Summing up, we can thus control the left-hand side of (51) by n≤O(1) 2−np2n2np/s. If one chooses s sufficiently close to 1, then this sum converges, and we are done. A similar argument handles πhl. The remaining paraproduct πhh then follows from (45) and H¨older’s inequality (which is still valid for r<1). One can modify the above argument to obtain the corresponding estimate for the bilinear Hilbert transform (in the Walsh model, at least); see e.g. [39], [53], [46]. A difficulty in that case is that the tiles are no longer lacunary, and one cannot guarantee the spatial disjointness of the trees Tin Tn. However one can still make the trees essentially disjoint in phase space, but then one can only use L2estimates to control T∈Tn|IT|instead of Lpestimates. Because of this, the above strategy only seems to work for the bilinear Hilbert transform when r>2/3; it appears that one needs very different techniques to handle the remaining case 1/2<r≤2/3. 304 Auscher, Hofmann, Muscalu, Tao, Thiele The weak boundedness condition (56) can actually be removed; see the remarks after Theorem 6.8. Informally, this theorem asserts that to prove the L2boundedness of an operator T it actually suffices to establish boundedness for a single function bPfor each interval IP, provided that bPis not degenerate (in the sense that its mean is large) and provided that T∗(1) is under control. In the next section we shall remove the condition on T∗(1), obtaining a “two-sided” version of this theorem. Proof: Again it suffices to show that T(1) is in BMO. By Lemma 3.1 it suffices to show that for every complete tree Twe have  P∈T\T∈TT |W(T(1))(P)|2|IT|(64) for some collection Tof disjoint convex trees in Twhose tops form a (1 −ε)-packing of Tfor some ε>0. Fix T. By Lemma 6.5 we can indeed find such a collection Twith the additional property that bis pseudo-accretive on T\T∈TT.The claim (64) then follows from the argument used to prove Corollary 6.4. The above arguments do not extend well to two-sided situations in which one controls T(b1) and T∗(b2) (unless one of b1,b2is close to a constant, e.g. in BMO norm). In order to handle the general case we need adapted Haar bases, to which we now turn. 6.2. Adapted Haar bases, and two-sided T(b) theorems. Let P be a collection of tiles, and let bbe a function which is strongly pseudoaccretive on P. For each P∈P, we define the adapted Haar wavelet φb P (introduced in [20]; see also [4]) by φb P:= |IP|−1/2[b]Pr [b]P χIPl−|IP|−1/2[b]Pl [b]P χIPr.(65) Observe that this collapses to φPif bis constant on IP. For nonconstant b,φb Pis no longer mean zero, but one can easily verify that φb P still obeys the weighted mean zero condition bφb P=0.(66) As a consequence we have the orthogonality property φb Pbφb Q= 0 for all distinct P,Q ∈P.(67) Trees, Extrapolation, T(b)305 From (65) we see that φb Pbφb P=[b]Pr[b]Pl [b]P =2 [b]−1 Pl+[b]−1 Pr .(68) In particular, from the strong pseudo-accretivity condition (59) we have the bound φb Pbφb P 1.(69) It is interesting that this bound uses only the strong pseudo-accretivity of b, and in particular does not require L∞control on b. Define the dual adapted Haar wavelet ψb Pby ψb P:= φb Pb φb Pbφb P . By (67), (69) we thus have that ψb P,φ b Q=δPQ where δis the Kronecker delta. In particular we have the representation formula f= P∈P Wbf(P)ψb P (70) whenever gis in the span of {ψb P:P∈P}, where the adapted wavelet coefficients Wbf(P) are defined by Wbf:= f,φb P. We have the following basic orthogonality property: Lemma 6.7. Let Tbe a convex tree, and let bbe a function which is pseudo-accretive on Tand obeys the mean bound |b|2mean∗(T)1.(71) Then for any function19 f∈Swe have  P∈T |Wbf(P)|21/2 f2.(72) In fact, the more general estimate  P∈T |Wb(bf)(P)|21/2 f2|b|21/2 mean∗(T) (73) holds for any f,b∈S. 19It can easily be seen, by aid of (70), that the estimate (72) can be reversed for all f in the span of the φb P, but we will not use this. 306 Auscher, Hofmann, Muscalu, Tao, Thiele Proof: From (71) and (6) we observe that |Wb|2size∗(T)1.(74) We first prove (72). From (65) and the identity |IP|−1/2Wb(P)= [b]Pl−[b]P=[b]P−[b]Pr, we obtain the identity Wbf(P)=Wf(P)−Wb(P) [b]P [f]P. If we replace Wbby Wbthen (72) follows from Bessel’s inequality and the orthonormality of the Haar wavelets φP. By the previous identity and the triangle inequality it thus suffices to show  P∈T Wb(P) [b]P [f]P 21/2 f2. We may discard [b]Pby pseudo-accretivity (58). The claim then follows from Carleson embedding (Lemma 5.1) and (74). Now we prove (73). Let Qdenote the collection of tiles in P+which are children of tiles in T, but are not in Titself. In order to ensure that the intervals {IQ:Q∈Q}partition ITwe will allow the tiles Qto have spatial intervals |IQ|=2 −M−1; the partition property then follows from the convexity of T. The function bf−Q∈Q[bf]QχIQhas mean zero on every interval IQ, and is thus orthogonal to φb Pfor every P∈T. We may thus freely replace bfby the averaged function Q∈Q[bf]QχIQin (73). By (72) it thus suffices to show that       Q∈Q [bf]QχIQ      2 2 f2 2|b|2mean∗(T). But from Cauchy-Schwarz we have |[bf]Q|2|IQ|f2 L2(IQ)|b|2mean(2Q)≤f2 L2(IQ)|b|2mean∗(T), and the claim follows by summing in Q. We can now give our main result, namely a dyadic local T(b) theorem. Trees, Extrapolation, T(b)307 Theorem 6.8 (Dyadic local T(b) theorem).Let Tbe a perfect Calder´on-Zygmund operator, and suppose that for each P∈P+we can find functions b1 P,b2 Pin S(IP)obeying the normalization [b1 P]P=[b2 P]P=1(75) and the bounds IP |b1 P|2+|Tb1 P|2+|b2 P|2+|T∗b2 P|2|IP|.(76) Then Tis bounded on L2. This theorem is a stronger version of the local T(b) theorem in [11] (but for the dyadic setting with perfect cancellation), which required L∞control in (76) instead of L2control. (This was generalized to BMO control and to non-doubling situations in [47].) Also it required the global T(b) theorem of David, Journ´e and Semmes [23] (which we instead deduce as a corollary of Theorem 6.8). We make some further remarks after the proof of the theorem. Proof: This proof is somewhat lengthy and so we split the argument into several stages. Step 0. Preliminary estimates: We begin with a basic lemma which already shows the importance of the normalization (75). Lemma 6.9 (b1 Pspans S(IP)/S0(IP)).For any tile P∈P+and any f∈S(IP), we have fL2(IP)f−[f]PL2(IP)+|IP|−1/2|f,b1 P|. Similarly for b2 P. Proof: Let hbe an arbitrary element of S(IP) with h2= 1. Then f,h=f,h−[h]Pb1 P+[h]Pf,b1 P=f−[f]P,h−[h]Pb1 P+[h]Pf,b1 P. By Cauchy-Schwarz and (76) we thus have |f,h| f−[f]PL2(IP)+|IP|−1/2|f,b1 P|. Taking suprema over all h, the claim follows. A useful application of the above lemma is the following convenient truncation property of the b1 Pand b2 P(already observed in [11]): 308 Auscher, Hofmann, Muscalu, Tao, Thiele Corollary 6.10. Let P,Qbe lacunary tiles with Q≤P. If we have the estimate IQ |Tb1 P|2+|b1 P|2K|IQ|(77) for some K1, then we have 2IQ |T(b1 PχIQ)|2K|IQ|. Similarly for b2 P(but with Treplaced by T∗). Proof: By (52) and (77) the portion of the integral on 2IQ\IQis acceptable, so it suffices to bound the integral on IQ. From Cauchy-Schwarz, (76) and (77) we have |T(b1 PχIQ),b 2 Q|=|b1 PχIQ,T∗b2 Q|≤b1 PL2(IQ)T∗b2 QL2(IQ)K1/2|IQ|. By Lemma 6.9 it thus suffices to show that T(b1 PχIQ)−[T(b1 PχIQ)]QL2(IQ)K1/2|IQ|1/2. Now observe that for every h∈S0(IQ)wehave T(b1 PχIQ),h=b1 PχIQ,T∗h=b1 P,T∗h=Tb1 P,h. By duality this implies that T(b1 PχIQ)−[T(b1 PχIQ)]Q=T(b1 P)−[T(b1 P)]Q on IQ. The claim then follows from (77). This corollary will be useful in estimating the operator T when acting on objects such as ψb1 P Qwhich can be expressed as linear combinations of truncated versions of b1 P. Similarly when estimating T∗on objects such as ψb2 P Q. Step 1. Overview of main argument: We now begin the main argument. Let Abe the best constant such that T∗χIPL1(IP)≤A|IP| for all tiles P∈P+. We claim that A=O(1); from this and the corresponding claim for TχIP(which is of course symmetric) the theorem will follow from the local T(1) theorem (Corollary 6.3). In fact we will show IP Tf ≤((1 −ε)A+O(1))|IP|f∞ (78) Trees, Extrapolation, T(b)309 for all tiles P∈P+and f∈S(IP), and some 0 <ε1 depending only on the implicit constant in (76). By duality this implies that A≤ (1 −ε)A+O(1), which will prove the desired bound on A. Fix P,f. We shall prove the estimate (78) in three steps. Firstly (in Step 2), we decompose fand reduce matters to proving a Carleson measure type estimate on the wavelet coefficients |T∗χIP,ψb1 P Q|2; this argument shall use stopping-time arguments (which we encapsulate as Lemma 6.11) based on b1 Pbut not on b2 P. Then (in Step 3), we decompose χIPand use stopping time arguments (again using Lemma 6.11) based on b2 Pbut not on b1 P. It will be important not to try to handle b1 Pand b2 Pat the same time as we will lose the crucial (1 −ε) packing property of the trees left out by the stopping time algorithm if we do so. The purpose of these stopping arguments is to impose some pseudoaccretivity and other regularity properties on the b1 Pand b2 P. Once we have enough regularity properties, we can then (in Step 4) do an elementary computation to estimate the wavelet coefficients |T∗χIP,ψb1 P Q| pointwise by the quantities which we know to be controlled by hypothesis (see (76) below). Step 2. Pruning the bad tiles of b1 P:We now begin the first of the three steps outlined above. We would like to break up finto linear combinations of the wavelets ψb1 P Q, but we cannot do this for all Qbecause we do not control the strong pseudo-accretivity of b1 P. However, by using Lemma 6.5 and some other selection algorithms we can find a large subtree of Tree(P) for which we can decompose fas desired, modulo acceptable errors: Lemma 6.11. Let P∈P+be a tile. Then we can partition Tree(P)=T1∪Pbuffer ∪ T∈T T where •Tis a collection of disjoint complete trees in Tree(P)whose tops form a (1 −ε)-packing of Tree(P)for some 0<ε1(depending only on the implicit constant in (76)); •T1is a tree with top Psuch that b1 Pis strongly pseudo-accretive on T1(with constants perhaps depending on ε); •Pbuffer is a 2-packing of Tree(P),T1∪Pbuffer is convex, b1 Pis pseudo-accretive on T1∪Pbuffer and we have the mean bounds |b1 P|2+|Tb1 P|2mean∗(T1∪Pbuffer)1(79) (with the implicit constant depending on ε). 310 Auscher, Hofmann, Muscalu, Tao, Thiele •We have the decomposition f=[f]Pb1 P+ Q∈T1 Wb1 Pf(Q)ψb1 P Q+ T∈T (fχIT−[f]PTb1 PT)+  Q∈Pbuffer ϕQ (80) whenever f∈S(IP), where the “buffer functions” ϕQare supported on IQ, have mean zero, and take the form ϕQ=aQb1 PχIQl+a Qb1 PχIQr+a Qb1 Ql+a Qb1 Qr where the coefficients aQ,a Q,a Q,a Qdepend on fand the b1 Pand (when |IQ|=2 −M) obey the bounds |aQ|+|a Q|+|a Q|+|a Q|f∞.(81) A similar statement holds with b1 Pand Tb1 Preplaced by b2 Pand T∗b2 P (but the sets T1,Pbuffer and Tare different then). The tree T1represents the “good” portion of the tree Tree(P), in which b1 Pis neither too large nor too small (so in particular the Wb1 P wavelet system is well-behaved on T1). The buffer tiles Pbuffer are those tiles immediately above T1(and are thus slightly less “good”), while the remaining trees Thave no good properties at all, except that they only occupy at most (1 −ε) of the tree Tree(P). This decomposition shares many features in common with Lemma 3.7 (for instance, the trees Tare formed from those intervals where bis too “heavy” or too “light”). In terms of the phase plane (but adapted to the Wb1 Pwavelet system instead of the Haar wavelet system), one can interpret the right-hand side of (80) as follows. The first term corresponds to the region of phase space below the tree T1. The second term corresponds to T1itself. The third term corresponds to the region above T1∪Pbuffer, while the last term is an error term corresponding to the region Pbuffer. In the model case b1 P=χIP, (80) simplifies to f=[f]PχIP+Π T1f+ T∈T ΠTf+ Q∈Pbuffer Wf(Q)φQ. Proof: We begin by applying Lemma 6.5 to Tree(P) to find a preliminary collection T0of disjoint convex trees in Tree(P) such that the tops of T0area(1−2ε)-packing, and such that b1 Pis pseudo-accretive (with constants depending on ε) on the tree T2:= Tree(P)\ T∈T0 T. Trees, Extrapolation, T(b)311 However we do not yet have (79). To obtain these bounds we let Q denote the set of all tiles Q∈T2for which |b1 P|2+|Tb1 P|2mean(Q)≥C/ε and which are maximal with respect to ≤. If the constant Cis chosen large enough, then Qis a ε-packing of Tree(P). Thus if we define T:= T0∪ Q∈Q (Tree(Q)∩T2) then we see that (79) holds on the tree T3:= Tree(P)\ T∈T T, while the tops of Tare still a (1 −ε)-packing of Tree(P). We now perform one minor modification to Tto make Tsiblingfree. If Tcontains two trees whose tops PT,PT are siblings, we can concatenate these trees and add a new tile 2PT=2PT to join these trees to a larger tree without affecting the (1 −ε)-packing nature of the tree tops. Repeating this process as often as necessary (it must terminate since Tree(P) only has a finite number of tiles) we can make Tsiblingfree. For similar reasons we may assume that the trees20 in Tare complete, since we can always replace an incomplete tree by the completion of that tree, absorbing any sub-trees that were also in Tif necessary. We now define21 Pbuffer to be the set of tiles Qin T3such that one or both22 of the children Ql,Qrof Qare not in T3. Since T3is a convex tree, the children of Qwho are not in T3must have disjoint spatial supports as Qvaries in Pbuffer. This implies that Pbuffer is a 2-packing. We now set T1:= T3\Pbuffer. Note that all children of tiles in T1lie in T1∪Pbuffer so that b1 Pis strongly pseudo-accretive on T1, but is merely pseudo-accretive on T1∪Pbuffer. 20Alternatively, we could avoid these modifications by combining the stopping time argument here with the one in Lemma 6.5. 21The algorithm here is extremely similar to the one used to prove Theorem 3.5. Indeed, one can even re-use Figure 3. The trees T0are the “light” trees where b1 P has too small a mean; the tiles Qcorrespond to the circled “heavy” tiles, where b1 P or Tb1 Phas too large an L2norm. The tiles Pbuffer are thus the buffer tiles, which are the ones just below the heavy or light tiles, as well as the tiles at the very finest scale. 22Of course, because we made Tsibling-free, the only way both the children of Qfail to be in T3is if Qis at the finest scale, i.e. if |IQ|=2 −M. 312 Auscher, Hofmann, Muscalu, Tao, Thiele The only property left to verify is the decomposition (80). If fis a constant multiple of b1 Pthen only the first term is non-zero (thanks to (66)) and the claim is easily verified. By subtracting off a constant multiple we may thus assume that fhas mean zero on IP. It will suffice to prove the identity assuming that [b1 P]Q,[b1 P]Ql,[b1 P]Qr= 0 for all Q∈Tree(P), since the general case then follows by an obvious limiting argument (the bounds (81) will not depend quantitatively on the above condition). In this case (70) applies23. Comparing this with (80) and using the mean zero condition, we reduce to showing that  Q∈Pbuffer Wb1 Pf(Q)ψb1 P Q+ T∈T Q∈T Wb1 Pf(Q)ψb1 P Q = T∈T (fχIT−[f]PTb1 PT)+  Q∈Pbuffer ϕQ for suitable ϕQ. Let Q∈Pbuffer. First suppose that neither child of Qis a top of a tree in T(since Tree(P) is complete, this can only happen when |IQ|=2 −M). In this case we simply set ϕQ:= Wb1 Pf(Q)ψb1 P Q. Now suppose that one child of Qis a top of a tree Tin T; without loss of generality we assume Qlis such a top. Since Tis sibling-free, Qr is not a top and must therefore lie in T1∪Pbuffer. In particular we have the lower bounds |[b1 P]Qr|,|[b1 P]Q|1.(82) We do not have good lower bounds on |[b1 P]Ql|, but fortunately we can take advantage of some “wiggle room” in the buffer, and avoid using this in our computations by exploiting the identity Wb1 Pf(Q)ψb1 P Q+ Q∈T Wb1 Pf(Q)ψb1 P Q = Q∈Tree(Q) Wb1 Pf(Q)ψb1 P Q− Q∈Tree(Qr) Wb1 Pf(Q)ψb1 P Q. 23One can easily verify (either by a dimension counting argument, or by inductively working from the finest scale upwards) that the wavelets ψb1 P Qfor Q∈Tree(P) span S0(P). Trees, Extrapolation, T(b)313 The function Q∈Tree(Q)Wb1 Pf(Q)ψb1 P Qclearly is supported on IQ and has mean zero, while f− Q∈Tree(Q) Wb1 Pf(Q)ψb1 P Q= Q∈Tree(P)\Tree(Q) Wb1 Pf(Q)ψb1 P Q is a constant multiple of b1 Pon IT. Thus we have  Q∈Tree(Q) Wb1 Pf(Q)ψb1 P Q=fχIQ−[f]Q [b1 P]Q b1 PχIQ. Similarly we have  Q∈Tree(Qr) Wb1 Pf(Q)ψb1 P Q=fχIQr−[f]Qr [b1 P]Qr b1 PχIQr. Subtracting the two we thus see that Wb1 Pf(Q)ψb1 P Q+ Q∈T Wb1 Pf(Q)ψb1 P Q=fχIT−[f]Q [b1 P]Q b1 PχIQ+[f]Qr [b1 P]Qr b1 PχIQr. If we thus define ϕQ:= [f]PTb1 PT−[f]Q [b1 P]Q b1 PχIQ+[f]Qr [b1 P]Qr b1 PχIQr we see that (81) follows; the mean zero condition can be seen by (75) and inspection. This completes the proof of Lemma 6.11. Roughly speaking, the above lemma states that we can find a large tree T1on which b1 Pis pseudo-accretive, on which b1 Pand Tb1 Pare effectively bounded, and for which we have a representation of the form (70). We now run an argument in the spirit of Lemma 3.1 to localize matters exclusively to this tree T1. We apply the above lemma first with the b1 P. We decompose fusing (80), thus estimating the left-hand side of (78) by the sum of the term below T1 IP [f]PTb1 P ,(83) the terms coming from T1  Q∈T1 Wb1 Pf(Q)IP Tψb1 P Q ,(84) 320 Auscher, Hofmann, Muscalu, Tao, Thiele Notice that b1,b2are only assumed to be in BMO27 rather than L∞. This generalization of the standard T(b) theorem appears to be new. Also observe that the above argument also works in the special case T(b1)=T ∗(b2) = 0 if we drop the para-accretivity and BMO hypotheses on b1,b2and instead impose the reverse H¨older conditions (1 |IP|IP|b1|2)1/2 |[b1]P|,(1 |IP|IP|b2|2)1/2 |[b2]P|1 on b1,b2. It is in fact likely that we can obtain a T(b)-type theorem for arbitrary (complex) dyadic A∞weights b1,b2(see [28]), but we will not attempt to give the most general statements here. References [1] P. Auscher, S. Hofmann, A. McIntosh and P. Tchamitchian, The Kato square root problem for higher order elliptic operators and systems on Rn, dedicated to the memory of Tosio Kato, J. Evol. Equ. 1(4) (2001), 361–385. [2] P. Auscher, S. Hofmann. M. Lacey, A. McIntosh and P. Tchamitchian, The solution of the Kato square root problem for second order elliptic operators on Rn,Ann. of Math. (2) (to appear). [3] P. Auscher, S. Hofmann, J. L. Lewis and P. Tchamitchian, Extrapolation of Carleson measures and the analyticity of Kato’s square-root operators, Acta Math. 187(2) (2001), 161–190. [4] P. Auscher and P. Tchamitchian, Bases d’ondelettes sur les courbes corde-arc, noyau de Cauchy et espaces de Hardy associ´es, Rev. Mat. Iberoamericana 5(3–4) (1989), 139–170. [5] P. Auscher and P. Tchamitchian, Square root problem for divergence operators and related topics, Ast´erisque 249 (1998), 172 pp. [6] J. Bergh and J. L¨ ofstr¨ om,“Interpolation spaces. An introduction”, Grundlehren der Mathematischen Wissenschaften 223, Springer-Verlag, Berlin-New York, 1976. [7] G. Beylkin, R. R. Coifman and V. Rokhlin, Fast wavelet transforms and numerical algorithms. I, Comm. Pure Appl. Math. 44(2) (1991), 141–183. 27For untruncated Calder´on-Zygmund operators in the limit M→∞it may well be necessary to require b1,b2to be in VMO rather than BMO. Trees, Extrapolation, T(b)321 [8] C. J. Bishop, L. Carleson, J. B. Garnett and P. W. Jones, Harmonic measures supported on curves, Pacific J. Math. 138(2) (1989), 233–236. [9] C. P. Calder´ on, On commutators of singular integrals, Studia Math. 53(2) (1975), 139–174. [10] L. Carleson, Lennart On convergence and growth of partial sumas of Fourier series, Acta Math. 116 (1966), 135–157. [11] M. Christ,AT(b) theorem with remarks on analytic capacity and the Cauchy integral, Colloq. Math. 60/61(2) (1990), 601–628. [12] M. Christ and J.-L. Journ´ e, Polynomial growth estimates for multilinear singular integral operators, Acta Math. 159(1–2) (1987), 51–80. [13] R. R. Coifman, A. McIntosh and Y. Meyer, L’int´egrale de Cauchy d´efinit un op´erateur born´e sur L2pour les courbes lipschitziennes, Ann. of Math. (2) 116(2) (1982), 361–387. [14] R. R. Coifman and Y. Meyer, On commutators of singular integrals and bilinear singular integrals, Trans. Amer. Math. Soc. 212 (1975), 315–331. [15] R. R. Coifman and Y. Meyer, Commutateurs d’int´egrales singuli`eres et op´erateurs multilin´eaires, Ann. Inst. Fourier (Grenoble) 28(3) (1978), 177–202. [16] R. R. Coifman and Y. Meyer, Fourier analysis of multilinear convolutions, Calder´on’s theorem, and analysis of Lipschitz curves, in: “Euclidean harmonic analysis” (Proc. Sem., Univ. Maryland, College Park, Md., 1979), Lecture Notes in Math. 779, Springer, Berlin, 1980, pp. 104–122. [17] R. R. Coifman and Y. Meyer, Au del`a des op´erateurs pseudodiff´erentiels, Ast´erisque 57 (1978), 185 pp. [18] R. R. Coifman and Y. Meyer, Nonlinear harmonic analysis, operator theory and P.D.E., in: “Beijing lectures in harmonic analysis” (Beijing, 1984), Ann. of Math. Stud. 112, Princeton Univ. Press, Princeton, NJ, 1986, pp. 3–45. [19] R. R. Coifman and Y. Meyer,“Ondelettes et op´erateurs. III. Op´erateurs multilin´eaires”, Actualit´es Math´ematiques, Hermann, Paris, 1991. [20] R. R. Coifman, P. W. Jones and S. Semmes, Two elementary proofs of the L2boundedness of Cauchy integrals on Lipschitz curves, J. Amer. Math. Soc. 2(3) (1989), 553–564. [21] G. David,“Wavelets and singular integrals on curves and surfaces”, Lecture Notes in Mathematics 1465, Springer-Verlag, Berlin, 1991. 322 Auscher, Hofmann, Muscalu, Tao, Thiele [22] G. David and J.-L. Journ´ e, A boundedness criterion for generalized Calder´on-Zygmund operators, Ann. of Math. (2) 120(2) (1984), 371–397. [23] G. David, J.-L. Journ´ e and S. Semmes,Op´erateurs de Calder´on-Zygmund, fonctions para-accr´etives et interpolation, Rev. Mat. Iberoamericana 1(4) (1985), 1–56. [24] G. David and S. Semmes,“Analysis of and on uniformly rectifiable sets”, Mathematical Surveys and Monographs 38, American Mathematical Society, Providence, RI, 1993. [25] C. Fefferman, Pointwise convergence of Fourier series, Ann. of Math. (2) 98 (1973), 551–571. [26] C. Fefferman and E. M. Stein, Some maximal inequalities, Amer. J. Math. 93 (1971), 107–115. [27] C. Fefferman and E. M. Stein,Hpspaces of several variables, Acta Math. 129(3–4) (1972), 137–193. [28] R. A. Fefferman, C. E. Kenig and J. Pipher, The theory of weights and the Dirichlet problem for elliptic equations, Ann. of Math. (2) 134(1) (1991), 65–124. [29] J. B. Garnett and P. W. Jones, BMO from dyadic BMO, Pacific J. Math. 99(2) (1982), 351–371. [30] L. Grafakos and N. J. Kalton, The Marcinkiewicz multiplier condition for bilinear operators, Studia Math. 146(2) (2001), 115–156. [31] S. Hofmann, M. Lacey and A. McIntosh, The solution of the Kato problem for divergence form elliptic operators with Gaussian heat kernel bounds, Ann. of Math. (2) (to appear). [32] S. Hofmann and J. L. Lewis, The Dirichlet problem for parabolic operators with singular drift terms, Mem. Amer. Math. Soc. 151(719) (2001), 113 pp. [33] S. Janson, On interpolation of multilinear operators, in: “Function spaces and applications” (Lund, 1986), Lecture Notes in Math. 1302, Springer, Berlin, 1988, pp. 290–302. [34] F. John, Quasi-isometric mappings, in: “Seminari 1962/63 Anal. Alg. Geom. e Topol.”, vol. 2, Ist. Naz. Alta Mat., Ediz. Cremonese, Rome, 1965, pp. 462–473. [35] P. W. Jones, Square functions, Cauchy integrals, analytic capacity, and harmonic measure, in: “Harmonic analysis and partial differential equations” (El Escorial, 1987), Lecture Notes in Math. 1384, Springer, Berlin, 1989, pp. 24–68. [36] P. W. Jones, Rectifiable sets and the traveling salesman problem, Invent. Math. 102(1) (1990), 1–15. Trees, Extrapolation, T(b)323 [37] N. H. Katz, Maximal operators over arbitrary sets of directions, Duke Math. J. 97(1) (1999), 67–79. [38] N. H. Katz, Remarks on maximal operators over arbitrary sets of directions, Bull. London Math. Soc. 31(6) (1999), 700–710. [39] M. Lacey and C. Thiele,Lpestimates on the bilinear Hilbert transform for 2 <p<∞,Ann. of Math. (2) 146(3) (1997), 693–724. [40] M. Lacey and C. Thiele, On Calder´on’s conjecture, Ann. of Math. (2) 149(2) (1999), 475–496. [41] M. Lacey and C. Thiele, A proof of boundedness of the Carleson operator, Math. Res. Lett. 7(4) (2000), 361–370. [42] J. L. Lewis and M. A. M. Murray, The method of layer potentials for the heat equation in time-varying domains, Mem. Amer. Math. Soc. 114(545) (1995), 157 pp. [43] A. McIntosh and Y. Meyer, Alg`ebres d’op´erateurs d´efinis par des int´egrales singuli`eres, C. R. Acad. Sci. Paris S´er. I Math. 301(8) (1985), 395–397. [44] Y. Meyer,“Ondelettes et op´erateurs. I. Ondelettes”, Actualit´es Math´ematiques, Hermann, Paris, 1990. [45] C. Muscalu, T. Tao and C. Thiele, Multi-linear operators given by singular multipliers, J. Amer. Math. Soc. 15(2) (2002), 469–496 (electronic). [46] C. Muscalu, T. Tao and C. Thiele,Lpestimates for the “biest”, submitted to Math. Ann. [47] F. Nazarov, S. Treil and A. Volberg, Accretive system Tbtheorems on nonhomogeneous spaces, Duke Math. J. (to appear). [48] F. Nazarov, S. Treil and A. Volberg,Tbtheorems on nonhomogeneous spaces, Acta Math. (to appear). [49] M. C. Pereyra, Lecture notes on dyadic harmonic analysis, in: “Second Summer School in Analysis and Mathematical Physics” (Cuernavaca, 2000), Contemp. Math. 289, Amer. Math. Soc., Providence, RI, 2001, pp. 1–60. [50] S. Semmes, Square function estimates and the T(b) theorem, Proc. Amer. Math. Soc. 110(3) (1990), 721–726. [51] E. M. Stein,“Harmonic analysis: real-variable methods, orthogonality, and oscillatory integrals”, Princeton Mathematical Series 43, Monographs in Harmonic Analysis III, Princeton University Press, Princeton, NJ, 1993. 324 Auscher, Hofmann, Muscalu, Tao, Thiele [52] C. Thiele, Time-frequency analysis in the discrete phase plane, in: “Topics in analysis and its applications”, World Sci. Publishing, River Edge, NJ, 2000, pp. 99–152. [53] C. Thiele, On the Bilinear Hilbert transform, Universit¨at Kiel, Habilitationsschrift (1998). [54] J. Verdera,L2boundedness of the Cauchy integral and Menger curvature, in: “Harmonic analysis and boundary value problems” (Fayetteville, AR, 2000), Contemp. Math. 277, Amer. Math. Soc., Providence, RI, 2001, pp. 139–158. P. Auscher: LAMFA, CNRS UMR 6140 Facult´edeMath´ematiques et d’Informatique Universit´e de Picardie-Jules Verne 80039 Amiens CEDEX 1 France E-mail address:[email protected] S. Hofmann: Mathematics Department University of Missouri Columbia MO, 65211 U.S.A. E-mail address:[email protected] C. Muscalu: Department of Mathematics UCLA Los Angeles CA 90095-1555 U.S.A. E-mail address:[email protected] T. Tao: Department of Mathematics UCLA Los Angeles CA 90095-1555 U.S.A. E-mail address:[email protected] C. Thiele: Department of Mathematics UCLA Los Angeles CA 90095-1555 U.S.A. E-mail address:[email protected] Trees, Extrapolation, T(b)325 Primera versi´o rebuda el 20 de setembre de 2001, darrera versi´o rebuda el 25 de juny de 2002.