Full text
Statistical Inference for Stochastic Processes (2025) 28:15 https://doi.org/10.1007/s11203-025-09333-w Adaptive exact recovery in sparse nonparametric models Natalia Stepanova1·Marie Turcicova2 Received: 24 July 2024 / Accepted: 18 September 2025 © The Author(s) 2025 Abstract We observe an unknown function of dvariables f(t),t∈[0,1]d, in the Gaussian white noise model of intensity ε>0. We assume that the function fis regular and that it is a sum of k-variate functions, where kvaries from 1 to s(1 ≤s≤d). These functions are unknown to us and only a few of them are nonzero. In this article, we address the problem of identifying the nonzero components of fin the case when d=dε→∞as ε→0andsis either fixed or s=sε→∞,s=o(d)as ε→∞. This may be viewed as a variable selection problem. We derive the conditions when exact variable selection in the model at hand is possible and provide a selection procedure that achieves this type of selection. The procedure is adaptive to a degree of model sparsity described by the sparsity parameter β∈(0,1).Wealsoderive conditions that make the exact variable selection impossible. Our results augment previous work in this area. Keywords Gaussian white noise ·Sparsity ·Exact selection ·Sharp selection boundary · Hamming risk ·Functional ANOVA model Mathematics Subject Classification Primary: 62G08 ·Secondary: 62H12 ·62G20 1 Introduction Analysis of high-dimensional data has recently become more and more important since large datasets have become increasingly available in every field of research – ranging from genomic sequencing and brain imaging to climate monitoring and social sciences. However, such analysis is often associated with specific phenomena that go beyond classical statistical theory. The main difficulty is that the dimension can be very large. Speaking in mathematical terms, it is natural to assume that the dimension tends to infinity. In this work, we address the problem of adaptive recovery of the sparsity pattern of a multivariate signal observed in the Gaussian white noise, and augment the results obtained earlier in Ingster and Stepanova (2014) and Stepanova and Turcicova (2025). Specifically, we assume that an unknown signal fof dvariables is observed in the Gaussian white noise BMarie Turcicova [email protected] 1School of Mathematics and Statistics, Carleton University, 1125 Colonel By Drive, K1S 5B6 Ottawa, ON, Canada 2Institute of Computer Science, Czech Academy of Sciences, Pod Vodárenskou vˇeží 271/2, 182 00 Prague, Czech Republic 0123456789().: V,-vol 123
15 Page 2 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 model dXε(t)=f(t)dt+εdW(t), t∈[0,1]d,(1) where dW is a d-parameter Gaussian white noise and ε>0 is the noise intensity. The signal fbelongs to a subspace of L2([0,1]d)=Ld 2with inner product (·,·)2and norm ·2that consists of regular enough functions, and we assume that d=dε→∞as ε→0. Consider an operator W:Ld 2→G0taking values in the set G0of centered Gaussian random variables such that if ξ0=W(g1)and η0=W(g2),whereg1,g2∈Ld 2,thencov(ξ0,η 0)=(g1,g2)2. The d-parameter Gaussian white noise dW in model (1) is defined through the operator W by W(g)=[0,1]d g(t)dW(t)∼N(0,g2 2), g∈Ld 2. In particular, if {g}∈Lis an orthonormal basis of Ld 2,thenW(g)∼N(0,1)for ∈L and, for any finite set {g}of the basis functions, the family {W(g)}forms a multivariate standard normal vector. Thus, the centered Gaussian measure on Ld 2determined by Whas a diagonal covariance operator (i.e., the identity operator). Furthermore, let Xε:Ld 2→G be an operator taking values in the set Gof Gaussian random variables such that if ξ= Xε(g1)and η=Xε(g2),whereg1,g2∈Ld 2,thenE(ξ) =(f,g1)2,E(η) =(f,g2)2,and cov(ξ, η) =ε2(g1,g2)2. By “observing the trajectory (1)” we mean observing a realization of the Gaussian field Xε(t),t∈[0,1]d, defined through the operator Xεby Xε(g)=[0,1]d g(t)dXε(t)∼N(f,g)2,ε 2g2 2,g∈Ld 2. In terms of the operators Wand Xε, the stochastic differential equation (1) can be expressed as Xε=f+εW,(2) and “observing the trajectory (2)” means that we observe all normal N(f,g)2,ε 2g2 2 random variables when gruns through Ld 2.Forany f∈Ld 2, the “observation” Xεin model (2) defines the Gaussian measure Pε, fon the Hilbert space Ld 2with mean function f and covariance operator ε2I,whereIis the identity operator (for references, see Giné and Nickl (2016), Ibragimov and Khasminskii (1997), Skorohod (1998)). In addition to regularity constraints, we assume that fhas a sparse structure and consider the problem of recovering exactly the sparsity pattern of fby using the asymptotically minimax approach. 1.1 Sparsity conditions In high-dimensional inference problems like the one under study, the quality of statistical procedures deteriorates very quickly as the dimension increases. One way to avoid the curse of dimensionality stemming from high-dimensional settings is to reduce the “working dimension” of the problem. In the present setup, this can be done by employing functional ANOVA-type decompositions for f. One common functional ANOVA model (see, for example, Lin (2000), Ingster and Suslina (2015)) assumes that the function of dvariablesisasum of functions of one variable (main effects), functions of two variables (two-way interactions), and so on. That is, the function fof dvariables is decomposed as f(t)= u⊆{1,...,d} fu(tu), t=(t1,...,td)∈[0,1]d,tu=(tj)j∈u, 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 3 of 26 15 where the sum is taken over all subsets u⊆{1,...,d}. Owen (1998) noted that the functional ANOVA decomposition is completely analogous to the one used in experimental statistics because the total variance σ2=[0,1]d(f(t)−I)2dt,whereI=[0,1]df(t)dt, can be written as σ2=u⊆{1,...,d}σ2 u,whereσ2 u=[0,1]kf2 u(tu)dtuif u=∅and σ2 ∅=0. In Ingster and Suslina (2015), this decomposition is used to set up nonparametric alternatives to the hypothesis of no signal in a signal detection problem. In Owen (1998)andLin(2000), the functional ANOVA decomposition is used to split the problem of integrating and estimating a multivariate function finto low-dimensional tasks. It is assumed that, in the above functional ANOVA decomposition, fu=const for u=∅ and, in order to guarantee uniqueness, that 1 0 fu(tu)dtj=0,for j∈u, that is, the terms fuare mutually orthogonal in Ld 2. Each function fu(tu)depends only on variables in tuand describes the “interaction” between these variables. Denote by #(u)the number of elements of a set uand let t−u=(tj)j/∈u. One example of the function fu(tu)that satisfies the above orthogonality condition is (see Owen (1998) for details) fu(tu)=[0,1]d−#(u)⎛ ⎜ ⎝f(t)− v⊂u v=u fv(tv)⎞ ⎟ ⎠dt−u. Another example is the function fu(tu)=j∈ufj(tj), where each fjintegrates to zero. In high-dimensional inference problems, alternatively, or in addition to the general functional ANOVA model, we may choose some s(1 ≤s≤d) and consider the expansion f(t)= u⊆{1,...,d},1≤#(u)≤s fu(tu), t∈[0,1]d, where, if sis small relative to d,f(t)is a sum of functions of a small number of variables. The model of interest in this work is obtained from the above orthogonal expansion by assuming additionally that f(t)has a sparse structure. To ensure that the recovery of the d-variate signal fin model (2) is feasible, the sparsity and regularity constraints on fare required. For 1 ≤k≤s,wheresis as above, define Uk,d={uk:uk⊆{1,...,d},#(uk)=k}and observe that #(Uk,d)=d k.If uk={j1,..., jk}∈Uk,d,1≤j1< ... < jk≤d, denote as before tuk=(tj1,...,tjk)∈[0,1]k, and assume that the signal fin model (2)hastheform f(t)= s k=1 uk∈Uk,d ηukfuk(tuk), t∈[0,1]d,(3) where the components fuk,uk∈Uk,d,1≤k≤s,satisfy the orthogonality condition 1 0 fuk(tuk)dtj=0,for j∈uk,(4) the nonrandom quantities ηuk∈{0,1}, called the index variables, mark the component fuk as “active” when ηuk=1 and as “inactive” when ηuk=0, and as ε→0 uk∈Uk,d ηuk=d k1−β (1+o(1)), 1≤k≤s,(5) 123
15 Page 4 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 where β∈(0,1)is an unknown model parameter called the sparsity index. Condition (5) is the sparsity condition that determines the sparsity structure of fin (3). This condition implies that if 1 ≤k<l≤s, the number of k-variate terms on the right-hand side of (3) that are active is less than the number of l-variate terms that are active. In a simpler setup, when instead of decomposition (3) in model (2)–(5), the signal f(t)is assumed to be a sum of functions, each depending on kvariables (1 ≤k≤d), of the form f(t)= uk∈Uk,d ηukfuk(tuk), t∈[0,1]d,(6) where ηk=(ηuk,uk∈Uk,d)is such that uk∈Uk,dηuk=d k1−β(1+o(1)) as ε→0, the problem of identifying exactly the active components of f(t)was studied in Ingster and Stepanova (2014)fork=1 and in Stepanova and Turcicova (2025)for1≤k≤d.Inthese articles, conditions under which exact identification of the nonzero components of f(t)as in (6) is possible and impossible were established, and adaptive selection procedures that, under some model assumptions, identify exactly all nonzero k-variate components fukwere proposed. In this work, we consider the two cases: (i) when sis fixed and (ii) when s=sε→∞, s=o(d)as ε→0. In both cases, the sparsity condition (5) yields s k=1 uk∈Uk,d ηuk= s k=1d k1−β (1+o(1)) =d s1−β (1+o(1)), ε →0, that is, only d s1−β(1+o(1)) =o(d s)orthogonal components fukof fin (3) are active and the remaining components are inactive, which implies that the function fis sparse.The values of βthat are close to 1 make the signal fin (3)highly sparse, whereas the values of βthat are close to 0 make it dense. Condition (5), in which 1 ≤s≤d, may be viewed as a natural extension of the sparsity condition uk∈Uk,d ηuk=d k1−β (1+o(1)) ε →0,for some 1 ≤k≤d, employed in Stepanova and Turcicova (2025) to recover the sparsity pattern of function f(t) as in (6) for a chosen k. Define the sets Hk β,d=Hk β,d(ε),1≤k≤s,andHs β,d=Hs β,d(ε) as follows: Hk β,d=ηk=(ηuk,uk∈Uk,d):ηuk∈{0,1},uk∈Uk,d,uk∈Uk,d ηuk=d k1−β (1+o(1)), Hs β,d={η=(η1,...,ηs):ηk∈Hk β,d,1≤k≤s}. Based on the “observation” Xεin model (2)–(5), we wish to identify, with high degree of accuracy, the sparsity pattern of f, that is, we wish to construct a good estimator ˆ η= (ˆ η1,..., ˆ ηs)of η=(η1,...,ηs)∈Hs β,dthat would tell us which terms fukin the sparse decomposition (3) are active. This may be viewed as a variable selection problem, and ˆ ηmay be named a selector. 1.2 Regularity conditions Attempting to provide an asymptotically minimax solution to the problem of recovering the sparsity pattern of a function fin model (2)–(5), we have to assume that the set of 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 5 of 26 15 signals fis not “too large". In this article, we will be interested in periodic Sobolev classes described by means of Fourier coefficients. Such classes are quite common in the literature on nonparametric estimation, signal detection, and variable selection. Namely, following the construction in Ingster and Suslina (2015), for uk∈Uk,d,1≤k≤ d,consider the set ˚ Zuk={=(l1,...,ld)∈Zd:lj=0for j/∈ukand lj= 0for j∈uk}. We also set ˚ Z∅=(0,...,0) d ,˚ Z=Z\{0},˚ Zk=˚ Z×...×˚ Z k ,and note that (for k=0we set Uk,d=∅) Zd=˚ Z∪{0}d= u⊆{1,...,d} ˚ Zu= 0≤k≤d uk∈Uk,d ˚ Zuk. Consider the Fourier basis {φ(t)}∈Zdof Ld 2defined as follows: φ(t)= d j=1 φlj(tj), =(l1,...,ld)∈Zd, φ0(t)=1,φ l(t)=√2cos(2πlt), φ−l(t)=√2sin(2πlt), l>0. (7) Observe that φ(t)=φ(tuk)for ∈˚ Zuk(for u=∅we set φ(tu)=1) and {φ(t)}∈Zd= u⊆{1,...,d}{φ(tu)}∈˚ Zu= 0≤k≤d uk∈Uk,d{φ(tuk)}∈˚ Zuk .(8) Next, let θ(uk)=(fuk,φ )Ld 2be the th Fourier coefficient of fukfor ∈˚ Zuk,uk∈Uk,d, 1≤k≤d. Then, for uk={j1,..., jk}∈Uk,d,where1≤j1< ... < jk≤d,1≤k≤s, the k-variate component fukon the right-hand side of (3) can be expressed as fuk(tuk)= ∈˚ Zuk θ(uk)φ(tuk), and the entire function fin (3) takes the form f(t)= s k=1 uk∈Uk,d ηuk ∈˚ Zuk θ(uk)φ(tuk). (9) Note that only those Fourier coefficients of fthat correspond to the orthogonal components fukare nonzero and that fuk2 2=(fuk,fuk)Ld 2=∈˚ Zuk θ2 (uk). For uk={j1,..., jk}∈Uk,d,where 1 ≤j1<...< jk≤d,1≤k≤s,weassumethat fukbelongs to the Sobolev class of k-variate functions with integer smoothness parameter σ≥1 for which the semi-norm ·σ,2is defined by fuk2 σ,2= k i1=1 ... k iσ=1 ∂σfuk ∂tji1...∂tjiσ 2 2 .(10) Under the periodic constraint, we can define the semi-norm · σ,2for the general case σ>0 in terms of the Fourier coefficients θ(uk),∈˚ Zuk. For this, assume that fuk(tuk) 123
15 Page 6 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 admits 1-periodic [σ]-smooth extension in each argument to Rk, i.e., for all derivatives f(n) uk of integer order 0 ≤n≤[σ],where f(0) uk=fuk, one has f(n) uk(tj1,...,tji−1,0,tji+1,...,tjk)=f(n) uk(tj1,...,tji−1,1,tji+1,...,tjk), 2≤i≤k−1, with obvious extension for i=1,k. Then, the expression in (10) corresponds to fuk2 σ,2= ∈˚ Zuk θ2 (uk)c2 ,c2 =⎛ ⎝ d j=1 (2πlj)2⎞ ⎠ σ =k i=1 (2πlji)2σ .(11) Next, denote by Fcukthe Sobolev ball of radius 1 with coefficients cuk=(c)∈˚ Zuk ,thatis, Fcuk=⎧ ⎪ ⎨ ⎪ ⎩ fuk(tuk)= ∈˚ Zuk θ(uk)φ(tuk), tuk∈[0,1]k: ∈˚ Zuk θ2 (uk)c2 ≤1⎫ ⎪ ⎬ ⎪ ⎭ and assume that every component fukof fin (3) belongs to this Sobolev ball, that is, fuk∈Fcuk,uk∈Uk,d,1≤k≤s.(12) Thus, the model under study is specified by equations (2)–(5)and(12). The collections cuk=(c)∈˚ Zuk ,uk∈Uk,d, are invariant with respect to a particular choice of the elements of uk∈Uk,dfor all 1 ≤k≤d. Hence, if uk={j1,j2,..., jk},1≤j1< ... < jk≤d,and∈˚ Zuk, we can write c=clj1,...,ljk,0,...,0. Thinking of a transformation from ˚ Zkto ˚ Zukthat maps =(l 1,...,l k)∈˚ Zkto =(lj1,...,ljk,0,...,0)∈˚ Zuk according to the rule lj=0for j/∈ukand ljp=lpfor p=1,...,k,we may, if needed, represent each collection cuk=(c)∈˚ Zuk as ck=(c)∈˚ Zk. Clearly, the Sobolev balls Fcuk are isomorphic for all ukof cardinality k(1 ≤k≤d). 1.3 Problem statement Based on the “observation” Xεin model (2)–(5)and(12), we wish to identify, with a high degree of accuracy, the active components fukof f. This leads to the problem of obtaining an estimator ˆ η=(ˆ η1,..., ˆ ηs)for η=(η1,...,ηs)∈Hs β,d,whichmaybeviewedasa variable selection problem, since ˆ ηselects important variables tj,j=1,...,d,andtheir relevant interactions in the ANOVA-type expansion (3). The estimator ˆ ηis named a selector since it serves to select the active components of fin decomposition (3). Identifying the active components of fin model (2)–(5)and(12) is feasible when the components fukof fare not “too small”, i.e., they are separated from zero in an appropriate way. Therefore, for a given uk∈Uk,d,1≤k≤sand r>0, we define the set ˚ Fcuk(r)={fuk∈Fcuk:fuk2≥r} and consider testing H0,uk:fuk=0vs.Hε 1,uk:fuk∈˚ Fcuk(rε,k), uk∈Uk,d,1≤k≤s,(13) where rε,k>0, 1 ≤k≤s,andmax 1≤k≤srε,k→0asε→0. To ensure that exact recovery of the sparsity pattern of f(in the sense defined precisely at the end of this section) is possible, we require the components fuk,uk∈Uk,d,1≤k≤s,tobe 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 7 of 26 15 selectable, that is, we require the family (indexed by ε)ofcollectionsrε={rε,k,1≤k≤s} in the hypothesis testing problem (13) to be above the so-called sharp selection boundary. Establishing the sharp selection boundary that makes the exact recovery of fin model (2)–(5) and (12) feasible is one of the goals in this article. For this purpose, the hypothesis testing problem (13) will be employed as an auxiliary problem. Since the Sobolev balls Fcukare isomorphic for all ukof cardinality k,1≤k≤s,the problem of testing s k=1d kpairs of the null and alternative hypotheses in (13) reduces to that of testing H0,uk:fuk=0vs.Hε 1,uk:fuk∈˚ Fcuk(rε,k), for some uk∈Uk,d,1≤k≤s,(14) and the number of tests one needs to carry out decreases from s k=1d kto s.Forevery chosen uk,1≤k≤s, the hypotheses H0,ukand Hε 1,uk,1≤k≤s, separate asymptotically (i.e., there exists a consistent test procedure for testing H0,ukversus Hε 1,uk,1≤k≤s)when the quantity min1≤k≤srε,kis not “too small”. The sharp selection boundary in the problem at hand could be described in terms of min1≤k≤srε,k, however, the construction of an adaptive selection procedure in Section 2suggests that it is more natural to describe the sharp selection boundary in terms of min1≤k≤saε,uk(rε,k)/%log d k,whereaε,uk(rε,k)is defined below by (22). Let rε={rε,k,1≤k≤s}be the same family of collections as in (13). Define the class of sparse multivariate functions of our interest by Fβ,σ s,d(rε)=&f:f(t)= s k=1 uk∈Uk,d ηukfuk(tuk), fuksatisfies (4), fuk∈˚ Fcuk(rε,k), uk∈Uk,d,ηk=(ηuk)uk∈Uk,d∈Hk β,d,1≤k≤s', where the dependence of Fβ,σ s,d(rε)on the smoothness parameter σis hidden in the coefficients cuk=(c)∈˚ Zuk defined in (11). First, we find the sharp selection boundary that allows us to verify whether the components of a signal f∈Fβ,σ s,d(rε)are all selectable. Next, under the assumption that the components fuk,uk∈Uk,d,1≤k≤s,areselectable,weconstructa selector ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤s)that recovers exactly the sparsity pattern of fin model (2)–(5)and(12) in the sense that for all β∈(0,1)and σ>0 lim sup ε→0 sup f∈Fβ,σ s,d(rε) Eε, f|ˆ η−η|=0,(15) where Eε, fis the expectation with respect to the measure Pε, f, and the expression Eε, f|ˆ η−η|=Eε, f⎛ ⎝ s k=1 uk∈Uk,d|ˆηuk−ηuk|⎞ ⎠ is the Hamming risk of ˆ ηas an estimator of η. The Hamming risk represents the average Hamming loss |ˆ η−η|:=s k=1uk∈Uk,d|ˆηuk−ηuk|, which counts the number of positions at which ˆ ηand ηdiffer. Finally, we show that if at least one of the components fukis not selectable, then exact recovery of the sparsity pattern of f∈Fβ,σ s,d(rε)is impossible. A related measure of risk that is often used in the literature on variable selection is the probability of wrong recovery,P ε, f(ˆ η= η), that determines the chance that a selector ˆ η 123
15 Page 8 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 will not agree with η. Most literature on variable selection in high dimensions focuses on constructing selectors such that the probability of wrong recovery is close to zero in some asymptotic sense (see, for example, Wainwright (2009), Wasserman and Roeder (2009), Zhang (2010), Zhao and Yu (2006)). Although the probability of wrong recovery has been used extensively by other authors, the Hamming risk is a more general measure of risk since, by means of Markov’s inequality, Pε, f(ˆ η= η)=Pε, f(|ˆ η−η|≥1)≤Eε, f|ˆ η−η|. In view of the last inequality, if the Hamming risk tends to zero, the probability of wrong recovery is guaranteed to tend to zero. For this reason, we only consider the Hamming risk as a measure of selector performance. 1.4 Sparse Gaussian sequence space model We shall study the recovery problem at hand in the sequence space of Fourier coefficients of fand name it the variable selection problem. A Gaussian sequence space model is equivalent to the corresponding Gaussian white noise model but is more convenient to deal with since it is written in terms of the Fourier coefficients. Let {φ(t)}∈Zdbe the orthonormal basis of Ld 2as in (7), and let θ(uk)=(fuk,φ )Ld 2 be the th Fourier coefficient of fukfor ∈˚ Zuk,uk∈Uk,d,1≤k≤d. Denote by X=Xε(φ)=[0,1]dφ(t)dXε(t),∈˚ Zuk,uk∈Uk,d,1≤k≤s,theth empirical Fourier coefficient. Then, in view of (4)and(8), the sequence space model that corresponds to model (2)–(5)and(12) takes the form X=(f,φ )2+εW(φ)= s k=1 vk∈Uk,d ηvk(fvk,φ )2+εW(φ) =ηukθ(uk)+εξ,∈˚ Zuk,uk∈Uk,d,1≤k≤s,(16) where ηk=(ηuk,uk∈Uk,d)∈Hk β,d,ξ=W(φ)=[0,1]dφ(t)dW(t)are iid standard normal random variables for all ∈˚ Zuk,uk∈Uk,d,1≤k≤s,θuk=(θ(uk), ∈˚ Zuk) consists of the Fourier coefficients θ(uk)=(fuk,φ )Ld 2and belongs to the ellipsoid cuk=&θuk=(θ(uk), ∈˚ Zuk)∈l2(Zd): ∈˚ Zuk θ2 (uk)c2 ≤1'. Next, for uk∈Uk,dand rε,k>0, define ˚ cuk(rε,k)=&θuk∈cuk: ∈˚ Zuk θ2 (uk)≥r2 ε,k',(17) and note that ˚ cuk(rε,k)=∅when rε,k>1/cε,0,where, recalling (11), cε,0:= inf∈˚ Zuc= (2π)σkσ/2. Therefore, we are only interested in the case when rε,k∈(0,(2π)−σk−σ/2). In the sequence space of Fourier coefficients, the hypothesis testing problem (14) becomes that of testing H0,uk:θuk=0vs.Hε 1,uk:θuk∈˚ cuk(rε,k), for some uk∈Uk,d,1≤k≤s.(18) 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 9 of 26 15 The collection of hypothesis testing problems in (18) will be used to establish conditions for the possibility of exact variable selection in model (16),andtodesignanadaptiveselection procedure that achieves this type of selection. For the family of collections rε={rε,k,1≤k≤s},rε,k>0, define the set σ s,d(rε)={θ=(θ1,...,θs):θk∈˚ σ k,d(rε,k), 1≤k≤s}, where ˚ σ k,d(rε,k)={θk=(θuk,uk∈Uk,d):θuk∈˚ cuk(rε,k), uk∈Uk,d},1≤k≤s. For uk∈Uk,d,1≤k≤s, denote Xuk={X,∈˚ Zuk}and let ˆ ηk=(ˆηuk,uk∈Uk,d), where ˆηuk=ˆηuk(Xuk)∈{0,1},beanestimatorofηk=(ηuk,uk∈Uk,d)∈Hk β,d.We have previously decided to call an aggregate estimator ˆ η=(ˆ η1,..., ˆ ηs)for η=(η1,...,ηs)∈ Hs β,daselector. When dealing with the problem of identifying nonzero ηukin model (16), the maximum Hamming risk of a selector ˆ ηcan be expressed as Rε,s(ˆ η):= sup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|ˆ η−η|, where Eθ,ηis the expectation with respect to the joint distribution of Xuk={X,∈˚ Zuk}, uk∈Uk,d,1≤k≤s, in model (16), and Eθ,η|ˆ η−η|=Eθ,η⎛ ⎝ s k=1 uk∈Uk,d|ˆηuk−ηuk|⎞ ⎠ is the Hamming risk of ˆ η. A “good” selector is the one that consistently chooses the nonzero ηukas ε→0. With this in mind, an exact selector ˆ ηis defined such that lim supε→0Rε,s(ˆ η)= 0forallβ∈(0,1)and σ>0. If such a selector exists, it is said that exact selection in model (16) is possible. 1.5 Summary of the main results Consider the problem of exact identification of the nonzero index variables ηukin the sequence space model (16), and recall the family of collections rε={rε,k,1≤k≤s}that determines the sets ˚ cuk(rε,k),uk∈Uk,d,1≤k≤s,asdefinedin(17). The main contribution of this work is the derivation of the sharp selection boundary, stated in terms of rε={rε,k,1≤ k≤s}, which defines when exact selection of the nonzero index variables ηukin model (16) is possible and when this type of selection is impossible. This boundary is determined by inequalities (36)and(40) bellow. Then, we construct an adaptive (independent of the sparsity index β) estimator ˆ ηof η∈Hs β,dattaining this boundary with the property lim sup ε→0 sup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|ˆ η−η|=0,(19) which holds for all β∈(0,1)and σ>0. We also show that if the family of collections rε={rε,k,1≤k≤s}falls below the selection boundary, one has, for β∈(0,1)and σ>0, lim inf ε→0inf ˜ ηsup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|˜ η−η|>0,(20) 123
15 Page 16 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 Theorem 1 Let s ∈{1,...,d},β∈(0,1), and σ>0be fixed numbers, and let d =dε→∞ and log d s=olog ε−1as ε→0. Let the family of collections rε={rε,k,1≤k≤s}, rε,k>0, be such that lim inf ε→0min 1≤k≤s aε,uk(rε,k) %log d k>√2(1+(1−β). (36) Then the selector ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤s)given by (33)and (34)satisfies lim sup ε→0 sup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|ˆ η−η|=0. Proof Denote the threshold in (33)bytε,k=*(2+)log d k+log Mk,1≤k≤s. When ηuk=0 (respectively ηuk=1), we shall write P0and E0(respectively Pθukand Eθuk) for Pθuk,ηukand Eθuk,ηuk. For all small enough ε, the maximum Hamming risk Rε,s(ˆ η)of the selector ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤s)satisfies Rε,s(ˆ η)=sup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|ˆ η−η|= sup η∈Hs β,d sup θ∈σ s,d(rε) s k=1 uk∈Uk,d Eθuk,ηuk|ˆηuk−ηuk| =sup η∈Hs β,d sup θ∈σ s,d(rε) s k=1⎛ ⎝ uk:ηuk=0 E0(ˆηuk)+ uk:ηuk=1 Eθuk(1−ˆηuk)⎞ ⎠ ≤sup η∈Hs β,d s k=1 uk:ηuk=0 E0max 1≤m≤Mk 1Suk,m>tε,k +sup η∈Hs β,d sup θ∈σ s,d(rε) s k=1 uk:ηuk=1 Eθuk1−max 1≤m≤Mk 1Suk,m>tε,k ≤ s k=1 sup ηk∈Hk β,d uk:ηuk=0 Mk m=1 P0Suk,m>tε,k + s k=1 sup ηk∈Hk β,d uk:ηuk=1 sup θuk∈˚ cuk(rε,k) Eθuk⎧ ⎨ ⎩ 1⎛ ⎝ Mk + m=1{Suk,m≤tε,k}⎞ ⎠⎫ ⎬ ⎭ ≤ s k=1d kMk m=1 P0Suk,m>tε,k +2 s k=1d k1−β sup θuk∈˚ cuk(rε,k) min 1≤m≤Mk PθukSuk,m≤tε,k.(37) By taking the index m0,k(1 ≤m0,k≤Mk) to be such that β∈(βm0,k,k,β m0,k+1,k]for 1≤k≤s, we can further estimate the second term on the right-hand side of (37) and arrive at 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 17 of 26 15 Rε,s(ˆ η)≤ s k=1d kMk m=1 P0Suk,m>tε,k+2 s k=1d k1−β sup θuk∈˚ cuk(rε,k) PθukSuk,m0,k≤tε,k =:I(1) ε,s+I(2) ε,s.(38) Similar to how this is done in the proof of Theorem 3.1 in Stepanova and Turcicova (2025), i.e., applying the exponential upper bound (31) to the probability P0Suk,m>tε,k,inview of (34) and the requirement log Mk=o(log d k),weobtainasε→0 I(1) ε,s≤ s k=1d kMkexp −(1+/2)&log d k+log Mk'(1+o(1)) =Os k=1&Mkd k'−/2=Omax 1≤k≤s&Mkd k'−/2=o(1). (39) Next, under the assumptions of the theorem, applying, as in the proof of Theorem 3.1 in Stepanova and Turcicova (2025), the exponential upper bound (32) to the probability PθukSuk,m0,k≤tε,k=PθukSuk,m0,k−Eθuk(Suk,m0,k)≤tε,k−Eθuk(Suk,m0,k)for those θuk∈˚ cuk(rε,k)for which Eθuk(Suk,m0,k)max ∈˚ Zuk ω(r∗ ε,k,m0,k)=o(1), ε →0, and Chebyshev’s inequality for those θuk∈˚ cuk(rε,k)for which the above limiting relation does not hold true, we obtain that I(2) ε,s=o(1)as ε→0. From this, (38), and (39), lim supε→0Rε,s(ˆ η)=0, which completes the proof. Theorem 1can be extended to the case of growing s. Theorem 2 Let β∈(0,1)and σ>0be fixed numbers, and let d =dε→∞and s =sε→ ∞be such that s =o(d),log log d=o(s), and log d s=olog ε−1as ε→0.Letthe family of collections rε={rε,k,1≤k≤s},r ε,k>0, be as in Theorem 1. Then the selector ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤s)given by (33)and (35)satisfies lim sup ε→0 sup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|ˆ η−η|=0. Proof The proof goes in part along the same lines as that of Theorem 1, with condition (35) used instead of condition (34). Observe that, when d→∞,k→∞,k=o(d),wehave d k∼dk k!and log d k∼klog(d/k). Therefore, using the conditions s=o(d)and log Mk=o(log d k), the geometric progression formula, and (35), we obtain as ε→0 I(1) ε,s≤ s k=1d kMkexp −(1+/2)&log d k+log Mk'(1+o(1)) ≤ s k=1 exp (−(/4)klog(d/s))= s k=1 (d/s)−(/4)k=O(d/s)−/4=o(1). 123
15 Page 18 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 Also, under the assumptions of Theorem 2, acting as in the proof of Theorem 1, we can obtain that I(2) ε,s=o(1)as ε→0, and hence lim supε→0Rε,s(ˆ η)≤lim supε→0I(1) ε,s+I(2) ε,s=0. The importance of condition log log d=o(s)when s=sε→∞,s=o(d)as ε→0 can be seen from the proof of Theorem 3.3 in Stepanova and Turcicova (2025). Theorems 1and 2show that if a positive family of collections rε={rε,k:1≤k≤s}is such that condition (36) holds, exact variable selection in the sequence space model (16)is possible; this is true for fixed sas well as for s=sε→∞,s=o(d)as ε→0. Theorems 1and 2may be viewed as extensions of Theorems 3.1 and 3.3 in Stepanova and Turcicova (2025), for more details we refer to the discussion of Section 4. 3.2 Lower bound on the minimax risk when sis fixed and when s→∞ We now turn to deriving the conditions when exact variable selection in model (16) is impossible, i.e. when the lower bound (20) holds true, which occurs if the family of collections rε={rε,k,1≤k≤s}in the hypothesis testing problem (18) falls below a certain level. The precise statement is as follows. Theorem 3 Let s ∈{1,...,d},β∈(0,1), and σ>0be fixed numbers, and let d =dε→∞ and log d s=olog ε−1as ε→0. Let the family of collections rε={rε,k,1≤k≤s}, rε,k>0, be such that lim sup ε→0 min 1≤k≤s aε,uk(rε,k) %log d k<√2(1+(1−β). (40) Then lim inf ε→0inf ˜ ηsup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|˜ η−η|>0, where the infimum is taken over all selectors ˜ ηof η∈Hs β,din model (16). Proof For every k(1 ≤k≤s), we choose some uk∈Uk,dandfixit.Letk=k(ε) be a map from (0,∞)to {1,...,s}defined as follows: k=argmin 1≤k≤s aε,uk(rε,k) %log d k. We have Rε,s:=inf ˜ ηsup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|˜ η−η| =inf ˜ ηsup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η⎛ ⎝ s k=1 uk∈Uk,d|˜ηuk−ηuk|⎞ ⎠ ≥inf ˜ ηk sup ηk∈Hk β,d sup θk∈˚ σ k,d(rε,k) Eθk,ηk⎛ ⎝ uk∈Uk,d |˜ηuk−ηuk|⎞ ⎠.(41) 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 19 of 26 15 Since log d s=olog ε−1implies log d k=olog ε−1, it follows from (41)and Theorem 3.2 in Stepanova and Turcicova (2025)that lim inf ε→0Rε,s>0 provided lim sup ε→0 aε,uk(rε,k) %log d k<√2(1+(1−β), which is true by the definition of kand condition (40). This completes the proof. The result of Theorem 3can be extended to the case of growing s. Theorem 4 Let β∈(0,1)and σ>0be fixed numbers, and let d =dε→∞and s =sε →∞be such that s =o(d),log log d=o(s), and log d s=olog ε−1as ε→0.Letthe family of collections rε={rε,k,1≤k≤s},r ε,k>0, be as in Theorem 3.Then lim inf ε→0inf ˜ ηsup η∈Hs β,d sup θ∈σ s,d(rε) Eθ,η|˜ η−η|>0, where the infimum is taken over all selectors ˜ ηof η∈Hs β,din model (16). Proof Using the same arguments as in the proof of Theorem 3, with reference to Theorem 3.4 of Stepanova and Turcicova (2025) instead of reference to Theorem 3.2 of Stepanova and Turcicova (2025), we arrive at the statement of Theorem 4. Theorems 3and 4show that if a positive family of collections rε={rε,k:1≤k≤s}is such that condition (40) holds, exact variable selection in the sequence space model (16)is impossible; this is true for fixed sas well as for s=sε→∞s=o(d)as ε→0.Theorems 3and 4may be viewed as extensions of Theorems 3.2 and 3.4 in Stepanova and Turcicova (2025), for more details we refer to the discussion in Section 4. 4 Discussion By using the notion of optimality of a statistical procedure from the minimax hypothesis testing theory, we arrive at the following important conclusions. When sis fixed, in view of Theorems 1and 3, the selector defined in (33)and(34) is seen to perform optimally with respect to a Hamming loss in the asymptotically minimax sense. Similarly, when s=sε →∞,s=o(d)as ε→0, by means of Theorems 2and 4, the selector defined in (33)and (35) is optimal with respect to a Hamming loss in the asymptotically minimax sense. We now explain our choice to state the sharp selection boundary in terms of the quantities aε,uk(rε,k),1≤k≤s, rather than in terms of the radii rε,k,1≤k≤s. To this end, consider the Gaussian vector model Xuk=μk,dηuk+εuk,uk∈Uk,d,1≤k≤s,(42) where μk,d>0 is the signal, the errors εukare iid standard normal random variables, the deterministic quantities ηuk∈{0,1}are as before and, in particular, satisfy the sparsity condition (5). For ηuk=1, the observation Xukis equal to the signal μk,dperturbed by random noise εuk.Forηuk=0, the observation Xukis merely random noise εuk.Dueto 123
15 Page 20 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 condition (5), only s k=1uk∈Uk,dηuk=s k=1d k1−β(1+o(1)) =d s1−β(1+o(1)) observations in model (42) contain signal, and thus the model is sparse. The sharp selection boundary in model (42) is given by, cf. inequalities (6) and (7) in Stepanova and Turcicova (2025), lim inf d→∞ min 1≤k≤s μk,d %log d k>√2(1+(1−β) and lim sup d→∞ min 1≤k≤s μk,d %log d k<√2(1+(1−β). By comparing the above inequalities to the sharp selection boundary determined by (36)and (40), we conclude that, in the variable selection problem at hand, the quantity aε,uk(rε,k)plays the same role as the signal μk,ddoes in the problem of recovering η=(ηk=(ηuk,uk∈ Uk,d), 1≤k≤s)in model (42). Additionally, the use of aε,uk(rε,k),1≤k≤s,for establishing the sharp selection boundary is natural for comparison purposes, as seen from the discussion below. Let us compare the main results of this article to those stated in Theorems 3.1 to 3.4 of Stepanova and Turcicova (2025) for a simpler model when, instead of decomposition (3) for function fin model (2)–(5)and(12), one has decomposition (6). The sharp selection boundary for the latter (simpler) model is stated in terms of aε,uk(rε,k)as in (22) and is given by the inequalities, cf. inequalities (36)and(40), lim inf ε→0 aε,uk(rε,k) %log d k>√2(1+(1−β) and lim sup ε→0 aε,uk(rε,k) %log d k<√2(1+(1−β). Here ukis an arbitrary element of Uk,d,andrε,kdetermines the set ˚ cuk(rε,k)as defined in (17), more precisely, rε,kis a radius of the centered l2-ball removed from the ellipsoid cuk. The above inequalities provide conditions on the l2-norm of (θ(uk))∈˚ Zuk ,uk∈Uk,d,that make exact variable selection in the Gaussian sequence space model studied in Stepanova and Turcicova (2025) possible (the first inequality) and impossible (the second inequality). In terms of the initial model (2)–(5)and(12), with (6) in place of (3), the quantity rε,k determines the set ˚ Fcuk(rε,k)={fuk∈Fcuk:fuk2≥rε,k}, and thus the above inequalities also provide conditions on the L2-norm of the components fukof f=uk∈Uk,dηukfukthat make exact selection of the active (nonzero) components possible and impossible. Note that the sharp selection boundary is the same for fixed kand for k→∞,k=o(d)as ε→0 provided the rate at which ktends to infinity is carefully regulated (see Theorems 3.1 to 3.4 in Stepanova and Turcicova (2025)). This is consistent with the results in Section 3of this article when sis fixed (see Theorems 1and 3)andwhens→∞,s=o(d)as ε→0 (see Theorems 2and 4). In case of growing s, the adaptive selector ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤ s),where ˆηukis defined by (33), is similar to that for fixed s, with a slight difference in the threshold *(2+)log d k+log Mk, as seen from conditions (34)and(35) imposed on . It is also of interest to make a bridge between the sharp selection boundary given by inequalities (36)and(40)andthesharp detection boundary established by Ingster and Suslina (2015) in the problem of testing 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 21 of 26 15 Table 1 The number Nk,dof active components of ffor β=0.87 and selected values of d and k k=1k=2k=3k=4 502345 d1002357 20023610 H0,s:f=0vs.Hε 1,s:f∈ 1≤k≤s uk∈Uk,d ˚ Fcuk(rε,k). As shown in Theorem 3.3 of Ingster and Suslina (2015), the sharp detection boundary is given by the inequalities lim inf ε→0min 1≤k≤s aε,uk(rε,k) %log d k>√2 and lim sup ε→0 min 1≤k≤s aε,uk(rε,k) %log d k<√2.(43) Specifically, when the first inequality in (43) holds, the hypotheses H0,sand Hε 1,sseparate asymptotically (i.e., there exists a consistent test procedure for testing H0,sversus Hε 1,s), whereas, when the second inequality in (43) holds, H0,sand Hε 1,smerge asymptotically (i.e., there is no consistent test procedure for testing H0,sversus Hε 1,s). By comparing the inequalities in (43) to the inequalities (36)and(40), we notice that the sharp selection boundary lies above the sharp detection boundary. This is not unusual because the problem of selecting nonzero components is more difficult than that of testing whether such nonzero components exit. In particular, in order to be selectable, the components need to be detectable. 5 Simulation study To showcase the exact selection procedure based on ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤s) in (33)–(34) in its capacity to recover the sparsity pattern of a signal f∈Fβ,σ s,d(rε)of the form f=s k=1uk∈Uk,dηukfuk, we conduct a small-scale simulation study. For this study, we choose σ=1, s=4, and d=50,100,200. We also choose β=0.87, which means that the signal fis highly sparse. The number of active components Nk,d:= uk∈Uk,dηuk of ffor each kis obtained from relation (5). The values of Nk,dare listed in Table 1. In this simulation study, we choose the active components of fas follows. First, consider the following nine functions defined on [0,1]: g1(t)=t2(2t−1−(t−0.5)2)exp(t)−0.5424, g2(t)=t2(2t−1−(t−1)5)−0.2887, g3(t)=1.5t22t−1cos(15t)−0.05011, g4(t)=t−1/2, g5(t)=5(t−0.7)3+0.29, g6(t)=2(t−0.4)2−0.1867, g7(t)=0.7(t2−0.1)3−0.0643, g8(t)=10(t2−0.5)5+0.068, g9(t)=3(t−0.8)4−0.1968. 123
15 Page 22 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 In order to simplify the notation in this section, instead of indexing the components fukof f by uk, (e.g., f{i,j}for u2={i,j}), we will index them by u(i) kfor i=1,...,d k.Byusing this notation, the active (and inactive) components fukof fon [0,1]k,k=1,...,4, are defined as follows: for k=1:fu(i) 1 (ti)=gi(ti), i=1,2, fu(i) 1 (tu(i) 1 )=0,i=3,...,d, for k=2:fu(i) 2 (ti,ti+1)=gi(ti)gi+1(ti+1), i=1,2,3, fu(i) 2 (tu(i) 2 )=0,i=4,...,d 2, for k=3:fu(i) 3 (t1,t2,ti+2)=g1(t1)g2(t2)gi+2(ti+2), i=1,...,4, for d=50 :fu(i) 3 (tu(i) 3 )=0,i=5,...,d 3, for d=100 :fu(5) 3 (t1,t2,t7)=g1(t1)g2(t2)g7(t7), fu(i) 3 (tu(i) 3 )=0,i=6,...,d 3, for d=200 :fu(i) 3 (t1,t2,ti+2)=g1(t1)g2(t2)gi+2(ti+2), i=5,6, fu(i) 3 (tu(i) 3 )=0,i=7,...,d 3, for k=4:fu(i) 4 (t1,t2,t3,ti+3)=g1(t1)g2(t2)g3(t3)gi+3(ti+3), i=1,...,5, for d=50 :fu(i) 4 (tu(i) 4 )=0,i=6,...,d 4, for d=100 :fu(i) 4 (t1,t2,t4,ti+2)=g1(t1)g2(t2)g4(t4)gi+2(ti+2), i=6,7, fu(i) 4 (tu(i) 4 )=0,i=8,...,d 4, for d=200 :fu(i) 4 (t1,t2,t4,ti+2)=g1(t1)g2(t2)g4(t4)gi+2(ti+2), i=6,7, fu(i) 4 (t1,t2,t5,ti−1)=g1(t1)g2(t2)g5(t5)gi−1(ti−1), i=8,9,10, fu(i) 4 (tu(i) 4 )=0,i=11,...,d 4, where tu(i) k=(tj)j∈u(i) k , as before. When the component functions fu(i) k areasabove,the orthogonality condition (4) holds true up to four decimal places for all chosen values of kand d. The active components are associated with u(i) k={i},i=1,2,for k=1, u(i) k={i,i+1},i=1,2,3,for k=2, u(i) k={1,2,i+2},i=1,...,N3,d,for k=3and d=50,100,200. For k=4, they are associated with u(i) k={1,2,3,i+3},i=1,...,5, when d=50,100,200;u(i) k={1,2,4,i+2},i=6,7,when d=100,200;and u(i) k= {1,2,5,i−1},i=8,9,10,when d=200. Recall the definition of the exact selector ˆ η=(ˆ ηk=(ˆηuk,uk∈Uk,d), 1≤k≤s)with components ˆηukgiven by (33)–(34): 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 23 of 26 15 ˆηuk=max 1≤m≤Mk 1Suk,m>)(2+)log d k+log Mk, where the weights of the statistics Suk,m=∈˚ Zuk ω(r∗ ε,k,m)(X/ε)2−1,1≤m≤Mk, are equal to ω(r∗ ε,k,m)=1 2ε2 (θ∗ (r∗ ε,k,m))2 aε,uk(r∗ ε,k,m),∈˚ Zuk, and the values of θ∗ (rε,k)and aε,uk(rε,k)are computed by using relations (24)and(26). Keeping in mind condition (34) and the requirements on Mk,wetake=1/%log d k, Mk=20, and choose the points βk,m,m=1,...,Mk, to be uniformly placed on the interval [0.001,0.999]. For each kand m,thevaluer∗ ε,k,mis found as a solution of equation (28) with aε,uk(rε,k)as in (26). In order to simulate the observations in model (16), we need to compute the true Fourier coefficients θ(uk), ∈˚ Zuk, in decomposition (9). In what follows, ukstands for u(i) kfor i=1,...,Nk,d. When k=1, we simply have u1={j}for j=1,2, tu1=tj,=ljand θ(u1)=fu1,φ ljL1 2=(gj,φ lj)L1 2. When k=2, we have u2={i,i+1}=:{j1,j2},fori=1,2,3, tuk=(tj1,tj2), and each index ∈˚ Zukhas only two nonzero elements. Denote these nonzero elements of by lj1 and lj2.Then θ(u2)=(fu2,φ lj1φlj2)L2 2=1 01 0 gj1(tj1)gj2(tj2)φlj1(tj1)φlj2(tj2)dtj1dtj2 =1 0 gj1(tj1)φlj1(tj1)dtj11 0 gj2(tj2)φlj2(tj2)dtj2=(gj1,φ lj1)L1 2(gj2,φ lj2)L1 2, and similarly for k=3 and 4. For k=1,2,3,4, the Fourier coefficients θ(u(i) k)with i>Nk,dare all equal to zero. For smooth Sobolev functions, the absolute values of their Fourier coefficients decay to zero at a polynomial rate. Therefore, although in theory = (l1,...,ld)∈˚ Zuk, we shall restrict ourselves to ∈n˚ Zuk def =˚ Zuk∩[−n,n]d,where n=622 for k=1, n=154 for k=2, n=65 for k=3, and n=36 for k=4.For each k,the chosen value of nensures that none of the nonzero coefficients θ∗ (r∗ ε,k,m), and hence none of the weights ω(r∗ ε,k,m), is missing in the evaluation of the statistics Suk,m,m=1,...,Mk. Thus, the model we are dealing with in this section is as follows: X=ηukθ(uk)+εξ,∈n˚ Zuk,1≤k≤4, with ε=5·10−5. The vector (ξ)∈n˚ Zuk consists of iid standard normal random variables ξ, and the component ηukof η=(ηk=(ηuk,uk∈Uk,d), 1≤k≤4)equals 1 if uk∈,u(1) k,u(2) k,...,u(Nk,d) k-and zero otherwise. For all choices of d,werunJ=15 independent cycles of simulations (using R software) and estimate the Hamming risk Eθ,ηs k=1uk∈Uk,d|ˆηuk−ηuk|by means of the quantity Err(ˆ η)=1 J J j=1 s k=1 uk∈Uk,d|ˆη(j) uk−ηuk|, 123
15 Page 24 of 26 Statistical Inference for Stochastic Processes (2025) 28:15 Table 2 Estimated Hamming risk Err(ˆ η)obtained from J=15 simulation cycles α d0.0001 0.0005 0.0009 0.001 0.0011 0.0012 0.005 0.5 1 50 1 1 0.86 0.80 0.53 0.20 0 0 0 100 1 1 0.93 0.93 0.67 0.47 0 0 0 200 1 1 1 1 0.93 0.67 0 0 0 where ˆη(j) ukis the value of ˆηukobtained in the j-th repetition of the experiment. The values of Err(ˆ η)for different dare seen in Table 2in the last column corresponding to α=1, and they are all zero. To illustrate the impact of a signal strength on the Hamming risk of ˆ η, we multiply the active component fu(1) 1 by α∈(0,1], while keeping the other active components unchanged. The values of the estimated Hamming risk Err(ˆ η)obtained for different choices of αare presented in Table 2. The selection procedure never detects a signal if there is none, and thus only active components may be misidentified as inactive. Since, out of all the active components, only one of them is multiplied by α∈(0,1], the entries in Table 2(i.e., the average numbers of misidentified components of the signal f=4 k=1uk∈Uk,dηukfuk) are all less than or equal to 1. In general, as seen from Table 2, the stronger the signal is, the smaller is the number of misidentified active components. It is also seen from Table 2that the estimated Hamming risk Err(ˆ η)increases as dincreases. This is not surprising since the model gets sparser as dgets larger, and the sparser the model is, the harder is to select correctly the active components of f. Overall, the numerical results of this section are in agreement with the statement of Theorem 1. 6 Concluding remarks The results obtained in this article provide the conditions for the possibility and impossibility of exact variable selection with respect to a Hamming loss in the Gaussian white noise model of intensity εwhen the d-variate signal fis the sum of a small number of k-variate functions with kvarying from 1 to s(1 ≤s≤d), for sbeing fixed and for s=sε→∞, s=o(d)as ε→0. The sharp selection boundary, which defines a precise demarcation between what is possible and impossible in this problem, is determined by inequalities (36) and (40) in terms of the quantities aε,uk(rε,k),1≤k≤s,definedby(22) and satisfying aε,uk(rε,k)∼c(σ, k)r2+k/(2σ) ε,kε−2as ε→0,where c(σ, k)is a known function of σand k (see relations (26)and(27)). Under the conditions when exact variable selection is possible, we proposed an adaptive selection procedure that achieves exact selection when the level of sparsity is unknown. In this article, the regularity constraints on the d-variate function fare expressed in terms of the Sobolev semi-norms for σ>0 and are imposed component-wise. In addition to the Sobolev balls Fcuk,uk∈Uk,d,1≤k≤s,ofσ-smooth functions fuk(tuk),where tuk=(tj1,...,tjk)∈Rk,uk={j1,..., jk},1≤j1< ... < jk≤d,asdefinedin Section 1.2, it might be of interest to consider the balls of analytic functions fuk(tuk),tuk= (tj1,...,tjk)∈Rk, with smoothness parameter σ>0 such that (a) fukis 1-periodic in each of its arguments, (b) fukcan be analytically continued from Rkto the k-dimensional strip 123
Statistical Inference for Stochastic Processes (2025) 28:15 Page 25 of 26 15 Sσ={zuk∈Ck:|Im zuk|≤σ/(2π)},and|f(zuk)|≤Mfor all zuk∈Sσand some M>0. A more general setup, when the regularity constraints are imposed on the whole d-variate function frather than on its k-variate components for 1 ≤k≤s, may also be addressed. Variable selection techniques for high-dimensional data are widely applied across numerous research fields. In the context of functional ANOVA model (in practice also called Smoothing Spline ANOVA decomposition, or SS-ANOVA), the problem of variable selection was studied, for example, in Wahba et al. (1995) with applications to data from health sciences. Therefore, it is also of interest to investigate how the selection method proposed in this article may be applied to health and well-being data. Acknowledgements The research of N. Stepanova was supported by a Discovery Grant of the Natural Sciences and Engineering Research Council of Canada. The research of M. Turcicova was partially supported by Discovery Grants of the Natural Sciences and Engineering Research Council of Canada during the author’s stays at Carleton University in the 2023–2024 academic year. Additionally, the author’s research was funded by the project “Research of Excellence on Digital Technologies and Wellbeing CZ.02.01.01/00/22_008/0004583" which is co-financed by the European Union. Author Contributions N.S. was responsible for conceptualization, methodology, investigation, simulation setting, writing - Original Draft but also Review & Editing, Supervision, Project administration, Funding acquisition; M.T. was responsible for methodology, investigation, simulation (setting and coding), writing - Original Draft but also Review & Editing. All authors reviewed the manuscript Funding Open access publishing supported by the institutions participating in the CzechELib Transformative Agreement. Data Availability No datasets were generated or analysed during the current study. Declarations Conflicts of Interest The authors confirm that there are no relevant financial or non-financial competing interests to report. Competing interests The authors declare no competing interests. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. References Belitser E, Nurushev N (2018) Uncertainty quantification for robust variable selection and multiple testing. Electron J Stat 16:5955–5979 Butucea C, Ndaoud M, Stepanova N, Tsybakov A (2018) Variable selection with Hamming loss. Ann Statist 46(5):1837–1875 Butucea C, Stepanova N (2017) Adaptive variable selection in nonparametric sparse models. Electron J Stat 11:2321–2357 Gao Z, Stove S (2020) Fundamental limits of exact support recovery in high dimensions. Bernoulli 26(4):2605– 2638 Genovese CR, Jin J, Wasserman L, Yao Z (2012) A comparison of the lasso and marginal regression. J Mach Learn Res 13:2107–2143 123