scieee AI-readable full text Open interactive document viewer

Lasso–Ridge Refitting: A Two-Stage Estimator for High-Dimensional Linear Regression

Liu, Guo

Abstract

This work introduces a novel Lasso–Ridge method. Theoretical results and simulations are provided. This version is a preliminary working note and is not fully polished. A later peer-reviewed article related to this note was published in Japanese Journal of Statistics and Data Science and is available at https://doi.org/10.1007/s42081-026-00345-1.

Full text

Lasso–Ridge Refitting: A Two-Stage Estimator for High-Dimensional Linear Regression Liu Guo1 1Department of Pure and Applied Mathematics, Waseda University August 22, 2025 Abstract The least absolute shrinkage and selection operator (Lasso) is a popular method for highdimensional statistics. However, it has limitations such as bias and prediction error. To address such disadvantages, many alternatives and refitting strategies have been proposed and studied. This work introduces a novel Lasso–Ridge method. We show that our estimator achieves better prediction performance than the standard Lasso, even when its tuning parameter is set at the optimal rate plog p/n. Moreover, the proposed method retains key advantages of the Lasso, including prediction consistency and reliable variable selection under standard conditions. Through extensive simulations, we further demonstrate that our estimator consistently outperforms the Lasso in both prediction and estimation accuracy, highlighting its potential as a powerful tool for high-dimensional linear regression. 1 Introduction High-dimensional linear regression, where the number of predictors pexceeds the number of observations n, commonly arises in a variety of scientific fields, including genomics, signal processing, and finance. In such settings, the goal is often to identify a sparse set of relevant variables while achieving good predictive and estimation performance. The Lasso is a widely used method for high-dimensional regression [1] that performs simultaneous variable selection and estimation by imposing an ℓ1-penalty. The theoretical properties of the Lasso have been extensively studied. Under suitable conditions such as the restricted eigenvalue condition, it satisfies oracle inequalities and achieves consistency in estimation and prediction [2]. However, a well-known drawback of the Lasso is that it introduces bias due to the ℓ1-penalty. Moreover, it has been shown that there exists a lower bound on the prediction error[3], and that the Lasso with a single choice of the penalty parameter cannot simultaneously achieve variable selection consistency and √n-consistency[4]. To address these drawbacks and achieve better performance, a variety of refitting strategies have been developed. These methods typically aim at modifying the initial Lasso estimate to reduce shrinkage bias and enhance both prediction and estimation accuracy. For instance, the OLS post-Lasso[5] applies ordinary least squares to the Lassoselected support to reduce bias. The Adaptive Lasso [6] uses a reweighted ℓ1penalty in a second-stage Lasso to better reflect the relative sizes of the coefficients. The Relaxed Lasso [7] adjusts the degree of penalization after variable selection, allowing greater flexibility in balancing sparsity and shrinkage. These approaches reflect a general two-stage refitting strategy: Lasso is first applied to obtain an initial result like variable selection, and then a second step is carried out based on that result. These refitting strategies has been proposed with theoretical analysis and experiments to address their advantages. Despite these successes, much of the existing research focuses primarily on consistency results and empirical simulations, with limited direct comparison of estimator performance. In this paper, we propose a novel Lasso–Ridge refitting approach. The intuition behind our method is to fill in the prediction gap introduced by the Lasso’s penalty by applying a ridge regression that brings a safe correction targeting the active set of the Lasso estimator. The main contributions of this work can be summarized as follows: 1 1. A theoretical guarantee of improved prediction under minimal assumptions on the design matrix and noise distribution. 2. Retention of key strengths of the Lasso, including variable selection and prediction consistency. 3. Substantially improved performance in both prediction and estimation across a variety of simulation settings. 2 Preliminaries and Background 2.1 Notation, Norms, and Model Assumptions We use ∥·∥1to denote the ℓ1-norm, i.e., ∥x∥1:= Pi|xi|, and ∥·∥2to denote the Euclidean (or ℓ2) norm, i.e., ∥x∥2:= pPix2 i. For a matrix A∈Rn×p, we denote the operator ∞-norm by ∥A∥∞:= max 1≤i≤n p X j=1 |Aij|, which corresponds to the maximum absolute row sum. The spectral norm, denoted ∥A∥2, is the largest singular value of Aand coincides with the operator norm induced by the Euclidean norm, i.e., ∥A∥2:= sup ∥x∥2=1 ∥Ax∥2. For a finite set E, we denote its cardinality by |E|. For any index set S⊆[p] := {1, . . . , p}, we write AS∈Rn×|S|to denote the submatrix of the design matrix A∈Rn×pconsisting of the columns indexed by S. That is, ASis formed by extracting the columns X·j for each j∈S. Similarly, for a vector β∈Rp, we write βS∈R|S|for the subvector containing the entries indexed by S. For completeness, we adopt standard conventions in high-dimensional statistics to ensure that all expressions remain well-defined, even in degenerate cases. For example, when S=∅, we interpret AS∈Rn×0. In our proofs, we explicitly separate cases involving empty index sets to simplify the argument structure. Throughout, care is taken to maintain mathematical rigor, and no ambiguity arises in definitions or proofs. We consider a standard linear regression model given by: y=Xβ0+ϵ, ϵ ∼N(0, σ2In), where y= (y1, . . . , yn)∈Rnis the response vector, X∈Rn×pis the design matrix, β0denotes the unknown coefficient vector, and ϵis the noise vector. This Gaussian assumption is made for simplicity, and all theoretical results below also hold when ε is a sub-Gaussian random vector with mean zero and sub-Gaussian parameter σ. We additionally assume that the columns of Xare normalized in such a way that for all j∈[p], we have ∥Xj∥2 2=n, where Xjrepresents the jth column of the matrix X. 2 2.2 The Lasso Estimator and Its Basic Properties The Lasso estimator is noted as ¯ β, with its definition as follows: ¯ β= arg min β∈Rp1 2n∥y−Xβ∥2 2+λL∥β∥1, Definition 2.1. We define the key quantities associated with the Lasso: •The equicorrelation set: E:=i∈ {1,2, . . . , p}:|XT i(y−X¯ β)/n|=λL •The equicorrelation signs: s:= sign XET(y−X¯ β)/n Assumption 2.2. The columns of the design matrix Xare in general position. Proposition 2.3 (Uniqueness and Equicorrelation Set of the Lasso Estimator).Under Assumption 2.1, the solution ¯ βis unique. Moreover, its selected support coincides with the E, to be more specific, ˆ βj= 0 if and only if j∈E. For more details on the structure and uniqueness of Lasso solutions, see [8]. Remark 2.4. This is a mild condition that holds with probability one when the entries of the design matrix Xare drawn from a continuous distribution. It ensures the uniqueness of Lasso solutions, thereby avoiding ambiguity when defining a refitting step based on the non-zero part of Lasso estimators. In the rest of the paper, we adopt this assumption without further mention. In addition, without this assumption, it also holds that every Lasso estimator ¯ βcorresponding to a fixed tuning parameter λLgives the same fitted value X¯ β, which ensures the uniqueness of the equicorrelation set Eand the equicorrelation signs s. Proposition 2.5. The KKT conditions for the Lasso estimator can be derived from subdifferential calculus [2]. Introducing ¯zas an element of the subdifferential of the ℓ1 norm at ¯ β, the conditions are: 1 nX⊤(y−X¯ β) = λL¯z, where ¯z∈∂∥¯ β∥1and is given componentwise by: ¯zi=(sign(¯ βi),if ¯ βi= 0, ui∈[−1,1],if ¯ βi= 0,for i= 1, . . . , p. Also, notice that: 1 nX⊤ E(y−X¯ β) = λLs. Here, sign(·)denotes the sign function. 3 3 Ridge Refitting for Lasso Estimator 3.1 Definition of a Ridge refitting strategy A new Lasso-Ridge estimator could be defined as: ˆ β= arg min β∈Rp,β−E=0 1 2n∥y−Xβ∥2 2+λ 2∥β−¯ β∥2 2, or equivalently: ˆ β=ˆ δ+¯ β, ˆ δ= arg min δ∈Rp,δ−E=0 1 2n∥(y−X¯ β)−Xδ∥2 2+λ 2∥δ∥2 2, where ¯ βrepresents the Lasso estimator, Erepresents the equicorrelation set of the Lasso estimator. Note that the refitting step is based on the non-zero sets of Lasso. For given y, Lasso selects no columns when λLis large enough. Also note that ϵcan also lead to no columns selected by lasso. In such cases, ˆ δis a zero vector and no refitting is made to the Lasso estimator. When Eis not an empty set, ˆ δEcould be regarded as a Ridge estimator: ˆ δE= arg min δ∈R|E|1 2n∥(y−X¯ β)−XEδE∥2 2+λ 2∥δE∥2 2. The tuning parameter λcontrols the second step. When λtends to infinity, the second step step becomes negligible, and no change is added to the Lasso. On the other hand, when λ= 0, this refitting strategy is the same as least square refitting after the Lasso. In comparison to the least square refitting, this new method aims to address the limitation left by Lasso tuning instead of trusting Lasso’s variable selection completely. In our theoretical analysis, the tuning parameter used in the second step depends on both the initial Lasso tuning parameter λLand the noise vector ε. For simplicity, we denote it simply by λthroughout the paper. The tuning parameter of the Lasso, denoted by λL, governs both variable selection and estimation simultaneously. However, previous works suggest that it may be difficult for one choice of λLto simultaneously yield reliable variable selection and strong predictive performance. This limitation motivates the use of refitting strategies. In many cases, λL is chosen conservatively to ensure sparsity, leaving room for a second-stage procedure to reduce bias or improve estimation. Notice that, this refitting strategy is to change the non-zero entries of the Lasso estimator to reduce squared loss on the selected variables. To be more specific, we have the following basic property: Proposition 3.1 (Reduction in Empirical Risk). ∥y−Xˆ β∥2≤ ∥y−X¯ β∥2. 4 4 Theoretical Results 4.1 Enhanced Prediction Performance Theorem 4.1 (Pointwise Prediction Improvement).Select the tuning parameter for the second step λ > 2∥G∥∞, where G=XE⊤XE/n. Then the Lasso-Ridge estimator satisfies the following inequality: 1 2n∥Xˆ β−Xβ0∥2 2≤1 2n∥X¯ β−Xβ0∥2 2−R(ε, X, β0), where R(ε, X, β0) = λ λL (λL 2−3 2∥1 nX⊤ε∥∞)∥ˆ δ∥2 2 is a random variable that is non-negative in many cases. Remark 4.2. One crucial idea behind the prediction improvement is to select the tuning parameter λ > 2∥G∥∞based on the support XEobtained from the Lasso step. It is worth noting that when the number of selected columns is small or the selected columns exhibit low correlation, a strong penalty in the second step is unnecessary, allowing greater flexibility for the second step. It is reasonable to consider cases where the support selected by Lasso is close to or smaller than the true support, which implies that the penalty need not be too large in order to preserve this property. While this tuning choice facilitates theoretical analysis, in practice there is more flexibility in selecting the penalty, and more reliable methods such as cross-validation can be used. In certain extreme cases where the Lasso tuning parameter λLis not set at an optimal rate for prediction or estimation, such as when λL≫plog p/n , substantial improvements in prediction accuracy can be observed, and this situation may occur when λLis chosen solely to ensure variable selection consistency. In the following theorems, we analyze the case where λLis chosen according to theoretical guidelines to achieve favorable prediction and estimation performance for Lasso. Corollary 4.3 (Probability of Prediction Superiority under Fixed Design).Select the tuning parameter for the second step λ > 2∥G∥∞. Consider the case where λL= 3σp2 log(2p/α)/n for some α∈(0,1) , we have: P(∥Xβ0−Xˆ β∥2≤ ∥Xβ0−X¯ β∥2)≥1−α. Remark 4.4. In fixed design, λL≥ ∥X⊤ε/n∥∞is often required for proving bounds for Lasso’s prediction and estimation error, where ∥X⊤ε/n∥∞is of the orderplog p/n. In practice, to avoid false alarm and to choose the most crucial support, there are cases where we need a stronger regularization. From these facts, it shows that the choice λL= 3σp2 log(2p/α)/n for some α∈(0,1) reflects a conservative but standard regime, where the refinement step corrects the prediction error from the strong regularization, and our Lasso–Ridge method improves Lasso’s prediction performance across a broad class of settings, requiring minimal assumptions on tuning or design. Corollary 4.5 (Asymptotic Prediction Improvement).Select the tuning parameter for the second step λ > 2∥G∥∞. For λL=Cσp2 log(2p/α)/n where Cis a moderate 5 constant, then P(∥Xβ0−Xˆ β∥2≤ ∥Xβ0−X¯ β∥2)→0as n→ ∞. Remark 4.6. For typical error distributions, a common requirement for the theoretical analysis of the Lasso is that the tuning parameter λLis of the orderplog p/n, ensuring that λL≥ ∥X⊤ε/n∥∞with high probability. We consider the case where Lasso’s tuning parameter is λlasso =C·σplog p/n with a constant C > 0. Although this constant might corresponds to relatively stronger regularization, the rate remains optimal from an asymptotic perspective, as it matches the standard order plog p/n used in highdimensional analysis. This lemma shows that, even λLis chosen theoretically reasonably, there are cases where the new estimator improve Lasso’s prediction with high probability. 4.2 Inheritance of Consistency Properties One might reasonably question whether the proposed Lasso–Ridge refitting procedure is simply designed to take advantage of specific structural features of the Lasso solution, potentially at the expense of its well-known consistency properties. However, as the following results demonstrate, this refitting approach retains the key consistency properties of Lasso. In this sense, the method should not be seen as an arbitrary adjustment, but rather as a safe refinement that builds on Lasso’s strengths while addressing some of its limitations. Theorem 4.7 (Sign Preservation of the New Estimator).If λ > 2∥G∥∞, we have sign(ˆ β) = sign(¯ β). Theorem 4.8 (Weak Prediction Consistency).If λL≍σplog p/n and λ > 2∥G∥∞, we have 1 2n∥Xˆ β−Xβ0∥2 2=Op σ∥β0∥1rlog p n!. Definition 4.9 (Bickel et al. [9]).We say that a matrix X∈Rn×psatisfies the Restricted Eigenvalue condition RE(c0, r), where c0>0 and r∈[p], if there exists a constant κ(c0, r)>0 such that for all index sets J⊆[p] with |J| ≤ r, the following inequality holds: 1 n∥Xv∥2 2≥κ(c0, r)∥vJ∥2 2,for all v∈Rpsuch that ∥vJc∥1≤c0∥vJ∥1. Theorem 4.10 (Existing Work: Prediction Consistency for Sign-consistency Refitting Strategies).If λL= 6σp2 log(p/δ)/n for some δ∈(0,1) and λ > 2∥G∥∞, under assumption 4.9, we have the following bound: 1 2n∥Xˆ β−Xβ0∥2 2+1 2n∥X¯ β−Xβ0∥2 2≤λ2 Ls(9 κ2(3, s)+16 κ2(1, s)). For details and a complete proof, see [10]. 6 5 Numerical Studies We adopt a simulation framework similar to that used in the Relaxed Lasso experiments[7] to compare the performance of our estimator with that of the Lasso under diverse conditions. The response variable follows the linear model: y=Xβ +ε, where the design matrix X∈Rn×pconsists of rows independently drawn from a multivariate normal distribution with zero mean and covariance matrix Σ, given by Σij =ρ|i−j|,with ρ∈ {0,0.5}. The noise vector ε∈Rnfollows a normal distribution N(0, σ2In), with σ∈ {0.5,1} representing different noise level. The response vector yis then generated according to the linear model above. We consider sample sizes n∈ {50,100,200}and numbers of predictors p∈ {100,200,400}, and the number of relevant (nonzero) coefficients is s∈ {5,10,30}. For k≤s, we set βk= 1, and for k > s, we set βk= 0. For the Lasso method, the tuning parameter λLis selected by cross-validation from 20 candidate values, spanning from 0 to ∥X⊤y∥∞/n. For the Lasso-Ridge method, the search grid expands to 20 ×10 candidates. All estimators use 5-fold cross-validation to select the regularization parameters. Each simulation setting is repeated 100 times, and the results are averaged over these repetitions. Performance is evaluated in terms of both prediction and estimation accuracy. In particular, we report the relative improvement of the Lasso–Ridge method over the standard Lasso, defined as Improvement = 100 ·MSELasso MSENew −1 where LLasso and LNew denote the error (e.g., prediction error or estimation error) associated with the Lasso and Lasso–Ridge estimators respectively. A positive value indicates improved performance of the Lasso–Ridge method in comparison to the standard Lasso. Overall, the simulation results show that the Lasso–Ridge estimator generally outperforms the Lasso. Except for a very few cases, where the performance of the new estimator is only slightly worse, it consistently achieves better results. In particular, for the very sparse setting, the new estimator exhibits a substantial improvement over the Lasso. This simulation demonstrates that the Lasso–Ridge method robustly improves the performance across a wide range of scenarios. 6 Conclusion and Further Studies In this work, we proposed a Lasso–Ridge estimator. We provided theoretical guarantees under minimal assumptions, and empirical evidence demonstrated consistent gains over the standard Lasso across various simulation regimes. Several directions remain open for future research. First, extending the theoretical analysis to more general noise structures and non-Gaussian designs would enhance the applicability of the method. Second, the integration of adaptive or data-driven tuning strategies within the refinement framework could further improve robustness in practice. 7 Table 1: Relative prediction improvement of Lasso–Ridge over Lasso for ρ= 0. Each group corresponds to a pair (s, σ). s= 5, σ = 0.5s= 5, σ = 1 n\p100 200 400 100 200 400 50 150 49 24 8 -22 -17 100 166 234 262 191 234 159 200 66 100 135 150 219 285 s= 10, σ = 0.5s= 10, σ = 1 n\p100 200 400 100 200 400 50 18 76 527 -4 55 167 100 126 132 102 97 78 -7 200 45 63 97 102 169 240 s= 30, σ = 0.5s= 30, σ = 1 n\p100 200 400 100 200 400 50 1156 488 691 314 206 371 100 2 59 1539 -2 20 442 200 4 4 1 24 31 16 Table 2: Relative prediction improvement of Lasso–Ridge over Lasso for ρ= 0.5. Each group corresponds to a pair (s, σ). s= 5, σ = 0.5s= 5, σ = 1 n\p100 200 400 100 200 400 50 138 229 173 70 93 57 100 131 193 166 139 210 194 200 98 141 176 110 169 223 s= 10, σ = 0.5s= 10, σ = 1 n\p100 200 400 100 200 400 50 79 58 41 11 7 -11 100 82 124 164 94 129 159 200 80 116 131 100 146 176 s= 30, σ = 0.5s= 30, σ = 1 n\p100 200 400 100 200 400 50 15 186 853 2 88 238 100 7 11 2 16 5 -14 200 25 46 56 38 69 91 8 2 Proofs of Main Results Proof of Proposition 3.1 Proof. Following from the definition of ˆ δas a solution to an optimization problem, we have: 1 2n∥y−X¯ β−Xˆ δ∥2 2+λ 2∥ˆ δ∥2 2≤1 2n∥y−X¯ β−X0∥2 2+λ 2∥0∥2 2.(1) Because ∥ˆ δ∥2≥0, we have: ∥y−Xˆ β∥2≤ ∥y−X¯ β∥2. Prove of theorem 4.1 Proof. According to inequality (1), we have: 1 2n∥y−X¯ β−Xˆ δ∥2 2+λ 2∥ˆ δ∥2 2≤1 2n∥y−X¯ β∥2 2. Substituting y=Xβ0+ε, we obtain: 1 2n∥Xβ0+ε−X¯ β+Xˆ δ∥2 2+λ 2∥ˆ δ∥2 2≤1 2n∥Xβ0+ε−X¯ β∥2 2. Expanding both sides, then: 1 2n∥Xβ0−X¯ β+Xˆ δ∥2 2+1 2nε⊤ε+1 nε⊤(Xβ0−X¯ β+Xˆ δ) + λ 2∥ˆ δ∥2 2 ≤1 2n∥Xβ0−X¯ β∥2 2+1 2nε⊤ε+1 nε⊤(Xβ0−X¯ β). Recall that ˆ β=ˆ δ+¯ β. By simplifying both sides of the inequality, we obtain: 1 2n∥Xβ0−Xˆ β∥2 2+λ 2∥ˆ δ∥2 2≤1 2n∥Xβ0−X¯ β∥2 2+1 nε⊤Xˆ δ, (2) It follows from dual-norm inequality that the right-hand side satisfies: 1 nε⊤Xˆ δ≤ ∥ˆ δ∥1∥1 nX⊤ε∥∞. Applying this and Lemma 1.6 (a norm inequality for ˆ δ), we derive: λ 2∥ˆ δ∥2 2−1 nε⊤Xˆ δ≥λ 2∥ˆ δ∥2 2−∥ˆ δ∥1∥1 nX⊤ε∥∞ ≥λ 2∥ˆ δ∥2 2−3λ 2λL∥1 nX⊤ε∥∞∥ˆ δ∥2 2 =λ λL (λL 2−3 2∥1 nX⊤ε∥∞)∥ˆ δ∥2 2. 15 Combining this inequality and (2), we arrive at the result: 1 2n∥X¯ β−Xβ0∥2 2−1 2n∥Xˆ β−Xβ0∥2 2≥λ λL (λL 2−3 2∥1 nX⊤ε∥∞)∥ˆ δ∥2 2.(3) Prove of corollary 4.2 and 4.3 Proof. By inequality (3), establishing the prediction improvement ∥X¯ β−Xβ0∥2≥ ∥Xˆ β− Xβ0∥2reduces to showing that λL/3≥ ∥X⊤ε/n∥∞. Thus we deduce that: P(∥Xβ0−Xˆ β∥2≤ ∥Xβ0−X¯ β∥2)≥P(∥1 nX⊤ε∥∞≤λL 3). Combining this inequality with Lemma 1.3, we obtain Corollaries 4.2 and 4.3. Prove of theorem 3 Proof. Following from the definition of Lasso estimator ¯ βas a solution to an optimization problem, we have 1 2n∥y−X¯ β∥2 2+λL∥¯ β∥1≤1 2n∥y−Xβ0∥2 2+λL∥β0∥1. Also substitute ˆ β=ˆ δ+¯ βin inequality (1), we get: 1 2n∥y−Xˆ β∥2 2+λ 2∥ˆ δ∥2 2≤1 2n∥y−X¯ β∥2 2. From these two inequalities, it follows that: 1 2n∥y−Xˆ β∥2 2+λ 2∥ˆ δ∥2 2+λL∥¯ β∥1≤1 2n∥y−Xβ0∥2 2+λL∥β0∥1. Replacing ∥ˆ δ∥2 2by Lemma 1.6 and splitting λL∥¯ β∥1, then we have: 1 2n∥y−Xˆ β∥2 2+1 3λL∥ˆ δ∥1+1 3λL∥¯ β∥1+2 3λL∥¯ β∥1≤1 2n∥y−Xβ0∥2 2+λL∥β0∥1. Simplifying it, we obtain: 1 2n∥y−Xˆ β∥2 2+1 3λL∥ˆ β∥1≤1 2n∥y−Xβ0∥2 2+λL∥β0∥1. Substituting y=Xβ0+ε, we obtain 1 2n∥Xβ0+ε−Xˆ β∥2 2+1 3λL∥ˆ β∥1≤1 2n∥Xβ0+ε−Xβ0∥2 2+λL∥β0∥1. Expanding the squared norm and rearranging terms, we obtain 1 2n∥Xβ0−Xˆ β∥2 2+1 3λL∥ˆ β∥1≤ε⊤X(β0−ˆ β) + λL∥β0∥1. 16 Using the Cauchy–Schwarz inequality, we obtain 1 2n∥Xβ0−Xˆ β∥2 2+1 3λL∥ˆ β∥1≤ ∥β0−ˆ β∥1∥1 nX⊤ε∥∞+λL∥β0∥1. Thus, on the event ∥1 nX⊤ε∥∞≤1 3λL, we have 1 2n∥Xβ0−Xˆ β∥2 2≤1 3λL∥β0−ˆ β∥1−1 3λL∥ˆ β∥1+λL∥β0∥1 ≤1 3λL∥β0∥1+λL∥β0∥1=4 3λL∥β0∥1. Using standard concentration bounds for sub-Gaussian noise, we have     1 nX⊤ε   ∞ =Op σrlog p n!. Setting λ≍σqlog p n, we derives the final result: 1 2n∥Xˆ β−Xβ0∥2 2=Op σ∥β0∥1rlog p n!. Acknowledgement This draft is not fully polished and is subject to further revision. It is uploaded to Zenodo as a record of ongoing research. Some parts of the text (e.g. introduction, phrasing or simulations) may later be revised or rewritten, and a polished version will be prepared for journal submission. At this stage the author is listed alone, but future versions may be published in collaboration with others. 17 Bibliography [1] Robert Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58(1):267–288, 1996. [2] Sara van de Geer. Estimation and Testing Under Sparsity. Seminar fur Statistik HGG 24.1 R, ETH Zentrum, Zurich, Switzerland, 2017. [3] Sara van de Geer. On tight bounds for the lasso. arXiv preprint arXiv:1807.05085, 2018. [4] Soumendra N. Lahiri. Necessary and sufficient conditions for variable selection consistency of the lasso in high dimensions. The Annals of Statistics, 49(2):820–844, 2021. [5] Alexandre Belloni and Victor Chernozhukov. Least squares after model selection in high-dimensional sparse models. Bernoulli, 19(2):521–547, 2013. [6] Hui Zou. The adaptive lasso and its oracle properties. Journal of the American Statistical Association, 101(476):1418–1429, 2006. [7] Nicolai Meinshausen. Relaxed lasso. Computational Statistics & Data Analysis, 52(1):374–393, 2007. [8] Ryan J. Tibshirani. The lasso problem and uniqueness. Electronic Journal of Statistics, 7:1456–1490, 2013. [9] Peter John Bickel., Ya’acov Ritov, and Alexandre B. Tsybakov. Simultaneous analysis of lasso and dantzig selector. The Annals of Statistics, 37(4):1705–1732, 2009. [10] Evgenii Chzhen, Mohamed Hebiri, and Joseph Salmon. On lasso refitting strategies. Bernoulli, 25(4A):3175–3200, 2019. [11] James E. Gentle. Matrix Algebra: Theory, Computations and Applications in Statistics. Springer International Publishing AG, Cham, third edition, 2024. 18