Skip to content
intermediate

Why Lasso Can Set Coefficients to Zero: An L1 Derivation

Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero.…

Published 2026-10-02Updated 2026-10-0410 min read
Ceremonial cavalry with Union Flags in a historical parade on The Mall, London.
Ceremonial cavalry with Union Flags in a historical parade on The Mall, London. Photo by Josh Withers on Pexels.

Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero. Fit ridge on the same data and every coefficient survives, small but alive.

The usual explanation — "the penalty pushes small numbers down until they round to zero" — is wrong, and it hides the actual mechanism. Nothing is rounding. The optimizer is choosing a point where a coefficient is exactly zero because that point is genuinely optimal under the L1 penalty. To see why, we need the objective, the constraint geometry, and one small case where you can watch a coefficient cross the line.

This is the derivation. I'll assume you already have OLS normal equations and ridge shrinkage in hand — if not, the prerequisite articles cover both. Here we start from the lasso objective and follow it to the zero.

Set Up the Objective and the Notation

Fix the notation first, because every later symbol needs a concrete referent.

Let XX be the n×pn \times p design matrix — nn observations, pp features. Let yy be the nn-vector of targets. Let β\beta be the pp-vector of coefficients we are solving for. The residual vector is

r=y−Xβ.r = y - X\beta.

The intercept is handled separately and is not penalized. Penalizing it would make the fit depend on where you happened to center the target, which is not a modeling choice you want the penalty making for you.

The lasso objective is

min⁡β  12n∥y−Xβ∥22+λ∥β∥1.\min_{\beta} \; \frac{1}{2n}\|y - X\beta\|_2^2 + \lambda\|\beta\|_1.

Read it term by term:

  • 12n∥y−Xβ∥22\frac{1}{2n}\|y - X\beta\|_2^2 is the squared error — the fit term. It wants β\beta to explain the data.
  • ∥β∥1=∑j=1p∣βj∣\|\beta\|_1 = \sum_{j=1}^{p}|\beta_j| is the L1 norm — the penalty term. It wants β\beta to be small, and it charges for every nonzero entry.
  • λ≥0\lambda \ge 0 is the single knob trading one against the other.

Ridge uses the same framework with one structural change: the penalty becomes λ∥β∥22\lambda\|\beta\|_2^2. Same fit term, same λ\lambda, different norm. That single substitution is the entire source of the behavioral difference, and the rest of this article is about why.

Three assumptions carry the derivation:

  1. Columns of XX are standardized (centered and scaled to unit variance). The L1 penalty charges absolute magnitude, so if one feature is measured in millimeters and another in kilometers, you are penalizing units, not signal.
  2. λ≥0\lambda \ge 0. A negative penalty would reward large coefficients, which is not regularization.
  3. No perfect collinearity in XX, so the solution is unique.

Knowledge check

Check your understanding

Answer this question before you continue.

Why does the derivation assume that feature columns are standardized?
Single Choice

Focus: Explain why lasso predictors should be standardized before applying the coefficient penalty.

Two Equivalent Views: Penalty and Constraint

The penalized form above is convenient for computation. The constrained form is convenient for geometry, and they describe the same solution set.

min⁡β  ∥y−Xβ∥22subject to∥β∥1≤t.\min_{\beta} \; \|y - X\beta\|_2^2 \quad \text{subject to} \quad \|\beta\|_1 \le t.

For every λ\lambda there exists some tt that yields the same solution, and vice versa. This is Lagrangian duality — I am stating it, not proving it here. The practical content is the direction of the knob: larger λ\lambda corresponds to smaller tt, which means more shrinkage and more zeros.

Now picture the two pieces.

The squared-error term draws ellipsoids around the OLS solution β^\hat{\beta}. Each ellipsoid is a contour of equal error: points on the same ellipse fit the data equally well. The center is β^\hat{\beta}, the unpenalized optimum.

The constraint draws a region around the origin. For L1 that region is a diamond in two dimensions; for L2 it is a circle.

The lasso solution is the first point where a growing error ellipsoid touches the constraint region. Start with a tiny ellipsoid near β^\hat{\beta} — it misses the constraint entirely. Inflate it. The moment it makes contact, you have found the constrained optimum. That contact point is the answer.

Everything that follows is about the shape of the thing the ellipsoid touches.

Knowledge check

Check your understanding

Answer this question before you continue.

As λ increases in the penalized formulation, what happens to the equivalent constraint radius t?
Comparison Reasoning

Focus: Relate the penalty strength in the penalized objective to the size of the equivalent L1 constraint.

Why the Diamond Has Corners and the Circle Does Not

Two coordinate plots compare constraint shapes and expanding error ellipses centered at the OLS solution. In the L1 plot, the first contact is at a diamond corner on an axis, so one coefficient is zero. In the L2 plot, contact is at a smooth, non-axis point, so both coefficients are nonzero.
The L1 ball’s corners provide natural contact points with zero coefficients; a smooth L2 boundary typically does not.

In two dimensions, the L1 ball {∣β1∣+∣β2∣≤t}\{|\beta_1| + |\beta_2| \le t\} is a diamond with vertices sitting on the axes at (t,0)(t, 0), (0,t)(0, t), (−t,0)(-t, 0), and (0,−t)(0, -t). In pp dimensions it becomes a cross-polytope, and its vertices and low-dimensional faces lie on coordinate subspaces — places where one or more coefficients are exactly zero.

The L2 ball {β12+β22≤t2}\{\beta_1^2 + \beta_2^2 \le t^2\} is a smooth sphere. No vertices. No flat faces. Its boundary curves everywhere, so an expanding ellipsoid generically meets it at a single tangent point in the interior of a face — a spot where every coordinate is nonzero. Coordinate-zero points do exist on the sphere (the axis points), but the smooth boundary gives the ellipsoid no reason to prefer them.

Here is the touch argument. An expanding ellipsoid generically first contacts a smooth surface at a tangent point in the interior of a face — a spot where every coordinate is nonzero. It can contact a corner exactly, and corners sit where some βj=0\beta_j = 0.

Be honest about the word "generically." The corner is not guaranteed. But it is a positive-probability event for L1 and a measure-zero event for L2. That asymmetry is the whole story: the diamond offers the ellipsoid a set of zero-coordinate landing pads, and the circle offers none.

Note: The corner is not the only way to get a zero. In higher dimensions the ellipsoid can land on an edge or a face of the cross-polytope, and any such contact pins at least one coefficient to zero. The vertex is just the easiest picture to hold in your head.

The Coordinate-Wise View: Soft Thresholding

The geometry is convincing, but I do not want you to take sparsity on faith from a picture. There is a second route to the same conclusion, and it is algebraic.

Take the special case where the design is orthonormal: X⊤X=IX^\top X = I. This is a teaching device, not the general case, but it strips the problem down to its mechanism. Under orthonormality the lasso solution reduces to

β^j=sign⁡ ⁣(β^jOLS)⋅max⁡ ⁣(∣β^jOLS∣−λ,  0).\hat{\beta}_j = \operatorname{sign}\!\left(\hat{\beta}_j^{\text{OLS}}\right) \cdot \max\!\left(\left|\hat{\beta}_j^{\text{OLS}}\right| - \lambda, \; 0\right).

Read that as a mechanism, not a formula. Take the OLS coefficient. Subtract λ\lambda from its magnitude. If nothing is left, the coefficient is exactly zero — not small, zero. This is soft thresholding, and it is the cleanest statement of lasso coefficient thresholding:

∣β^jOLS∣≤λ  ⟹  β^j=0.\left|\hat{\beta}_j^{\text{OLS}}\right| \le \lambda \;\Longrightarrow\; \hat{\beta}_j = 0.

Compare ridge under the same orthonormal design:

β^j=β^jOLS1+λ.\hat{\beta}_j = \frac{\hat{\beta}_j^{\text{OLS}}}{1 + \lambda}.

Ridge divides. Lasso subtracts. Division by a finite number never reaches zero, no matter how large λ\lambda grows. Subtraction crosses zero and stops there. That is the difference between shrinkage and selection, and it is the same fact the diamond was showing you.

Common mistake: Treating the orthonormal result as the general rule. With correlated predictors, the threshold is no longer a clean per-coefficient comparison against λ\lambda. The next section shows what actually happens.

Knowledge check

Check your understanding

Answer this question before you continue.

For an orthonormal design, a feature has an OLS coefficient of −0.4 and λ is 0.4. What is its lasso coefficient?
Scenario Interpretation

Focus: Apply the orthonormal-design soft-thresholding condition to determine when a lasso coefficient is exactly zero.

A Small Worked Case: Shrinkage, Then a Zero

Let's make it concrete. Two standardized predictors, and suppose the OLS solution is

β^OLS=(2.0,  0.5).\hat{\beta}^{\text{OLS}} = (2.0, \; 0.5).

I am choosing these numbers so the arithmetic stays hand-checkable. In the orthonormal case, soft thresholding gives the lasso solution directly:

β^(λ)=(sign⁡(2.0)max⁡(2.0−λ,0),  sign⁡(0.5)max⁡(0.5−λ,0)).\hat{\beta}(\lambda) = \left(\operatorname{sign}(2.0)\max(2.0 - \lambda, 0), \; \operatorname{sign}(0.5)\max(0.5 - \lambda, 0)\right).

Walk λ\lambda upward and watch what happens.

λ\lambdaβ^1\hat{\beta}_1β^2\hat{\beta}_2
0.02.00.5
0.21.80.3
0.51.50.0
1.01.00.0
2.00.00.0

At λ=0.5\lambda = 0.5 the second coefficient hits zero, because ∣β^2OLS∣=0.5≤λ|\hat{\beta}_2^{\text{OLS}}| = 0.5 \le \lambda. The threshold condition is satisfied exactly, and the coefficient does not pass through zero into negative territory — it stops. For any λ>0.5\lambda > 0.5, feature 2 is gone from the model. At λ=2.0\lambda = 2.0, feature 1 follows.

That moment at λ=0.5\lambda = 0.5 is what "selection" means operationally: the active feature set changes. Before it, two features carry weight. After it, one does. The model did not merely get smaller; its structure changed.

One more piece of context worth knowing: the full solution path β^(λ)\hat{\beta}(\lambda) is piecewise linear in λ\lambda. Between threshold crossings, each coefficient moves along a straight line. That is why path algorithms can compute the entire sequence of solutions efficiently rather than refitting from scratch at every λ\lambda — but that is an implementation detail, not part of the derivation.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked orthonormal case with OLS coefficients (2.0, 0.5), what is the lasso solution at λ = 0.2?
Output Prediction

Focus: Compute the two lasso coefficients in the article's orthonormal worked case at a specified penalty strength.

What the Derivation Does Not Promise

A zero coefficient is a statement about the objective at a particular λ\lambda. It is not a statement about the world. Five limits matter.

Correlated predictors. With near-duplicate columns, lasso tends to keep one and zero the other, and which one it picks can flip across resamples. The predictions may be stable while the selection is not. If your features come in correlated groups, this instability is the norm, not an edge case.

Sparse is not important. A zero means "this coefficient was not worth its L1 cost at this λ\lambda." It does not mean the variable is irrelevant. A feature can be genuinely predictive and still be zeroed because a correlated neighbor already carries the signal.

Scaling matters. Penalizing raw coefficients on unstandardized features penalizes measurement units. Standardize first, or the penalty is doing something you did not intend.

Recovery has conditions. Exact support recovery — getting the true nonzero set back — requires enough samples, low noise, and a design that is not too correlated. When those conditions fail, L1 selection can behave close to random. The number of samples needs to be sufficiently large relative to the number of true nonzero coefficients, the noise level, and the correlation structure of XX.

λ\lambda is a modeling decision. Cross-validation tends to under-penalize, keeping a few irrelevant features because they do not hurt prediction. BIC-style criteria tend to over-penalize. Neither is a truth oracle; they optimize different things.

When to Reach for Lasso, and When Not To

Convert the mechanism into a decision rule.

Reach for lasso when pp is large relative to nn, you suspect only a subset of features matter, and you want a sparse, inspectable model. The corners of the L1 ball are doing useful work for you.

Prefer ridge when predictors are many and correlated and you mainly care about prediction stability rather than selection. Ridge will not hand you a feature list, but it will not hand you an unstable one either.

Consider elastic net when you want sparsity but need correlated groups to enter or leave together. It blends the L1 and L2 penalties, which is the natural extension when the diamond's corners are too sharp for your data.

And do not treat a lasso fit as a causal or scientific conclusion. Treat it as a hypothesis generator. The zero coefficients tell you which features might deserve attention; held-out validation tells you whether that attention pays off.

Tip: Put scaling and the penalty inside a pipeline so the same transformation is applied at fit and predict time. A scaler fitted on training data and forgotten at prediction time is a quiet, recurring bug.

Where to Go Next

The zero coefficients come from the corners of the L1 ball, and the soft-thresholding formula is the same fact written in algebra. Two views, one mechanism.

Here is the next move I would make. Take a small standardized dataset, fit a lasso, and sweep λ\lambda across a wide range. Plot the coefficient path — λ\lambda on a log scale against each coefficient. Watch a line dive to zero and stay there. That picture will do more for your intuition than any derivation, because you will have produced it yourself.

Then apply the decision rule: use the sparsity as a hypothesis about which features deserve attention, and verify it out of sample before you believe it.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A lasso fit assigns a feature a zero coefficient, and that feature is correlated with another selected feature. Which interpretation is supported by the article?
Question 1 of 2Misconception Check

Focus: Distinguish a zero lasso coefficient from proof that a feature has no predictive relevance.

A team has many correlated predictors and mainly wants stable predictions, not a feature list. Which choice best matches the article's guidance?
Question 2 of 2Scenario Interpretation

Focus: Choose among lasso, ridge, and elastic net based on sparsity, correlated predictors, and the modeling goal described in the article.

References

  1. 1.13. Feature selection — scikit-learn 1.9.0 documentationscikit-learn.org
  2. Sparsity and smoothness via the fused lassodept.stat.lsa.umich.edu
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.