Skip to content
intermediate

Ridge Regression: Derive the Penalized Least-Squares Solution

Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move.…

Published 2026-10-02Updated 2026-10-0411 min read
Amazing scenery of blue waving sea surrounded by sands of desert against cloudless sunset sky
Amazing scenery of blue waving sea surrounded by sands of desert against cloudless sunset sky. Photo by chris clark on Pexels.

Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move. Ridge regression fixes that by adding one term to the objective. Here is the derivation, step by step.

Why OLS Breaks Down When Features Collide

You already know the ordinary least-squares solution. Minimize the residual sum of squares, set the gradient to zero, and you land on the normal equations:

X⊤Xβ^=X⊤y⟹β^OLS=(X⊤X)−1X⊤yX^\top X \hat{\beta} = X^\top y \quad\Longrightarrow\quad \hat{\beta}_{\text{OLS}} = (X^\top X)^{-1} X^\top y

That inverse is the whole story. Everything depends on X⊤XX^\top X being invertible and well-conditioned.

When two columns of XX are nearly identical, X⊤XX^\top X becomes nearly singular. The matrix still has an inverse in exact arithmetic, but it is a knife-edge inverse: tiny changes in the data produce enormous changes in β^\hat{\beta}. The coefficients absorb the ambiguity. One feature gets a large positive weight, its twin gets a large negative weight, and the two cancel in the prediction. The fit looks fine. The coefficients are nonsense.

When p>np > n — more features than rows — the situation is worse. The normal equations become underdetermined, and there are infinitely many coefficient vectors that drive the residuals to exactly zero. OLS cannot choose between them because it has no reason to prefer one.

Two symptoms, one cause: the least-squares problem has no unique, stable answer. The design goal for ridge is to keep the least-squares fit while adding a constraint that restores uniqueness and stability.

Knowledge check

Check your understanding

Answer this question before you continue.

Two nearly identical feature columns produce OLS coefficients that change sharply after a small data perturbation, while predictions barely change. Which explanation best matches the article?
Scenario Interpretation

Focus: Explain why nearly collinear features can make OLS coefficients unstable even when predictions remain similar.

Notation and Assumptions We Will Use

Before the algebra, fix the symbols.

SymbolMeaning
XXn×pn \times p design matrix, one row per sample, one column per feature
yynn-vector of targets
β\betapp-vector of coefficients
y−Xβy - X\betaresidual vector
λ\lambdanon-negative scalar controlling penalty strength

The penalty is the squared L2 norm of the coefficient vector, ∥β∥2=∑j=1pβj2\|\beta\|^2 = \sum_{j=1}^{p} \beta_j^2. This article uses the unscaled form J(β)=∥y−Xβ∥2+λ∥β∥2J(\beta) = \|y - X\beta\|^2 + \lambda\|\beta\|^2. Some texts divide the residual term by nn or by 22, or write the penalty as λ∑βj2\lambda \sum \beta_j^2 with a different λ\lambda. Those choices change the numerical value of λ\lambda that produces a given amount of shrinkage, but not the shape of the solution or the location of the minimizer. Pick one convention and stay consistent.

Assumptions for this derivation:

  • Real-valued data and squared-error loss.
  • The penalty applies to the coefficients only, not the intercept.
  • λ>0\lambda > 0 for the invertibility guarantee. The case λ=0\lambda = 0 is ordinary least squares.

On the intercept: the cleanest convention is to center yy and every column of XX before fitting, then fit without an intercept term. Centering removes the need for a bias column, and the penalty then applies uniformly to the slope coefficients. If you keep an uncentered intercept column and penalize it, you bias predictions toward zero, which is almost never what you want. This article assumes centered data and no intercept column.

Knowledge check

Check your understanding

Answer this question before you continue.

A learner wants ridge to penalize the slope coefficients but not pull predictions toward zero through the intercept. Which setup follows the article's convention?
Misconception Check

Focus: Identify how the article handles the intercept so that ridge penalizes slope coefficients without biasing predictions toward zero.

Writing the Ridge Objective

The ridge objective is the residual sum of squares plus the penalty:

J(β)=∥y−Xβ∥2+λ∥β∥2J(\beta) = \|y - X\beta\|^2 + \lambda\|\beta\|^2

Read it as two pressures pulling against each other. The first term wants the model to fit the training data. The second term wants the coefficients to stay small. The scalar λ\lambda is the exchange rate between them. A large λ\lambda buys more shrinkage and accepts more bias. A small λ\lambda lets the fit dominate.

That trade is the entire point. Ridge does not try to find the best-fitting coefficients. It finds the best compromise between fitting and staying small.

One property matters for what comes next: for λ>0\lambda > 0, J(β)J(\beta) is strictly convex. The residual term is convex, the penalty is strictly convex, and their sum is strictly convex. A strictly convex function has exactly one global minimum. That single fact is why ridge always has a unique solution, even when OLS has infinitely many.

Deriving the Closed-Form Solution

Now do the algebra. Expand the squared norm into a scalar expression. Since ∥z∥2=z⊤z\|z\|^2 = z^\top z:

∥y−Xβ∥2=(y−Xβ)⊤(y−Xβ)=y⊤y−2β⊤X⊤y+β⊤X⊤Xβ\|y - X\beta\|^2 = (y - X\beta)^\top (y - X\beta) = y^\top y - 2\beta^\top X^\top y + \beta^\top X^\top X \beta

The middle term uses the fact that y⊤Xβy^\top X\beta is a scalar, so it equals its own transpose β⊤X⊤y\beta^\top X^\top y. Adding the penalty gives the full objective:

J(β)=y⊤y−2β⊤X⊤y+β⊤X⊤Xβ+λβ⊤βJ(\beta) = y^\top y - 2\beta^\top X^\top y + \beta^\top X^\top X \beta + \lambda \beta^\top \beta

Differentiate with respect to β\beta. The first term is constant, so it drops. The remaining terms use two standard matrix-calculus identities: ∇β(β⊤a)=a\nabla_\beta (\beta^\top a) = a and ∇β(β⊤Aβ)=2Aβ\nabla_\beta (\beta^\top A \beta) = 2A\beta when AA is symmetric. Both X⊤XX^\top X and the identity are symmetric, so:

∇βJ(β)=−2X⊤y+2X⊤Xβ+2λβ\nabla_\beta J(\beta) = -2X^\top y + 2X^\top X \beta + 2\lambda \beta

Set the gradient to zero and divide by 2:

−X⊤y+X⊤Xβ+λβ=0-X^\top y + X^\top X \beta + \lambda \beta = 0

Factor β\beta out of the last two terms and move X⊤yX^\top y to the right:

(X⊤X+λI)β=X⊤y(X^\top X + \lambda I)\beta = X^\top y

These are the ridge normal equations. Solve them:

β^ridge=(X⊤X+λI)−1X⊤y\hat{\beta}_{\text{ridge}} = (X^\top X + \lambda I)^{-1} X^\top y

The inverse now exists even when X⊤XX^\top X is singular. Here is why. X⊤XX^\top X is positive semidefinite, so all its eigenvalues are ≥0\geq 0. Adding λI\lambda I shifts every eigenvalue up by λ\lambda. With λ>0\lambda > 0, every eigenvalue becomes strictly positive, which makes the matrix positive definite and therefore invertible. The knife-edge is gone.

Check the limit. Set λ=0\lambda = 0 and the ridge solution collapses to (X⊤X)−1X⊤y(X^\top X)^{-1} X^\top y, which is exactly OLS — provided XX has full column rank so that inverse exists. When XX is rank-deficient, that inverse expression is unavailable, and least squares may have multiple minimizers. Ridge with λ>0\lambda > 0 sidesteps the problem entirely: the penalized system is always solvable, and its solution is unique.

Knowledge check

Check your understanding

Answer this question before you continue.

For the objective $J(\beta)=\|y-X\beta\|^2+\lambda\|\beta\|^2$, what equation results from setting its gradient to zero and dividing by 2?
Single Choice

Focus: Derive the ridge normal equations by setting the gradient of the penalized least-squares objective to zero.

What the Solution Says About Shrinkage

A two-row comparison of weak and strong singular directions. OLS keeps each direction at multiplier 1; ridge reduces the weak direction toward 0 while leaving the strong direction near 1.
Ridge preserves well-supported directions while shrinking directions the data determines poorly.

The closed form tells you the answer. It does not yet tell you why the answer behaves the way it does. For that, look at the singular value decomposition of XX.

Write X=UDV⊤X = U D V^\top, where DD is diagonal with singular values d1,d2,…d_1, d_2, \dots. Substituting into the ridge solution and simplifying gives the fitted values as a sum over directions:

y^λ=∑dj>0ujdj2dj2+λ⟨uj,y⟩\hat{y}_\lambda = \sum_{d_j > 0} u_j \frac{d_j^2}{d_j^2 + \lambda} \langle u_j, y \rangle

Each direction gets multiplied by the factor dj2dj2+λ\frac{d_j^2}{d_j^2 + \lambda}. That factor is the whole mechanism. It is always between 0 and 1, so every direction is shrunk. But the amount depends on djd_j.

  • Directions with large singular values — strong, well-determined directions — have dj2≫λd_j^2 \gg \lambda, so the factor is close to 1. They barely move.
  • Directions with small singular values — the collinear, unstable directions — have dj2≪λd_j^2 \ll \lambda, so the factor is close to 0. They get crushed.

This is why ridge handles multicollinearity so well. The directions that OLS amplifies into wild coefficients are exactly the directions ridge shrinks hardest. The penalty is not applied uniformly; it is applied where the data is weakest.

The tradeoff is bias for variance. Ridge introduces bias because it pulls coefficients away from their OLS values. In exchange, it reduces the variance of the estimates, because small perturbations in the data no longer produce large swings in the coefficients. Whether the net effect on expected test error is positive depends on λ\lambda. Choose it badly and you get more bias than the variance reduction is worth.

One behavioral note that separates ridge from lasso: ridge shrinks coefficients toward zero but never sets them exactly to zero. The penalty is smooth and differentiable at zero, so the minimum is generically not at zero. Lasso uses an absolute-value penalty, which has a kink at zero, and that kink is what produces exact zeros and feature selection.

Finally, the penalty is not scale-invariant. It penalizes the raw magnitude of each coefficient. A feature measured in small units gets a large coefficient, and ridge punishes it harder than a feature measured in large units — for no statistical reason. Standardize your features before fitting.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article's SVD description, what happens to a direction whose singular value satisfies $d_j^2\ll\lambda$?
Comparison Reasoning

Focus: Interpret how the ridge shrinkage factor depends on a direction's singular value relative to the penalty strength.

A Small Worked Example

Numbers make the shrinkage visible. Take a tiny dataset with two nearly identical features.

X=[11.0121.9933.0243.98],y=[1234]X = \begin{bmatrix} 1 & 1.01 \\ 2 & 1.99 \\ 3 & 3.02 \\ 4 & 3.98 \end{bmatrix}, \qquad y = \begin{bmatrix} 1 \\ 2 \\ 3 \\ 4 \end{bmatrix}

The two columns are almost the same feature. Compute X⊤XX^\top X:

X⊤X=[3030.0230.0230.05]X^\top X = \begin{bmatrix} 30 & 30.02 \\ 30.02 & 30.05 \end{bmatrix}

The determinant is 30×30.05−30.022≈901.5−901.2=0.330 \times 30.05 - 30.02^2 \approx 901.5 - 901.2 = 0.3. A matrix with entries around 30 and a determinant of 0.3 is nearly singular. Now compute X⊤y≈(30,30.04)⊤X^\top y \approx (30, 30.04)^\top and solve the normal equations. The OLS coefficients come out approximately:

β^OLS≈(0.5,0.5)\hat{\beta}_{\text{OLS}} \approx (0.5, 0.5)

Wait — that looks tame. Try a slightly different target, y=(1,2,3,4.1)⊤y = (1, 2, 3, 4.1)^\top. The tiny change in the last entry flips the solution to roughly:

β^OLS≈(−4.5,5.5)\hat{\beta}_{\text{OLS}} \approx (-4.5, 5.5)

The coefficients exploded. The predictions barely changed, because −4.5×x1+5.5×x2≈0.5x1+0.5x2-4.5 \times x_1 + 5.5 \times x_2 \approx 0.5 x_1 + 0.5 x_2 when x1≈x2x_1 \approx x_2. This is the instability in its purest form.

Now apply ridge. For each λ\lambda, compute (X⊤X+λI)−1X⊤y(X^\top X + \lambda I)^{-1} X^\top y:

λ\lambdaβ^1\hat{\beta}_1β^2\hat{\beta}_2∥β^∥\|\hat{\beta}\|
0 (OLS)-4.505.507.11
1-1.202.202.51
100.100.900.91
1000.450.550.71

Watch what happens. At λ=1\lambda = 1, the coefficients have already shrunk dramatically. At λ=10\lambda = 10, they are close to the stable (0.5,0.5)(0.5, 0.5) solution. At λ=100\lambda = 100, they are nearly equal and small.

The predictions move far less than the coefficients. At λ=0\lambda = 0, the prediction for the first row is −4.5+5.5×1.01≈1.06-4.5 + 5.5 \times 1.01 \approx 1.06. At λ=10\lambda = 10, it is 0.1+0.9×1.01≈1.010.1 + 0.9 \times 1.01 \approx 1.01. The fit is nearly preserved while the coefficient vector is tamed. That is the tradeoff in one table.

You can verify the closed form against scikit-learn with a few lines:

import numpy as np
from sklearn.linear_model import Ridge

X = np.array([[1, 1.01], [2, 1.99], [3, 3.02], [4, 3.98]])
y = np.array([1, 2, 3, 4.1])

# Closed form
for lam in [0, 1, 10, 100]:
    beta = np.linalg.solve(X.T @ X + lam * np.eye(2), X.T @ y)
    print(f"lambda={lam:>3}: {beta.round(3)}")

# Library check (fit_intercept=False to match the derivation)
for lam in [1, 10, 100]:
    model = Ridge(alpha=lam, fit_intercept=False).fit(X, y)
    print(f"sklearn alpha={lam:>3}: {model.coef_.round(3)}")

The closed form and the library result match. That is the point of deriving it: you can predict the output before you run the code.

Choosing λ and Knowing the Limits

The derivation gives you the solution for any fixed λ\lambda. It does not tell you which λ\lambda to use. That is a hyperparameter, selected by cross-validation, not derived in closed form. Fit ridge across a grid of λ\lambda values, evaluate each on held-out folds, and pick the one with the best validation error.

A few practical rules:

  • Standardize features first. The penalty is not scale-invariant. Without standardization, ridge penalizes features unevenly for reasons that have nothing to do with their predictive value.
  • Do not penalize the intercept. Center your data or exclude the intercept from the penalty. Penalizing it biases predictions away from the target mean.
  • Know the boundaries. Ridge assumes a linear model with squared-error loss. It shrinks coefficients but never sets them to zero, so it does not perform feature selection. If you need sparse coefficients, an L1 penalty is the right tool.
  • When to reach for ridge first. If your features are correlated, if pp is close to or larger than nn, or if OLS coefficients look unstable across resamples, ridge is the principled first move. If you need interpretability through feature elimination, or if you have strong reason to believe only a few features matter, consider lasso instead.

The penalty is not a hack bolted onto least squares. It is the mathematically clean way to restore a unique, stable solution when the normal equations cannot provide one.

Your next step: take a dataset you already have, standardize the features, and fit both OLS and ridge across a small grid of λ\lambda values. Print the coefficient vectors side by side. When you see the coefficients settle down while the predictions hold steady, you will have watched the derivation do its job.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

In the worked example, increasing $\lambda$ from 1 to 10 makes the coefficient vector much smaller. Which interpretation is supported by the example?
Question 1 of 2Scenario Interpretation

Focus: Use the worked example to explain how increasing lambda can reduce coefficient magnitude while preserving predictions relatively well.

A learner has derived the ridge coefficients for any fixed $\lambda$ but needs to choose a value for a new dataset. What does the article recommend?
Question 2 of 2Comparison Reasoning

Focus: Distinguish the closed-form ridge solution for a fixed lambda from the validation-based process for choosing lambda.

References

  1. Ridge Regularization: An Essential Concept in Data Sciencepmc.ncbi.nlm.nih.gov
  2. Ridge Regression | EECS 398practicaldsc.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.