Skip to content
intermediate

Least-Squares Residuals as Projections: Why Orthogonality Appears

You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a…

Published 2026-10-02Updated 2026-10-0411 min read
High voltage power lines stretching across a clear deep blue sky.
High voltage power lines stretching across a clear deep blue sky. Photo by Kris Møklebust on Pexels.

You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a certificate: the model is correct. It is not. The cleanliness you are seeing is partly a geometric guarantee that least squares produces no matter what data you feed it — even data a straight line has no business describing.

That guarantee is worth understanding precisely, because it tells you exactly where to stop trusting your residual plot and where to start. By the end of this article you will be able to derive the projection representation of least squares, explain why the residual vector is orthogonal to the fitted values and to every column of the design matrix, and separate those algebraic identities from the statistical assumptions they say nothing about.

The Setup: Design Matrix, Coefficients, Residuals

Fix the notation first, so every symbol later maps to something you can point at in your own data.

You have nn observations. Stack the target values into a vector y∈Rny \in \mathbb{R}^n. Stack the predictors — including a column of ones if you have an intercept — into a design matrix X∈Rn×pX \in \mathbb{R}^{n \times p}, where pp is the number of coefficients. The unknown coefficients form β∈Rp\beta \in \mathbb{R}^p.

The model is written

y=Xβ+εy = X\beta + \varepsilon

and the fitted form is

y^=Xβ^.\hat{y} = X\hat{\beta}.

The residual vector is what the model failed to reproduce:

r=y−Xβ^.r = y - X\hat{\beta}.

Least squares chooses β^\hat{\beta} to minimize the squared Euclidean norm of that residual:

β^=arg⁡min⁡β∥y−Xβ∥2.\hat{\beta} = \arg\min_{\beta} \|y - X\beta\|^2.

Two things this derivation needs: XX has full column rank (no redundant columns), and the objective is the squared Euclidean norm. Two things it does not need yet: any assumption that the errors are independent, homoscedastic, or even that the true relationship is linear. Those are statistical questions. The geometry below holds regardless.

If you have already worked through fitting a line and reading its error, you know residuals as the vertical gaps between points and the fitted line. What changes here is that we stop treating them as a list of numbers and start treating them as a single vector living in the same nn-dimensional space as yy.

Knowledge check

Check your understanding

Answer this question before you continue.

In the stated setup, which expression gives the residual vector after fitting?
Single Choice

Focus: Identify the residual vector in the article's least-squares notation.

Why the Column Space Is the Whole Story

Every candidate prediction the model can produce has the form XβX\beta. As β\beta ranges over all of Rp\mathbb{R}^p, the vector XβX\beta sweeps out the column space of XX — the set of all linear combinations of the columns. Call it C(X)\mathcal{C}(X).

That subspace is where the entire game happens. It has dimension at most pp, while yy lives in Rn\mathbb{R}^n. Unless p=np = n and the columns span everything, yy generally sits outside C(X)\mathcal{C}(X). The part of yy that no linear combination of columns can reach is exactly the error you are stuck with.

This reframes the optimization. Minimizing ∥y−Xβ∥2\|y - X\beta\|^2 over β\beta is the same as finding the point in the subspace C(X)\mathcal{C}(X) closest to the fixed point yy. Picture Rn\mathbb{R}^n as ordinary space, C(X)\mathcal{C}(X) as a plane through the origin, and yy as a point floating off that plane. Least squares drops a perpendicular from yy onto the plane and calls the foot of that perpendicular y^\hat{y}.

Once you see it that way, orthogonality stops being a fact you memorize. It becomes the definition of "closest point."

Knowledge check

Check your understanding

Answer this question before you continue.

Why is minimizing $\|y-X\beta\|^2$ equivalent to a closest-point problem?
Comparison Reasoning

Focus: Explain how least-squares minimization becomes a closest-point problem in the design matrix's column space.

The Projection Argument: Closest Point Implies Perpendicular

A schematic subspace plane labeled C(X) contains the fitted vector y-hat from the origin. The observation vector y ends above the plane, and the residual r runs from y-hat to y at a right angle to the plane. The vectors y-hat and r form the two perpendicular parts of y.
Least squares selects the closest point in the column space, making the residual perpendicular to that space.

Let us derive it rather than assert it. Suppose β^\hat{\beta} minimizes the objective. Pick any direction u∈Rpu \in \mathbb{R}^p and consider moving along it: β^+tu\hat{\beta} + tu. If β^\hat{\beta} is truly the minimizer, then the objective at t=0t = 0 cannot be beaten, so the function

f(t)=∥y−X(β^+tu)∥2f(t) = \|y - X(\hat{\beta} + tu)\|^2

has a minimum at t=0t = 0. Write r=y−Xβ^r = y - X\hat{\beta}. Then

f(t)=∥r−tXu∥2=∥r∥2−2t (Xu)⊤r+t2∥Xu∥2.f(t) = \|r - tXu\|^2 = \|r\|^2 - 2t\,(Xu)^\top r + t^2\|Xu\|^2.

This is a parabola in tt opening upward. Its minimum sits where the derivative vanishes:

f′(0)=−2 (Xu)⊤r=0.f'(0) = -2\,(Xu)^\top r = 0.

Since uu was arbitrary, (Xu)⊤r=0(Xu)^\top r = 0 for every uu, which is the same as

X⊤r=0.X^\top r = 0.

Read that equation geometrically. X⊤rX^\top r is the vector of dot products between rr and each column of XX. Setting it to zero says rr is orthogonal to every column, and therefore to every linear combination of columns — that is, to the entire column space.

The same condition falls out of the normal equations. Setting the gradient of ∥y−Xβ∥2\|y - X\beta\|^2 to zero gives X⊤Xβ=X⊤yX^\top X \beta = X^\top y, and rearranging yields X⊤(y−Xβ^)=0X^\top(y - X\hat{\beta}) = 0. The normal equations are the algebraic shadow of the geometric fact. The geometry is the thing to remember; the algebra is how you compute it.

Knowledge check

Check your understanding

Answer this question before you continue.

At a least-squares minimum, the objective cannot decrease for a small move in any coefficient direction. What condition does this imply for $r=y-X\hat{\beta}$?
Single Choice

Focus: Connect minimization along arbitrary coefficient directions to residual orthogonality with the design columns.

The Hat Matrix and the Two Orthogonal Pieces

The projection has a name and a matrix. Define

H=X(X⊤X)−1X⊤.H = X(X^\top X)^{-1}X^\top.

Then y^=Hy\hat{y} = Hy and r=(I−H)yr = (I - H)y. This HH is the hat matrix, because it puts the hat on yy.

Two properties do all the work. HH is symmetric (H⊤=HH^\top = H) and idempotent (H2=HH^2 = H). Idempotence is the algebraic statement that projecting twice is the same as projecting once — once you are on the plane, projecting again does not move you. The same holds for I−HI - H, which projects onto the orthogonal complement.

Now the identity readers usually see quoted falls out in one line:

r⊤y^=y⊤(I−H)⊤Hy=y⊤(I−H)Hy=y⊤(H−H2)y=0.r^\top \hat{y} = y^\top (I - H)^\top H y = y^\top (I - H) H y = y^\top (H - H^2) y = 0.

Fitted values and residuals are orthogonal because H(I−H)=0H(I - H) = 0: the two projection matrices target complementary subspaces. Nothing about your data entered that calculation.

Because the two pieces are perpendicular, the Pythagorean theorem applies to the vectors:

∥y∥2=∥y^∥2+∥r∥2.\|y\|^2 = \|\hat{y}\|^2 + \|r\|^2.

That is the vector-level version of the split between explained and unexplained variation you may have seen as sums of squares.

Common mistake: Reading HH as a property of your model's quality. HH depends only on XX. Change yy to anything at all — random noise, a completely different process — and the identities still hold exactly. Orthogonality is a property of the fitting procedure, not of the fit.

Knowledge check

Check your understanding

Answer this question before you continue.

For ordinary least squares, with $\hat{y}=Hy$ and $r=(I-H)y$, what explains $r^\top\hat{y}=0$?
Misconception Check

Focus: Use the hat-matrix projection identities to explain why fitted values and residuals are orthogonal.

Worked Example: Watch the Residuals Go Perpendicular

Let us make this concrete with numbers you can reproduce in a few lines of NumPy.

import numpy as np

# Four points, one predictor plus intercept
x = np.array([0.0, 1.0, 2.0, 3.0])
y = np.array([1.0, 3.0, 2.0, 5.0])

X = np.column_stack([np.ones_like(x), x])
beta, *_ = np.linalg.lstsq(X, y, rcond=None)

y_hat = X @ beta
r = y - y_hat

print("beta:", beta)
print("X^T r:", X.T @ r)      # should be ~[0, 0]
print("r . y_hat:", r @ y_hat)  # should be ~0

Run it and you will see X^T r print two values at machine precision — essentially zero — and r . y_hat likewise. The residuals are orthogonal to the intercept column and to the predictor column, and therefore to the fitted values.

Now the part that matters. Replace y with data that a straight line clearly cannot describe — say, points along a curve:

y_curved = np.array([1.0, 4.0, 9.0, 16.0])  # quadratic, not linear
beta, *_ = np.linalg.lstsq(X, y_curved, rcond=None)
y_hat = X @ beta
r = y_curved - y_hat

print("X^T r:", X.T @ r)      # still ~[0, 0]
print("r . y_hat:", r @ y_hat)  # still ~0

The identities hold exactly, to machine precision, even though the model is wrong. Plot the residuals against the fitted values and you will see a clear U-shaped pattern — the signature of a missing quadratic term. The orthogonality is intact; the model is still bad.

That is the central demonstration of this article. The angle tells you nothing. The shape tells you everything.

What Orthogonality Does Not Prove

This is where the geometry ends and statistics begins, and confusing the two is the most common error I see in intermediate learners.

The statement r⊥C(X)r \perp \mathcal{C}(X) is a theorem about a minimization problem. It is true for any yy, any full-rank XX, and any solver that actually finds the minimum. It says nothing about:

  • whether the true relationship between predictors and target is linear,
  • whether the error variance is constant across observations,
  • whether the errors are independent,
  • whether you have omitted a relevant variable,
  • whether the errors are normally distributed.

Those are claims about ε\varepsilon, the unobservable noise term, and they require assumptions you impose on top of the geometry. The geometric result is about rr, the observable residual. They are different objects, and the identity connecting them only holds under a correctly specified model.

Note: When XX is rank-deficient — for example, two identical predictor columns — β^\hat{\beta} is not unique. The fitted values y^\hat{y} and the residuals rr still are, because they depend only on the column space, not on which basis you chose for it. Orthogonality survives; coefficient interpretation does not.

The practical consequence is the one your residual-analysis workflow already told you: you must look at the shape of the residuals, not just their correlation with y^\hat{y}. That correlation is zero by construction. Plot rr against each predictor, against y^\hat{y}, and against observation order, and ask what structure survived the fit.

Where the Geometry Changes: Weighted and Nonlinear Least Squares

The clean perpendicular picture is a special case, and knowing when it stops applying keeps you from misapplying the intuition.

Weighted least squares minimizes ∥W1/2(y−Xβ)∥2\|W^{1/2}(y - X\beta)\|^2 for a weight matrix WW. The residuals are orthogonal to the columns in the transformed space, not the original one. If you plot raw residuals against raw fitted values, the orthogonality you expect may not appear — and that is correct behavior, not a bug.

Nonlinear least squares minimizes ∥y−f(β)∥2\|y - f(\beta)\|^2 where ff is nonlinear in β\beta. There is no global column space. At the solution, orthogonality holds between the residual and the columns of the Jacobian — the local linearization. Solvers like SciPy's leastsq expose this directly: the gtol parameter controls the desired orthogonality between the function vector and the Jacobian columns, and it is used as a convergence criterion. Orthogonality becomes something the algorithm enforces rather than something the geometry guarantees.

Total least squares measures the distance from each point to the fitted line along the shortest direction, not vertically. The residual is perpendicular to the fitted curve in the original coordinate space, which is a genuinely different geometry from ordinary least squares.

The takeaway: before you claim orthogonality, name the norm and name the space. The identity is only as meaningful as the objective it came from.

Reading Residuals With the Right Mental Model

Here is how I want you to use this.

Treat orthogonality as a sanity check on your computation, not as evidence about your model. If you compute X⊤rX^\top r and it is not essentially zero, something is wrong with your solver, your design matrix, or your arithmetic. That is a useful alarm. It is not a quality score.

When residuals show structure — curvature, a funnel, clusters — the fix is almost never a different solver. It is a different feature, a transformation, an interaction term, or a different model form. The solver did its job perfectly; it found the closest point in the subspace you gave it. If that subspace cannot represent the truth, no amount of numerical care will save you.

When residuals look random, remember that randomness is consistent with many unmodeled problems: omitted variables, heteroscedasticity, dependence between observations. A clean residual plot narrows the space of things that are obviously wrong. It does not certify that anything is right.

The angle is guaranteed. The shape is the evidence. Go refit a model you already have, verify X⊤r=0X^\top r = 0 numerically, then plot rr against y^\hat{y} and against each predictor. Ask what structure survived the projection — that is the question the geometry was never going to answer for you.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A straight-line fit to curved data has $X^\top r$ numerically near zero, while its residual plot shows a U-shaped pattern. Which interpretation fits the article?
Question 1 of 2Scenario Interpretation

Focus: Distinguish least-squares residual orthogonality from evidence that the model is correctly specified.

You switch from ordinary to weighted least squares. Which statement best describes the orthogonality geometry?
Question 2 of 2Comparison Reasoning

Focus: Identify the space in which weighted least-squares residual orthogonality applies.

References

  1. leastsq — SciPy v1.16.2 Manualdocs.scipy.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.