Skip to content
advanced

When Is Least Squares Unbiased? Assumptions Behind OLS Guarantees

A coefficient can be unbiased on average and still be wrong in your one dataset. Worse, it can be unbiased and still predict badly, and unbiased and still…

Published 2026-10-02Updated 2026-10-0411 min read
Aerial view of a sunny beach in Corpus Christi with waves and coastal buildings.
Aerial view of a sunny beach in Corpus Christi with waves and coastal buildings. Photo by Hameen Reynolds on Pexels.

A coefficient can be unbiased on average and still be wrong in your one dataset. Worse, it can be unbiased and still predict badly, and unbiased and still carry no causal meaning. Those are three different failures, and they come from three different assumptions.

You can already fit a least-squares line. You can read residuals, compute an R², and print a coefficient table. Somewhere along the way you absorbed the sentence "OLS is unbiased" and filed it as a property of your fitted model. It is not. It is a property of a procedure under a contract, and if you cannot name the clauses of that contract, you cannot tell when you have voided it.

This article fixes the target of the guarantee, states the classical setup, derives where unbiasedness comes from, derives where the variance result comes from, and then shows why neither property survives a misspecified model or licenses a causal claim.

What "Unbiased" Actually Claims

The OLS estimate is a random variable. That sentence is the whole game, and it is the one most learners skip.

Imagine the data-generating process as a machine that emits a fresh dataset of size n every time you press a button. You press it, fit OLS, and get a coefficient vector. You press it again and get a slightly different one. Press it a thousand times and you have a thousand coefficient vectors — a sampling distribution.

Unbiasedness is a statement about the mean of that distribution:

E[β^]=β\mathbb{E}[\hat{\beta}] = \beta

The average of the estimates across repeated samples equals the true parameter. That is all it says.

It does not say your estimate is close to β. It does not say your estimate is correct. It says that if you could run the experiment forever, the errors would cancel. In any single sample — the only sample you will ever have — an unbiased estimator can be far off.

The error of an estimator decomposes into two pieces:

E[(β^−β)2]=(E[β^]−β)2⏟bias2+E[(β^−E[β^])2]⏟variance\mathbb{E}[(\hat{\beta} - \beta)^2] = \underbrace{(\mathbb{E}[\hat{\beta}] - \beta)^2}_{\text{bias}^2} + \underbrace{\mathbb{E}[(\hat{\beta} - \mathbb{E}[\hat{\beta}])^2]}_{\text{variance}}

Unbiasedness zeroes the first term. It says nothing about the second. A procedure can be perfectly unbiased and so noisy that any given estimate is useless. Hold that frame: the rest of this article is about which assumptions zero the bias term, which assumptions control the variance term, and why neither term answers the prediction or causal question.

The Setup and Notation

Write the classical linear model as

y=Xβ+εy = X\beta + \varepsilon

where yy is an n×1n \times 1 vector of outcomes, XX is an n×pn \times p design matrix whose rows are observations and columns are regressors, β\beta is the p×1p \times 1 vector of true parameters, and ε\varepsilon is the n×1n \times 1 vector of errors.

The OLS estimator is

β^=(X⊤X)−1X⊤y\hat{\beta} = (X^\top X)^{-1} X^\top y

which requires X⊤XX^\top X to be invertible — equivalently, XX must have full column rank. This is the normal-equation solution you already derived; I will not re-derive it here. What matters now is the distinction between two objects that look similar and behave nothing alike:

  • ε=y−Xβ\varepsilon = y - X\beta is the true error. You never observe it. It is a property of the data-generating process.
  • e=y−Xβ^e = y - X\hat{\beta} is the residual. You compute it every time you fit. It is a property of your sample and your estimate.

Conflating them is the source of most confusion about OLS assumptions. The assumptions below are about ε\varepsilon and XX. They are not about yy alone, and they are not about ee.

Assumptions That Buy Unbiasedness

A four-level flowchart starts with full column rank, which makes OLS defined. Adding zero conditional mean leads to unbiased coefficients. Adding homoscedastic and uncorrelated errors leads to the classical variance formula and BLUE. A separate normality branch points to exact small-sample t and F tests.
Separate the condition for unbiasedness from the stronger conditions behind variance and small-sample inference.

Substitute the model into the estimator:

β^=(X⊤X)−1X⊤(Xβ+ε)=β+(X⊤X)−1X⊤ε\hat{\beta} = (X^\top X)^{-1} X^\top (X\beta + \varepsilon) = \beta + (X^\top X)^{-1} X^\top \varepsilon

Take the conditional expectation given XX:

E[β^∣X]=β+(X⊤X)−1X⊤E[ε∣X]\mathbb{E}[\hat{\beta} \mid X] = \beta + (X^\top X)^{-1} X^\top \mathbb{E}[\varepsilon \mid X]

The second term vanishes whenever E[ε∣X]=0\mathbb{E}[\varepsilon \mid X] = 0, which gives E[β^∣X]=β\mathbb{E}[\hat{\beta} \mid X] = \beta. This is the zero conditional mean assumption, and it is the standard sufficient condition for conditional unbiasedness.

Read the algebra carefully, because the direction of the implication matters. Zero conditional mean is enough to guarantee unbiasedness. It is not the only way the second term can vanish — a nonzero conditional error mean could in principle be annihilated by the particular matrix weighting (X⊤X)−1X⊤(X^\top X)^{-1} X^\top. But that weaker moment condition is not something you can verify or reason about in practice. Zero conditional mean is the condition you actually check, and it is the one the rest of this article assumes.

Operationally, the assumption says that knowing XX does not systematically shift the average error. It does not say the errors are independent of XX, and it does not say the error distribution is the same at every value of XX. The conditional mean is zero; the conditional variance can still move with XX, as the next section shows.

In practice the assumption fails whenever:

  • A relevant variable is omitted and correlated with an included regressor.
  • A regressor is measured with systematic error.
  • The outcome and a regressor are jointly determined, so the regressor absorbs part of the error.

A worked example

Suppose the true process is y=2+3x1+4x2+εy = 2 + 3x_1 + 4x_2 + \varepsilon with E[ε∣x1,x2]=0\mathbb{E}[\varepsilon \mid x_1, x_2] = 0. You fit only x1x_1:

y=α+β1x1+u,u=4x2+εy = \alpha + \beta_1 x_1 + u, \quad u = 4x_2 + \varepsilon

The composite error uu now contains x2x_2. If x2x_2 is correlated with x1x_1 — say Cov(x1,x2)=c≠0\text{Cov}(x_1, x_2) = c \neq 0 — then E[u∣x1]≠0\mathbb{E}[u \mid x_1] \neq 0, and the omitted-variable formula gives

E[β^1]=3+4⋅Cov(x1,x2)Var(x1)\mathbb{E}[\hat{\beta}_1] = 3 + 4 \cdot \frac{\text{Cov}(x_1, x_2)}{\text{Var}(x_1)}

With c=0.5c = 0.5 and Var(x1)=1\text{Var}(x_1) = 1, the expected estimate is 3+2=53 + 2 = 5. Not close. Not noisy. Systematically wrong, and no sample size fixes it — the bias term does not shrink as nn grows. More data sharpens your estimate of the wrong number.

Knowledge check

Check your understanding

Answer this question before you continue.

Assuming the design matrix has full column rank, what does E[ε | X] = 0 guarantee for OLS?
Question 1 of 2Misconception Check

Focus: Connect the zero conditional mean assumption to conditional unbiasedness of OLS.

In the worked example, Cov(x₁, x₂) = 0.5 and Var(x₁) = 1. What is the expected estimate of the coefficient on x₁ when x₂ is omitted?
Question 2 of 2Scenario Interpretation

Focus: Apply the omitted-variable formula to determine the expected coefficient when a correlated regressor is omitted.

Where the Variance Result Comes From

Unbiasedness is half the story. The other half is how much β^\hat{\beta} moves across samples. Under spherical errors — constant variance and zero covariance across observations — the sampling variance is

Var(β^∣X)=σ2(X⊤X)−1\text{Var}(\hat{\beta} \mid X) = \sigma^2 (X^\top X)^{-1}

Two assumptions enter here and only here:

  • Homoscedasticity: Var(εi∣X)=σ2\text{Var}(\varepsilon_i \mid X) = \sigma^2 for all ii. Constant error variance.
  • No autocorrelation: Cov(εi,εj∣X)=0\text{Cov}(\varepsilon_i, \varepsilon_j \mid X) = 0 for i≠ji \neq j. Errors are uncorrelated across observations.

Neither is required for unbiasedness. This is the split that matters: violate homoscedasticity and β^\hat{\beta} is still unbiased, but the usual standard errors are wrong. Confidence intervals get the wrong width. t-tests get the wrong p-values. You will report a precise-looking result that is not precise, or miss a real effect because the interval is too wide.

The (X⊤X)−1(X^\top X)^{-1} term carries a second lesson. When columns of XX are nearly linearly dependent, X⊤XX^\top X approaches singularity and its inverse inflates. Coefficients become highly sensitive to small perturbations in yy. This is multicollinearity: variance inflation without bias. The estimate is still centered on β; it just wanders further from it in any given sample.

One more boundary: normality is not required for unbiasedness or for the Gauss-Markov variance result. It is required for exact small-sample t and F distributions. In large samples, the central limit theorem makes those tests approximately valid without it.

Knowledge check

Check your understanding

Answer this question before you continue.

Suppose zero conditional mean still holds, but error variance changes with X. Which conclusion follows?
Comparison Reasoning

Focus: Distinguish the effect of heteroscedasticity on OLS unbiasedness from its effect on usual standard errors.

Gauss-Markov: What BLUE Does and Does Not Promise

The Gauss-Markov theorem states that under the classical assumptions — linearity in parameters, full rank, E[ε∣X]=0\mathbb{E}[\varepsilon \mid X] = 0, homoscedasticity, and no autocorrelation — OLS is the Best Linear Unbiased Estimator.

Unpack each word, because the theorem is narrower than its reputation:

  • Linear: restricted to estimators that are linear functions of yy.
  • Unbiased: restricted to estimators with zero bias.
  • Best: minimum variance within that restricted class.

"Best" is a within-class claim. It does not say OLS beats every estimator. It says OLS beats every linear unbiased estimator. Step outside the class and the guarantee evaporates.

Ridge regression is the concrete counterexample. By accepting a small amount of bias, ridge can achieve lower mean squared error than OLS — the bias term grows, but the variance term shrinks faster. That trade is exactly the bias-variance decomposition from the first section, made operational. BLUE is a reason to trust OLS when the assumptions hold. It is not a reason to stop checking them.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the stated Gauss-Markov assumptions, what does “Best” in BLUE mean?
Single Choice

Focus: Interpret the scope of the Gauss-Markov theorem's minimum-variance claim.

Unbiased Coefficients, Bad Predictions

Estimating β and predicting yy at new values of XX are different problems. Unbiasedness answers the first. It says nothing about the second.

Expected squared prediction error at a new point x0x_0 decomposes into irreducible noise, squared bias, and variance:

E[(y0−y^0)2]=σ2+Bias2+Var(y^0)\mathbb{E}[(y_0 - \hat{y}_0)^2] = \sigma^2 + \text{Bias}^2 + \text{Var}(\hat{y}_0)

An unbiased estimator contributes zero to the middle term and can still contribute heavily to the last one. High-variance coefficients produce high-variance predictions. The model is right on average and unstable in practice.

Two failure modes deserve names:

Extrapolation. A model unbiased on the training distribution can be badly wrong outside the observed range of XX. The linear form is an assumption, not a fact, and nothing in the data constrains behavior beyond the range you observed.

Large irreducible noise. Even a correctly specified conditional mean gives poor point predictions when σ2\sigma^2 is large. The signal is thin; the noise dominates.

The practical signal is simple: compare training error to held-out error, and inspect where the model is being asked to predict relative to the data it saw. If the new XX values sit outside the training range, treat the prediction as an extrapolation, not an estimate.

Unbiased Coefficients, Wrong Causal Story

Here is the trap that catches experienced practitioners. Unbiasedness is defined relative to the model you wrote down. It says nothing about whether that model matches the causal structure of the world.

E[ε∣X]=0\mathbb{E}[\varepsilon \mid X] = 0 can hold perfectly for a purely predictive specification while the coefficient has no causal interpretation. The regression is doing its job — estimating a conditional mean — and you are asking it a question it was never built to answer.

The threats that survive a clean regression:

  • Confounding: an unmeasured variable drives both XX and yy.
  • Simultaneity: XX and yy determine each other.
  • Selection: the sample was chosen in a way correlated with the outcome.
  • Measurement error in regressors: attenuates coefficients in ways that mimic real effects.

The same coefficient can be unbiased for a descriptive conditional mean and meaningless as an effect estimate. Causal claims require a design or identification argument — randomization, an instrument, a discontinuity, a defensible set of controls — that regression mechanics cannot supply. The math will not warn you. It will hand you a tight confidence interval around a number that means nothing.

Diagnosing Violations in Practice

Turn each assumption into a check you can run.

AssumptionWhat it protectsHow to checkWhat to do when it fails
E[ε∣X]=0\mathbb{E}[\varepsilon \mid X] = 0UnbiasednessResiduals vs. each regressor; domain reasoning about omitted variablesAdd controls, change design, or abandon causal claims
HomoscedasticityStandard errorsResidual variance vs. fitted valuesHeteroscedasticity-consistent standard errors
No autocorrelationStandard errorsResidual correlation for time-ordered or grouped dataCluster-robust or Newey-West standard errors
Full rank / low collinearityCoefficient stabilityCondition number or variance inflation factorsDrop, combine, or regularize regressors
Normality (optional)Exact small-sample testsResidual QQ plotRely on asymptotic tests, or bootstrap

The decision rule splits by goal:

  • If the goal is inference, fix the standard errors first, then interrogate the design. A heteroscedasticity-robust or cluster-robust standard error costs nothing and removes a whole class of false confidence.
  • If the goal is prediction, validate on held-out data and consider regularized or more flexible models. Unbiasedness is not the metric you care about; out-of-sample error is.

Common mistake: Treating a high R² or a significant p-value as evidence that the model is correct. Both can survive severe assumption violations. Neither is a substitute for checking E[ε∣X]=0\mathbb{E}[\varepsilon \mid X] = 0.

The One Rule to Carry Forward

Unbiasedness is a conditional promise about the average of an estimator across repeated samples. It is bought by the zero-conditional-mean assumption and paired with a variance result that depends on spherical errors. Neither property survives a misspecified model. Neither licenses a causal claim. Neither guarantees a good prediction.

The next move is concrete. Take a model you have already fit. Plot the residuals against each regressor, one at a time. Look for structure — curvature, fanning, clusters. Then ask the question that actually determines whether you can trust the coefficient: is it plausible that E[ε∣X]=0\mathbb{E}[\varepsilon \mid X] = 0 here? If you cannot answer yes with a reason, you have a descriptive fit, not an estimate you can build a decision on.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A model's coefficient estimator is unbiased under the training data-generating process, but a prediction is requested at an x-value well outside the training range. What can you conclude from unbiasedness alone?
Question 1 of 2Scenario Interpretation

Focus: Explain why coefficient unbiasedness does not ensure accurate predictions outside the observed range of regressors.

A regression satisfies its stated zero-conditional-mean condition for a predictive specification. What else is needed before interpreting its coefficient as a causal effect?
Question 2 of 2Misconception Check

Focus: Recognize that unbiasedness relative to a predictive model does not, by itself, identify a causal effect.

References

  1. Key Assumptions of OLS: Econometrics Reviewwww.albert.io
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.