Skip to content
intermediate

Derive the Bias–Variance–Noise Decomposition for Squared Error

Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…

Published 2026-10-02Updated 2026-10-0411 min read
Breathtaking view of the ocean under a bright blue sky, ideal for serene stock imagery.
Breathtaking view of the ocean under a bright blue sky, ideal for serene stock imagery. Photo by Jeffrey Eisen on Pexels.

Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not a bug, and it is not noise in the data. It is a third thing, and the decomposition is the accounting that separates all three.

By the end of this article, you will be able to write the expectation, name every symbol, and prove the split yourself instead of reciting it.

Why One Error Number Hides Three Causes

A single test error conflates three distinct sources. A model that is systematically wrong. A model that is unstable across training sets. And randomness in the target itself that no model can remove.

The bias–variance decomposition is an accounting identity, not a new model. It does not predict anything on its own. What it does is tell you which cause you are actually paying for when your error is high.

If you have read the intuition-level treatment of the bias–variance tradeoff, you already have the diagnostic categories. Here we ask a sharper question: what do those words mean as exact mathematical objects? The shape of the result is:

Expected squared error = squared bias + variance + noise.

Each term has a different owner. Squared bias belongs to the model class and the fitting procedure. Variance belongs to the training data. Noise belongs to the data-generating process. Keep those owners in mind; they are the whole point.

Notation and the Data-Generating Assumptions

Before any algebra, fix the symbols. Every derivation mistake I have seen comes from losing track of what is random and what is fixed.

Fix a query point x. Define the target as

y=f(x)+εy = f(x) + \varepsilon

where f(x)f(x) is the true conditional mean, E[y∣x]=f(x)\mathbb{E}[y \mid x] = f(x), and ε\varepsilon is noise with mean zero and variance σ2\sigma^2. The noise is independent of the training set.

Define the learned predictor as f^(x;D)\hat{f}(x; D). It is a random function because it depends on the random training set DD, drawn i.i.d. from the joint distribution P(x,y)P(x, y).

Define the expectation ED[⋅]\mathbb{E}_D[\cdot] over repeated draws of training sets, and the mean prediction:

fˉ(x)=ED[f^(x;D)].\bar{f}(x) = \mathbb{E}_D[\hat{f}(x; D)].

The assumptions the derivation depends on:

  • Squared loss.
  • Additive noise with zero mean.
  • Noise independent of the training data.
  • A fixed query point xx.

Note: The noise term is irreducible only under these assumptions. Change them and the label becomes misleading, as we will see later.

One more distinction matters. In the fixed-design setting, the inputs are held fixed and all randomness comes from the noise εi\varepsilon_i. In the random-design setting, the inputs are also resampled. The fixed-design version tends to show larger bias and smaller variance, because less randomness lands in the variance term. Both are valid; they answer slightly different questions. I will use the random-design version because it matches how most people actually resample data.

Knowledge check

Check your understanding

Answer this question before you continue.

At a fixed query point x, what does the article’s \(\bar{f}(x)\) represent?
Single Choice

Focus: Distinguish the training-set-averaged predictor from the true conditional mean and a single fitted predictor.

Deriving the Decomposition Step by Step

Start from the expected squared error at the fixed point xx:

Ey,D[(y−f^(x;D))2].\mathbb{E}_{y,D}\left[(y - \hat{f}(x; D))^2\right].

The expectation runs over both the noise in yy and the randomness in the training set DD.

Step 1: Insert and subtract the mean prediction. Add and subtract fˉ(x)\bar{f}(x) inside the square:

y−f^(x;D)=(y−fˉ(x))+(fˉ(x)−f^(x;D)).y - \hat{f}(x; D) = \big(y - \bar{f}(x)\big) + \big(\bar{f}(x) - \hat{f}(x; D)\big).

Step 2: Expand the square. This gives three terms:

E[(y−fˉ(x))2]+E[(fˉ(x)−f^(x;D))2]+2 E[(y−fˉ(x))(fˉ(x)−f^(x;D))].\mathbb{E}\left[\big(y - \bar{f}(x)\big)^2\right] + \mathbb{E}\left[\big(\bar{f}(x) - \hat{f}(x; D)\big)^2\right] + 2\,\mathbb{E}\left[\big(y - \bar{f}(x)\big)\big(\bar{f}(x) - \hat{f}(x; D)\big)\right].

Step 3: Kill the cross term. This is the step most readers skip and most mistakes hide in. The factor (fˉ(x)−f^(x;D))\big(\bar{f}(x) - \hat{f}(x; D)\big) has mean zero over DD, because fˉ(x)\bar{f}(x) is defined as the mean of f^(x;D)\hat{f}(x; D). And it is independent of the noise in yy. So the expectation of the product factors into a product of expectations, and one factor is zero:

ED[fˉ(x)−f^(x;D)]=fˉ(x)−fˉ(x)=0.\mathbb{E}_D\left[\bar{f}(x) - \hat{f}(x; D)\right] = \bar{f}(x) - \bar{f}(x) = 0.

The cross term vanishes.

Step 4: Split the remaining two terms. The first term involves only yy and fˉ(x)\bar{f}(x). Write y=f(x)+εy = f(x) + \varepsilon:

Ey[(y−fˉ(x))2]=(f(x)−fˉ(x))2+σ2.\mathbb{E}_y\left[\big(y - \bar{f}(x)\big)^2\right] = \big(f(x) - \bar{f}(x)\big)^2 + \sigma^2.

The second term is the spread of predictions around their own mean:

ED[(fˉ(x)−f^(x;D))2]=Var⁡D[f^(x;D)].\mathbb{E}_D\left[\big(\bar{f}(x) - \hat{f}(x; D)\big)^2\right] = \operatorname{Var}_D\left[\hat{f}(x; D)\right].

Step 5: Write the identity.

Ey,D[(y−f^(x;D))2]=(f(x)−fˉ(x))2⏟Bias2+Var⁡D[f^(x;D)]⏟Variance+σ2⏟Noise.\mathbb{E}_{y,D}\left[(y - \hat{f}(x; D))^2\right] = \underbrace{\big(f(x) - \bar{f}(x)\big)^2}_{\text{Bias}^2} + \underbrace{\operatorname{Var}_D\left[\hat{f}(x; D)\right]}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Noise}}.

In one plain sentence: average error = systematic miss + instability + unavoidable noise.

Knowledge check

Check your understanding

Answer this question before you continue.

Why does the cross term involving \(\bar{f}(x)-\hat{f}(x;D)\) vanish in the derivation?
Misconception Check

Focus: Explain why the cross term vanishes in the squared-error decomposition under the stated assumptions.

A Small Worked Example You Can Check by Hand

A number line shows predictions 8, 9, 11, and 12 centered around both the true value and mean prediction, 10. Below it, four compact tiles show squared bias 0, variance 2.5, noise 1, and total 3.5.
The predictions average to the truth, so bias is zero; their spread and the noise account for the total error.

Abstract labels become real once you push numbers through them. Pick a query point xx where the true value is f(x)=10f(x) = 10, and let the noise variance be σ2=1\sigma^2 = 1.

Suppose your learner is a toy predictor whose output depends on which training set it saw. Enumerate four hypothetical training sets:

Training setPrediction f^(x;D)\hat{f}(x; D)
D1D_18
D2D_212
D3D_39
D4D_411

Mean prediction: fˉ(x)=(8+12+9+11)/4=10\bar{f}(x) = (8 + 12 + 9 + 11)/4 = 10.

Squared bias: (f(x)−fˉ(x))2=(10−10)2=0(f(x) - \bar{f}(x))^2 = (10 - 10)^2 = 0.

Variance: the average squared deviation from the mean:

(8−10)2+(12−10)2+(9−10)2+(11−10)24=4+4+1+14=2.5.\frac{(8-10)^2 + (12-10)^2 + (9-10)^2 + (11-10)^2}{4} = \frac{4 + 4 + 1 + 1}{4} = 2.5.

Noise: σ2=1\sigma^2 = 1.

Total: 0+2.5+1=3.50 + 2.5 + 1 = 3.5.

Now verify directly. For each training set, the expected squared error at xx is the squared deviation from f(x)f(x) plus the noise variance:

  • D1D_1: (10−8)2+1=5(10-8)^2 + 1 = 5
  • D2D_2: (10−12)2+1=5(10-12)^2 + 1 = 5
  • D3D_3: (10−9)2+1=2(10-9)^2 + 1 = 2
  • D4D_4: (10−11)2+1=2(10-11)^2 + 1 = 2

Average: (5+5+2+2)/4=3.5(5 + 5 + 2 + 2)/4 = 3.5. The arithmetic closes. That closing is the proof.

Two sanity checks worth internalizing. Change the model and the noise term stays at 1. Change the true function and the variance term stays at 2.5. Each term answers to a different master.

If you want to verify this over many simulated training sets rather than four, the same arithmetic scales directly:

import numpy as np

rng = np.random.default_rng(0)
f_x = 10.0
sigma = 1.0
n_sets = 10_000

# Toy predictor: true value plus a training-set-dependent offset
offsets = rng.normal(0, np.sqrt(2.5), n_sets)
predictions = f_x + offsets

bias_sq = (f_x - predictions.mean()) ** 2
variance = predictions.var()
noise = sigma ** 2
print(bias_sq + variance + noise)

The printed value converges to roughly 3.5. The code confirms the hand calculation; it does not replace it.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, the squared bias is 0, the prediction variance is 2.5, and the noise variance is 1. What total expected squared error does the decomposition give?
Output Prediction

Focus: Compute the decomposition’s total expected squared error from the example’s squared bias, variance, and noise.

What Each Term Actually Owns

Squared bias is owned by the model class and the fitting procedure. It is the gap between the average prediction and the truth. The decomposition is evaluated for a specified learning procedure and a specified training-set size. Change either one and the bias term can move. A more flexible model class can reduce bias by representing more of the structure in f(x)f(x); a larger training set can also shift the expected fitted predictor, which is why "more data never helps bias" is too strong a claim.

Variance is owned by the learner's sensitivity to the particular training sample. It is the spread of predictions across resampled datasets. A model that memorizes its training set will have high variance: small changes in the data produce large changes in the fit.

Noise is owned by the data-generating process. No model, no amount of data, and no amount of tuning removes it under these assumptions.

The tradeoff is a statement about how a change in model flexibility moves two of the three terms in opposite directions. It is not a law that the sum must always improve. Sometimes adding flexibility cuts bias more than it adds variance; sometimes it does the reverse. The decomposition tells you which is happening; it does not tell you the answer in advance.

Common mistake: Assuming the clean three-way split transfers to classification. It does not. The decomposition is a squared-loss result. For 0-1 loss, cross-entropy, or classification accuracy, the terms do not simply add, and the "main prediction" is defined differently. Treat the decomposition as a reasoning tool for squared-error regression, and as a hypothesis generator elsewhere.

Knowledge check

Check your understanding

Answer this question before you continue.

A learner is retrained on resampled datasets, and its predictions at the same x move substantially from one fit to another. Which term describes this spread?
Scenario Interpretation

Focus: Identify prediction variance as sensitivity to the particular training sample.

From the Theorem to Training and Validation Curves

The theorem averages over infinitely many training sets. In practice you have one dataset, so you cannot compute ED[f^(x;D)]\mathbb{E}_D[\hat{f}(x; D)] directly. This is where most people confuse the idealized expectation with the quantities they can actually measure.

Training error and validation error are single-sample estimates. They are not bias and variance. A large train-validation gap is evidence about variance, not a measurement of it. The gap tells you the model fits its particular training sample much better than new data; it does not tell you how much of the gap is variance versus noise versus bias.

Resampling methods give you a practical proxy. Repeated splits, cross-validation, and bagging-style ensembles let you observe the spread of predictions across training sets. Averaging predictions over many resampled models is the closest practical analogue of fˉ(x)\bar{f}(x). This is one reason ensembles can reduce the variance term: they approximate the mean prediction, and the mean prediction has no variance around itself. The size of that reduction depends on how the members are constructed and how correlated their errors are; averaging highly correlated models buys less than averaging diverse ones.

Common mistake: Reading a high validation error as proof of high bias. A high validation error does not tell you which term is large. You need a comparison across model complexity or across resamples to attribute the error. One number, one dataset, one model — that is not enough to separate the causes.

Assumptions That Break the Clean Split

The derivation is clean because the assumptions are strong. Here is where it stops being a reliable guide.

Nonzero-mean noise. If E[ε]≠0\mathbb{E}[\varepsilon] \neq 0, the bias term absorbs the noise mean, and the "irreducible" label becomes misleading. The noise is still there; it is just hiding inside bias.

Residual dependence that breaks the cross-term argument. The cross term vanishes because the residual y−fˉ(x)y - \bar{f}(x) has mean zero at the fixed xx and is independent of the training-set randomness. What matters is that the held-out outcome's residual is mean-zero at xx and appropriately independent of the training procedure. A residual that is correlated with the input xx is not by itself a failure: as long as f(x)f(x) is the conditional mean, the residual can vary with xx and the pointwise decomposition still holds. The decomposition breaks when the residual's mean at xx is not zero, or when the training procedure and the held-out outcome are no longer independent in the way the cross-term argument requires.

Non-squared loss. Classification losses need different machinery. The decomposition changes form, and the terms do not simply add.

Distribution shift. If the data distribution changes between training and deployment, the expectation is taken over the wrong distribution. The decomposition describes a world that no longer exists.

The practical rule: use the decomposition as a reasoning tool for squared-error regression. Treat it as a hypothesis generator elsewhere, not a measurement.

Where to Go Next

Before you add complexity or collect more data, ask which of the three terms you are actually attacking. Then name the experiment that would show it.

A concrete next step: take a low-flexibility model and a high-flexibility model on your own regression data. Run both through repeated train-validation splits. Record the spread of predictions across resamples, not just the mean error. If the high-flexibility model's predictions swing widely across resamples while its average prediction barely moves, you are watching variance. To attribute bias, you need a known or defensibly estimated target function — as in a simulation or a controlled example where you set f(x)f(x) yourself. On ordinary data, the true conditional target is not directly observed, so a stable distance from one observed label cannot by itself identify bias. Repeated resampling tells you about prediction variability and model stability; bias attribution requires additional assumptions or comparisons.

That experiment turns the algebra into a debugging routine. The intuition-level tradeoff article gives you the diagnostic categories; the learning-curve article gives you the workflow. This derivation gives you the reason the categories are real.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

One fitted model has much lower training error than validation error. What conclusion is justified by the article?
Question 1 of 2Comparison Reasoning

Focus: Distinguish practical train-validation diagnostics from the theoretical bias and variance terms.

If the additive noise has nonzero mean, what does the article say happens to the clean interpretation of the decomposition?
Question 2 of 2Misconception Check

Focus: Explain how nonzero-mean noise affects the interpretation of the clean decomposition.

References

  1. Single estimator versus bagging: bias-variance decomposition — scikit-learn 1.9.0 documentationscikit-learn.org
  2. Bias/Variance Tradeoffocw.mit.edu
  3. Rethinking Bias-Variance Trade-off for Generalization of ...proceedings.mlr.press
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.