Skip to content
advanced

Why Averaging Trees Helps: Derive the Random-Forest Variance Effect

Everyone repeats that a forest reduces variance. Almost nobody writes down the formula that governs the reduction — and that formula contains a term that…

Published 2026-10-02Updated 2026-10-048 min read
Stunning aerial photo of Liloan's pristine island coastline with lush greenery.
Stunning aerial photo of Liloan's pristine island coastline with lush greenery. Photo by Ariel Raz on Pexels.

Everyone repeats that a forest reduces variance. Almost nobody writes down the formula that governs the reduction — and that formula contains a term that never goes away, no matter how many trees you add.

This is the derivation behind that claim. We will assume every tree has the same marginal variance and that any two trees share one common pairwise correlation, derive the variance of their average, and then work a small numerical example. The result is a model of the mechanism, not a proof about your dataset. Bias is out of scope, and the formula will not predict your held-out error.

What the Variance Claim Actually Says

The quantity we care about is the variance of the average of BB predictors, not the variance of a single tree. Those are different objects, and conflating them is where most hand-waving starts.

If you have already read why a forest beats a single tree, and how bagging differs from boosting, you have the intuition. This article supplies the algebra. Boosting is sequential error correction; we are only handling the parallel-averaging case here.

Two assumptions do all the work:

  1. Equal marginal variance. Every tree's prediction has the same variance σ2\sigma^2.
  2. Common pairwise correlation. Any two distinct trees have the same correlation ρ\rho.

Both are modeling conveniences, not properties of real trees. We will state exactly where they break later. For now, accept them and watch what falls out.

Notation and the Two Assumptions

Fix an input point xx and fix the training sample. Let

T1,T2,…,TBT_1, T_2, \dots, T_B

be the predictions of BB trees at xx. Each TiT_i is a random variable because the tree depends on a bootstrap resample of the fixed training data and on the random feature subsets chosen at each split. The training sample itself is held constant; the randomness we are modeling is the tree-construction randomness, not the randomness of drawing a new dataset.

The ensemble prediction is the arithmetic mean:

Tˉ=1B∑i=1BTi\bar{T} = \frac{1}{B} \sum_{i=1}^{B} T_i

We assume:

Var⁡(Ti)=σ2for every i\operatorname{Var}(T_i) = \sigma^2 \quad \text{for every } i Corr⁡(Ti,Tj)=ρfor every i≠j\operatorname{Corr}(T_i, T_j) = \rho \quad \text{for every } i \neq j

Equal variance is a convenience. Real trees in a forest differ in depth, in which features they split on, and in how much of the data they see, so their variances are not identical. A single correlation parameter is a stronger simplification still: real trees have heterogeneous, data-dependent correlations, and the correlation between two trees depends on where in feature space you evaluate them.

That last point matters. The correlation here is between predictions at a fixed input xx, not between errors. It is a local quantity. Two trees may agree closely in one region of the input space and disagree sharply in another. The averaging argument is therefore local, not global — it describes what happens at one point, and you would have to integrate over xx to say anything about overall performance.

Knowledge check

Check your understanding

Answer this question before you continue.

In the derivation's setup, what source of randomness makes each tree prediction T_i a random variable?
Single Choice

Focus: Distinguish tree-construction randomness from variation caused by drawing a new training dataset in the setup.

Deriving the Variance of the Average

Start from the definition of variance for a sum, with the constant factor pulled out:

Var⁡(Tˉ)=Var⁡ ⁣(1B∑i=1BTi)=1B2Var⁡ ⁣(∑i=1BTi)\operatorname{Var}(\bar{T}) = \operatorname{Var}\!\left(\frac{1}{B}\sum_{i=1}^{B} T_i\right) = \frac{1}{B^2} \operatorname{Var}\!\left(\sum_{i=1}^{B} T_i\right)

The variance of a sum expands into a double sum over all ordered pairs:

Var⁡ ⁣(∑i=1BTi)=∑i=1B∑j=1BCov⁡(Ti,Tj)\operatorname{Var}\!\left(\sum_{i=1}^{B} T_i\right) = \sum_{i=1}^{B}\sum_{j=1}^{B} \operatorname{Cov}(T_i, T_j)

Split that double sum into two parts. When i=ji = j, the covariance is just the variance. When i≠ji \neq j, it is the covariance between two distinct trees.

=∑i=1BVar⁡(Ti)⏟diagonal+∑i≠jCov⁡(Ti,Tj)⏟off-diagonal= \underbrace{\sum_{i=1}^{B} \operatorname{Var}(T_i)}_{\text{diagonal}} + \underbrace{\sum_{i \neq j} \operatorname{Cov}(T_i, T_j)}_{\text{off-diagonal}}

Now count the terms. There are BB diagonal entries and B(B−1)B(B-1) off-diagonal entries. Substitute the shared variance and the shared covariance. Since Cov⁡(Ti,Tj)=ρσ2\operatorname{Cov}(T_i, T_j) = \rho\sigma^2 for i≠ji \neq j:

Var⁡ ⁣(∑i=1BTi)=Bσ2+B(B−1)ρσ2\operatorname{Var}\!\left(\sum_{i=1}^{B} T_i\right) = B\sigma^2 + B(B-1)\rho\sigma^2

Divide by B2B^2:

Var⁡(Tˉ)=Bσ2+B(B−1)ρσ2B2=σ2B+B−1Bρσ2\operatorname{Var}(\bar{T}) = \frac{B\sigma^2 + B(B-1)\rho\sigma^2}{B^2} = \frac{\sigma^2}{B} + \frac{B-1}{B}\rho\sigma^2

Factor out σ2\sigma^2 and you get the closed form:

Var⁡(Tˉ)=ρσ2+1−ρBσ2\boxed{\operatorname{Var}(\bar{T}) = \rho\sigma^2 + \frac{1-\rho}{B}\sigma^2}

Read the two terms separately.

The first term, ρσ2\rho\sigma^2, does not contain BB. It is the irreducible floor set by the correlation between trees. If every tree made the same mistake, ρ=1\rho = 1, and the floor equals the variance of a single tree — averaging buys you nothing.

The second term, 1−ρBσ2\frac{1-\rho}{B}\sigma^2, is the part averaging can actually remove. It shrinks as 1/B1/B.

Now take the limit:

lim⁡B→∞Var⁡(Tˉ)=ρσ2\lim_{B \to \infty} \operatorname{Var}(\bar{T}) = \rho\sigma^2

The second term vanishes. The first survives. Averaging can only remove the uncorrelated part of the variance, and it can never remove the correlated part.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's assumptions, what is Var(T̄) when B = 4, σ² = 2, and ρ = 0.25?
Output Prediction

Focus: Apply the derived variance formula to a specified ensemble size, marginal variance, and common correlation.

A Small Numerical Example

A curve for correlation 0.2 falls steeply at first, then levels toward a horizontal variance floor at 0.2 as the number of trees increases; points mark 1, 10, 100, and 1000 trees.
With correlation fixed at 0.2, adding trees removes the 1/B component but cannot lower variance below 0.2.

Pick σ2=1\sigma^2 = 1 and ρ=0.2\rho = 0.2. Then:

Var⁡(Tˉ)=0.2+0.8B\operatorname{Var}(\bar{T}) = 0.2 + \frac{0.8}{B}
Trees BBSecond term 0.8B\frac{0.8}{B}Ensemble variance
10.8001.000
20.4000.600
50.1600.360
100.0800.280
500.0160.216
1000.0080.208
10000.00080.2008

The decay is fast at first and then stalls. Going from 1 tree to 10 cuts variance from 1.0 to 0.28 — a 72% reduction. Going from 10 to 1000 cuts it from 0.28 to 0.2008 — another 28%, but you paid a hundredfold in compute. The floor is 0.20.2, and no number of trees will push you below it.

Now hold B=100B = 100 fixed and vary ρ\rho:

Correlation ρ\rhoEnsemble variance
0.00.008
0.10.109
0.20.208
0.50.505
0.90.901

Changing ρ\rho moves the floor. It does not change the decay rate, which is always 1/B1/B. This is the practical reading: past a modest number of trees, more trees buy almost nothing, and only decorrelation lowers the floor.

Knowledge check

Check your understanding

Answer this question before you continue.

For σ² = 1 and B = 100, what ensemble variance does the formula give if ρ = 0.5?
Output Prediction

Focus: Calculate ensemble variance in the numerical example when the common correlation is changed.

From the Formula to Forest Diversity

The formula tells you where to push. If the floor is ρσ2\rho\sigma^2, then the only lever that lowers it is ρ\rho.

Bootstrap sampling and per-split feature subsampling are both attempts to push ρ\rho down. Bootstrap sampling shows each tree a different resample of the data, so trees disagree about which observations matter. Feature subsampling forces trees to consider different candidate splits, so they disagree about which features matter. Both mechanisms inject disagreement, and disagreement is what lowers ρ\rho.

There is a tradeoff. More aggressive randomization can raise the variance of individual trees — a tree that only sees a random third of the features at each split may be a worse tree than one that sees all of them. It can also change bias. The formula captures only the variance side, and only under the equal-variance assumption. If randomization makes each tree noisier, σ2\sigma^2 goes up even as ρ\rho goes down, and the net effect on the ensemble is not obvious from the formula alone.

Keep the causal claim cautious. The formula motivates the design of random forests. It does not prove that any particular randomization scheme helps on a given dataset. That is an empirical question, and it is answered by held-out evaluation, not by algebra.

Knowledge check

Check your understanding

Answer this question before you continue.

A practitioner increases feature subsampling and sees less agreement among trees. Which conclusion best matches the article's reasoning?
Scenario Interpretation

Focus: Explain how randomization motivates forest diversity while recognizing the variance tradeoff and limits of the simplified result.

Where the Derivation Stops Being True

The formula is a mechanism explanation, and it has clear failure modes.

Real trees are not exchangeable. Their variances differ, and their correlations depend on the input point and on the training data. The equal-variance assumption is a modeling convenience; the common-correlation assumption is a stronger one. When trees have different variances, the algebra still works but the closed form becomes a weighted expression with no single ρ\rho to point at.

The formula says nothing about bias. Randomization can change bias as well as variance. A variance-only account is incomplete, and treating the formula as a full bias-variance decomposition will mislead you. There is active research suggesting that feature randomization can reduce bias relative to plain bagging, not just variance — which the formula above cannot see.

Correlation is not a fixed constant you can read off a fitted forest. It is a modeling assumption, not a measured quantity. You cannot compute ρ\rho from a trained model and plug it in. You can estimate pairwise prediction correlations on held-out data, but that estimate is local to the data you measured it on.

The derivation is conditional on the training sample. It describes the variance of the ensemble's prediction at a fixed input, given the data. It does not describe how the ensemble behaves across resampled datasets. Those are different sources of variability, and the formula only addresses one.

Use the formula to reason about direction and limits, not to predict a number.

What to Take Away

The variance of an average of BB correlated predictors is:

Var⁡(Tˉ)=ρσ2+1−ρBσ2\operatorname{Var}(\bar{T}) = \rho\sigma^2 + \frac{1-\rho}{B}\sigma^2

A floor set by correlation, plus a term that averaging can remove. That is the whole structure.

The practical rule follows directly. Increase ensemble size until the second term is negligible relative to the first, then stop adding trees and attack correlation if you need more. In scikit-learn, that means tuning max_features and the sampling scheme rather than n_estimators.

The honest boundary: the math explains why averaging helps. It does not tell you how much it will help on your data. To find that out, run a controlled experiment. Hold the data split fixed, vary ensemble size and randomization settings, and compare held-out behavior. The formula tells you what to expect in direction. The experiment tells you what actually happened.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Under the simplified formula, if the term that shrinks with ensemble size is already negligible, which change most directly targets any remaining variance floor, assuming marginal variance is held fixed?
Question 1 of 2Comparison Reasoning

Focus: Choose the variance-reduction lever suggested by the formula once the ensemble-size term is already small.

A learner uses the derived expression as a direct prediction of a forest's held-out error. Which correction is most accurate?
Question 2 of 2Misconception Check

Focus: Identify the scope limits of the variance derivation and distinguish conditional variance from a complete account of predictive performance.

References

  1. Randomization Can Reduce Both Bias and Variancewww.jmlr.org
  2. Chapter 11 Random Forests | Hands-On Machine Learning with Rbradleyboehmke.github.io
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.