Skip to content
intermediate

Why a Bootstrap Sample Leaves Observations Out of Bag

A random forest hands you an out-of-bag score, and someone nearby says, "Each tree leaves out about a third of the data." That number feels like a setting.…

Published 2026-10-02Updated 2026-10-048 min read
Close-up of a yellow Ethernet cable with connectors on a blue background.
Close-up of a yellow Ethernet cable with connectors on a blue background. Photo by Ann H on Pexels.

A random forest hands you an out-of-bag score, and someone nearby says, "Each tree leaves out about a third of the data." That number feels like a setting. It is not. It is a fixed consequence of one sampling rule, and once you derive it, you stop treating OOB as magic and start reading it as a coverage count.

Here is the question worth answering precisely: if I draw a bootstrap sample of size nn from nn training rows, what is the probability that one specific row — call it observation ii — never gets picked at all?

What a Bootstrap Sample Actually Does

Start with the mechanism, because the probability has nothing to attach to until you can picture the draw.

You have nn training observations. A bootstrap sample draws nn items with replacement. Each draw picks one of the nn rows uniformly, so any single row has probability 1/n1/n of being chosen on that draw. You repeat this nn times. The result is a sample of the same size as the original, but it is a multiset: some rows appear twice, some three times, and some never appear.

That last group is the point. The out-of-bag set for one tree is the set of training rows that never appeared in that tree's bootstrap sample. It is not a separate holdout split you carved off. It is a by-product of the resampling itself.

Bagging trains each base learner on its own resample and aggregates the predictions. The OOB set falls out of that process for free. So the real question is not "how many rows are left out" — that varies per sample — but "what is the probability that this specific row is left out." That is a marginal probability, and it is stable.

Knowledge check

Check your understanding

Answer this question before you continue.

Which description correctly explains why a row can be out of bag for a tree?
Misconception Check

Focus: Explain how replacement sampling creates an out-of-bag set for an individual tree.

Deriving the Omission Probability

A fixed row i is shown outside each of n independent draw slots, each with miss probability 1 minus 1 over n. The all-miss path is labeled (1 − 1/n)^n; its complement, 1 − (1 − 1/n)^n, is the probability the row appears at least once.
Multiplying the per-draw miss chance across n independent draws gives the omission probability; inclusion is its complement.

Let's build it one draw at a time. Define nn as the number of training observations, and fix your attention on a single observation ii.

Step 1: one draw. On a single draw, the probability that ii is not selected is the complement of being selected:

P(not selected on one draw)=1−1nP(\text{not selected on one draw}) = 1 - \frac{1}{n}

Step 2: all nn draws. The draws are independent and identically distributed. Each draw has the same 1/n1/n chance of hitting ii, and one draw's outcome tells you nothing about the next. For independent events, the probability that all of them miss is the product:

P(i omitted)=(1−1n)nP(i \text{ omitted}) = \left(1 - \frac{1}{n}\right)^n

Step 3: the complement. The probability that ii appears at least once is:

P(i included)=1−(1−1n)nP(i \text{ included}) = 1 - \left(1 - \frac{1}{n}\right)^n

That is the whole derivation. Two lines, one assumption.

Step 4: name the assumption. The result depends on the draws being independent and identically distributed, with each row carrying equal probability 1/n1/n. If your resampling scheme weights rows unequally, or draws are somehow dependent, this formula no longer describes your sample. The math is exact only under the standard bootstrap.

Step 5: read the scope. This is a per-observation marginal probability. It does not say "exactly 37% of rows are left out." It says each row, considered on its own, has this chance of being omitted. The actual count of distinct omitted rows in any one sample is a random variable with its own distribution.

Common mistake: Treating (1−1/n)n(1 - 1/n)^n as the fraction of rows omitted in a given sample. It is the expected omission probability per row, not a guarantee about any single resample.

Knowledge check

Check your understanding

Answer this question before you continue.

A bootstrap sample draws 4 times from 4 equally likely rows, with replacement. What is the probability that one specified row is omitted from all 4 draws?
Single Choice

Focus: Apply the independent-draw derivation to calculate a fixed row's omission probability.

The Large-Sample Limit and Why It Is Roughly 37%

Now push nn upward and watch what the formula does.

As nn grows, (1−1n)n\left(1 - \frac{1}{n}\right)^n approaches 1/e1/e, which is about 0.3680.368. So the omission probability approaches roughly 36.8%, and the inclusion probability approaches about 63.2%.

The convergence is fast enough that the rule of thumb holds at modest sizes. Here are worked numbers:

nn(1−1/n)n(1 - 1/n)^nOmission
100.3487~34.9%
1000.3660~36.6%
10000.3677~36.8%

At n=10n = 10 the gap from 1/e1/e is real — about two percentage points — and worth computing rather than assuming. By n=100n = 100 you are already close.

The limit is a constant, not a function of dataset size. That is why "about a third left out" survives across wildly different datasets. It is baked into the sampling rule, not into your data.

One interpretation worth holding onto: roughly 63% of unique rows appear in a given resample, so the expected number of distinct rows is about 0.632n0.632n. The rest are out of bag.

Note: 1/e1/e is a limit, not an exact value. At small nn, compute the actual probability instead of rounding to 37%.

Knowledge check

Check your understanding

Answer this question before you continue.

For a bootstrap sample of size 10 drawn from 10 rows, which interpretation best matches the article's calculation?
Comparison Reasoning

Focus: Interpret the large-sample limit for omission and distinguish it from the finite-sample value.

From Omission Probability to OOB Coverage

Now scale from one tree to a forest. This is where the per-sample probability turns into something you can observe in a model.

Let p=(1−1/n)np = (1 - 1/n)^n be the omission probability for a single tree, and let BB be the number of trees. For a fixed observation ii, each tree omits ii independently with probability pp. This independence holds because the trees' bootstrap samples are drawn independently under the same sampling rule — that is the standard bagging setup, and it is the condition that makes the count below a clean binomial. The number of trees that can predict ii is therefore Binomial(B,p)(B, p).

The expected number of OOB trees for a given row is about 0.368B0.368B. With B=100B = 100, that is roughly 37 trees. This is why OOB estimates stabilize as the forest grows: more trees means more OOB votes per row, and the average settles down.

The probability that a row is OOB for at least one tree is:

1−(1−p)B1 - (1 - p)^B

This rises quickly with BB, but it is never exactly 1 for finite BB. There is always some chance a row is in-bag for every single tree.

That is the mechanism behind a behavior you may have already hit: scikit-learn's oob_decision_function_ can contain NaN when n_estimators is small, because a data point was never left out during the bootstrap. The math predicted it before the code showed it.

Warning: The derivation describes coverage, not accuracy. It tells you how many trees can vote on a row. It says nothing about whether those votes are any good.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's standard independent-bootstrap setup, a forest has 100 trees. About how many trees are expected to have a fixed row out of bag?
Scenario Interpretation

Focus: Relate per-tree omission probability to the expected number of OOB predictions across a forest.

What the Probability Does Not Promise

The tempting overreach is to conclude that "about 37% left out" means the OOB score is a free, unbiased test score. It is not that simple, and the distinction matters.

Coverage is not independence. For each tree that contributes to a row's OOB prediction, that row was genuinely excluded from training. But the contributing trees are correlated with one another — they were grown on overlapping resamples of the same data — and the OOB estimate is an internal resampling estimate, not a separate external evaluation. The row is held out of the trees that vote on it; the estimate as a whole still lives inside the training data.

OOB error has known weak regimes. It has been reported to behave poorly in settings with small sample sizes, many predictors, weak effects, and balanced classes. Treat it as a useful approximation, not a guarantee.

The arithmetic is configuration-dependent. Change max_samples, and the omission probability changes. Set bootstrap=False, and there is no OOB set at all — scikit-learn only exposes oob_score when bootstrap=True. The ~37% figure assumes the default: nn draws of size nn with replacement.

My rule: use OOB as a cheap internal signal while you are developing. Keep a reserved evaluation set for the decisions that actually matter.

Check the Arithmetic Yourself

The derivation earns its keep when it predicts what you observe. Verify it in a few lines.

Compute the limit directly:

import numpy as np

for n in [10, 100, 1000]:
    print(n, (1 - 1/n)**n, 1/np.e)

Then simulate the draw and compare the empirical omission rate to the formula:

import numpy as np

rng = np.random.default_rng(0)
n, trials = 100, 20_000
omitted = sum(0 not in rng.integers(0, n, size=n) for _ in range(trials))
print(omitted / trials, (1 - 1/n)**n)

The two numbers should land close together. That is the point: the formula predicts the output before you run it.

If you want to connect it to a real model, fit a small RandomForestClassifier with oob_score=True and inspect how many trees actually cover each row. Compare the observed counts to the binomial expectation of about 0.368B0.368B. Keep this as verification of the derivation, not a replacement for it.

Where This Leaves You

The ~37% figure is a fixed consequence of sampling nn times with replacement. It tells you how much OOB evidence exists per row — not how trustworthy that evidence is. Those are two different questions, and conflating them is how people end up trusting an OOB score they should have double-checked.

The next practical step: watch how OOB estimates behave as the forest grows, and compare them against a reserved evaluation set. If the two diverge in a way that surprises you, you have found a regime where coverage was never the whole story.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A researcher observes that each row receives many OOB votes in a large forest. What conclusion is supported by that coverage information alone?
Question 1 of 2Misconception Check

Focus: Distinguish OOB prediction coverage from the accuracy or reliability of the resulting estimate.

With a small number of trees, a row has no OOB prediction. Which explanation and response best fit the article?
Question 2 of 2Scenario Interpretation

Focus: Explain why a row may lack an OOB prediction in a small forest and choose a cautious interpretation of OOB results.

References

  1. RandomForestClassifier — scikit-learn 1.9.1 documentationscikit-learn.org
  2. (out-of-bag estimates) - UC Berkeley Statisticswww.stat.berkeley.edu
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.