Skip to content
intermediate

Why the Best Validation Score Is Often Too Optimistic

You run a grid search. The leaderboard prints a winner at 0.91. You ship it, and the first honest measurement lands near 0.86. Nothing broke. The model did…

Published 2026-10-02Updated 2026-10-049 min read
A captivating sunrise over a foggy landscape with vibrant warm colors and silhouettes.
A captivating sunrise over a foggy landscape with vibrant warm colors and silhouettes. Photo by Anton Kudryashov on Pexels.

You run a grid search. The leaderboard prints a winner at 0.91. You ship it, and the first honest measurement lands near 0.86. Nothing broke. The model did not degrade overnight. The number was never a measurement of your model — it was the maximum of a noisy set, and maxima run high.

The Score You Report Is a Maximum, Not an Average

Beginners collapse two different quantities into one. The first is the score of a specific candidate: train this configuration, evaluate it, read the number. The second is the score of the candidate that won a comparison: run many configurations, sort the results, and report the top of the list.

Those are not the same random variable. The first is a measurement. The second is the output of a selection operator applied to a set of measurements.

Selection is an operation on numbers. If you hand me twenty estimates and ask for the best, I return the maximum. The maximum of a set behaves differently from any single member of that set — it leans upward, and it leans harder as the set grows. That lean is what we call selection-induced optimism, and it is the specific failure this article is about.

To be precise about scope: this is not data leakage. It is not distribution shift. It is not the mechanics of nested cross-validation. You already know validation data must stay separate from training data. The question here is narrower and sharper: what happens when the same validation data is used to choose among candidates?

Note: Leakage contaminates the estimate itself. Selection optimism is a property of taking a maximum over honest but noisy estimates. The estimates can be perfectly clean and the reported score can still be too high.

Knowledge check

Check your understanding

Answer this question before you continue.

A search evaluates several candidates on the same validation data and reports the winner. What quantity is that reported score?
Misconception Check

Focus: Distinguish a score measured for one candidate from the score reported after selecting among candidates.

Notation and the Assumptions Behind the Illustration

Before the arithmetic, define the symbols. I want you to be able to judge exactly how far this model generalizes.

Let there be K candidates, indexed by k=1,…,Kk = 1, \dots, K. Each candidate has a true generalization score μk\mu_k — the number you would get if you could evaluate it on the entire population of future data. You never observe μk\mu_k. You observe a validation estimate sks_k.

Model the estimate as the truth plus an error term:

sk=μk+eks_k = \mu_k + e_k

where eke_k is a zero-mean error with standard deviation σ\sigma, and the errors are independent across candidates.

Four assumptions are doing real work here, and I want them labeled as assumptions rather than facts:

  • Finite candidate set. You compared a known, countable number of configurations.
  • Roughly unbiased individual estimates. Each sks_k is centered on its own μk\mu_k. No systematic contamination.
  • Comparable noise scale. Every candidate's estimate carries roughly the same σ\sigma.
  • Independent errors. The eke_k do not move together.

The quantity you report is max⁡ksk\max_k s_k. The quantity you want is the true score of the selected candidate, μk∗\mu_{k^*}, where k∗=arg⁡max⁡kskk^* = \arg\max_k s_k.

The first assumption that breaks in practice is independence. Candidates in a real grid share folds, share features, share preprocessing. Their errors tend to move together. When that shared component is positive, comparisons become less independent and the extra-max effect shrinks relative to the independent case. How much it shrinks depends on the correlation structure, not on a single fixed factor. The direction, though, does not flip: correlated noise still pushes the maximum upward.

Knowledge check

Check your understanding

Answer this question before you continue.

If candidate k* is chosen because it has the largest validation estimate, which quantity represents the true score you want to know for that selected candidate?
Single Choice

Focus: Identify the true-score quantity corresponding to the candidate selected by its validation estimate.

Why the Maximum Drifts Upward: A Worked Example

A horizontal score axis marks the shared true score at 0.80. Candidate estimates cluster around it, with the rightmost point circled at about 0.837 to show the upward gap created by selecting the maximum.
When equally good candidates have noisy estimates, choosing the highest score selects an upward fluctuation.

Now make it concrete. Suppose you compare K=20K = 20 candidates that are genuinely identical. Every μk\mu_k equals the same value, say 0.800.80. No candidate is actually better than any other. The only thing separating them is noise.

Give each estimate a noise scale of σ=0.02\sigma = 0.02.

If the errors are roughly normal, each sks_k is a draw from a distribution centered at 0.80 with standard deviation 0.02. Now ask: what does the maximum of twenty such draws look like?

The maximum of KK independent draws from a distribution sits above the mean by roughly σ\sigma times a factor that grows with KK. For twenty draws from a normal distribution, that factor is approximately 1.87. So the expected maximum is roughly:

0.80+1.87×0.02≈0.8370.80 + 1.87 \times 0.02 \approx 0.837

Read that number carefully. No candidate is better than any other. The truth is 0.80 for all of them. Yet the winner's validation score lands near 0.84 — an optimism of about four points, manufactured entirely by the act of taking a maximum.

Two consequences follow immediately.

First, the gap scales with σ\sigma. Noisier estimates mean a larger reported score, even with the same underlying models. Small validation sets have large σ\sigma, which is exactly where beginners have the least data to spare.

Second, the gap scales with KK. Try forty candidates instead of twenty and the factor rises to about 2.07, pushing the expected maximum to roughly 0.841. The more configurations you try, the higher your reported score climbs — even when the models are indistinguishable.

Picture a dot plot: twenty points scattered around a single marked true value at 0.80, and the rightmost dot circled. The circle is your "winner." The distance from the circle to the true value is the optimism. It is not a property of the winning model. It is a property of the circle.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article’s example, 20 identical candidates each have true score 0.80 and noise scale 0.02; the approximate normal-maximum factor is 1.87. What expected winning validation score does this imply?
Output Prediction

Focus: Use the article’s bounded noisy-estimate example to estimate how selection changes the reported score.

What the Illustration Does and Does Not Claim

Under the stated assumptions, the example establishes one thing: the expected reported score exceeds the expected true score of the selected candidate. That is the whole claim.

It does not say the selected model is bad. Selection may still hand you the best available choice. The problem is the number attached to it, not the choice itself.

Where the assumptions fail, the magnitude changes:

  • Correlated candidates shrink the effect relative to the independent case, because shared noise does not stack the same way independent noise does. The exact reduction depends on how strongly the errors move together.
  • Unequal noise across candidates complicates the picture — a high-variance candidate is more likely to win by luck.
  • Small validation sets inflate σ\sigma and therefore inflate the optimism, right where you can least afford it.

Keep this distinct from leakage. Leakage means the held-out data indirectly informed the model. Selection optimism means the held-out data honestly ranked a set of honest estimates, and you reported the top of the ranking as if it were a single measurement.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the illustration’s stated assumptions, which conclusion does it support?
Comparison Reasoning

Focus: State the claim supported by the illustration without overstating what selection optimism implies.

Separated Evaluation: Paying Once for an Honest Number

The fix is structural, not statistical sleight of hand. The data used to select must not be the data used to report.

In the notation above, the bias comes from reporting max⁡ksk\max_k s_k when the sks_k are the same draws you used to pick k∗k^*. On fresh data, the selected candidate's estimate is no longer a maximum over those noise draws. The upward bias is removed, not merely reduced.

This maps onto a workflow you already have. Selection happens on validation folds. Reporting happens once, on held-out data. The held-out number is allowed to be lower than the validation winner — that is the point, not a failure.

The cost is honest: separated evaluation spends data and compute, and it gives you one number instead of a leaderboard. That is the trade. A leaderboard is a ranking tool; a held-out score is a performance claim. Confusing the two is how the 0.91 becomes a 0.86 in production.

Decision rule: Treat the winning validation score as a ranking signal, not a performance claim, whenever the same data informed the selection. The more candidates you compared and the noisier the estimate, the wider the gap you should expect between the winner and the truth.

Common Mistakes That Recreate the Bias

The theory is only useful if you can spot it in your own experiments. These patterns recreate the maximum:

  • Reporting the best cross-validation score from a large grid as if it were the model's expected performance. The grid searched; the score is the search's output.
  • Re-running the search after seeing the test score. This converts the test set into a selection set and reintroduces the same maximum on data you promised not to touch.
  • Comparing candidates across different folds, seeds, or preprocessing steps. Extra noise gets amplified by the maximum.
  • Choosing a winner from a difference smaller than the noise scale. When the gap is under σ\sigma and the candidates are otherwise comparable, the ranking is weak evidence — the winner may simply be the luckiest draw, not the better model.
  • Treating one lucky run as progress. If the gap does not survive a change of seed or fold split, it was never a gap.

A Practical Rule for Reading Your Own Leaderboard

Before you trust a winning score, ask three questions:

  1. How many candidates were compared?
  2. How noisy is the estimate?
  3. Was the reporting data ever used for selection?

If the answer to the third question is yes, the number is optimistic by construction. If the candidate count is large and the noise is high, expect the gap to be wide. If the candidates are strongly correlated, expect a narrower gap than the independent-noise illustration predicts — but do not expect zero.

I would rather ship an honest 0.86 I can act on than an optimistic 0.91 I cannot. The honest number tells me where the model actually stands, which is the only place improvement can start.

Your next move: audit your last experiment. Count how many candidates you compared, and check whether the number you reported came from data that saw the selection. Then look forward to the next step in this curriculum — evaluating the entire selection procedure rather than a single chosen model, and quantifying uncertainty around the final estimate so you know how much of the gap is signal and how much is the circle.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A team evaluates its chosen model on a test set, sees an unexpectedly low score, then reruns the search and keeps the candidate with the best score on that same test set. What has the team done?
Question 1 of 2Scenario Interpretation

Focus: Recognize when a supposedly held-out test set has become part of model selection.

A team compares many candidates using noisy validation estimates, and the candidates are strongly correlated. Which interpretation best matches the article?
Question 2 of 2Scenario Interpretation

Focus: Interpret how candidate count, estimate noise, and candidate correlation affect expected selection optimism.

References

  1. 3.5. Validation curves: plotting scores to evaluate models — scikit-learn 1.9.0 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.