Skip to content
advanced

How Certain Is a Cross-Validation Score? Derive the Limits of Fold-Based Uncertainty

Two models. Mean cross-validation scores of 0.842 and 0.847. The instinct is immediate: the second one wins.

Published 2026-10-02Updated 2026-10-049 min read
Vibrant close-up of network cable connectors with colorful lighting.
Vibrant close-up of network cable connectors with colorful lighting. Photo by Nic Wood on Pexels.

Two models. Mean cross-validation scores of 0.842 and 0.847. The instinct is immediate: the second one wins.

That instinct rests on a hidden assumption — that the fold scores are independent draws, so their spread can be divided by k\sqrt{k} and read as a standard error. The folds are not independent. The training sets overlap heavily. And the number you get from that division is a lower bound on your ignorance, not an estimate of it.

Let's derive where that dependence comes from, and what it does and does not license you to conclude.

What a Fold Score Actually Measures

Fix notation first, because every uncertainty claim depends on what object you are describing.

Take a dataset of nn observations. Split it into kk folds. Let BjB_j be the held-out block for fold jj, and let TjT_j be the training set — the complement of BjB_j. Fit a model on TjT_j, evaluate it on BjB_j, and record the score:

sj=metric(f^Tj, Bj)s_j = \text{metric}\big(\hat{f}_{T_j},\, B_j\big)

The cross-validation estimate is the mean:

sˉ=1k∑j=1ksj\bar{s} = \frac{1}{k}\sum_{j=1}^{k} s_j

Each sjs_j is a random variable. Its randomness comes from the split, from the data, and from any algorithmic randomness such as initialization or subsampling. It is not a draw from a fixed population parameter with a known sampling distribution. The mean sˉ\bar{s} is a summary of one particular resampling procedure — this dataset, this kk, this splitter — not an estimate of a single well-defined quantity whose distribution you can look up.

That distinction is the whole article. If you already know k-fold mechanics and the grouped or temporal constraints that force different splitters, hold that as background. The question here is narrower: given the fold scores you already have, what can their spread tell you?

Picture the kk training sets as overlapping blocks. Fold 1 trains on blocks 2 through kk. Fold 2 trains on blocks 1, 3, 4, …, kk. The shared observations are not a rounding error. They are most of the data.

Knowledge check

Check your understanding

Answer this question before you continue.

What does the mean fold score, as defined here, summarize?
Single Choice

Focus: Interpret the cross-validation mean as a summary of the specified resampling procedure rather than as a universal performance quantity.

Why the Folds Are Not Independent Samples

A three-row by three-column matrix shows three folds and three data blocks. In each row, one block is held out and the other two are used for training; the held-out block rotates across rows. The shared training blocks are highlighted, and the resulting fold scores are visually linked.
Rotating the held-out block changes each fold, but the training data overlap—so the resulting scores are dependent.

Consider two folds jj and j′j'. Their training sets TjT_j and Tj′T_{j'} each contain k−1k-1 of the kk blocks. They share k−2k-2 of those blocks. The overlap fraction is:

k−2k−1\frac{k-2}{k-1}

For k=10k=10, that is 8/9≈89%8/9 \approx 89\% of the training data. Two models trained on 89% identical data will make correlated errors on their respective held-out blocks. When one fold's model happens to fit the idiosyncrasies of the shared data well, the other fold's model tends to as well. The errors move together.

This correlation is not incidental. It is induced by the resampling design, and it grows as kk grows. Larger kk means more shared training data per pair, which means tighter coupling between fold scores.

There is a second dependence channel. Each observation is used for training in k−1k-1 folds and for testing in exactly one. The fold scores are not draws from a common independent pool of errors; they are computed from a single dataset that every fold has already seen in some role.

Now state the assumption the naive formula needs. The standard error sd(sj)/k\text{sd}(s_j)/\sqrt{k} is valid when the sjs_j are independent and identically distributed. That assumption is false here. Positive correlation means the effective number of independent pieces of information is smaller than kk — often much smaller.

Knowledge check

Check your understanding

Answer this question before you continue.

With ten-fold cross-validation, two folds' training sets share eight of their nine training blocks. What does this overlap imply for interpreting the fold scores?
Misconception Check

Focus: Explain how overlapping training sets induce dependence among fold scores.

A Small Numerical Example: What the Naive Standard Error Misses

Make it concrete. Suppose you run k=10k=10 folds and get these scores:

s = [0.81, 0.86, 0.83, 0.88, 0.79, 0.85, 0.84, 0.87, 0.82, 0.85]

The mean is sˉ=0.840\bar{s} = 0.840. The sample standard deviation is roughly 0.0280.028. The naive standard error is:

senaive=0.02810≈0.0089\text{se}_{\text{naive}} = \frac{0.028}{\sqrt{10}} \approx 0.0089

A naive 95% interval would be sˉ±1.96⋅0.0089\bar{s} \pm 1.96 \cdot 0.0089, roughly [0.823,0.857][0.823, 0.857]. That interval looks tight. It suggests you know the score to within about a percentage point.

The interval is too narrow. The direction of the error is always optimistic: the naive formula treats kk correlated numbers as kk independent numbers, so it divides by a denominator that is too large. The true sampling variance of sˉ\bar{s} is larger than the naive estimate suggests, because the fold scores carry less independent information than their count implies.

I want to be careful here. I am not going to hand you a universal correction factor. The right adjustment depends on the stability of the learning algorithm and the dependence structure of the data. A highly stable learner — one whose fitted function barely changes when you swap a few training points — produces more correlated fold scores, so the naive formula understates uncertainty more severely. An unstable learner produces noisier, less correlated scores, and the gap narrows. There is no single multiplier that fixes this across algorithms and datasets.

What the example establishes is the direction and the mechanism, not a precise number. The naive interval is a floor, not a ceiling.

Knowledge check

Check your understanding

Answer this question before you continue.

For the ten scores in the example, the sample standard deviation is about 0.028 and the naive standard error is about 0.0089. Why should that standard error not be read as a reliable measure of uncertainty?
Scenario Interpretation

Focus: Use the numerical example to explain why dividing fold-score spread by the square root of the fold count can understate uncertainty.

Fold Spread Is Not a Confidence Interval

This is the distinction most practitioners blur, so let's be exact.

A confidence interval needs two things: a target quantity and a coverage guarantee. The standard deviation of fold scores provides neither by itself. It answers a different question — how much did the split move the score? — not how far is the mean from the truth?

Two targets get conflated constantly. The first is the performance of the specific model fitted on the full dataset — the one you would actually ship. The second is the expected performance of the learning algorithm across hypothetical new datasets drawn from the same distribution. Ordinary cross-validation does not directly measure the first: each fold evaluates a model trained on a different subset, not the full-data model. And it estimates the second only under assumptions about how the data were generated and how the algorithm responds to training-set size. Neither target is universally "more reliable" to estimate; which one you care about determines which validation design is appropriate.

And low variance does not imply low bias. Small fold spread can coexist with a badly biased estimate — if the folds all share the same systematic flaw, they will agree with each other while all being wrong. Agreement is not accuracy.

Common mistake: Reading a tight fold spread as evidence that the mean is close to the true generalization error. The spread measures sensitivity to the split. It says nothing about whether the split itself is representative.

My rule: report fold spread as a stability diagnostic, and treat any interval built from it as a lower bound on uncertainty unless a dependence-aware method was used.

Knowledge check

Check your understanding

Answer this question before you continue.

A practitioner observes a small standard deviation among fold scores. What conclusion is warranted by the article?
Misconception Check

Focus: Distinguish fold-to-fold spread as a stability diagnostic from an interval with a coverage guarantee.

Comparing Two Models Without Overclaiming

Back to 0.842 versus 0.847. Here is what most comparisons get wrong.

Fold scores for two models on the same splits are paired, not independent samples. Model A and Model B are evaluated on the identical held-out blocks. Comparing their means as if they were independent throws away the pairing and misstates the variance of the difference. The correct object of study is the per-fold difference dj=sj(A)−sj(B)d_j = s_j^{(A)} - s_j^{(B)}, not the two means in isolation.

Even with pairing handled, a difference smaller than the fold-to-fold spread is not evidence of a real difference. It is consistent with split noise. The honest question is whether the observed gap survives resampling of the data and the split — which is why a dependence-aware test or a properly designed resampling procedure is the appropriate tool.

Practical guidance:

  • Compare per-fold differences, not aggregate means, when the splits are shared.
  • Watch for selection effects. Picking the best of many candidates by mean CV score inflates the winner's apparent advantage. That is a separate problem from dependence, but it compounds it.
  • When the decision hinges on a small gap, treat the comparison as inconclusive until you have evidence that accounts for dependence, not just a tighter-looking number.

When This Level of Caution Is Worth the Cost

Caution should not become paralysis. Here is where I draw the line.

The naive reading is fine when the gap between candidates is large relative to fold spread, when the decision is low-stakes, or when the model will be retrained on all available data anyway. In those cases, the extra computation buys you little.

It is not fine when the decision is high-stakes, when candidates are close, when the dataset is small, or when the result feeds a claim about generalization.

Two interactions matter:

With kk: Larger kk reduces the bias of the estimate but increases fold dependence. The naive standard error becomes more misleading, not less, as you increase kk. This is counterintuitive and worth internalizing.

With the learning algorithm: Highly stable learners produce more correlated fold scores, so the naive formula understates uncertainty more severely. Stability is usually a virtue; here it quietly widens the gap between your reported interval and reality.

When the stakes justify it, the neighboring tools are nested evaluation for selection and bootstrap resampling for uncertainty. Both are designed to account for dependence in ways that dividing a standard deviation by k\sqrt{k} does not.

What to Do With 0.842 and 0.847

Return to the opening comparison. The fold mean is a point estimate from one resampling procedure. The fold spread is a stability diagnostic. Neither is a confidence interval unless the dependence was handled explicitly.

So here is the durable rule: do not declare a winner on a gap smaller than the noise the split itself introduces.

Concretely: rerun the comparison with repeated k-fold, record the spread of the repeated means, and treat that spread as a diagnostic of split sensitivity — not as a confidence interval. Repeated k-fold scores are still dependent, because the same observations keep appearing across folds and repetitions. If the gap dissolves into the spread, you have learned something more valuable than a winner — you have learned that your data cannot distinguish these two models, and that is a finding, not a failure. If the gap survives, you still need a procedure justified for your target and assumptions before claiming a real difference.

Run the experiment. Inspect the spread. Then decide — and be honest about what the spread does and does not prove.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Two candidate models were evaluated on the same folds. Which comparison best respects the design described in the article?
Question 1 of 2Comparison Reasoning

Focus: Choose the appropriate comparison quantity when two models are evaluated on identical cross-validation splits.

A comparison reports mean scores of 0.842 and 0.847, and the gap dissolves into the spread of repeated k-fold means. What is the most defensible conclusion?
Question 2 of 2Scenario Interpretation

Focus: Interpret a small cross-validation score gap cautiously and identify what repeated-fold variability can and cannot establish.

References

  1. Cross-validation Confidence Intervals for Test Errorlucasjanson.fas.harvard.edu
  2. Cross-validation: what does it estimate and how well ... - PMCpmc.ncbi.nlm.nih.gov
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.