Skip to content
intermediate

How Stacking Uses Out-of-Fold Predictions to Train a Meta-Model

Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing…

Published 2026-10-02Updated 2026-10-049 min read
A breathtaking sunrise over a vast mountainous landscape with clear skies.
A breathtaking sunrise over a vast mountainous landscape with clear skies. Photo by Adriana FT on Pexels.

Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing crashed. No warning fired. The meta-model simply learned to trust a signal that will not exist at prediction time.

That signal is the in-sample prediction. And the fix is the out-of-fold matrix.

The Meta-Model Needs Honest Predictions

Stacking has a simple contract: several base models each produce a prediction, and a second-level meta-model learns how to combine those predictions into a final answer. If you have seen voting versus stacking before, you already know the difference — voting fixes the combination rule in advance, stacking learns it.

Here is the hidden requirement that trips people up. The meta-model is trained on meta-features: the base models' outputs. For the meta-model to learn a useful combination rule, those meta-features must behave the same way during training as they will during inference. At inference, every base model predicts a row it has never seen. So during training, the meta-features must also come from base models predicting rows they never saw.

That is the entire idea. Everything below is the mechanism that makes it true.

The failure we are going to explain: if you train the meta-model on predictions the base models made for their own training rows, the meta-model learns from base outputs that are systematically easier than the ones it will see in production.

Notation and the Two-Level Setup

Let the training set be (X,y)(X, y) with nn rows. We split it into KK folds. We have MM base estimators h1,…,hMh_1, \dots, h_M and one meta-model gg.

Define the fold assignment k(i)k(i) as the fold containing row ii. Each row belongs to exactly one held-out fold.

The object we are building is the out-of-fold prediction matrix ZZ, of shape (n,M)(n, M). Row ii holds each base model's prediction for row ii, produced by a copy of that model that never saw row ii:

Z[i,m]=hm(−k(i))(xi)Z[i, m] = h_m^{(-k(i))}(x_i)

Here hm(−k)h_m^{(-k)} is the copy of base model mm trained on every fold except fold kk. The superscript is doing real work: it says this prediction came from a model that did not train on this row.

One assumption carries the whole construction: rows are exchangeable and folds are representative. This is the same assumption ordinary k-fold cross-validation already relies on. When it fails — grouped data, temporal data, imbalanced classes — the naive fold assignment breaks, and the construction must inherit the correct splitter instead.

Building the Out-of-Fold Matrix, Fold by Fold

Three fold lanes show models trained on the other two folds predicting their held-out rows. Arrows place those predictions into the corresponding rows of Z, with separate columns for two base models.
Each row of Z receives predictions from models that excluded that row's fold.

The construction is a loop. For each fold kk:

  1. Fit a fresh copy of every base estimator on the complement of fold kk.
  2. Predict fold kk with each of those copies.
  3. Write those predictions into the rows of ZZ that belong to fold kk.

After KK iterations, every row of ZZ is filled by a model that did not train on it. The assignment rule from above is exactly this loop written in one line.

The cost is the part people underestimate. Each base model is refit KK times, so you pay K×MK \times M fits, not MM fits. With 5 folds and 4 base models, that is 20 training runs before the meta-model even starts.

A six-row worked example

Take n=6n = 6 rows, K=3K = 3 folds, M=2M = 2 base models. Fold 1 holds rows 1–2, fold 2 holds rows 3–4, fold 3 holds rows 5–6.

Fold held outRows used to trainRows predictedPredictions written to ZZ
13, 4, 5, 61, 2Z[1,:],Z[2,:]Z[1,:], Z[2,:]
21, 2, 5, 63, 4Z[3,:],Z[4,:]Z[3,:], Z[4,:]
31, 2, 3, 45, 6Z[5,:],Z[6,:]Z[5,:], Z[6,:]

Row 3's meta-features come from base models trained on rows 1, 2, 5, and 6 — never on row 3. That is the guarantee. ZZ has exactly one prediction per training row per base model, and its row order matches yy. That alignment is what makes ZZ a legitimate feature matrix: you fit gg on (Z,y)(Z, y) with its own validation discipline, exactly as you would fit any model on any feature matrix.

Knowledge check

Check your understanding

Answer this question before you continue.

In the six-row example, which rows train the base-model copies that produce row 3's entries in Z?
Output Prediction

Focus: Trace which training rows produce a particular row's out-of-fold meta-features.

Why In-Sample Predictions Are a Shortcut

Now the tempting shortcut. Fit each base model once on all of XX. Predict XX. Use those predictions as meta-features.

It looks fine. The meta-model still trains on real rows and real labels. The pipeline runs without error. And that is precisely why the bug survives so long — nothing in the code tells you it is wrong.

The core problem is a distribution mismatch. The meta-model is trained on base outputs generated under one condition — the base model saw the row during fitting — but it will be applied to base outputs generated under a different condition: the base model has never seen the row. The meta-model has no way to know which condition it is in. It learns a combination rule tuned to the wrong input distribution.

What does that mismatch look like in practice? It depends on the base models. A flexible model that memorizes its training rows will produce in-sample predictions that are optimistically accurate and often overconfident. The meta-model sees a base learner that appears near-perfect and may learn to over-weight it. On new data, that same model's errors are the largest, and the ensemble degrades. That is one common failure pattern — not the only one. Even a well-regularized base model produces in-sample predictions that are systematically different from its out-of-sample predictions, and the meta-model will still learn from the wrong distribution.

PropertyIn-sample meta-featuresOut-of-fold meta-features
What the prediction measuresHow the base model fits rows it trained onHow the base model behaves on unseen rows
Relationship to inference-time inputsMismatched distributionApproximate match
What gg learnsA combination rule tuned to easier base outputsA combination rule tuned to realistic base outputs
Behavior on new dataUnreliable — the mismatch may or may not surface as over-trustMore likely to hold

The general rule underneath: any feature derived from the target must be generated without the row's own label influencing it. This is the same principle behind target encoding and feature selection inside cross-validation. Stacking leakage is not a special case — it is that principle wearing a different hat.

Common mistake: Checking that the pipeline runs and the training score looks good. Both are true for the broken version. The only reliable check is asking, for each meta-feature, which rows trained the model that produced it?

Knowledge check

Check your understanding

Answer this question before you continue.

Why can training the meta-model on base models' predictions for the same rows used to fit those base models create a problem?
Misconception Check

Focus: Explain why in-sample base predictions provide unsuitable training features for the meta-model.

From Training to Inference: What Actually Gets Kept

Here is the loop most learners leave open. The fold models are scaffolding. They exist only to manufacture honest meta-features, and then they are discarded.

After ZZ is built, refit each base estimator once on the full training set. These full-data models are what serve new cases. The trace for a new row xx:

  1. Pass xx through each full-data base model.
  2. Assemble the MM predictions into a meta-feature vector.
  3. Feed that vector to gg.
  4. Return gg's output.

The shape contract matters: the meta-feature vector at inference must have the same column order and meaning as ZZ. If column 1 was a random forest's probability during training, column 1 must be that same probability at inference.

For classifiers, you also choose between predicted labels and predicted probabilities as meta-features. Probabilities usually carry more information, but the choice must be consistent between training and inference. If you pass raw features through to the meta-model, that changes the input contract and must be applied identically at both stages.

Knowledge check

Check your understanding

Answer this question before you continue.

After building Z and fitting the meta-model, how should a new case be processed in the described stack?
Scenario Interpretation

Focus: Describe which fitted models generate meta-features for new cases after out-of-fold training.

Where This Construction Breaks

The construction is sound, but it has boundaries worth knowing before you trust it.

Nested validation. ZZ is training data for gg. If you evaluate the whole stack on the same rows used to build ZZ, your reported score is optimistic. Evaluating the stack honestly requires an outer split.

Preprocessing leakage. Scaling, imputation, and target encoding must be fit inside each fold, not on the full training set before splitting. A scaler fit on all of XX has already seen the held-out fold.

Splitter mismatch. Grouped, temporal, or imbalanced data needs the matching splitter, and the meta-model must respect the same structure. Shuffled folds on time-series data leak the future into the past.

Small data. Each fold model trains on (K−1)/K(K-1)/K of the data, so it is weaker than the final full-data version. The meta-model partially absorbs this bias, but with little data the gap is real.

Diminishing returns. Stacking adds K×MK \times M fits and a second tuning problem. It earns its cost when base models are genuinely diverse and individually strong — not as a default upgrade.

Knowledge check

Check your understanding

Answer this question before you continue.

A team builds Z with out-of-fold predictions, then reports the stack's score on those same rows. What change does the article identify as necessary for an honest evaluation?
Comparison Reasoning

Focus: Distinguish constructing out-of-fold meta-features from honestly evaluating the complete stack.

A Decision Rule for Using Stacking

Use stacking when you have several diverse, reasonably strong base models and a validation setup you trust. Skip it when one model clearly dominates, when data is small, or when you cannot afford an outer validation loop.

The one-line test: if you cannot explain how each meta-feature was generated without the row's own label, the stack is not trustworthy yet.

My advice for the next hour: build the out-of-fold matrix by hand on a small dataset — six rows, three folds, two base models — using plain scikit-learn clone and KFold. Print ZZ. Then fit a StackingClassifier on the same data and compare its internal meta-features against yours. When the two matrices line up, the mechanism stops being a library detail and becomes something you own.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Which procedure constructs the out-of-fold matrix Z as defined in the article?
Question 1 of 2Single Choice

Focus: Summarize how each row receives one out-of-fold prediction from each base estimator.

A dataset contains time-ordered observations, so randomly shuffled folds could let later information influence predictions for earlier rows. What is the appropriate response when constructing stacking features?
Question 2 of 2Scenario Interpretation

Focus: Apply the article's splitter-matching principle when row exchangeability does not hold.

References

  1. Combine predictors using stacking — scikit-learn 1.9.1 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.