How Stacking Uses Out-of-Fold Predictions to Train a Meta-Model
Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing…

Key topics
Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing crashed. No warning fired. The meta-model simply learned to trust a signal that will not exist at prediction time.
That signal is the in-sample prediction. And the fix is the out-of-fold matrix.
The Meta-Model Needs Honest Predictions
Stacking has a simple contract: several base models each produce a prediction, and a second-level meta-model learns how to combine those predictions into a final answer. If you have seen voting versus stacking before, you already know the difference — voting fixes the combination rule in advance, stacking learns it.
Here is the hidden requirement that trips people up. The meta-model is trained on meta-features: the base models' outputs. For the meta-model to learn a useful combination rule, those meta-features must behave the same way during training as they will during inference. At inference, every base model predicts a row it has never seen. So during training, the meta-features must also come from base models predicting rows they never saw.
That is the entire idea. Everything below is the mechanism that makes it true.
The failure we are going to explain: if you train the meta-model on predictions the base models made for their own training rows, the meta-model learns from base outputs that are systematically easier than the ones it will see in production.
Notation and the Two-Level Setup
Let the training set be with rows. We split it into folds. We have base estimators and one meta-model .
Define the fold assignment as the fold containing row . Each row belongs to exactly one held-out fold.
The object we are building is the out-of-fold prediction matrix , of shape . Row holds each base model's prediction for row , produced by a copy of that model that never saw row :
Here is the copy of base model trained on every fold except fold . The superscript is doing real work: it says this prediction came from a model that did not train on this row.
One assumption carries the whole construction: rows are exchangeable and folds are representative. This is the same assumption ordinary k-fold cross-validation already relies on. When it fails — grouped data, temporal data, imbalanced classes — the naive fold assignment breaks, and the construction must inherit the correct splitter instead.
Building the Out-of-Fold Matrix, Fold by Fold
The construction is a loop. For each fold :
- Fit a fresh copy of every base estimator on the complement of fold .
- Predict fold with each of those copies.
- Write those predictions into the rows of that belong to fold .
After iterations, every row of is filled by a model that did not train on it. The assignment rule from above is exactly this loop written in one line.
The cost is the part people underestimate. Each base model is refit times, so you pay fits, not fits. With 5 folds and 4 base models, that is 20 training runs before the meta-model even starts.
A six-row worked example
Take rows, folds, base models. Fold 1 holds rows 1–2, fold 2 holds rows 3–4, fold 3 holds rows 5–6.
| Fold held out | Rows used to train | Rows predicted | Predictions written to |
|---|---|---|---|
| 1 | 3, 4, 5, 6 | 1, 2 | |
| 2 | 1, 2, 5, 6 | 3, 4 | |
| 3 | 1, 2, 3, 4 | 5, 6 |
Row 3's meta-features come from base models trained on rows 1, 2, 5, and 6 — never on row 3. That is the guarantee. has exactly one prediction per training row per base model, and its row order matches . That alignment is what makes a legitimate feature matrix: you fit on with its own validation discipline, exactly as you would fit any model on any feature matrix.
Knowledge check
Check your understanding
Answer this question before you continue.
Why In-Sample Predictions Are a Shortcut
Now the tempting shortcut. Fit each base model once on all of . Predict . Use those predictions as meta-features.
It looks fine. The meta-model still trains on real rows and real labels. The pipeline runs without error. And that is precisely why the bug survives so long — nothing in the code tells you it is wrong.
The core problem is a distribution mismatch. The meta-model is trained on base outputs generated under one condition — the base model saw the row during fitting — but it will be applied to base outputs generated under a different condition: the base model has never seen the row. The meta-model has no way to know which condition it is in. It learns a combination rule tuned to the wrong input distribution.
What does that mismatch look like in practice? It depends on the base models. A flexible model that memorizes its training rows will produce in-sample predictions that are optimistically accurate and often overconfident. The meta-model sees a base learner that appears near-perfect and may learn to over-weight it. On new data, that same model's errors are the largest, and the ensemble degrades. That is one common failure pattern — not the only one. Even a well-regularized base model produces in-sample predictions that are systematically different from its out-of-sample predictions, and the meta-model will still learn from the wrong distribution.
| Property | In-sample meta-features | Out-of-fold meta-features |
|---|---|---|
| What the prediction measures | How the base model fits rows it trained on | How the base model behaves on unseen rows |
| Relationship to inference-time inputs | Mismatched distribution | Approximate match |
| What learns | A combination rule tuned to easier base outputs | A combination rule tuned to realistic base outputs |
| Behavior on new data | Unreliable — the mismatch may or may not surface as over-trust | More likely to hold |
The general rule underneath: any feature derived from the target must be generated without the row's own label influencing it. This is the same principle behind target encoding and feature selection inside cross-validation. Stacking leakage is not a special case — it is that principle wearing a different hat.
Common mistake: Checking that the pipeline runs and the training score looks good. Both are true for the broken version. The only reliable check is asking, for each meta-feature, which rows trained the model that produced it?
Knowledge check
Check your understanding
Answer this question before you continue.
From Training to Inference: What Actually Gets Kept
Here is the loop most learners leave open. The fold models are scaffolding. They exist only to manufacture honest meta-features, and then they are discarded.
After is built, refit each base estimator once on the full training set. These full-data models are what serve new cases. The trace for a new row :
- Pass through each full-data base model.
- Assemble the predictions into a meta-feature vector.
- Feed that vector to .
- Return 's output.
The shape contract matters: the meta-feature vector at inference must have the same column order and meaning as . If column 1 was a random forest's probability during training, column 1 must be that same probability at inference.
For classifiers, you also choose between predicted labels and predicted probabilities as meta-features. Probabilities usually carry more information, but the choice must be consistent between training and inference. If you pass raw features through to the meta-model, that changes the input contract and must be applied identically at both stages.
Knowledge check
Check your understanding
Answer this question before you continue.
Where This Construction Breaks
The construction is sound, but it has boundaries worth knowing before you trust it.
Nested validation. is training data for . If you evaluate the whole stack on the same rows used to build , your reported score is optimistic. Evaluating the stack honestly requires an outer split.
Preprocessing leakage. Scaling, imputation, and target encoding must be fit inside each fold, not on the full training set before splitting. A scaler fit on all of has already seen the held-out fold.
Splitter mismatch. Grouped, temporal, or imbalanced data needs the matching splitter, and the meta-model must respect the same structure. Shuffled folds on time-series data leak the future into the past.
Small data. Each fold model trains on of the data, so it is weaker than the final full-data version. The meta-model partially absorbs this bias, but with little data the gap is real.
Diminishing returns. Stacking adds fits and a second tuning problem. It earns its cost when base models are genuinely diverse and individually strong — not as a default upgrade.
Knowledge check
Check your understanding
Answer this question before you continue.
A Decision Rule for Using Stacking
Use stacking when you have several diverse, reasonably strong base models and a validation setup you trust. Skip it when one model clearly dominates, when data is small, or when you cannot afford an outer validation loop.
The one-line test: if you cannot explain how each meta-feature was generated without the row's own label, the stack is not trustworthy yet.
My advice for the next hour: build the out-of-fold matrix by hand on a small dataset — six rows, three folds, two base models — using plain scikit-learn clone and KFold. Print . Then fit a StackingClassifier on the same data and compare its internal meta-features against yours. When the two matrices line up, the mechanism stops being a library detail and becomes something you own.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


