Validate Scikit-Learn Inference Inputs Before Prediction
A missing column throws. A reordered DataFrame does not. That asymmetry is where confident wrong numbers come from.

Key topics
A missing column throws. A reordered DataFrame does not. That asymmetry is where confident wrong numbers come from.
You already have a fitted pipeline. You can call predict. What you do not have is a guard between raw incoming data and the model — and the model will not build one for you. It will accept whatever you hand it, run it through the same arithmetic it learned during training, and return a number with no opinion about whether that number means anything.
So the job is yours: read the contract off the fitted artifact, enforce it in code, and make the enforcement the only door into predict.
What the Model Actually Remembers About Its Inputs
Before writing a single check, get the ground truth. A fitted pipeline is not a black box about its inputs — it stores what it learned, and you can read it back.
If you fitted on a DataFrame with named columns, scikit-learn records them on the estimator as feature_names_in_. If your pipeline starts with a ColumnTransformer, the fitted transformers know which columns they were assigned and what they produced. get_feature_names_out() gives you the transformed feature names in order.
import joblib
pipeline = joblib.load("model.joblib")
# Names the estimator saw at fit time, in order
print(pipeline.feature_names_in_)
# What each preprocessing branch consumed and emitted
pre = pipeline.named_steps["preprocess"]
for name, transformer, columns in pre.transformers_:
print(name, columns)
Note:
feature_names_in_only exists when you fit on a DataFrame with string column names. Fit on a NumPy array and the model has no memory of names at all — which is exactly why fitting on named frames is worth the small overhead.
Here is the trap. The moment you retype that column list into a validation script by hand, you have created a second source of truth. Two lists, one model. They agree today. They will disagree after the next refit, and the disagreement will be silent because nothing compares them.
Derive the expected columns from the artifact. Then the contract cannot drift away from the model, because it is the model's contract.
There is a second distinction that matters more than the first. Hard failures — a missing column, a numeric column arriving as strings — announce themselves. Silent semantic changes — the same column name, the same dtype, a different meaning or unit — do not. A amount column that switches from dollars to cents passes every structural check you can write and still poisons every prediction downstream. Structural checks catch the loud failures. Semantic checks catch the expensive ones.
Knowledge check
Check your understanding
Answer this question before you continue.
Capture the Contract at Fit Time
The validator needs two things: the expected column order and the expected dtype for each column. Both must come from training, not from the incoming frame. If you infer expected dtypes from the data you are about to validate, you have validated the data against itself.
So capture the contract once, when you fit, and save it next to the artifact.
import json
import pandas as pd
def capture_contract(X_train):
contract = {
"columns": list(X_train.columns),
"dtypes": {c: str(X_train[c].dtype) for c in X_train.columns},
}
with open("input_contract.json", "w") as f:
json.dump(contract, f, indent=2)
return contract
contract = capture_contract(X_train)
The contract is plain JSON: a column list and a dtype map. It travels with the model, it diffs cleanly in version control, and it does not require unpickling anything to inspect. When the upstream schema changes, you update the contract and redeploy the check — no refit required.
Knowledge check
Check your understanding
Answer this question before you continue.
A Minimal Validator You Can Run in Ten Lines
Start with the smallest thing that works. This function checks presence, then order, then dtype, then missing values — and each check fails with a message specific enough to act on.
import pandas as pd
def validate_inputs(df, contract):
expected_columns = contract["columns"]
expected_dtypes = contract["dtypes"]
missing = [c for c in expected_columns if c not in df.columns]
if missing:
raise ValueError(f"Missing columns: {missing}")
extra = [c for c in df.columns if c not in expected_columns]
if extra:
raise ValueError(f"Unexpected columns: {extra}")
if list(df.columns) != expected_columns:
raise ValueError(
f"Column order changed.\n"
f" expected: {expected_columns}\n"
f" received: {list(df.columns)}"
)
for col, dtype in expected_dtypes.items():
if str(df[col].dtype) != dtype:
raise TypeError(
f"Column '{col}' has dtype {df[col].dtype}, expected {dtype}"
)
nulls = df.isna().sum()
if nulls.any():
raise ValueError(f"Null values present:\n{nulls[nulls > 0]}")
return df.copy()
Three design choices are doing real work here.
First, the function rejects reordered input rather than silently reordering it. That is a deliberate policy: if the incoming column order differs from training, something upstream changed, and you want to know about it before the model does. Silently reordering would hide the change and let a broken pipeline keep producing numbers.
Second, the function returns a copy instead of mutating the caller's DataFrame. Silent mutation of caller state is its own bug class — you fix one path and break another that was relying on the original order.
Third, the checks run in a deliberate sequence. Presence first, because a missing column makes every later check meaningless. Then order, then dtype, then nulls. Each failure produces a different exception with a different fix.
On a good frame, the call returns quietly:
clean = validate_inputs(raw, contract)
preds = pipeline.predict(clean)
On a bad one, you get text you can grep for:
TypeError: Column 'amount' has dtype object, expected float64
That message is the success criterion. If your validator cannot tell you which column and what it expected, it is not finished.
Warning: dtype inference and feature-name behavior have shifted across pandas and scikit-learn releases. Pin your versions, and treat a dtype mismatch as information about the environment as well as the data.
Knowledge check
Check your understanding
Answer this question before you continue.
Wiring the Check Into the Prediction Path
A standalone function is a suggestion. A single entry point is a rule.
Wrap load-and-predict so there is exactly one door into the model:
def predict_checked(df):
pipeline = joblib.load("model.joblib")
with open("input_contract.json") as f:
contract = json.load(f)
clean = validate_inputs(df, contract)
return pipeline.predict(clean)
Two rules make this hold up.
Validate before the pipeline sees the data. Do not bury the check inside a custom transformer. A transformer that raises mid-fit_transform produces confusing stack traces, and worse, it can be bypassed by anyone who calls the underlying steps directly.
Keep the validator separate from the pipeline artifact. The contract is metadata about the model, not part of it. That separation is what lets you inspect the contract without loading the model, and update it without retraining.
The call site is then one line you can drop into a script, a notebook cell, or a batch job:
preds = predict_checked(incoming_df)
Catching the Silent Semantic Change
Structural checks are necessary and insufficient. Here is the failure they cannot see: amount stays named amount, stays float64, and switches from dollars to cents. Every check passes. Every prediction is wrong by a factor of one hundred.
The only defense is comparison against a baseline recorded at training time. Capture the min, max, mean, and null rate for each numeric column when you fit, store them next to the artifact, and check incoming data against a tolerance band.
def check_ranges(df, training_stats, tolerance=0.5):
warnings = []
for col, stats in training_stats.items():
lo, hi = df[col].min(), df[col].max()
if lo < stats["min"] * (1 - tolerance) or hi > stats["max"] * (1 + tolerance):
warnings.append(f"{col}: range [{lo:.2f}, {hi:.2f}] "
f"vs training [{stats['min']:.2f}, {stats['max']:.2f}]")
return warnings
For categorical columns, check the levels themselves. An unseen category, a renamed level, or a case change that maps to a different code all pass structural validation and all change the prediction.
Common mistake: treating these as hard failures. A range warning is evidence, not a verdict. You decide whether to reject the row, quarantine it for review, or predict with a logged flag. The check makes the decision visible; it does not make the decision for you.
This is not ceremony. Errors in the data fed to a model can nullify whatever accuracy you gained in training — the input is a production asset, and it deserves the same scrutiny as the weights.
Knowledge check
Check your understanding
Answer this question before you continue.
Failure Modes and Debugging Signals
Keep this table close. When something breaks at prediction time, the error usually tells you which layer failed.
| Symptom | Likely cause | What it usually means |
|---|---|---|
KeyError or feature-name mismatch | Missing column | Upstream schema change or renamed field |
ValueError: Column order changed | Reordered columns | Upstream reordered the frame; the validator caught it |
object dtype reaching a numeric transformer | String-typed numerics | Unparsed input or a silent coercion to NaN |
| Extra columns ignored or fatal | Schema drift | Depends on the pipeline — decide deliberately |
| Prediction from all-null input | Imputation defaults | The model is predicting about nothing |
The reordered-columns row is the dangerous one. When a pipeline selects by name, order does not matter and nothing raises. When it selects by position, order matters and nothing raises either — you just get a different number. That is the case your validator exists to catch, and it is the reason the order check is not optional.
One Experiment: Break the Contract on Purpose
Reading about checks is not the same as watching them fire. Take one known-good row, record its prediction, then mutate exactly one thing at a time.
baseline = predict_checked(good_row)[0]
# Drop a column
print(predict_checked(good_row.drop(columns=["amount"])))
# Swap two columns
swapped = good_row[["b", "a", "c"]]
print(predict_checked(swapped))
# Cast a numeric to string
stringy = good_row.assign(amount=good_row["amount"].astype(str))
print(predict_checked(stringy))
# Shift a value by 100x
shifted = good_row.assign(amount=good_row["amount"] * 100)
print(predict_checked(shifted)[0], "vs baseline", baseline)
Run it and build your own table: which mutations raise, which pass silently, and which change the prediction. The first three should raise — missing column, reordered columns, wrong dtype. The last one should pass silently and change the prediction. That contrast is the lesson: the validator catches structural breakage, and the range check is what catches the semantic shift.
Then decide, from what you observed, which checks belong in hard-fail mode and which belong in warn-and-log mode. That decision is the actual deliverable — the code is just how you make it real.
The Decision Rule
Structural checks fail hard. Semantic checks warn and log. Every prediction path goes through one validated entry point.
That is the whole discipline, and it fits in three sentences because the enforcement is what matters, not the philosophy.
The next practical move is to record training-time input statistics — per-column min, max, mean, null rate, and category levels — and save them alongside the model artifact. Without that baseline, your semantic checks have nothing to compare against and degrade into guesses. With it, you have the boundary where validation stops and monitoring begins: validation asks whether this input is allowed, monitoring asks whether the stream of allowed inputs is drifting away from what the model learned.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


