Skip to content
intermediate

Inspect Regression Residuals With a Scikit-Learn Diagnostic Experiment

A respectable R² can sit on top of a model that is wrong in a structured way. The score summarizes average fit; it says nothing about where the fit fails.…

Published 2026-10-02Updated 2026-10-047 min read
Networking equipment with connected cables, showcasing modern technology infrastructure.
Networking equipment with connected cables, showcasing modern technology infrastructure. Photo by Vladimir Srajber on Pexels.

A respectable R² can sit on top of a model that is wrong in a structured way. The score summarizes average fit; it says nothing about where the fit fails. A residual plot is how you ask that question.

This is a short, runnable diagnostic experiment. You will fit a linear model, plot residuals against fitted values, add one second view, name a visible pattern, and choose exactly one follow-up check. The goal is not a prettier plot. The goal is a better next question.

What This Experiment Will and Will Not Prove

A residual plot is evidence about structure, not a verdict on your model. It tells you where to look next. It does not tell you the model is broken, and a clean plot does not tell you the model is correct.

Here is what "done" looks like for this experiment:

  • A fitted regression model.
  • A residual-versus-fitted scatter plot.
  • One additional diagnostic view.
  • A named pattern, written in one sentence.

Three assumptions hold the experiment together. First, the model is linear in its parameters. Second, residuals are computed on data the model did not memorize — in-sample residuals can look better than they are, so we will hold some data back. Third, the feature set stays fixed for now; we are diagnosing the fit, not redesigning the pipeline.

If you have already read about residuals as the leftover signal after the model takes its share, this is where that idea becomes a repeatable check.

Set Up a Small, Honest Regression Example

We need a dataset where structure is visible rather than buried in noise. A small synthetic set with one feature and a deliberately curved relationship works well, because you already know the truth and can verify the plot is telling you something real.

import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)
X = np.linspace(0, 10, 200).reshape(-1, 1)
y = 2.0 * X.ravel() + 0.4 * X.ravel()**2 + rng.normal(0, 2.0, size=200)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=0
)

model = LinearRegression().fit(X_train, y_train)
y_hat = model.predict(X_test)
residuals = y_test - y_hat

print("coef:", model.coef_, "intercept:", model.intercept_)
print("test R^2:", model.score(X_test, y_test))

Expected output will look roughly like a coefficient near 2 plus a positive intercept, and a test R² that is high but not perfect — the squared term is real, and a straight line cannot capture it. That gap between "high R²" and "wrong shape" is the whole point.

Note: Computing residuals = y_test - y_hat explicitly matters. When you hide the subtraction behind a library call, you stop owning the quantity you are plotting, and you stop noticing when the sign convention flips.

Knowledge check

Check your understanding

Answer this question before you continue.

For a held-out observation with `y_test = 12` and prediction `y_hat = 15`, what residual does this experiment compute?
Output Prediction

Focus: Compute and interpret residuals using the article's stated sign convention.

Plot Residuals Against Fitted Values

This is the primary view. Residuals go on the vertical axis; fitted values go on the horizontal.

plt.figure(figsize=(7, 5))
plt.scatter(y_hat, residuals, s=30, alpha=0.7)
plt.axhline(0, color="red", linestyle="--", linewidth=1.2)
plt.xlabel("Fitted values")
plt.ylabel("Residuals")
plt.title("Residuals vs. fitted")
plt.show()

Fitted values are the right x-axis because they expose structure the model failed to capture across its own prediction range. Plotting residuals against the raw target or a single feature can hide the same structure behind the scale of the input.

A healthy residual-versus-fitted plot looks like a shapeless band centered near zero with roughly constant vertical spread. No curve, no funnel, no cluster. The red reference line is the visual anchor: it turns a cloud of points into a readable signal by giving your eye a baseline to compare against.

On this dataset you should see a clear U-shape or arch. The model under-predicts at the low and high ends and over-predicts in the middle. That is the squared term announcing itself.

Common mistake: Plotting residuals against the raw target and concluding "the residuals look fine." The raw target can absorb the curvature and mask it.

Knowledge check

Check your understanding

Answer this question before you continue.

To make the primary residual diagnostic described in the article, which quantities belong on the horizontal and vertical axes, respectively?
Single Choice

Focus: Select the axes that reveal residual structure across the model's prediction range.

Add a Second View: Spread or Order

One plot shows shape. A second plot makes spread or dependence visible. Pick one and justify it.

If your concern is changing variance, use a scale-location style view: the square root of the absolute residuals against fitted values. The square root compresses large values so a widening funnel becomes easier to see.

plt.figure(figsize=(7, 5))
plt.scatter(y_hat, np.sqrt(np.abs(residuals)), s=30, alpha=0.7)
plt.xlabel("Fitted values")
plt.ylabel("sqrt(|residuals|)")
plt.title("Scale-location")
plt.show()

A flat, even band suggests roughly constant spread. A funnel — narrow on one side, wide on the other — suggests the error spread changes with the prediction level, which affects how much you should trust errors at the extremes.

If your rows carry meaning, such as time order, plot residuals against observation order instead. That view catches dependence the fitted-value plot cannot see. Either way, the second view narrows the hypothesis; it does not confirm it.

Knowledge check

Check your understanding

Answer this question before you continue.

You suspect residual spread grows as predictions increase, and the rows have no meaningful order. Which second view best matches the article's recommendation?
Scenario Interpretation

Focus: Choose an additional diagnostic view suited to investigating changing residual spread.

Read the Pattern Without Overclaiming

Turn the plots into a small vocabulary, then apply discipline.

PatternWhat it suggestsNext check
Curvature (U, arch, wave)Missing nonlinear term or interactionAdd a squared or interaction feature, re-plot
Funnel (changing spread)Error variance depends on prediction levelTry a transformation of the target, re-plot
A few far-out pointsOutliers or high-leverage observationsInspect those rows before changing the model
Flat bandNo visible structure in this viewDo not conclude normality — that needs a Q-Q plot

The discipline is simple: a pattern is a hypothesis about the data or the model. The next step is a test, not a rewrite. A flat residual plot does not confirm normally distributed errors; that is a different question with a different view.

Choose the Next Diagnostic, Not the Next Rewrite

A flowchart shows fitting on training data, inspecting validation residuals, making one targeted change and checking again in a loop, then using the untouched test set for a final check.
Use validation residuals to guide one revision; keep the test set untouched until the final check.

Map each pattern to one follow-up check.

  • Curvature: add a squared term and re-plot.
  • Funnel: try a log or square-root transform of the target and re-plot.
  • Far-out points: inspect those observations directly before touching the model.

Change one thing at a time. If you add a squared term and a transform simultaneously, you cannot attribute the result to either.

Here is the boundary that keeps this honest. The test split you used to inspect the original model is now spent for diagnostic purposes. If you compare a revised model against those same residuals, you are tuning to the diagnostic, and any apparent improvement may not generalize. So the comparison below uses a validation split carved from the training data, and the test set stays untouched until the end.

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline

X_fit, X_val, y_fit, y_val = train_test_split(
    X_train, y_train, test_size=0.3, random_state=1
)

poly_model = make_pipeline(
    PolynomialFeatures(degree=2, include_bias=False),
    LinearRegression(),
).fit(X_fit, y_fit)

val_residuals = y_val - poly_model.predict(X_val)

Re-plot val_residuals against the new fitted values. If the curvature disappears on validation data, you learned something real about the data. If it persists, the problem is likely elsewhere — a different feature, a different functional form, or noise.

Only after you have settled on a revision do you touch the test set once, as a final check:

print("final test R^2:", poly_model.score(X_test, y_test))

Warning: Do not tune the model until the residuals look pretty. A model adjusted to satisfy the diagnostic is overfit to the diagnostic itself, and you will have learned nothing that generalizes.

Knowledge check

Check your understanding

Answer this question before you continue.

A residual plot shows curvature, and you want to evaluate adding a squared feature. Which next step follows the article's procedure?
Debugging

Focus: Choose a targeted follow-up for curvature without reusing the diagnostic test split for tuning.

Where This Fits in Your Evaluation Workflow

Residual inspection answers questions about structure and spread. It does not replace held-out evaluation, learning curves, or metric choice. Those are separate diagnostics with their own evidence, and they answer different questions.

Skip this experiment when the target is categorical, when the model is not a regression model, or when the real problem is upstream in the data pipeline. A residual plot cannot diagnose a broken join.

The loop is short: fit, plot, name the pattern, run one targeted check — on validation data, not the test set. Run it on your own dataset this week. Write the pattern you see in one sentence before you change anything. Then pick exactly one follow-up check and let the output tell you whether your hypothesis was right.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A residual-versus-fitted plot looks like a flat band. What conclusion is warranted by the article?
Question 1 of 2Misconception Check

Focus: Distinguish absence of visible residual structure from evidence of normally distributed errors.

Which statement best matches how the article places residual inspection in an evaluation workflow?
Question 2 of 2Comparison Reasoning

Focus: Explain how residual inspection relates to other evaluation tools and the final test-set check.

References

  1. 19 Residual-diagnostics Plots | Explanatory Model Analysisema.drwhy.ai
  2. Residual Plots Explained: Check Regression Assumptions | DataCampwww.datacamp.com
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.