Skip to content
beginner

Plot and Interpret Learning Curves With Scikit-Learn

You have seen the classic two-line plot. One line starts high and drifts down. The other starts low and climbs. Somewhere on the right, they meet, and the…

Published 2026-10-02Updated 2026-10-048 min read
Breathtaking sunrise over rolling hills with fields and rural landscape, featuring mist and vibrant colors.
Breathtaking sunrise over rolling hills with fields and rural landscape, featuring mist and vibrant colors. Photo by S. von Hoerst on Pexels.

You have seen the classic two-line plot. One line starts high and drifts down. The other starts low and climbs. Somewhere on the right, they meet, and the article you were reading says "now you know whether you need more data or a better model."

Then you call learning_curve yourself, and the questions arrive all at once. Which scoring metric? Which cv? Why are the shaded bands so wide? And what, exactly, are you supposed to change based on this picture?

This tutorial answers those questions in order. We will generate the curves, read them honestly, and pick one experiment to run next. I assume you already know what training and validation scores are and have seen the high-bias and high-variance shapes. If those ideas are still fuzzy, read the intuition piece on learning curves first, then come back. Here we focus on producing the plot and turning it into a decision.

What the learning_curve Function Actually Returns

Before you run anything, build a precise mental model of the function. It is simpler than it looks, and most confusion comes from expecting it to behave like fit.

learning_curve takes an estimator, the full training set (X, y), and a few configuration arguments. It returns three arrays:

  • train_sizes: the training-set sizes it actually used, in the same units you passed in.
  • train_scores: shape (n_ticks, n_cv_folds).
  • test_scores: shape (n_ticks, n_cv_folds) — the validation scores.

That shape is the key detail. For every training size, the function runs cross-validation and records one score per fold. You get a grid of scores, not a single line. Averaging across folds is your job.

A few arguments deserve attention:

train_sizes accepts fractions or absolute counts. A float in (0, 1] is treated as a fraction of the maximum training set size. The default is five evenly spaced fractions from 0.1 to 1.0. Absolute integers are interpreted as sample counts.

cv controls the splitting strategy and defaults to 3-fold. For classifiers, scikit-learn uses stratified folds by default. If your data is grouped or time-ordered, you must pass an explicit splitter — a random split will quietly lie to you.

scoring defaults to the estimator's own score method. Set it explicitly to the metric that matches your problem. The default is convenient, not correct.

One behavior surprises almost everyone: the function does not save or return a fitted model. It trains many temporary models, scores them, and discards them. If you expected to reuse the estimator afterward, you cannot. That is by design — the curve is a diagnostic, not a training run.

If you also pass return_times=True, you get fit_times and score_times. That is useful when the curve is slow rather than wrong.

Knowledge check

Check your understanding

Answer this question before you continue.

What does each row of `train_scores` or `test_scores` represent?
Single Choice

Focus: Interpret the dimensions and meaning of the arrays returned by `learning_curve`.

A Minimal Runnable Example

Let us get to a working plot fast. The example uses only scikit-learn, NumPy, and Matplotlib, plus a small built-in dataset, so it is deterministic and needs no downloads.

import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import learning_curve
from sklearn.linear_model import LogisticRegression

X, y = load_breast_cancer(return_X_y=True)

estimator = LogisticRegression(max_iter=5000)

train_sizes, train_scores, val_scores = learning_curve(
    estimator,
    X,
    y,
    train_sizes=np.linspace(0.1, 1.0, 8),
    cv=5,
    scoring="accuracy",
    n_jobs=-1,
)

train_mean = train_scores.mean(axis=1)
train_std = train_scores.std(axis=1)
val_mean = val_scores.mean(axis=1)
val_std = val_scores.std(axis=1)

plt.figure(figsize=(8, 5))
plt.plot(train_sizes, train_mean, "o-", label="Training score")
plt.fill_between(train_sizes, train_mean - train_std, train_mean + train_std, alpha=0.2)

plt.plot(train_sizes, val_mean, "s-", label="Validation score")
plt.fill_between(train_sizes, val_mean - val_std, val_mean + val_std, alpha=0.2)

plt.xlabel("Training examples")
plt.ylabel("Accuracy")
plt.title("Learning curve")
plt.legend(loc="best")
plt.grid(True)
plt.show()

Expected output: a two-line plot. The training curve starts high and drifts down as more data is added. The validation curve starts low and rises. The shape is the success criterion here, not any specific number.

Why average across folds? Because a single fold's score is noisy. The mean is your best estimate; the standard deviation is the honest picture of how much that estimate wobbles. Plotting the band as mean ± std shows both.

One detail to internalize: the training score is measured on the same data the model just fit. It is optimistic by construction. A high training score is not evidence of a good model — it is evidence that the model can memorize.

Knowledge check

Check your understanding

Answer this question before you continue.

In the example, what does `train_scores.mean(axis=1)` produce?
Output Prediction

Focus: Explain what averaging cross-validation scores along the fold axis produces for plotting.

Reading the Shape Without Overclaiming

Two side-by-side learning-curve plots show training and validation scores against training examples. The high-bias curves meet at a mediocre plateau; the high-variance curves remain separated while validation rises. Suggested first experiments are more model capacity or better features for high bias, and more data for high variance.
Compare curve shape and choose one first experiment; treat each pattern as a hypothesis, not a diagnosis.

You will see two patterns most of the time.

High bias. Both curves converge to a similar, mediocre score. The validation curve has flattened before the data runs out. Within the range you measured, adding more rows did not improve validation performance.

High variance. A wide, persistent gap between a high training score and a lower validation score. The validation curve is still climbing at the largest training size. More data is likely to help.

Here is the honest reading: these shapes are consistent with a cause, not a diagnosis. A wide gap can come from too much model flexibility, too little data, noisy labels, or a leaky evaluation setup. The curve narrows the hypothesis space; it does not name the culprit.

The band width matters as much as the line. Wide bands mean the estimate itself is unstable. Any conclusion you draw from a small difference between two curves is fragile. Before you interpret a shape, confirm the split is valid for your data — grouped, temporal, or stratified as needed. A broken split produces a beautiful, meaningless curve.

Knowledge check

Check your understanding

Answer this question before you continue.

A curve shows a high training score, a lower validation score, and a validation line that is still climbing at the largest measured size. Which interpretation best fits the article?
Scenario Interpretation

Focus: Recognize a curve pattern consistent with high variance without treating it as a definitive diagnosis.

Three Ways to Get a Misleading Curve

Most wrong conclusions come from one of three setup mistakes. Each has a debugging signal.

Leakage through preprocessing. If you scale, impute, or encode on the full dataset before splitting, information from the validation folds bleeds into training. Validation scores inflate and the curve flattens. The fix is a Pipeline so every transform is fit inside each fold. Signal: validation score suspiciously close to the training score.

Wrong split for the data. Random folds on grouped or time-ordered data let near-duplicate or future information into training. The curve looks healthy; the deployed model is not. Signal: validation score drops when you fix the split.

Metric mismatch. Accuracy on an imbalanced target makes both curves look flat and high while the minority class is ignored. The curve is not wrong — it is answering a different question than the one that matters. Signal: a curve that barely moves across training sizes.

A fourth, quieter problem: too few folds or too few points on the x-axis. Three folds and five sizes produce a jagged line that invites over-reading. More folds and more sizes cost time but buy a stable picture.

Knowledge check

Check your understanding

Answer this question before you continue.

A learning curve has suspiciously similar, high training and validation scores. You discover that scaling was fit on the full dataset before cross-validation. What is the best fix?
Debugging

Focus: Choose a fold-safe preprocessing strategy when preprocessing can leak validation information.

From Curve to One Justified Experiment

Now the payoff. Convert the shape into a single, falsifiable next step.

For the high-bias shape: test added model capacity or a feature change first. A plateau over the sizes you measured is a hypothesis, not a verdict — it says the model stopped improving across the range you plotted, not that more data is useless forever. If both curves sit at a mediocre score and the validation line has gone flat, a richer model or a better feature is the cheapest experiment to run before you collect more rows.

For the high-variance shape: more data, stronger regularization, or a simpler model. If the validation curve is still climbing at the largest size, more data is the cheapest first bet.

Pick one change. Rerun the same curve with the same cv, scoring, and train_sizes. Compare shapes, not single numbers.

A successful experiment narrows the gap or pushes the validation curve past its old plateau. A failed experiment leaves the curve unchanged — and that is itself evidence that your guessed cause was wrong. Keep a short log: curve shape, change made, new shape. That log is the reusable asset that turns curve reading into a habit.

Where Learning Curves Stop Being Useful

Curves are a data-and-complexity diagnostic, not a hyperparameter search. To choose a specific value for a parameter, use a validation curve over that parameter instead.

Curves say nothing about calibration, threshold choice, or the cost of specific error types. They cannot separate label noise from model flexibility — both produce a persistent gap, and only inspecting the data or the residuals can tell them apart. And they assume your evaluation split reflects deployment. If the world shifts after training, a healthy curve is a historical artifact, not a forecast.

Skip the curve entirely when your dataset is so small that each fold is meaningless, or when a single held-out set already answers your question.

What to Do Next

Generate the curve with an explicit cv and scoring. Read the shape and the band width together. Name the most plausible cause. Then change exactly one thing and rerun.

The curve is a hypothesis generator, not a verdict. When it is healthy but your model still fails, move to the neighboring diagnostics: a validation curve to tune a specific parameter, or residual and error analysis to find structure the score is hiding.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Both curves level off at a mediocre score, and the validation curve has flattened across the measured training sizes. Which first experiment best follows the article's advice?
Question 1 of 2Scenario Interpretation

Focus: Select one justified follow-up experiment for a high-bias-shaped curve while preserving a controlled comparison.

A model has a persistent gap between high training scores and lower validation scores. What conclusion is warranted from the curve alone?
Question 2 of 2Misconception Check

Focus: Recognize that a learning-curve pattern narrows hypotheses but does not uniquely establish a cause.

References

  1. learning_curve — scikit-learn 1.9.0 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.