Skip to content
beginner

A Decision Tree Experiment: Inspect Splits and Observe Overfitting

A decision tree can score 100% on the data it was trained on and still be wrong about almost everything else. That gap is not a bug. It is the lesson.

Published 2026-10-02Updated 2026-10-0410 min read
Close-up of a person analyzing a colorful graph chart with a pen in a modern office setting.
Close-up of a person analyzing a colorful graph chart with a pen in a modern office setting. Photo by https://kaboompics.com/ on Pexels.

A decision tree can score 100% on the data it was trained on and still be wrong about almost everything else. That gap is not a bug. It is the lesson.

This is a short, runnable experiment. The whole point is to make that gap appear on purpose, then shrink it by turning one dial: how deep the tree is allowed to grow. By the end, you will have a small script that fits a DecisionTreeClassifier at several depths and prints training accuracy next to validation accuracy, so you can watch the two numbers pull apart.

I like this experiment because it is the cheapest way I know to feel overfitting in your hands instead of reading about it. You fit a model, look inside it, change one number, and the output tells you whether you understood.

What This Experiment Will Show You

Before any code, here is what "done" looks like:

  • A script that fits a decision tree at several max_depth values.
  • A table with two columns: training accuracy and validation accuracy.
  • A visible pattern: training accuracy climbs toward 1.0, while validation accuracy rises, peaks, and then falls.

That falling tail is the thing you are hunting. When you see it, you have observed overfitting directly rather than taking someone's word for it.

You will need Python with scikit-learn, NumPy, and pandas installed. matplotlib is optional, and only if you want to look at a picture of the tree.

Two assumptions carry over from earlier reading. First, you already know that a tree splits data into regions and that depth controls how many splits stack up. Second, you have seen the idea of a train/validation boundary. We are not re-deriving impurity here, and we are not re-teaching why validation exists. We are going to use both.

Why max_depth? Because it is the single most direct complexity control on a tree. It is one integer, it is easy to reason about, and its effect on the model is immediate. If you want one dial to turn while you watch generalization move, this is the dial.

Set Up the Data and the Split

We will use a small, familiar dataset so you can reason about the features by name instead of staring at anonymous columns. The Iris dataset ships with scikit-learn and is small enough to fit in your head.

import numpy as np
import pandas as pd
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

iris = load_iris()
X = pd.DataFrame(iris.data, columns=iris.feature_names)
y = pd.Series(iris.target, name="species")

X_train, X_val, y_train, y_val = train_test_split(
    X, y, test_size=0.3, random_state=42, stratify=y
)

print("train shape:", X_train.shape, "validation shape:", X_val.shape)
print("class balance (train):")
print(y_train.value_counts(normalize=True).round(3))

The random_state=42 matters more than it looks. Without a fixed seed, scikit-learn can break ties between equally good splits differently on each run, and your numbers will drift between executions. Fix the seed once, and every later comparison is fair.

Each split has one job. The training data fits the tree. The validation data judges it. That is the whole contract, and it is worth saying plainly because it is easy to violate by accident.

Warning: Do not use the validation set to pick your depth and then report that same validation score as your final result. The moment validation influences a decision, it stops being a clean judge. If you want an honest final number, hold back a separate test set, or use cross-validation on the training data to choose the depth.

Printing the shapes and class balance first is not busywork. If one class dominates, plain accuracy will flatter your model later, and you want to know that before you start trusting scores.

Knowledge check

Check your understanding

Answer this question before you continue.

Why does the experiment keep `random_state=42` fixed while comparing tree depths?
Single Choice

Focus: Explain why a fixed random seed makes depth comparisons more interpretable.

Fit One Tree and Read Its Structure

Start with an unconstrained tree. No depth limit, no minimum leaf size. Let it grow until it is satisfied.

from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)

train_acc = accuracy_score(y_train, tree.predict(X_train))
val_acc = accuracy_score(y_val, tree.predict(X_val))

print(f"train accuracy: {train_acc:.3f}")
print(f"validation accuracy: {val_acc:.3f}")
print("nodes:", tree.tree_.node_count)
print("depth:", tree.get_depth())

On a dataset this small and clean, you will likely see training accuracy at or near 1.000 and validation accuracy noticeably lower. That is the gap, and it is the entire subject of this article.

Now look inside the tree. A fitted tree is not a black box; it is a data structure you can read.

from sklearn.tree import plot_tree
import matplotlib.pyplot as plt

plt.figure(figsize=(12, 6))
plot_tree(tree, max_depth=2, feature_names=iris.feature_names,
          class_names=iris.target_names, filled=True, fontsize=9)
plt.show()

Notice the max_depth=2 in the plotting call. That is deliberate. A full-depth tree on Iris is a wall of boxes, and a wall of boxes teaches nothing at a glance. Truncating the picture to two levels keeps it readable while still showing you the top splits, the thresholds, and the class counts in each node.

To see a prediction as a sequence of comparisons rather than a magic number, walk one sample down the tree:

sample = X_val.iloc[[0]]
path = tree.decision_path(sample)
leaf = tree.apply(sample)

print("predicted class:", iris.target_names[tree.predict(sample)[0]])
print("leaf node id:", leaf[0])
print("nodes visited:", path.indices)

Every prediction is a chain of "is this feature less than or equal to that threshold?" questions, ending at a leaf that holds a class distribution. The tree does not know anything you cannot read off its nodes.

Here is the mechanism behind the gap: an unconstrained tree keeps splitting until its leaves are pure, meaning every training row in a leaf shares a label. Pure leaves are perfect on training data by construction. They are also where the memorization lives.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article's plotting call, what does `plot_tree(..., max_depth=2)` do?
Misconception Check

Focus: Distinguish limiting a tree visualization from limiting the fitted tree.

Sweep max_depth and Watch the Gap

A schematic chart with tree depth increasing along the horizontal axis and accuracy on the vertical axis. The training curve rises and levels off near the top; the validation curve rises, peaks at a moderate depth, then dips.
Compare both curves across depths: a widening gap and a validation downturn can signal overfitting.

Now the core experiment. Fit a fresh tree at each depth, and record both scores side by side.

results = []

for depth in range(1, 16):
    model = DecisionTreeClassifier(max_depth=depth, random_state=42)
    model.fit(X_train, y_train)
    results.append({
        "max_depth": depth,
        "train_acc": accuracy_score(y_train, model.predict(X_train)),
        "val_acc": accuracy_score(y_val, model.predict(X_val)),
    })

df = pd.DataFrame(results)
df["gap"] = df["train_acc"] - df["val_acc"]
print(df.round(3).to_string(index=False))

Your exact numbers will depend on the split, but the shape is what matters. It looks roughly like this:

max_depthtrain_accval_accgap
10.670.640.03
20.950.910.04
30.980.930.05
51.000.910.09
101.000.890.11
151.000.890.11

Read the columns in order. Training accuracy rises and then pins at 1.000. Validation accuracy rises, peaks somewhere in the middle, and then turns down. The gap widens as depth increases.

That turning point is a pattern that suggests overfitting on this split: the tree stops capturing signal and starts fitting noise. The mechanism is simple arithmetic. Each additional level roughly doubles the number of regions the tree can carve out. A shallow tree can only draw a few broad boundaries. A deep tree can isolate smaller and smaller groups of training rows until each group is its own private leaf. When a leaf contains two or three training examples, it has learned those examples, not the pattern behind them.

Keep random_state=42 fixed across the whole sweep. If you let it vary, differences in the table could come from split randomness rather than from depth, and you would be reading noise as signal.

Knowledge check

Check your understanding

Answer this question before you continue.

As `max_depth` grows, training accuracy reaches 1.00 while validation accuracy peaks and then declines. What does this pattern suggest on this split?
Scenario Interpretation

Focus: Interpret rising training accuracy and falling validation accuracy as evidence of overfitting.

Read the Curve, Not the Single Number

The best depth is the one with the best validation score, not the best training score. Training accuracy is a fit report. It tells you how well the tree memorized. It does not tell you whether the tree is any good.

Two signatures are worth memorizing:

  • Large train-minus-validation gap: overfitting. The model fits the training data much better than new data.
  • Both scores low: underfitting. The model is too simple to capture the pattern at all.

There is a trap in reading these curves. With a small validation set, one lucky depth can look best by chance. A single spike in the validation column is weak evidence. A flat plateau where several neighboring depths score similarly is much more trustworthy, because it means the result is not sensitive to the exact depth you picked.

That is also why the downturn in the table is a clue, not a verdict. A single split gives you one estimate of the curve. To check whether the peak is stable, run the same sweep with cross-validation on the training data and see whether the best depth moves. If it does, the exact turning point was partly noise. If it holds, you have a pattern worth trusting.

Depth is also not your only control. min_samples_leaf and min_samples_split limit how finely a node may be carved, and they often produce smoother generalization behavior than depth alone. A tree with min_samples_leaf=5 cannot create a leaf holding two rows, no matter how deep it goes.

Knowledge check

Check your understanding

Answer this question before you continue.

Two neighboring depths have similar validation scores across a flat plateau, while one other depth has a single higher spike. Which evidence is more trustworthy for choosing a depth, according to the article?
Comparison Reasoning

Focus: Prefer a stable validation-score plateau over an isolated peak when judging a depth sweep.

When the Experiment Misleads You

The experiment is simple, which means it is easy to fool yourself. Here are the failure modes I see most often.

Non-determinism. Without a fixed random_state, ties between equally good splits get broken differently on each run. Your table shifts, and you chase a difference that was never real.

Leakage. If a feature encodes the label, or was computed using the full dataset before you split, validation accuracy will look excellent and mean nothing. The tree is reading the answer key.

Class imbalance. Plain accuracy can look high while the tree quietly ignores the minority class. Check per-class behavior before you trust the headline number.

Tiny validation sets. A handful of rows makes every accuracy jump look dramatic. Always report the number of validation samples next to the score, so the reader knows how much weight the number can carry.

High-dimensional data with few rows. Trees overfit easily when features outnumber samples. On data like that, the depth sweep will turn down earlier than you expect, and the peak validation score will be lower.

Common mistake: Treating the best validation score as a property of the algorithm. It is a property of this dataset, this split, and this seed. Change any of them and the winning depth can move.

Your Next Experiment: Change a Second Dial

You have one dial working. Now add a second and watch how the curve responds.

Re-run the sweep with min_samples_leaf set to a few values, and compare the shape of the validation curve against the depth-only version:

for leaf_size in [1, 3, 5, 10]:
    scores = []
    for depth in range(1, 16):
        model = DecisionTreeClassifier(
            max_depth=depth, min_samples_leaf=leaf_size, random_state=42
        )
        model.fit(X_train, y_train)
        scores.append(accuracy_score(y_val, model.predict(X_val)))
    best = max(scores)
    print(f"min_samples_leaf={leaf_size:>2}  best val_acc={best:.3f}")

Then try two more modifications. Swap the single validation split for cross-validation and see whether the best depth moves. And compare your chosen tree against a trivial baseline that always predicts the majority class, so you learn to ask whether the model beats doing nothing at all.

Record two numbers when you finish: the winning depth and the gap at that depth. That pair is the honest summary of the experiment. A single accuracy figure without its gap is incomplete evidence, and you should treat it that way.

The same tension you just watched — complexity up, training fit up, generalization eventually down — is what drives the ensemble methods that come next. A random forest does not escape this tradeoff. It manages it by combining many constrained trees instead of trusting one deep one. Run the sweep once more with min_samples_leaf, and you will have the intuition you need to understand why.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A tree gets an unexpectedly excellent validation score. You discover that one input feature was computed using the full dataset, including validation rows. What is the central problem to address?
Question 1 of 2Debugging

Focus: Recognize label leakage as a reason an apparently excellent validation score may be meaningless.

You rerun the depth sweep with `min_samples_leaf=5`. Which outcome is guaranteed by the constraint described in the article?
Question 2 of 2Scenario Interpretation

Focus: Predict how a minimum leaf-size constraint changes the splits a decision tree may make.

References

  1. 1.10. Decision Trees — scikit-learn 1.9.1 documentationscikit-learn.org
  2. Understanding the decision tree structure — scikit-learn 1.9.0 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.