How Decision Trees Score Classification Splits: Gini, Entropy, and Gain
You fit a tree, print it, and there it is: the root splits on worst radius, not mean texture. The library made a choice. You can't say why.

Key topics
You fit a tree, print it, and there it is: the root splits on worst radius, not mean texture. The library made a choice. You can't say why.
That gap matters. "The tree finds the best split" is not a mental model — it's a placeholder for one. The actual mechanism is a scoring loop: every candidate split is graded by how much class uncertainty it removes, weighted by how many samples land on each side. Learn to run that loop by hand and the library's choice stops being magic.
If you already read a tree as a sequence of partitions, you have the prerequisite. Here we stop asking what the tree looks like and start asking why each cut was chosen.
What a Split Score Actually Measures
A node's impurity is one number describing how mixed its class labels are. A node with 90% one class is nearly pure; a 50/50 node is maximally confused.
A split is scored by comparing the parent's impurity to the size-weighted average of the children's impurity. Whatever impurity disappears in that comparison is the gain. Bigger gain, cleaner separation.
Two impurity families dominate classification trees: Gini and entropy. They answer the same question with different curves, and most of the time they agree on which split wins.
Notation and the Two Impurity Functions
Before any formula, fix the symbols. For a node containing samples from classes, let be the fraction of samples in that node belonging to class . The fractions sum to 1: .
That vector of class probabilities is the entire input to both impurity functions.
Gini impurity
Derive it from its plain-language meaning. Pick a random sample from the node, then assign it a random label drawn from the node's own class distribution. What's the probability you mislabel it?
The probability you draw class is . The probability you assign class is also . So the probability of a correct assignment is , and the misclassification probability is . That is Gini impurity.
For a binary problem, write for the fraction of the positive class:
At or (pure node), . At (balanced), . Gini peaks at the balanced node.
Knowledge check
Check your understanding
Answer this question before you continue.
Entropy
Entropy measures expected surprise. A pure node surprises you never — every sample has the same label — so entropy is 0. A balanced binary node is maximally uncertain, and with the base-2 convention the unit is a bit: one yes/no question's worth of uncertainty.
For the binary case:
At or , . At , . Entropy peaks at the balanced node too — just on a different scale.
Both functions are concave and symmetric about . That shared shape is why they usually rank splits the same way. The scales differ: Gini lives in for binary problems, entropy in .
Note: These formulas describe the empirical label distribution in the node — the fractions you actually counted. They say nothing about the true data-generating distribution. That distinction becomes important when you start trusting split scores as evidence about the world.
Knowledge check
Check your understanding
Answer this question before you continue.
Weighted Impurity Decrease, Step by Step
Now the scoring rule. For a candidate split that sends a fraction of the parent's samples left and right:
Read each term:
- is the impurity before the split.
- and are the children's impurities.
- and are the fractions of the parent's samples reaching each child, with .
The weights are the whole point. Without them, a split that isolates one sample into a pure leaf would score perfectly — impurity 0 on one side, and who cares about the other. Weighting by child size means a tiny pure child can't dominate the score. The gain has to come from cleaning up a meaningful share of the data.
scikit-learn writes the same quantity in raw counts:
where is the number of samples at the current node, and are the counts in the left and right children, and is the total. When you pass sample_weight, all four counts are weighted sums.
That same expression powers the min_impurity_decrease parameter: a split is only accepted if its decrease clears your threshold. Set it to zero and the tree splits whenever any gain exists — which is exactly how you grow a tree that memorizes noise.
Common mistake: Treating the gain as a global optimization. The tree maximizes gain at this node only, with no lookahead. A split that looks mediocre now might have enabled a brilliant split two levels down. Greedy search never finds out.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: Scoring Two Candidate Splits
Here is a small binary-class table. Fifteen rows, one label column, two candidate features.
| Row | Feature A | Feature B | Label |
|---|---|---|---|
| 1 | low | small | 0 |
| 2 | low | small | 0 |
| 3 | low | small | 1 |
| 4 | low | medium | 0 |
| 5 | low | medium | 1 |
| 6 | low | medium | 1 |
| 7 | low | large | 1 |
| 8 | low | large | 1 |
| 9 | high | small | 0 |
| 10 | high | small | 0 |
| 11 | high | medium | 0 |
| 12 | high | medium | 0 |
| 13 | high | large | 1 |
| 14 | high | large | 1 |
| 15 | high | large | 1 |
Parent node. Count the labels: eight 0s, seven 1s. So and .
Gini: .
Entropy: .
Candidate 1 — split on Feature A. Left child (A = low): rows 1–8, labels 0,0,1,0,1,1,1,1 → three 0s, five 1s. Right child (A = high): rows 9–15, labels 0,0,0,0,1,1,1 → four 0s, three 1s.
Left Gini: .
Right Gini: .
Weighted child sum: .
Gain: .
Candidate 2 — split on Feature B. Left child (B = small): rows 1,2,3,9,10 → three 0s, two 1s. Right child (B = medium or large): rows 4–8 and 11–15 → five 0s, five 1s.
Left Gini: .
Right Gini: .
Weighted child sum: .
Gain: .
Feature A wins under Gini. Now repeat with entropy.
Candidate 1, entropy. Left child (3/8, 5/8): . Right child (4/7, 3/7): .
Weighted child sum: .
Gain: .
Candidate 2, entropy. Left child (3/5, 2/5): . Right child (5/10, 5/10): .
Weighted child sum: .
Gain: .
| Candidate | Parent impurity | Left child | Right child | Weighted child sum | Gain |
|---|---|---|---|---|---|
| A (Gini) | 0.498 | 0.468 | 0.490 | 0.479 | 0.019 |
| B (Gini) | 0.498 | 0.480 | 0.500 | 0.493 | 0.005 |
| A (entropy) | 0.997 | 0.954 | 0.985 | 0.969 | 0.028 |
| B (entropy) | 0.997 | 0.971 | 1.000 | 0.991 | 0.006 |
Both criteria pick Feature A. The magnitudes differ — 0.019 versus 0.028 — because the scales differ, not because entropy "found more signal." The ranking is what matters, and here the two criteria agree.
They don't always agree. When two candidates are close in gain, the different curves can flip the winner. That's the tie-breaking regime, and it's where criterion choice actually shows up.
Knowledge check
Check your understanding
Answer this question before you continue.
Gini vs Entropy in Practice
Given how similar the two are, how should you choose?
Start with Gini. It's the scikit-learn default, it's cheaper to compute (squares versus logarithms), and on large datasets with many candidate thresholds that cost is real. For most classification problems, the tree you get is effectively the same.
Switch to entropy when you have a reason. If you're comparing against ID3- or C4.5-style behavior, or diagnosing how split selection interacts with feature cardinality, entropy is the criterion those algorithms use.
The cardinality issue is worth understanding. Entropy-based gain is more prone to favoring features with many distinct values or categories. A noisy high-cardinality feature offers more thresholds to try, so it has more chances to score well by luck. Gini's bias in that direction is smaller, which is why CART defaults to it and why C4.5 introduced gain ratio as a normalization.
Tip: If your tree's best splits keep landing on an ID-like column, suspect split bias before you suspect a real signal. Drop the column and refit. If the tree's structure barely changes, the column was noise.
scikit-learn exposes both through the criterion parameter, and log_loss is the same family as entropy for binary classification.
What a Good Split Score Does Not Tell You
A high-gain split is a statement about your training sample. Nothing more.
It doesn't guarantee generalization. Greedy local optimization can build a deep tree that fits noise perfectly. Every split looked good at the time. The tree still fails on held-out data.
It isn't causal. Impurity decrease is not feature importance in any causal sense. A feature can score well because it's correlated with a proxy for the label, or simply because it has many distinct values to try.
It can stall on patterns no single split can separate. The XOR pattern is the classic case: no axis-aligned cut reduces impurity, so a greedy tree stops early even though the classes are perfectly separable by a diagonal boundary.
It can be fooled by cardinality. A noisy high-cardinality feature can outscore a genuinely informative low-cardinality one. Watch for this when a suspicious column keeps winning.
The procedure is reproducible: compute parent impurity, compute weighted child impurity for each candidate, subtract, pick the largest gain. Then treat that number as a local score, not a verdict.
Fit a tree in scikit-learn, print the tree_.impurity array, and check whether the splits the library chose match the ones you scored by hand. When they diverge, you've found either a bug in your arithmetic or a subtlety in the algorithm — and both are worth chasing.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


