Compare Random Forests and Extra Trees With a Controlled Experiment
One score is a snapshot. A distribution is a measurement. If you fit Random Forest and Extra Trees once, watch Extra Trees edge ahead by 0.01 accuracy, and…

Key topics
One score is a snapshot. A distribution is a measurement. If you fit Random Forest and Extra Trees once, watch Extra Trees edge ahead by 0.01 accuracy, and declare it the winner, you have measured noise and called it quality.
This tutorial builds a fair comparison. You will fit RandomForestClassifier and ExtraTreesClassifier under identical conditions, then sweep the seed to turn a single number into a distribution. The deliverable is the experimental design, not the score.
What a Fair Comparison Actually Requires
You already know the mechanism from the prerequisite comparison: Random Forest draws a bootstrap sample for each tree and searches for the best split among a random feature subset, while Extra Trees uses the full training set and picks a random threshold for each candidate feature before choosing among those random splits. One sentence each, because we are not re-teaching it here. We are testing what it produces.
A controlled experiment holds everything constant except the thing you are studying. Here is the contract:
| Held constant | Manipulated |
|---|---|
| Dataset | Ensemble class (Random Forest vs Extra Trees) |
| Train/test split | random_state seed (swept, not fixed) |
n_estimators | |
max_features | |
| Evaluation metric |
The seed sweep is the important part. When you refit both models across many seeds, you get a distribution of scores per model, and distributions can be compared honestly. That is the success criterion: not one number, but a spread.
Note: "Controlled" means one variable moves. If you change
n_estimatorsfor one model and not the other, you are no longer comparing ensembles — you are comparing hyperparameters.
Knowledge check
Check your understanding
Answer this question before you continue.
Load Data and Fix the Evaluation Split
We need a small, well-understood dataset so runtime stays short and results are reproducible. A scikit-learn built-in classification set is ideal.
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
data = load_breast_cancer()
X, y = data.data, data.target
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.25,
stratify=y,
random_state=42,
)
print("train shape:", X_train.shape)
print("test shape: ", X_test.shape)
print("train class counts:", np.bincount(y_train))
print("test class counts: ", np.bincount(y_test))
Expected output is a shape tuple and two class-count arrays. Eyeball them before trusting any score. stratify=y keeps the class balance stable across the split, which matters when one class is rarer.
Two assumptions are baked in here. First, no leakage: the test set is untouched until scoring. Second, no preprocessing that peeks at the test set. If you add scaling later, fit the scaler on training data only. Tree ensembles do not need feature scaling, which is one reason they are a clean first experiment.
Common mistake: Splitting, then scaling on the full dataset, then splitting again. The scaler has now seen the test set. Keep the pipeline honest from the first line.
Knowledge check
Check your understanding
Answer this question before you continue.
Fit Both Ensembles Under Identical Settings
Here is the smallest useful runnable comparison. Both models get matched n_estimators, max_features, and random_state.
from sklearn.ensemble import RandomForestClassifier, ExtraTreesClassifier
from sklearn.metrics import accuracy_score
def fit_and_score(model, X_train, y_train, X_test, y_test):
model.fit(X_train, y_train)
preds = model.predict(X_test)
return accuracy_score(y_test, preds)
rf = RandomForestClassifier(
n_estimators=200, max_features="sqrt", random_state=42, n_jobs=-1
)
et = ExtraTreesClassifier(
n_estimators=200, max_features="sqrt", random_state=42, n_jobs=-1
)
rf_score = fit_and_score(rf, X_train, y_train, X_test, y_test)
et_score = fit_and_score(et, X_train, y_train, X_test, y_test)
print(f"RandomForestClassifier: {rf_score:.4f}")
print(f"ExtraTreesClassifier: {et_score:.4f}")
You will get two scores within a small margin of each other. That closeness is the point, and it is not a coincidence. Both models average many trees, and that shared averaging mechanism dominates the difference between them. The margin alone proves nothing yet — it is one draw from a noisy distribution.
Tip: Set
n_jobs=-1so trees build in parallel. Each tree is independent, so this is free speed on a multicore machine.
Knowledge check
Check your understanding
Answer this question before you continue.
Sweep the Seed to See Variability
Now the experimental upgrade. Loop over a range of random_state values, refit both models each time, and collect held-out scores.
seeds = range(20)
rf_scores, et_scores = [], []
for seed in seeds:
rf = RandomForestClassifier(
n_estimators=200, max_features="sqrt", random_state=seed, n_jobs=-1
)
et = ExtraTreesClassifier(
n_estimators=200, max_features="sqrt", random_state=seed, n_jobs=-1
)
rf_scores.append(fit_and_score(rf, X_train, y_train, X_test, y_test))
et_scores.append(fit_and_score(et, X_train, y_train, X_test, y_test))
rf_scores = np.array(rf_scores)
et_scores = np.array(et_scores)
print(f"RandomForest: mean={rf_scores.mean():.4f} std={rf_scores.std():.4f}")
print(f"ExtraTrees: mean={et_scores.mean():.4f} std={et_scores.std():.4f}")
Expected output is two summary lines with mean and standard deviation, and — most likely — visible overlap between the two distributions.
Here is where mechanism earns its keep. Before running, predict the spread. Extra Trees randomizes split thresholds, which tends to reduce variance across draws. Random Forest's bootstrap sampling tends to reduce bias. So I would expect Extra Trees to show a slightly tighter spread, and the two means to sit close together. If your output disagrees, that is evidence about this dataset, not a refutation of the mechanism — the effect is conditional on data size, feature redundancy, and budget.
Note: The standard deviation here is across seeds, not across folds. It measures how much the result depends on the particular randomness draw. That is exactly the quantity a single split hides.
Knowledge check
Check your understanding
Answer this question before you continue.
Read the Result Without Declaring a Winner
If the distributions overlap, the honest conclusion is: no reliable difference on this data at this budget. Not "Extra Trees wins." Not "Random Forest is more accurate." The comparison is conditional on the experimental setup.
Name the confounds before you generalize:
- Dataset size and feature redundancy. Redundant features give random thresholds more room to work; small datasets make every split noisier.
n_estimators. More trees shrink variance for both models, compressing the difference.max_features. Fewer features per split increases tree diversity and shifts both means and spreads.- Metric choice. Accuracy hides class imbalance that precision or recall would expose.
A universal winner claim is unsupported because none of those knobs were held fixed across all possible datasets. You tested one configuration on one dataset.
So how do you actually choose? Use a decision rule, not a 0.01 gap:
- Speed matters? Extra Trees is typically faster because it skips the optimal-split search.
- Variance tolerance low? Prefer the model with the tighter spread in your sweep.
- Interpretability needed? Both are ensembles; neither gives you a single readable tree. Reach for feature importances or a surrogate tree instead.
Common mistake: Reporting the best seed as if it were the model's true performance. That is cherry-picking. Report the mean and the spread, or report nothing.
One Modification Worth Running Next
A comparison you cannot extend is a dead end. Change one knob and rerun the sweep. My first pick: drop max_features from "sqrt" to a smaller value, or raise n_estimators from 200 to 500.
Predict the direction before you run. More trees should shrink variance for both models. Fewer features per split should increase tree diversity, which may move the mean and the spread together. Then compare the new distributions against your baseline: did the knob move the mean, the spread, or both?
rf = RandomForestClassifier(
n_estimators=500, max_features=0.3, random_state=seed, n_jobs=-1
)
One runtime note: Extra Trees usually stays faster even at higher n_estimators, because it never searches for the optimal threshold. That speed is a real, measurable difference — and it is the kind of tradeoff a controlled experiment lets you see instead of assume.
The reusable asset here is the template: fixed split, matched settings, seed sweep, distribution comparison. Point it at your own dataset, change one knob at a time, and let the spread tell you what the single score never could.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


