Skip to content
beginner

Run a Bounded Hyperparameter Search With Scikit-Learn

A search is not a machine that finds the best number. It is a bounded experiment whose output is a comparison.

Published 2026-10-02Updated 2026-10-048 min read
Close-up of a car dashboard with illuminated gauges displaying speed and RPM.
Close-up of a car dashboard with illuminated gauges displaying speed and RPM. Photo by Jae Park on Pexels.

A search is not a machine that finds the best number. It is a bounded experiment whose output is a comparison.

You write a search space with twelve values per parameter, hit run, and twenty minutes later you still cannot say whether the winning setting is real or lucky. The problem is rarely the code. It is the mental model. By the end of this tutorial you will have a runnable search, a results table you can actually read, and a test set that has not been touched.

What a Search Object Actually Does

Before you configure anything, hold the right picture in your head. A scikit-learn hyperparameter search bundles five decisions into one object:

  • an estimator (the model being tuned)
  • a parameter space (the candidate values)
  • a sampling method (grid or random)
  • a cross-validation scheme (how folds are built)
  • a score function (what "better" means)

For every candidate setting, the search fits the estimator on each training fold, scores it on the held-out fold, and averages the scores. The winner is the setting with the best average, not the best single fold. That distinction is the whole point of cross-validation hyperparameter tuning: one lucky fold should not crown a model.

After fitting, the search object behaves like the estimator it wraps. It exposes predict and score, and by default it refits the winning setting on the full development set, so best_estimator_ is ready to use.

One boundary worth naming: only hyperparameters belong in the search space. These are values you set before training, like C or max_depth. Parameters learned during training, such as coefficients or split thresholds, are not yours to search.

If you have already read about train/validation/test boundaries, you know why the test set must stay out of this loop. Here we only need the operational consequence: the search sees the development set and nothing else.

Knowledge check

Check your understanding

Answer this question before you continue.

Why does the search select a candidate using its average cross-validation score rather than its score on one fold?
Misconception Check

Focus: Explain how cross-validation scores determine the winning candidate in a hyperparameter search.

Set Up the Development and Test Boundary First

A dataset splits into development data and a held-out test set. Development data flows into cross-validated search and a refitted best model; only that model then flows to the test set for one final score. No path leads from the test set back to the search.
Use development data to compare candidates; reserve the test set for one final evaluation.

Make the split the first line of code, before you write a single parameter value. Once you have seen a test score, you cannot unsee it, and every later decision is quietly informed by it.

import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)

X_dev, X_test, y_dev, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

The environment assumption is plain: scikit-learn, NumPy, pandas, and SciPy, with a small built-in dataset so the example runs anywhere. SciPy matters because the search space below samples from scipy.stats. The stratify=y argument keeps the class balance similar in both parts, which matters on small datasets.

From here on, the search code references X_dev and y_dev. It never mentions X_test. That is not a stylistic choice; it is the boundary that keeps your final estimate honest.

Note: A single split is a coarse boundary. If your dataset is small, the development set itself will be cross-validated inside the search, which is exactly what the next section does.

Knowledge check

Check your understanding

Answer this question before you continue.

In the example split, what does `stratify=y` help preserve between the development and test sets?
Single Choice

Focus: Describe the purpose of stratifying the development/test split in the example.

Build a Bounded Search Space

Bounded means you can state, in one sentence, why each value is in the list and what range it covers. If you cannot, you are guessing with extra steps.

I reach for two habits here. First, use logarithmic spacing for scale-sensitive parameters such as C or alpha. Linear spacing wastes candidates in a region that behaves almost identically. Second, keep the space small enough that you can explain it.

Grid search enumerates the full cartesian product, so cost multiplies fast. Random search samples a fixed number of candidates and is the better default when only a few parameters actually matter. For random search, continuous distributions from scipy.stats let a larger n_iter keep refining the same range instead of re-testing the same discrete points.

from scipy.stats import loguniform
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

pipe = Pipeline([
    ("scaler", StandardScaler()),
    ("clf", LogisticRegression(max_iter=1000)),
])

param_dist = {
    "clf__C": loguniform(1e-3, 1e2),
    "clf__penalty": ["l2"],
}

A practical starting budget for a beginner: a handful of values per parameter, two or three parameters, and five folds. Do the arithmetic before you run. Two parameters with four values each and five folds is 4 × 4 × 5 = 80 fits. That number is your cost, and it should be visible before you press enter.

Common mistake: Fitting the scaler outside the search leaks validation information into every fold. Wrapping scaling in a Pipeline means it is refit inside each fold, which is what keeps the estimate honest.

Knowledge check

Check your understanding

Answer this question before you continue.

A grid has two parameters with four values each and uses five-fold cross-validation. How many fits does the search require?
Output Prediction

Focus: Calculate the fit budget for a grid search from candidate counts and cross-validation folds.

Run the Search and Read the Results

Fit the search on the development data only, with a fixed random_state so the run is repeatable.

from sklearn.model_selection import RandomizedSearchCV

search = RandomizedSearchCV(
    estimator=pipe,
    param_distributions=param_dist,
    n_iter=20,
    cv=5,
    scoring="roc_auc",
    random_state=42,
    n_jobs=-1,
)

search.fit(X_dev, y_dev)

print("Best params:", search.best_params_)
print("Best mean CV score:", round(search.best_score_, 4))

Inspect best_params_ and best_score_ first, then open cv_results_ for the full picture: mean test score, standard deviation across folds, and rank for every candidate.

import pandas as pd

results = pd.DataFrame(search.cv_results_)
cols = ["rank_test_score", "mean_test_score", "std_test_score", "params"]
print(results[cols].sort_values("rank_test_score").head())

You should see a small table sorted by rank, with a mean score and a spread for each candidate. That table is your success criterion. If it looks like a wall of near-identical numbers, the search space was too narrow to matter.

The standard deviation matters as much as the mean. Two settings with nearly identical means but very different spreads are not equally trustworthy. A setting that wins by 0.002 with a spread of 0.05 is noise wearing a crown.

Because refit=True is the default, the winning setting is refit on the whole development set, which is why search.best_estimator_ is ready to predict.

Knowledge check

Check your understanding

Answer this question before you continue.

After fitting the search with `scoring="roc_auc"` and `cv=5`, what does `search.best_score_` report?
Scenario Interpretation

Focus: Interpret the meaning of the fitted search object's `best_score_` result.

Finish With One Honest Test-Set Evaluation

The search is done. The development-set comparisons are done. Now you spend the test set, once, and never feed the result back into tuning.

from sklearn.metrics import roc_auc_score

final_model = search.best_estimator_
test_score = roc_auc_score(y_test, final_model.predict_proba(X_test)[:, 1])

print("Final test ROC AUC:", round(test_score, 4))

This is the number you report. It is not a candidate to compare against other settings, and it is not a reason to rerun the search with a different range. The moment you use it to choose anything, it stops being a test set and becomes a validation set you happened to call by another name.

Expect the test score to differ from best_score_. The search's mean CV score is an average over folds on the development data; the test score is a single measurement on data the model has never seen. A small gap is normal. A large gap usually means the search overfit the validation folds, the split was unlucky, or the two sets are not drawn from the same distribution.

Warning: If you find yourself running the test evaluation, changing a parameter, and running it again, stop. That loop is tuning on the test set, and it quietly converts your honest estimate into an optimistic one.

Failure Modes You Will Actually Hit

The winning setting sits at the edge of the grid. That is a signal, not a result. The search is telling you the useful range lies outside what you gave it. Widen the range and rerun.

Preprocessing fitted outside the search. If you scale or encode before the search, every fold sees information from the others. Wrap it in a pipeline.

A typo in a parameter name. Scikit-learn raises an error at fit time and names the offending key. Read the message instead of guessing.

Scores that swing wildly across folds. Usually too few samples, too many folds, or a metric that is unstable on small validation sets.

Tuning against the test set, even once. That converts it into a validation set and destroys your only honest estimate.

One Modification Worth Trying

Change exactly one thing and compare. Swap grid search for random search with the same budget, or widen one parameter's range. Keep the test set untouched through the comparison; the experiment is about the search, not the final number.

If the winner barely moves, the extra search budget bought little. If it moves a lot, your original range was the bottleneck. Record the seed, the search space, and the selected settings so the comparison is reproducible rather than remembered.

Where This Leaves You

Run the search on the development set. Read the mean and the spread together. Touch the test set exactly once, at the end, and report that number without letting it steer anything.

The natural next question is whether the selected score itself is optimistic, since you chose it by looking at validation folds. That is what nested cross-validation exists to answer. For now, take one concrete action: rerun your own search with a narrower range and a fixed seed, then compare the selected settings against your first run. The difference will tell you whether your search space or your budget was the real constraint.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You have completed the development-set search and evaluated the selected model on the test set. The test score is lower than `best_score_`. What is the article's recommended next step?
Question 1 of 2Scenario Interpretation

Focus: Use an untouched test set for one final evaluation without feeding its result back into tuning.

You widen one parameter's range while keeping the test set untouched. If the selected setting changes substantially, what does the article suggest this indicates?
Question 2 of 2Comparison Reasoning

Focus: Interpret how a controlled search modification can reveal whether the original range or budget constrained the search.

References

  1. API design for machine learning software: experiences from the scikit-learn projectarxiv.org
  2. bergstra12a.dviwww.jmlr.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.