
Run a Bounded Hyperparameter Search With Scikit-Learn
A search is not a machine that finds the best number. It is a bounded experiment whose output is a comparison.
Read tutorialMachine-learning situations in which information from outside the permitted training boundary contaminates features, preprocessing, model selection, or evaluation.
Tagged articles
36 articles in this tag.

A search is not a machine that finds the best number. It is a bounded experiment whose output is a comparison.
Read tutorial
A stacking classifier can score beautifully on your test set for a reason that has nothing to do with skill: the meta-model was trained on predictions the…
Read tutorial
Swap OrdinalEncoder for OneHotEncoder, rerun the script, and the score moves. That single number feels like a verdict: one encoder is better. It isn't. You…
Read tutorial
A model with a respectable AUC can still make the wrong decision every single day. The scores are fine. The cutoff is the problem.
Read tutorial
You have three pipelines. Their scores are close enough that the ranking feels like a coin flip. So you run them against the test set to break the tie —…
Read tutorial
You run SelectKBest, get a tidy list of "important" features, then rerun with a different random seed and the list changes. That surprise is not a bug in…
Read tutorial
Fill the missing values, split the data, train the model, and watch your score climb. It feels like progress. It is usually a leak.
Read tutorial
Same model. Same features. Two validation schemes. One score says 0.94, the other says 0.61.
Read tutorial
You fit ridge, you fit lasso, lasso wins by a hair, and you quietly file away "lasso is better." That conclusion is a coin flip wearing a lab coat. One…
Read tutorial
You split your data once. Your model scores 0.91. You re-run the same split with a different random seed, and suddenly it scores 0.84. Same model. Same…
Read tutorial
You train a model. The validation score comes back at 0.98. You feel like a genius. Then the model goes into the real world and performs like a coin flip.
Read tutorial
A feature idea is a hypothesis. Change one thing, measure the same way twice, and let the output decide.
Read tutorial
You've tried different algorithms. You've tuned parameters until your eyes glaze over. And still, your model's score sits at a plateau, stubbornly refusing…
Read tutorial
You're staring at a wide table of columns, and the model is underperforming. The fix could mean deleting half the columns, reshaping the ones that remain,…
Read tutorial
Every supervised learning problem is a question about a table. The hard part is knowing which columns the model is allowed to look at.
Read tutorial
You train a model. Cross-validation gives you a confident accuracy score. You deploy. The model underperforms in ways your validation never hinted at.
Read tutorial
Your best tuning score is not your model's true performance. It is the score of the luckiest experiment you ran.
Read tutorial
A model that fits is not yet a model you can trust. Fitting takes one line. Trust takes a split, a reading of the numbers, and a look at where the errors…
Read tutorial
Your model is only as smart as the table you hand it. Most beginners skip straight to fitting an algorithm, then blame the model when results disappoint.…
Read tutorial
Every beginner hits the same wall: you load a real dataset, and it is full of holes. Columns you need are dotted with NaN. Your model refuses to run. So…
Read tutorial
Here is the trap I see beginners fall into constantly: they try five algorithms, compare the scores, keep the winner, then tweak it, compare again, and…
Read tutorial
The most dangerous bug in a beginner text classifier is not in the model. It is in the order of operations. Fit your vectorizer on the whole dataset before…
Read tutorial
You run a grid search, watch the cross-validation score climb, and feel good about the model you are about to ship. Then the model lands on real data and…
Read tutorial
You tuned a model, watched the cross-validation score climb, and reported the best number. Then new data arrived and the model performed worse than…
Read tutorial
Your data is clean. Your rows are ready. Then you feed the model a column of colors—"red", "green", "blue"—and it refuses to train. The error message is…
Read tutorial
You fit PCA, print the explained variance ratio, and see that PC1 captures 92% of the variance. That looks like a strong result. It might be an artifact of…
Read tutorial
You have a wide table of features, a model to build, and a quiet suspicion that half those columns are noise. The instinct is to shrink the dataset. But…
Read tutorial
Fit degree 1 through 9, watch training error fall to almost zero, and pick degree 9. That is the trap. Training error measures how well a model remembers…
Read tutorial
You scale the whole table, split it, train a model, and get 0.94. You ship it, and real data hands you 0.71. Nothing broke in production. The break…
Read tutorial
You fit a scaler on your full dataset, split into training and test sets, train a model, and watch it score beautifully. Then the model meets real data and…
Read tutorial
Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing…
Read tutorial
Same data. Same SVC call. One run scores 0.85, the next 0.55, and the only thing that changed was whether you scaled the features first.
Read tutorial
Target encoding is one of the most seductive traps in feature engineering. You encode a high-cardinality feature like city or merchant_id, watch your…
Read tutorial
Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…
Read tutorial
You've trained a model. The accuracy looks great. You're ready to ship it. Then someone asks the question that stops every beginner cold: How do you know…
Read tutorial
More models do not automatically mean a better model. An ensemble only wins when its members make different mistakes—and the way you combine them…
Read tutorial