
Estimate a Statistic’s Uncertainty With a Reproducible Bootstrap Experiment
You have one sample, one number, and no honest way to say how much that number would move if you collected the data again. The bootstrap does not create…
Read tutorialArticles whose authoritative article variant is practice and whose primary outcome is a runnable implementation or experiment.
Tagged articles
49 articles in this tag.

You have one sample, one number, and no honest way to say how much that number would move if you collected the data again. The bootstrap does not create…
Read tutorial
A search is not a machine that finds the best number. It is a bounded experiment whose output is a comparison.
Read tutorial
A stacking classifier can score beautifully on your test set for a reason that has nothing to do with skill: the meta-model was trained on predictions the…
Read tutorial
Swap OrdinalEncoder for OneHotEncoder, rerun the script, and the score moves. That single number feels like a verdict: one encoder is better. It isn't. You…
Read tutorial
A model with a respectable AUC can still make the wrong decision every single day. The scores are fine. The cutoff is the problem.
Read tutorial
Your clusters looked clean until you changed the random seed. Now the labels shuffled, the boundaries moved, and you are not sure whether you found…
Read tutorial
You have three pipelines. Their scores are close enough that the ranking feels like a coin flip. So you run them against the test set to break the tie —…
Read tutorial
Two detectors, one dataset, two different alert lists — and no labels to referee the disagreement.
Read tutorial
You run SelectKBest, get a tidy list of "important" features, then rerun with a different random seed and the list changes. That surprise is not a bug in…
Read tutorial
Fill the missing values, split the data, train the model, and watch your score climb. It feels like progress. It is usually a leak.
Read tutorial
One score is a snapshot. A distribution is a measurement. If you fit Random Forest and Extra Trees once, watch Extra Trees edge ahead by 0.01 accuracy, and…
Read tutorial
Same model. Same features. Two validation schemes. One score says 0.94, the other says 0.61.
Read tutorial
You fit ridge, you fit lasso, lasso wins by a hair, and you quietly file away "lasso is better." That conclusion is a coin flip wearing a lab coat. One…
Read tutorial
Change one number in a DBSCAN call and the whole story of your data can flip: two clusters become one, or a clean grouping dissolves into a field of noise.…
Read tutorial
A decision tree can score 100% on the data it was trained on and still be wrong about almost everything else. That gap is not a bug. It is the lesson.
Read tutorial
Your model is live. Labels arrive in three weeks. Right now, the only thing you can actually see is the input stream — and it looks different from what you…
Read tutorial
A feature idea is a hypothesis. Change one thing, measure the same way twice, and let the output decide.
Read tutorial
Two representations. One corpus. One split. One classifier. The only honest way to know what hashing costs you is to measure it against the vocabulary you…
Read tutorial
Every scikit-learn tutorial shows you the same two lines. Almost none of them stop to show you what changed inside the object between them.
Read tutorial
A hard label is a decision. A posterior probability is a confession about how close that decision was.
Read tutorial
A single accuracy number tells you almost nothing. A controlled experiment tells you whether that number means anything at all.
Read tutorial
You set a learning rate, hit run, and the loss curve does something strange. Maybe it barely moves. Maybe it bounces. Maybe it turns into nan. The number…
Read tutorial
You fit the model, plot a tidy dendrogram, cut it at the biggest gap, and report three clusters. Then you change one argument — the linkage — and get a…
Read tutorial
Most people stop at fit(). They get a label array, scatter-plot it, and call the job done. But the label array is not the result — it is the beginning of…
Read tutorial
Same data. Same code. Two accuracy scores that disagree, and nothing in the output explains why.
Read tutorial
You already know noisy labels hurt models. Knowing it changes nothing. The moment you flip a known fraction of labels yourself, retrain, and watch the…
Read tutorial
Your model passed validation. Then it went to production and started predicting one class far more often than it should. The features look normal. The…
Read tutorial
You have seen the classic two-line plot. One line starts high and drifts down. The other starts low and climbs. Somewhere on the right, they meet, and the…
Read tutorial
A model that fits is not yet a model you can trust. Fitting takes one line. Trust takes a split, a reading of the numbers, and a look at where the errors…
Read tutorial
You call predict(), and you get back a tidy row of zeros and ones. Clean. Decisive. And completely silent about the thing you actually wanted to see: the…
Read tutorial
The most dangerous bug in a beginner text classifier is not in the model. It is in the order of operations. Fit your vectorizer on the whole dataset before…
Read tutorial
You tuned a model, watched the cross-validation score climb, and reported the best number. Then new data arrived and the model performed worse than…
Read tutorial
You fit PCA, print the explained variance ratio, and see that PC1 captures 92% of the variance. That looks like a strong result. It might be an artifact of…
Read tutorial
You fit a model, call a feature-importance attribute, and get a tidy ranked bar chart. It looks like a verdict. It is not. That chart is a measurement of…
Read tutorial
A count target is not a permission slip. It is a hypothesis about how the mean and variance move together — and you can test it.
Read tutorial
Fit degree 1 through 9, watch training error fall to almost zero, and pick degree 9. That is the trap. Training error measures how well a model remembers…
Read tutorial
You scale the whole table, split it, train a model, and get 0.94. You ship it, and real data hands you 0.71. Nothing broke in production. The break…
Read tutorial
oob_score=True looks like a free validation set. It is not free, and it is not quite a validation set. It is a running estimate computed on the same rows…
Read tutorial
A respectable R² can sit on top of a model that is wrong in a structured way. The score summarizes average fit; it says nothing about where the fit fails.…
Read tutorial
One row can steer a line. Add a single extreme observation to an otherwise clean dataset, refit, and watch the slope tilt — even though every other point…
Read tutorial
You resample your data to fix the imbalance, rerun the same model, and watch ROC-AUC barely move while PR-AUC collapses. Nothing about the model changed.…
Read tutorial
You fit the pipeline, you save it, you load it back, and the predictions look wrong. Nothing crashed. Nothing warned you. The numbers are just quietly…
Read tutorial
Your model reports 0.91 accuracy. That number is probably correct, and it is probably useless for deciding what to fix next.
Read tutorial
A model can score 94% accuracy and still be quietly failing one class. Here is how to see it.
Read tutorial
A model can rank every sample correctly and still lie about the number attached to each one. This experiment isolates that lie, then measures whether a…
Read tutorial
Here is the uncomfortable truth about missing data: in a real dataset, you never see the values that went missing. You cannot check whether your imputation…
Read tutorial
Same data. Same SVC call. One run scores 0.85, the next 0.55, and the only thing that changed was whether you scaled the features first.
Read tutorial
Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…
Read tutorial
A missing column throws. A reordered DataFrame does not. That asymmetry is where confident wrong numbers come from.
Read tutorial