
Classical Anomaly Detection: Find Unusual Data Without Labels
Anomaly detection is not "find the weird rows." It is a decision about what counts as normal, made before you ever run an algorithm. Get that decision…
Read tutorialMeasuring predictive behavior and interpreting quality with metrics, error analysis, diagnostics, and task-appropriate evidence.
Tagged articles
78 articles in this tag.

Anomaly detection is not "find the weird rows." It is a decision about what counts as normal, made before you ever run an algorithm. Get that decision…
Read tutorial
If you already understand decision trees, you know the dilemma: a deep tree memorizes the training data and fails on new data, while a shallow tree is too…
Read tutorial
Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorial
Two models can post nearly identical test scores and still fail for opposite reasons. One misses because it never learned the pattern. The other misses…
Read tutorial
Your model scores 0.82 accuracy on the test set. That feels like a fact—something solid you can report and defend. Run the experiment again on a different…
Read tutorial
Two forecasters can post the same Brier score and deserve opposite verdicts. One is honest but says almost nothing. The other is sharp but lies about its…
Read tutorial
A stacking classifier can score beautifully on your test set for a reason that has nothing to do with skill: the meta-model was trained on predictions the…
Read tutorial
A model with a respectable AUC can still make the wrong decision every single day. The scores are fine. The cutoff is the problem.
Read tutorial
Your spam filter just flagged an email with a score of 0.55. Is it spam? The model isn't telling you—it's asking you to decide. Somewhere between the…
Read tutorial
Clustering always returns groups. Even on random noise, even on data with no structure worth naming, every algorithm will happily partition your points and…
Read tutorial
Your clusters looked clean until you changed the random seed. Now the labels shuffled, the boundaries moved, and you are not sure whether you found…
Read tutorial
You have three pipelines. Their scores are close enough that the ranking feels like a coin flip. So you run them against the test set to break the tie —…
Read tutorial
Two detectors, one dataset, two different alert lists — and no labels to referee the disagreement.
Read tutorial
You run SelectKBest, get a tidy list of "important" features, then rerun with a different random seed and the list changes. That surprise is not a bug in…
Read tutorial
Fill the missing values, split the data, train the model, and watch your score climb. It feels like progress. It is usually a leak.
Read tutorial
One score is a snapshot. A distribution is a measurement. If you fit Random Forest and Extra Trees once, watch Extra Trees edge ahead by 0.01 accuracy, and…
Read tutorial
Same model. Same features. Two validation schemes. One score says 0.94, the other says 0.61.
Read tutorial
You fit ridge, you fit lasso, lasso wins by a hair, and you quietly file away "lasso is better." That conclusion is a coin flip wearing a lab coat. One…
Read tutorial
A model returns 0.4. The default cutoff turns that into "negative," and nobody asked whether that was the right call.
Read tutorial
You split your data once. Your model scores 0.91. You re-run the same split with a different random seed, and suddenly it scores 0.84. Same model. Same…
Read tutorial
Two models. Mean cross-validation scores of 0.842 and 0.847. The instinct is immediate: the second one wins.
Read tutorial
Your model aced the test set. Clean evaluation, strong metrics, confident you. Then you deploy it, and the real world quietly moves on without telling you.
Read tutorial
You train a model. The validation score comes back at 0.98. You feel like a genius. Then the model goes into the real world and performs like a coin flip.
Read tutorial
Your decision tree scores 98% on training data. On new data, it drops to 74%. The tree didn't fail because it wasn't smart enough. It failed because it was…
Read tutorial
A decision tree can score 100% on the data it was trained on and still be wrong about almost everything else. That gap is not a bug. It is the lesson.
Read tutorial
Your model is live. Labels arrive in three weeks. Right now, the only thing you can actually see is the input stream — and it looks different from what you…
Read tutorial
You improved the model twice. The validation score went up, then up again. Now every change you try makes it worse or does nothing. You are not out of…
Read tutorial
A feature idea is a hypothesis. Change one thing, measure the same way twice, and let the output decide.
Read tutorial
You fit a model to something like house prices or income. Most values bunch on the left, and a long tail stretches to the right. The model behaves…
Read tutorial
A single accuracy number tells you almost nothing. A controlled experiment tells you whether that number means anything at all.
Read tutorial
Your best tuning score is not your model's true performance. It is the score of the luckiest experiment you ran.
Read tutorial
A model that predicts "no fraud" for every transaction can score 98% accuracy while catching zero fraud. The number feels trustworthy. It is not. Accuracy…
Read tutorial
Same data. Same code. Two accuracy scores that disagree, and nothing in the output explains why.
Read tutorial
You already know noisy labels hurt models. Knowing it changes nothing. The moment you flip a known fraction of labels yourself, retrain, and watch the…
Read tutorial
You've added features. You've tuned hyperparameters. You've tried a more flexible model. And still, validation performance sits at the same stubborn…
Read tutorial
Your model passed validation. Then it went to production and started predicting one class far more often than it should. The features look normal. The…
Read tutorial
If you are like most beginners, you start guessing. Maybe more data will fix it. Maybe a bigger model. Maybe fancier features. You try one, see little…
Read tutorial
You have seen the classic two-line plot. One line starts high and drifts down. The other starts low and climbs. Somewhere on the right, they meet, and the…
Read tutorial
You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a…
Read tutorial
You have probably heard the advice: "Just use a random forest. Trees handle everything." It sounds practical. It is also how many beginners end up with a…
Read tutorial
Most introductions to linear regression make it sound like drawing a line through some dots. That description is true and almost useless. Any line can be…
Read tutorial
A model that fits is not yet a model you can trust. Fitting takes one line. Trust takes a split, a reading of the numbers, and a look at where the errors…
Read tutorial
You call predict(), and you get back a tidy row of zeros and ones. Clean. Decisive. And completely silent about the thing you actually wanted to see: the…
Read tutorial
You train your first classifier. It hits 85% accuracy on the test set. You feel great—until you realize that a rule which ignores your data entirely would…
Read tutorial
A model that scores 90% accurate can still fail at the exact job it was built for. The headline number looks impressive, but it may be hiding a model that…
Read tutorial
Most beginners treat machine learning as a collection of algorithms to memorize. They collect names like trading cards—logistic regression, random forest,…
Read tutorial
You run a grid search. The leaderboard prints a winner at 0.91. You ship it, and the first honest measurement lands near 0.86. Nothing broke. The model did…
Read tutorial
Here is the trap I see beginners fall into constantly: they try five algorithms, compare the scores, keep the winner, then tweak it, compare again, and…
Read tutorial
Most people assume multiclass classification is just binary classification with more labels. It is not. The real problem is that many classical models only…
Read tutorial
Most classification tutorials teach you to pick one answer. A news article is politics or finance. A movie is action or comedy. A support ticket is billing…
Read tutorial
The most dangerous bug in a beginner text classifier is not in the model. It is in the order of operations. Fit your vectorizer on the whole dataset before…
Read tutorial
You run a grid search, watch the cross-validation score climb, and feel good about the model you are about to ship. Then the model lands on real data and…
Read tutorial
You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…
Read tutorial
A rating scale looks like it should be either a multiclass problem or a regression problem. Both instincts quietly discard the information that matters…
Read tutorial
Your model just aced its training data. Ninety-eight percent accuracy. Beautiful curves. Then you hand it new examples, and it stumbles like it never saw…
Read tutorial
You fit a model, call a feature-importance attribute, and get a tidy ranked bar chart. It looks like a verdict. It is not. That chart is a measurement of…
Read tutorial
A count target is not a permission slip. It is a hypothesis about how the mean and variance move together — and you can test it.
Read tutorial
Fit degree 1 through 9, watch training error fall to almost zero, and pick degree 9. That is the trap. Training error measures how well a model remembers…
Read tutorial
You fit a straight line to data that clearly curves, and the line misses everything. Not by a little—systematically. It sits above the points in one…
Read tutorial
A model can keep the same weights, the same features, and the same threshold, and still report precision 0.9 on your test set and precision 0.3 in…
Read tutorial
A regression model hands you one number. The person acting on that number needs to know how much weight it can bear. A prediction interval is how you show…
Read tutorial
A model can rank every case perfectly and still hand you probabilities you cannot spend. Calibration is the difference between a score that orders risk and…
Read tutorial
oob_score=True looks like a free validation set. It is not free, and it is not quite a validation set. It is a running estimate computed on the same rows…
Read tutorial
A single decision tree can feel like a confident liar: it nails your training data, then wobbles on new data, and reshapes its entire structure when one…
Read tutorial
A respectable R² can sit on top of a model that is wrong in a structured way. The score summarizes average fit; it says nothing about where the fit fails.…
Read tutorial
You have data. You have a prediction goal. And you're stuck on a question that feels like it should be simple: should I use regression or classification?
Read tutorial
You tune a model, retrain, and the score moves. The question is whether that movement came from your edit or from luck. Most beginners cannot tell the…
Read tutorial
A regression model reports a solid R-squared, so you call it done. Then you plot the predictions against the actual values and notice something…
Read tutorial
One row can steer a line. Add a single extreme observation to an otherwise clean dataset, refit, and watch the slope tilt — even though every other point…
Read tutorial
You resample your data to fix the imbalance, rerun the same model, and watch ROC-AUC barely move while PR-AUC collapses. Nothing about the model changed.…
Read tutorial
You train a fraud detector on a dataset where only 0.5% of transactions are fraudulent. The ROC AUC comes back at 0.95. You feel great. Then you look at…
Read tutorial
Your model reports 0.91 accuracy. That number is probably correct, and it is probably useless for deciding what to fix next.
Read tutorial
A model can score 94% accuracy and still be quietly failing one class. Here is how to see it.
Read tutorial
A model can rank every sample correctly and still lie about the number attached to each one. This experiment isolates that lie, then measures whether a…
Read tutorial
Same data. Same SVC call. One run scores 0.85, the next 0.55, and the only thing that changed was whether you scaled the features first.
Read tutorial
Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…
Read tutorial
You've trained a model. The accuracy looks great. You're ready to ship it. Then someone asks the question that stops every beginner cold: How do you know…
Read tutorial
You have a tabular dataset, a prediction to make, and a list of algorithm names that reads like a menu in a language you have not learned yet. Random…
Read tutorial