
Classical Anomaly Detection: Find Unusual Data Without Labels
Anomaly detection is not "find the weird rows." It is a decision about what counts as normal, made before you ever run an algorithm. Get that decision…
Read tutorialArticles explicitly planned for readers at the intermediate difficulty level.
Tagged articles
113 articles in this tag.

Anomaly detection is not "find the weird rows." It is a decision about what counts as normal, made before you ever run an algorithm. Get that decision…
Read tutorial
If you already understand decision trees, you know the dilemma: a deep tree memorizes the training data and fails on new data, while a shallow tree is too…
Read tutorial
Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorial
Two models can post nearly identical test scores and still fail for opposite reasons. One misses because it never learned the pattern. The other misses…
Read tutorial
You have a thousand bootstrap replicates sitting in a NumPy array. You have a histogram. What you do not have is the two numbers you actually need to…
Read tutorial
A random forest hands you an out-of-bag score, and someone nearby says, "Each tree leaves out about a third of the data." That number feels like a setting.…
Read tutorial
Your model scores 0.82 accuracy on the test set. That feels like a fact—something solid you can report and defend. Run the experiment again on a different…
Read tutorial
You have one sample, one number, and no honest way to say how much that number would move if you collected the data again. The bootstrap does not create…
Read tutorial
Two forecasters can post the same Brier score and deserve opposite verdicts. One is honest but says almost nothing. The other is sharp but lies about its…
Read tutorial
A stacking classifier can score beautifully on your test set for a reason that has nothing to do with skill: the meta-model was trained on predictions the…
Read tutorial
A model with a respectable AUC can still make the wrong decision every single day. The scores are fine. The cutoff is the problem.
Read tutorial
Your spam filter just flagged an email with a score of 0.55. Is it spam? The model isn't telling you—it's asking you to decide. Somewhere between the…
Read tutorial
Clustering always returns groups. Even on random noise, even on data with no structure worth naming, every algorithm will happily partition your points and…
Read tutorial
K-means does not discover the real categories hiding in your data. It draws geometric boundaries around points that happen to sit close together. Those…
Read tutorial
Your clusters looked clean until you changed the random seed. Now the labels shuffled, the boundaries moved, and you are not sure whether you found…
Read tutorial
Every clustering algorithm will happily partition pure noise into tidy groups. Run k-means on random data, and it returns clean circles of assigned points.…
Read tutorial
Two detectors, one dataset, two different alert lists — and no labels to referee the disagreement.
Read tutorial
You run SelectKBest, get a tidy list of "important" features, then rerun with a different random seed and the list changes. That surprise is not a bug in…
Read tutorial
One score is a snapshot. A distribution is a measurement. If you fit Random Forest and Extra Trees once, watch Extra Trees edge ahead by 0.01 accuracy, and…
Read tutorial
Same model. Same features. Two validation schemes. One score says 0.94, the other says 0.61.
Read tutorial
A model returns 0.4. The default cutoff turns that into "negative," and nobody asked whether that was the right call.
Read tutorial
You split your data once. Your model scores 0.91. You re-run the same split with a different random seed, and suddenly it scores 0.84. Same model. Same…
Read tutorial
A satisfaction rating of 4 when the truth was 5 is not the same kind of mistake as predicting 1. Most models cannot tell the difference. This one can — but…
Read tutorial
You add more columns to your dataset expecting a better model. Instead, your k-nearest neighbors accuracy drops, your clusters turn to mush, and every…
Read tutorial
Your model aced the test set. Clean evaluation, strong metrics, confident you. Then you deploy it, and the real world quietly moves on without telling you.
Read tutorial
You train a model. The validation score comes back at 0.98. You feel like a genius. Then the model goes into the real world and performs like a coin flip.
Read tutorial
You learned K-means, and it felt clean: pick a number of groups, let centroids pull points inward, and read off the labels. Then you hit a dataset with…
Read tutorial
Two points sit the same distance from a dense blob. DBSCAN labels one a cluster member and the other noise. That looks arbitrary until you stop asking how…
Read tutorial
Change one number in a DBSCAN call and the whole story of your data can flip: two clusters become one, or a clean grouping dissolves into a field of noise.…
Read tutorial
A fully grown decision tree can score near-perfect on the data it was trained on and still lose to a single split on data it has never seen. The usual…
Read tutorial
You fit a tree, print it, and there it is: the root splits on worst radius, not mean texture. The library made a choice. You can't say why.
Read tutorial
You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from.…
Read tutorial
Your model is live. Labels arrive in three weeks. Right now, the only thing you can actually see is the input stream — and it looks different from what you…
Read tutorial
A model degrades in production. Three engineers offer three diagnoses: "the inputs drifted," "the class balance changed," "the meaning changed." All three…
Read tutorial
You improved the model twice. The validation score went up, then up again. Now every change you try makes it worse or does nothing. You are not out of…
Read tutorial
A feature idea is a hypothesis. Change one thing, measure the same way twice, and let the output decide.
Read tutorial
You picked 2^18 buckets. You trained the model. The metrics came back slightly worse than your explicit-vocabulary baseline, and now you are staring at the…
Read tutorial
Feature hashing trades a little information and a lot of interpretability for something genuinely rare in machine learning: a categorical encoding that…
Read tutorial
Two representations. One corpus. One split. One classifier. The only honest way to know what hashing costs you is to measure it against the vocabulary you…
Read tutorial
You fit a linear model. The coefficients look sensible. Each feature seems to contribute its own share, and the numbers align with your intuition. Then you…
Read tutorial
Two houses sit side by side in your table. One is 1,800 square feet with 3 bedrooms. The other is 2,000 square feet with 5 bedrooms. To a k-NN model, the…
Read tutorial
You're staring at a wide table of columns, and the model is underperforming. The fix could mean deleting half the columns, reshaping the ones that remain,…
Read tutorial
Picture a point sitting exactly between two well-separated blobs of data. K-means will force it into one group or the other, even when the data could…
Read tutorial
A hard label is a decision. A posterior probability is a confession about how close that decision was.
Read tutorial
Linear regression draws a straight line through continuous numbers. Logistic regression draws an S-curve that outputs probabilities. They look like…
Read tutorial
Naive Bayes and logistic regression can classify the same dataset with nearly the same accuracy, yet they behave very differently on small samples, missing…
Read tutorial
Random forests grow many trees independently and average their votes. Gradient boosting grows trees one at a time, each one aimed at the mistakes the…
Read tutorial
A single accuracy number tells you almost nothing. A controlled experiment tells you whether that number means anything at all.
Read tutorial
"Gradient descent works" and "gradient descent is guaranteed to converge" sound like the same sentence. They are not. One is a hope backed by experience.…
Read tutorial
You train a model. Cross-validation gives you a confident accuracy score. You deploy. The model underperforms in ways your validation never hinted at.
Read tutorial
You fit the model, plot a tidy dendrogram, cut it at the biggest gap, and report three clusters. Then you change one argument — the linkage — and get a…
Read tutorial
Change one parameter, get a different tree. Same points, same distance metric, same greedy loop — yet the dendrogram reorganizes itself. That parameter is…
Read tutorial
Most beginners ask the wrong question first. They ask, "Which clustering algorithm is better?" The real question is, "Which question am I actually trying…
Read tutorial
Your best tuning score is not your model's true performance. It is the score of the luckiest experiment you ran.
Read tutorial
You train a model, split your data randomly, and the test score looks great. Then the model meets real-world data and quietly falls apart. The usual…
Read tutorial
A model that predicts "no fraud" for every transaction can score 98% accuracy while catching zero fraud. The number feels trustworthy. It is not. Accuracy…
Read tutorial
You load a saved model, feed it a fresh batch, and get predictions back. No error. No warning. Just numbers that look perfectly reasonable—except a column…
Read tutorial
"The model says education matters" sounds like a complete explanation. It is not. It is the beginning of a much narrower claim: the model leaned on…
Read tutorial
You fit an Isolation Forest, call decision_function, and get a column of numbers. You sort them, pick a cutoff, and ship the alerts. But if someone asks…
Read tutorial
Fit K-means twice on the same data with the same k, and you can get two different clusterings. The second run often reports a lower cost than the first.…
Read tutorial
Most people stop at fit(). They get a label array, scatter-plot it, and call the job done. But the label array is not the result — it is the beginning of…
Read tutorial
Fit KNN regression with k=1 and the prediction curve threads through every training point. Set k=50 and the same curve flattens toward a horizontal line.…
Read tutorial
Two models trained on the same data can look at a new point and give you sharply different answers. Neither one is broken. Each is answering a different…
Read tutorial
Your model passed validation. Then it went to production and started predicting one class far more often than it should. The features look normal. The…
Read tutorial
Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero.…
Read tutorial
If you are like most beginners, you start guessing. Maybe more data will fix it. Maybe a bigger model. Maybe fancier features. You try one, see little…
Read tutorial
You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a…
Read tutorial
You fit LinearDiscriminantAnalysis, plot the boundary, and get a straight line. Then someone asks why it is straight, and the honest answer is "because the…
Read tutorial
A point can sit closer to its neighbors than anything else in the dataset and still be the anomaly. That is the contradiction Local Outlier Factor was…
Read tutorial
A logistic regression prediction is not a label. It is a probability—and the label is a decision you make on top of that probability.
Read tutorial
Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The…
Read tutorial
You run a grid search. The leaderboard prints a winner at 0.91. You ship it, and the first honest measurement lands near 0.86. Nothing broke. The model did…
Read tutorial
A multiclass model hands you K raw scores. The obvious move — divide each by their sum — collapses the first time a score goes negative. Here is the fix,…
Read tutorial
Most classification tutorials teach you to pick one answer. A news article is politics or finance. A movie is action or comedy. A support ticket is billing…
Read tutorial
The intuition article told you Naive Bayes "combines prior and evidence." Then you tried to write $P(x_1, x_2, \dots, x_n \mid y)$ for real data, and the…
Read tutorial
You run a grid search, watch the cross-validation score climb, and feel good about the model you are about to ship. Then the model lands on real data and…
Read tutorial
You tuned a model, watched the cross-validation score climb, and reported the best number. Then new data arrived and the model performed worse than…
Read tutorial
You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…
Read tutorial
A rating scale looks like it should be either a multiclass problem or a regression problem. Both instincts quietly discard the information that matters…
Read tutorial
Most people assume PCA "removes unimportant features." It does not. It rotates your data into new directions ranked by variance, then drops the quietest…
Read tutorial
You fit PCA, print the explained variance ratio, and see that PC1 captures 92% of the variance. That looks like a strong result. It might be an artifact of…
Read tutorial
You can call PCA(), read explained_variance_ratio_, and get a working result. Then someone asks why the principal directions are eigenvectors of the…
Read tutorial
You have a wide table of features, a model to build, and a quiet suspicion that half those columns are noise. The instinct is to shrink the dataset. But…
Read tutorial
A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.
Read tutorial
A count target is not a permission slip. It is a hypothesis about how the mean and variance move together — and you can test it.
Read tutorial
A model can keep the same weights, the same features, and the same threshold, and still report precision 0.9 on your test set and precision 0.3 in…
Read tutorial
A regression model hands you one number. The person acting on that number needs to know how much weight it can bear. A prediction interval is how you show…
Read tutorial
A model can rank every case perfectly and still hand you probabilities you cannot spend. Calibration is the difference between a score that orders risk and…
Read tutorial
Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.
Read tutorial
oob_score=True looks like a free validation set. It is not free, and it is not quite a validation set. It is a running estimate computed on the same rows…
Read tutorial
A single decision tree can feel like a confident liar: it nails your training data, then wobbles on new data, and reshapes its entire structure when one…
Read tutorial
"Extra Trees" does not mean more trees. It means extra randomness — and that single distinction changes how the ensemble learns, how fast it trains, and…
Read tutorial
A respectable R² can sit on top of a model that is wrong in a structured way. The score summarizes average fit; it says nothing about where the fit fails.…
Read tutorial
A classification tree asks which split makes the labels purest. A regression tree has no labels to purify — only numbers that scatter. So what is it…
Read tutorial
Regularization is not a magic switch that makes every model better. It is a deliberate trade: you give up a little accuracy on your training data to gain…
Read tutorial
You tune a model, retrain, and the score moves. The question is whether that movement came from your edit or from luck. Most beginners cannot tell the…
Read tutorial
Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move.…
Read tutorial
You resample your data to fix the imbalance, rerun the same model, and watch ROC-AUC barely move while PR-AUC collapses. Nothing about the model changed.…
Read tutorial
You train a fraud detector on a dataset where only 0.5% of transactions are fraudulent. The ROC AUC comes back at 0.95. You feel great. Then you look at…
Read tutorial
Your model reports 0.91 accuracy. That number is probably correct, and it is probably useless for deciding what to fix next.
Read tutorial
A model can rank every sample correctly and still lie about the number attached to each one. This experiment isolates that lie, then measures whether a…
Read tutorial
One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with…
Read tutorial
Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing…
Read tutorial
Most explanations of support vector machines stop at the picture: two clouds of points, a line between them, a gap on either side. That picture is correct,…
Read tutorial
Same data. Same SVC call. One run scores 0.85, the next 0.55, and the only thing that changed was whether you scaled the features first.
Read tutorial
Most machine learning models are hoarders. They keep every training point around and let each one vote on the prediction. A support vector machine is the…
Read tutorial
Target encoding is one of the most seductive traps in feature engineering. You encode a high-cardinality feature like city or merchant_id, watch your…
Read tutorial
Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…
Read tutorial
You run PCA on your data and get one picture. You run t-SNE on the same data and get a completely different one. One shows two messy blobs. The other shows…
Read tutorial
A missing column throws. A reordered DataFrame does not. That asymmetry is where confident wrong numbers come from.
Read tutorial
More models do not automatically mean a better model. An ensemble only wins when its members make different mistakes—and the way you combine them…
Read tutorial
XGBoost has a reputation problem. It gets treated like a secret weapon—a mysterious algorithm that wins competitions and powers production systems while…
Read tutorial
Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by…
Read tutorial