
Classical Anomaly Detection: Find Unusual Data Without Labels
Anomaly detection is not "find the weird rows." It is a decision about what counts as normal, made before you ever run an algorithm. Get that decision…
Read tutorialMethods that investigate structure in data without using a target label as the learning outcome.
Tagged articles
17 articles in this tag.

Anomaly detection is not "find the weird rows." It is a decision about what counts as normal, made before you ever run an algorithm. Get that decision…
Read tutorial
Clustering always returns groups. Even on random noise, even on data with no structure worth naming, every algorithm will happily partition your points and…
Read tutorial
K-means does not discover the real categories hiding in your data. It draws geometric boundaries around points that happen to sit close together. Those…
Read tutorial
Your clusters looked clean until you changed the random seed. Now the labels shuffled, the boundaries moved, and you are not sure whether you found…
Read tutorial
Every clustering algorithm will happily partition pure noise into tidy groups. Run k-means on random data, and it returns clean circles of assigned points.…
Read tutorial
Two detectors, one dataset, two different alert lists — and no labels to referee the disagreement.
Read tutorial
You learned K-means, and it felt clean: pick a number of groups, let centroids pull points inward, and read off the labels. Then you hit a dataset with…
Read tutorial
Two points sit the same distance from a dense blob. DBSCAN labels one a cluster member and the other noise. That looks arbitrary until you stop asking how…
Read tutorial
You can write the Gaussian mixture density in a single line. The trouble starts the moment you take the log.
Read tutorial
Picture a point sitting exactly between two well-separated blobs of data. K-means will force it into one group or the other, even when the data could…
Read tutorial
A hard label is a decision. A posterior probability is a confession about how close that decision was.
Read tutorial
Most beginners ask the wrong question first. They ask, "Which clustering algorithm is better?" The real question is, "Which question am I actually trying…
Read tutorial
You fit an Isolation Forest, call decision_function, and get a column of numbers. You sort them, pick a cutoff, and ship the alerts. But if someone asks…
Read tutorial
Fit K-means twice on the same data with the same k, and you can get two different clusterings. The second run often reports a lower cost than the first.…
Read tutorial
Most people stop at fit(). They get a label array, scatter-plot it, and call the job done. But the label array is not the result — it is the beginning of…
Read tutorial
You have heard both terms. You can even recite rough definitions. Then you open your own dataset, stare at the columns, and freeze. Is this supervised or…
Read tutorial
You run PCA on your data and get one picture. You run t-SNE on the same data and get a completely different one. One shows two messy blobs. The other shows…
Read tutorial