
A Practical Clustering Workflow: From Unlabeled Data to Defensible Structure
Clustering always returns groups. Even on random noise, even on data with no structure worth naming, every algorithm will happily partition your points and…
Read tutorialUnsupervised methods for grouping observations and assessing whether discovered structure is stable, useful, or meaningful.
Tagged articles
15 articles in this tag.

Clustering always returns groups. Even on random noise, even on data with no structure worth naming, every algorithm will happily partition your points and…
Read tutorial
K-means does not discover the real categories hiding in your data. It draws geometric boundaries around points that happen to sit close together. Those…
Read tutorial
Your clusters looked clean until you changed the random seed. Now the labels shuffled, the boundaries moved, and you are not sure whether you found…
Read tutorial
Every clustering algorithm will happily partition pure noise into tidy groups. Run k-means on random data, and it returns clean circles of assigned points.…
Read tutorial
You learned K-means, and it felt clean: pick a number of groups, let centroids pull points inward, and read off the labels. Then you hit a dataset with…
Read tutorial
Two points sit the same distance from a dense blob. DBSCAN labels one a cluster member and the other noise. That looks arbitrary until you stop asking how…
Read tutorial
Change one number in a DBSCAN call and the whole story of your data can flip: two clusters become one, or a clean grouping dissolves into a field of noise.…
Read tutorial
You can write the Gaussian mixture density in a single line. The trouble starts the moment you take the log.
Read tutorial
Picture a point sitting exactly between two well-separated blobs of data. K-means will force it into one group or the other, even when the data could…
Read tutorial
A hard label is a decision. A posterior probability is a confession about how close that decision was.
Read tutorial
You fit the model, plot a tidy dendrogram, cut it at the biggest gap, and report three clusters. Then you change one argument — the linkage — and get a…
Read tutorial
Change one parameter, get a different tree. Same points, same distance metric, same greedy loop — yet the dendrogram reorganizes itself. That parameter is…
Read tutorial
Most beginners ask the wrong question first. They ask, "Which clustering algorithm is better?" The real question is, "Which question am I actually trying…
Read tutorial
Fit K-means twice on the same data with the same k, and you can get two different clusterings. The second run often reports a lower cost than the first.…
Read tutorial
Most people stop at fit(). They get a label array, scatter-plot it, and call the job done. But the label array is not the result — it is the beginning of…
Read tutorial