Unsupervised Learning
Unsupervised learning finds structure in data that has NO answer key. It groups similar things (clustering), compresses many features into a few (dimensionality reduction), flags the weird ones (anomaly detection), and produces embeddings. The hard part is proving a result is good when there is no label to check against.
Unsupervised Method Families 🧩
Without labels, the objective is intrinsic structure: group, compress, flag, or associate. Evaluation shifts from accuracy to internal coherence and downstream usefulness.
01.The Problem: Structure Without an Answer Key
Supervised learning (topic 40) always hands you the right answer for each example. What if there is no right answer?
Imagine dumping 10,000 songs into a room with no genre tags. Nobody told you which are jazz and which are metal.
Can you still sort them into piles that "feel" alike — and spot the one song that belongs to no pile?
That is unsupervised learning: find structure using only the inputs, with no target labels.
Formally you get only inputs x₁ … xₙ with no y. Instead of minimizing prediction error, the algorithm optimizes an objective defined inside the data itself:
- Similarity — points in the same group should be close together (clustering).
- Reconstruction / variance — a low-dimensional representation should preserve as much of the covariance as possible (PCA, autoencoders).
- Density — high-probability regions describe normal behavior; low-density regions are anomalies.
- Co-occurrence — items appearing together imply association rules.
02.The Trade That Cuts Both Ways
No answer key is a superpower and a curse at the same time.
The good news
You can now use every unlabeled record you own — usually 10 to 1000 times more data than a supervised pipeline can afford to label.
And you can discover segments nobody thought to ask about.
The bad news
There is no ground truth to say the result is correct.
So quality gets judged three ways instead:
- Internal coherence — are the piles actually tight and well-separated?
- Stability — do you get the same piles if you restart the algorithm?
- Downstream usefulness — the honest standard: does it improve a real decision?
This one sentence changes everything. In supervised learning you ask "what is my error?" In unsupervised learning you ask "what structure do I believe exists, and can I defend it?"
03.A Tiny Worked Example: k-Means in 4 Points
K-Means is the "round-robin grouping" algorithm. You say how many piles you want (that number is k), and it finds them.
Watch it on points plotted along a line, with k = 2:
codedata: 1 2 3 10 11 12 step 1: guess centers at 5 and 8 → assign each point to nearest center pile A = {1,2,3} pile B = {10,11,12} step 2: move each center to its pile mean → centers become 2 and 11 step 3: reassign — nothing moves. STOP.
Two clean piles, and the centers (2 and 11) barely move on the next pass.
Each loop does two jobs, forever until stable:
- Assignment: every point joins its nearest center.
- Update: every center moves to the mean of its members.
The score it drives down is inertia — the within-cluster sum of squares (each point's squared distance to its center, added up).
In one line
K-means"assign to nearest center, recenter, repeat" until nothing moves.
04.Visual Intuition: The Shape the Algorithm Believes
Every clusterer assumes a shape. Pick the wrong one and you get nonsense.
codeK-Means believes in round blobs: DBSCAN finds ANY shape + noise: ○ ○ ● = noise ◜◝ the thin arc is ONE (tight spherical piles, roughly ◟◞ cluster; far-off equal size, convex) . . dots are noise ● . . ● scattered evenly . .
- K-Means (Lloyd's algorithm) minimizes within-cluster sum of squares (inertia). It is
O(n·k·d·iterations), scales to hundreds of millions of points via MiniBatchKMeans or FAISS, and is the default for coarse segmentation. Its assumptions are strict: clusters are roughly equal-sized, isotropic (spherical), and convex. It is sensitive to feature scaling, so standardize first.kmust be supplied; pick it with the elbow on inertia, the silhouette peak, or stability across seeds rather than intuition. - Gaussian Mixture Models replace hard assignment with soft responsibilities and fit ellipsoidal clusters with covariance, yielding per-point membership probabilities and a likelihood you can compare across
kwith BIC/AIC. - DBSCAN / HDBSCAN are density-based: clusters are contiguous high-density regions separated by low-density gaps. They find arbitrary shapes, need no
k, and label sparse points as noise — exactly what you want for geospatial, log, and event data. Classical DBSCAN degrades in high dimensions and is sensitive to eps, which is why HDBSCAN (with a minimum-cluster-size parameter only) is the practical choice. - Agglomerative / hierarchical clustering builds a merge dendrogram, letting you cut at multiple granularities — useful when the natural taxonomy depth is unknown.
05.The Analogy: Sorting a Box of Mixed Buttons
Carry one picture through the whole topic: dump a box of mixed buttons on the table and sort them with no names.
- Clustering is making piles by look-and-feel — big red ones here, small blue there. You invented the groups; nobody gave you the categories.
- Outlier / anomaly detection is holding up the one square, glowing button and saying "that one fits nowhere."
- Dimensionality reduction (PCA) is noticing you can describe most buttons with just two numbers — size and darkness — and ignore the other ten features without losing the pattern.
Now the hard part of the analogy.
A friend sorts the same box differently. Who is right?
There is no answer key. You check three things instead: are the piles tight, would you sort the same way tomorrow (stability), and do the piles actually help you do something (pick the right thread for each button). Unsupervised results earn trust the same way — internal coherence plus downstream usefulness — never by matching a label, because no label exists.
06.Dimensionality Reduction and Anomalies in Practice
PCA finds orthogonal directions of maximum variance via eigendecomposition of the covariance matrix (equivalently SVD of the centered data). Projecting onto the top d components removes collinearity, denoises, and makes distance metrics meaningful again — the standard cure for the curse of dimensionality before clustering. Two practical caveats: PCA is linear, so it cannot unroll manifolds; and components are only interpretable through their loadings, so name them only after inspection. In predictive pipelines, fit PCA on the training fold only, then transform validation data.
t-SNE preserves local neighborhoods for visualization by converting distances into conditional probabilities and minimizing KL divergence between the high- and low-dimensional graphs. Perplexity sets the effective neighborhood size. Distances and cluster sizes between blobs in a t-SNE plot are not meaningful, and re-running with a different seed or perplexity can reshape the picture — treat it as a hypothesis generator, never as evidence.
UMAP builds a fuzzy topological simplicial complex and optimizes its low-dimensional layout. It is much faster than t-SNE, handles tens of millions of points, and supports transform() for new data, which makes it usable for production embedding maps and exploratory analytics dashboards.
Anomaly detection assumes normality dominates: learn what "typical" looks like, then flag what deviates. Options — Isolation Forest isolates points with random splits (anomalies need fewer splits, giving an O(n·t) score that works well in moderate dimensions); Mahalanobis distance and elliptic envelope flag points outside a Gaussian shell; KDE / density methods flag low-likelihood regions; autoencoder reconstruction error flags unusual sequences. Compare all of these against the supervised alternative (topic 40) when a handful of confirmed incidents exist.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
import numpy as np
Xs = StandardScaler().fit_transform(X)
results = {}
for k in range(2, 11):
km = KMeans(n_clusters=k, n_init=10, random_state=0).fit(Xs)
results[k] = silhouette_score(Xs, km.labels_) # higher = better separated
best_k = max(results, key=results.get)
print(f"k={best_k} silhouette={results[best_k]:.3f}")
# Silhouette on 1M+ rows: score a random sample, not the full matrix07.How to Trust a Result With No Label
Validation is the hard part, and it has a hierarchy.
- Internal consistency — silhouette, Davies-Bouldin, BIC, stability across seeds/bootstrap resamples.
- Held-out structure — cluster on 70%, assign the remaining 30%, measure assignment stability.
- Label correlation — check whether discovered segments separate a known outcome (churn, revenue, LTV) they were never trained on. This is the strongest cheap evidence.
- Downstream A/B — if segmentation drives a campaign, lift versus the old segmentation is the verdict.
Always test against a null. Run the same algorithm on shuffled or synthetic uniform data. If your "clusters" appear there too, they are artifacts of the metric, not the data.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Uses all available data with zero labeling cost — usually 10-1000x more rows than supervised pipelines.
- Reveals segments, drivers, and anomalies that no stakeholder had a name for.
- Embeddings and reduced representations improve every downstream model and cut storage/compute.
- Density and isolation methods detect novel failure modes without waiting for labeled incidents.
Trade-offs & Constraints
- No ground truth: results are interpretable but not verifiable, so disagreements are unresolvable analytically.
- Sensitive to scaling, distance metric, and hyperparameters (k, eps, perplexity) that quietly change conclusions.
- Cluster labels are not stable across retrains, which breaks naive production contracts.
- High-dimensional Euclidean distances concentrate, making results meaningless without dimensionality reduction.
Audio encoders produce fixed-size embeddings for ~43M tracks; self-supervised and metadata objectives train the representation, then clustering over those vectors yields compact, human-inspectable style groups used for radio generation, similarity search, and cold-start recommendation — no genre labels required.
Staff+ Engineering Takeaways
- Unsupervised learning optimizes intrinsic structure — similarity, variance preservation, density, or co-occurrence — instead of prediction error.
- K-means assumes spherical, equal-scale clusters; GMMs add soft assignments; DBSCAN/HDBSCAN find arbitrary density shapes and noise.
- PCA compresses linearly by variance; t-SNE and UMAP are neighborhood-preserving visualization tools, with UMAP faster and incremental.
- Feature scaling and distance choice are not tuning details in clustering — they change which structure you find.
- Credibility comes from internal metrics, assignment stability, correlation with untouched known outcomes, and downstream A/B lift.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Customer clusters must be recomputed monthly, and the product team notices cluster "3" means something different every run. The most robust fix is:
How clear and actionable was this distributed systems breakdown?