Overview
Curated: · Written: · Reviewed:
Key takeaways
- Supervised learning estimates a relationship from inputs to a defined target using labeled examples; classification predicts categories or probabilities, while regression predicts continuous quantities.
- Unsupervised learning finds structure without a target supplied for each training example. Clusters, components, embeddings, densities, and anomalies are model outputs whose business meaning must be validated separately.
- The decisive question is not which family sounds advanced, but what decision, feedback, label, and intervention exist at prediction time.
- Evaluation must reproduce deployment boundaries. Split by time, entity, group, geography, or experiment assignment when random rows would leak future or related information.
- Preprocessing, feature selection, dimensionality reduction, clustering, and pseudo-labeling must be fit inside the training portion of every evaluation fold.
- Semi-supervised and self-supervised approaches can use unlabeled data, but noisy assumptions and pseudo-label errors can amplify bias rather than create free ground truth.
1. Start from the decision and target
Supervised learning needs an observable outcome connected to the intended decision: fraud confirmed after investigation, delivery time, churn within a declared horizon, or defect category. Define prediction time, label window, eligible population, unit of analysis, action, error costs, and when truth becomes available. A convenient historical column is not necessarily a valid target; it may encode a past policy, human bias, post-outcome information, or a proxy that changes after deployment.
Classification outputs a class, score, or probability. The operating threshold belongs to the product decision and should be chosen using false-positive/false-negative costs, capacity, calibration, and subgroup behavior—not defaulted to 0.5. Regression outputs a continuous value or distribution; choose error metrics that match units and asymmetric consequences. Always compare against simple rules and dummy baselines before crediting model complexity.
2. What unsupervised learning does—and does not do
Clustering groups examples according to a chosen representation, distance or affinity, and algorithmic objective. K-means minimizes within-cluster squared distance and favors roughly convex, similar-variance geometry; density, hierarchical, graph, or mixture approaches encode different assumptions. A cluster ID is not automatically a customer segment, diagnosis, risk level, or causal group. Test stability across samples and seeds, inspect exemplars, obtain domain interpretation, and measure whether the downstream use improves an outcome.
Dimensionality reduction compresses or transforms features for visualization, denoising, speed, or modeling. PCA finds linear directions of variance after centering, but does not scale features automatically and high variance need not be predictive or causal. Nonlinear visualization can distort global distance and density. Fit transformations only on training data and validate the complete downstream pipeline on held-out deployment-like data.
Anomaly and novelty detection rank observations that differ under a representation and reference distribution. Rare is not the same as harmful, and known labeled fraud may make supervised ranking more appropriate. Novelty detection assumes training data largely represents normal behavior; contaminated training sets and changing seasonality can invalidate the score. Set review thresholds using operational capacity and measured precision, and maintain a feedback path.
3. Evaluation and leakage
Create a final untouched test set representing the future population and use training/validation or nested cross-validation for selection. Choose splits from the data-generating process: time splits for forecasting and policy drift, group splits when one customer/device/patient creates several rows, and spatial or site splits when local similarity would inflate performance. Random row splitting can put near duplicates or the same entity on both sides.
Every learned transform belongs inside the fold: imputation, scaling, vocabulary, target encoding, feature selection, PCA, sampling, cluster assignment, and pseudo-label generation. Fitting them once on all data leaks validation distribution into the model. Tune hyperparameters only on validation evidence and report uncertainty across folds or time windows; repeated test-set inspection turns the test into another training signal.
Supervised metrics require trustworthy labels and match the operating decision. For imbalance, accuracy can hide failure; use precision-recall, class-specific recall, calibration, cost, capacity, and threshold curves as appropriate. Unsupervised internal metrics measure an algorithmic shape, not business value. When external labels exist, keep them out of fitting but use them for evaluation; otherwise combine stability, domain review, downstream experiments, and qualitative failure analysis.
4. Limited labels and hybrid strategies
Semi-supervised learning combines a small labeled set with unlabeled examples under assumptions such as local smoothness or cluster consistency. Self-training accepts high-confidence model predictions as pseudo-labels; label propagation spreads label information through a similarity graph. Both can reinforce early mistakes or majority bias. Calibrate confidence, isolate a human-labeled evaluation set, control class balance, audit subgroup error, and compare against spending the same effort on more representative labels.
Self-supervised representation learning constructs a training signal from the data itself, then adapts the representation to a supervised or unsupervised downstream task. Active learning asks humans to label selected informative cases. Weak supervision combines noisy programmatic or heuristic label sources. These are label-acquisition and representation strategies, not permission to skip target definition, provenance, consent, leakage controls, or deployment evaluation.
5. Production feedback and monitoring
Log feature/representation version, model version, score, threshold, action, and eventual outcome with privacy controls. Monitor input coverage, missingness, representation and score distributions, cluster sizes/stability, anomaly-review yield, calibration, label delay, and decision outcomes. Changes may come from population drift, instrumentation, policy, adversaries, seasonality, or the model's own intervention.
Feedback is selective: only flagged cases may be reviewed, and treatment changes the observed outcome. Do not naively retrain on this biased stream. Preserve exploration or randomized audit samples where safe, track propensity and label source, delay training until labels mature, and investigate silent populations. Retraining is a governed release with leakage-safe evaluation, fairness and capacity checks, canarying, rollback, and post-deployment verification.
6. Worked example: the same 50,000 customers, two different questions
Given one table of 50,000 customers, a supervised and an unsupervised framing answer different questions and fail in different ways.
Supervised: predict churn in the next 30 days, trained on 50,000 rows where 4,100 churned (8.2% prevalence). The label is the definition of the task, and it is auditable — precision at the top decile is 0.34 against a 0.082 base rate, a 4.1x lift, and that number is checkable against what actually happened.
Unsupervised: segment the same customers with k-means on 14 standardized behavioral features. There is no label, so nothing is right or wrong; there is only whether the structure is stable and useful:
| k | Inertia | Silhouette | Smallest cluster | Churn rate spread across clusters |
|---|---|---|---|---|
| 3 | 412,000 | 0.31 | 9,800 | 6.1% - 11.4% |
| 5 | 318,000 | 0.28 | 2,300 | 3.2% - 19.7% |
| 8 | 265,000 | 0.19 | 410 | 1.1% - 24.0% |
Inertia falls monotonically with k and says nothing about the right answer. Silhouette peaks at k=3, but k=5 separates churn far more sharply (3.2% to 19.7%) while keeping every cluster above 2,000 members. At k=8 the smallest cluster holds 410 customers and will not survive a re-run on next month's data.
The error to avoid is treating cluster 4 as a class:
labels = KMeans(n_clusters=5, random_state=0).fit_predict(X)
# Wrong: these ids are arbitrary. Re-running with a different seed permutes them,
# so "cluster 4 is the at-risk segment" is not a statement that survives a refit.
df["segment"] = labels
# What the labels are actually good for: a hypothesis to test with a real outcome.
df.groupby(labels)["churned_30d"].mean() # 0.032, 0.058, 0.081, 0.121, 0.197
The churn rates on the last line are supervised evidence about an unsupervised grouping. That combination — unsupervised structure, supervised validation — is usually stronger than either framing alone, and it is the one an interviewer is listening for.
