Skip to content
Tech Interview Prep home
Technical interview guide

Supervised vs. Unsupervised Learning

Learning from labeled examples versus finding structure in unlabeled data — and where semi-supervised and reinforcement learning fit.

Read
25 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: scikit-learn stable documentation and Google Machine Learning Crash Course accessed 2026-08-30..

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Compare supervised and unsupervised learning using a product example.

QA-2

How do you define a supervised-learning target safely?

QA-3

Design an evaluation split for repeat customers over time.

QA-4

How do you prevent preprocessing leakage in an ML pipeline?

QA-5

How do classification and regression differ in outputs and evaluation?

QA-6

How would you evaluate a rare-event classifier?

QA-7

How should a production classification threshold be chosen?

QA-8

How do you decide whether clustering is appropriate?

QA-9

Explain k-means assumptions and common failure modes.

QA-10

How do you validate customer segments produced by clustering?

QA-11

Explain PCA and when it is useful.

QA-12

Design an anomaly-detection system when labels are scarce.

QA-13

How would you use semi-supervised learning safely?

QA-14

Compare self-training and label propagation.

QA-15

How do self-supervised, semi-supervised, active, and weak supervision differ?

QA-16

How do you choose between collecting labels and using an unsupervised proxy?

QA-17

How do you evaluate dimensionality reduction for visualization?

QA-18

How do you choose the number of clusters?

QA-19

How should label quality be measured and improved?

QA-20

How do you detect and prevent selective-label feedback loops?

QA-21

Design monitoring for a supervised model with delayed labels.

QA-22

Design monitoring for a deployed clustering system.

QA-23

How would you decide when to retrain an ML model?

QA-24

How do you choose between supervised anomaly classification and unsupervised anomaly detection?

QA-25

How would you present an unsupervised finding to stakeholders?