Machine Learning Interview Questions — Distilled From 7 Tutorials
This is the interview distillation of the 7-part AI/ML tutorial track: the questions ML interviewers actually ask, compressed to the answers that score. If you've done the tutorials, this is your revision sheet; if not, each answer tells you exactly what to learn next.
1. What are the assumptions of linear regression — and which one breaks most?
Answer: linearity, independent errors, homoscedasticity (constant error variance), no severe multicollinearity, roughly normal residuals. The one that breaks most in practice: multicollinearity — correlated features make coefficients unstable and uninterpretable. Diagnose with VIF; fix by dropping/combining features or using regularization.
2. Precision vs recall vs F1 — when do you optimize each?
from sklearn.metrics import precision_score, recall_score, f1_score # cancer screening: missing a case is catastrophic → recall # spam filter: false alarms annoy → precision # need one number: F1 = 2PR/(P+R)
Answer: precision = of my positive predictions, how many were right; recall = of the real positives, how many did I catch. Optimize recall when missing positives is costly (fraud, disease); precision when false alarms are costly (spam). F1 balances both. And never report accuracy on imbalanced data — 99% accuracy on 1% fraud is a useless model.
3. What is ROC-AUC really telling you?
Answer: the probability that a randomly chosen positive ranks above a randomly chosen negative — a threshold-independent measure of ranking quality. 0.5 = coin flip, 1.0 = perfect separation. The trap: AUC can look great on imbalanced data while precision is terrible — pair it with a precision-recall curve when positives are rare.
4. Your model scores 99% on train, 70% on test. Diagnose and fix.
Answer: textbook overfitting — high variance. Fixes, in order: more training data; simpler model; regularization (L1/Lasso zeroes features, L2/Ridge shrinks them); cross-validation for honest estimates; early stopping; dropout for neural nets. The follow-up: "And if both train and test are bad?" — underfitting: richer model, better features, less regularization.
5. L1 vs L2 regularization — when does it matter?
Answer: L1 (Lasso) drives coefficients to exactly zero → automatic feature selection. L2 (Ridge) shrinks everything smoothly → better when all features matter a bit. ElasticNet mixes both. Interview one-liner: "L1 selects, L2 shrinks."
6. How do you detect data leakage?
Answer: leakage = the model sees information at train time it won't have at prediction time (a classic: including a "days_to_churn" column when predicting churn). Symptoms: suspiciously perfect validation scores. Prevention: split before any preprocessing, fit scalers/imputers on train only, and think causally — "would I know this value at prediction time?"
7. Train/test split vs cross-validation?
Answer: a single split is fast but noisy — your score depends on the lucky split. K-fold cross-validation trains K times on different folds and averages: a stabler estimate, at K× the cost. Use CV for model selection and small datasets; a single holdout is fine for final evaluation on big data. And the test set is touched once — tuning on it is just overfitting with extra steps.
8. How do you handle imbalanced classes?
- Resampling: SMOTE/oversample the minority, or undersample the majority.
- Class weights:
class_weight='balanced'— penalize minority errors more. - Threshold tuning: 0.5 is a default, not a law — move it along the precision-recall curve.
- Right metric: PR-AUC / F1, never accuracy.
Say all four and you've covered the standard answer completely.
9. Feature engineering vs more data vs better model?
Answer: the experienced ordering: more/better data first (garbage in, garbage out), then features (domain-informed features beat algorithm swaps), then model tuning last. The interview signal is humility about modeling: "I'd spend the week on the data pipeline, not the hyperparameters." Concrete wins: target encoding for high-cardinality categoricals, interaction features the domain suggests, and ruthless removal of leaky features.
10. Walk me through your last ML project end to end
Answer template: problem framing (what business metric moved?) → data (sources, size, quality issues found) → iteration (what you tried, what the validation said) → evaluation (metric matched to the business cost) → deployment reality (latency, monitoring, retraining). Interviewers score the story arc — decisions and trade-offs — more than the final accuracy number.
In this series
- Top 50 Python DSA Problems by Pattern (With Solutions) — two pointers to intervals.
- Data Science Interview Questions: Statistics, pandas, SQL — the three muscles.
- Machine Learning Interview Questions — Distilled From 7 Tutorials (this post).
Full Python roadmap: Python Learning Roadmap 2026 — all 50 tutorials across 8 tracks, including the 7-part AI/ML track this post distills.
Comments
Post a Comment