Model Evaluation Done Right: Train-Test Split and Cross-Validation
Part 4 of the Python for AI/ML track. Last updated: September 2026.
A model that scores 100% on its training data can still be worthless. Evaluation — measuring performance honestly on data the model never saw — is what separates ML from numerology. Get this wrong and you ship a model that fails on day one.
The trap: overfitting
An unconstrained decision tree memorizes the training set — every quirk, every noise point. Watch the gap between train and test scores:
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
iris = load_iris()
X_train, X_test, y_train, y_test = train_test_split(
iris.data, iris.target, test_size=0.3, random_state=1)
deep = DecisionTreeClassifier(random_state=1).fit(X_train, y_train) # no limits
shallow = DecisionTreeClassifier(max_depth=2, random_state=1).fit(X_train, y_train)
print(f"Deep tree: train={deep.score(X_train, y_train):.3f} test={deep.score(X_test, y_test):.3f}")
# Deep tree: train=1.000 test=0.956 ← memorized, then stumbled
print(f"Shallow tree: train={shallow.score(X_train, y_train):.3f} test={shallow.score(X_test, y_test):.3f}")
# Shallow tree: train=0.962 test=0.956 ← honest, generalizes
The deep tree's perfect training score is a lie — it learned the noise. The shallow tree scores lower on training but the same on unseen data. The test score is the only score that counts.
train_test_split done properly
- random_state — fixes the shuffle so your results are reproducible. Always set it.
- test_size=0.2–0.3 — hold out 20–30% for the final verdict.
- stratify=y for classification — keeps class proportions identical in both splits, critical for small or imbalanced datasets.
- Never touch the test set until the final evaluation — no tuning on it, no peeking.
X_train, X_test, y_train, y_test = train_test_split(
iris.data, iris.target, test_size=0.2, random_state=42, stratify=iris.target)
Cross-validation: one split is noisy
A single train/test split can be lucky or unlucky. Cross-validation rotates the test fold — with cv=5, the data is split into 5 parts, the model trains on 4 and tests on the 5th, five times:
from sklearn.model_selection import cross_val_score
model = DecisionTreeClassifier(random_state=1)
scores = cross_val_score(model, iris.data, iris.target, cv=5)
print("Fold scores:", [f"{s:.3f}" for s in scores])
# Fold scores: ['0.967', '0.967', '0.900', '1.000', '1.000']
print(f"Mean: {scores.mean():.4f}") # Mean: 0.9667
Report the mean ± std across folds — "96.7% ± 3.6%" is an honest claim; a single 100% fold is not. Use cross-validation to compare models and tune settings, then confirm the winner once on the held-out test set.
Picking the right metric for the job
- Regression: MAE when errors cost linearly, RMSE when big errors are expensive, R² for a quick "variance explained" summary.
- Balanced classification: accuracy is fine.
- Imbalanced classification (fraud, disease): accuracy lies — optimize precision/recall/F1 or ROC-AUC instead.
- Ranking/recommendation: precision@k, NDCG — how good are the top results?
Rule of thumb: the metric is the business goal translated into math. If false alarms cost support tickets, optimize precision. If misses cost lives, optimize recall. Say the tradeoff out loud before you train.
Key takeaways
- Overfitting = memorizing training noise; the train/test gap exposes it.
- Split with random_state for reproducibility and stratify for classification.
- Cross-validation (cv=5) gives an honest mean ± std — use it to compare models.
- Never tune on the test set; it gets one final look, at the end.
- Choose the metric from the business cost of errors, not habit.
Next in this series: Feature Engineering Essentials for Machine Learning.
Comments
Post a Comment