Model Evaluation Done Right: Train-Test Split and Cross-Validation

Part 4 of the Python for AI/ML track. Last updated: September 2026.

A model that scores 100% on its training data can still be worthless. Evaluation — measuring performance honestly on data the model never saw — is what separates ML from numerology. Get this wrong and you ship a model that fails on day one.

The trap: overfitting

An unconstrained decision tree memorizes the training set — every quirk, every noise point. Watch the gap between train and test scores:

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier

iris = load_iris()
X_train, X_test, y_train, y_test = train_test_split(
    iris.data, iris.target, test_size=0.3, random_state=1)

deep = DecisionTreeClassifier(random_state=1).fit(X_train, y_train)          # no limits
shallow = DecisionTreeClassifier(max_depth=2, random_state=1).fit(X_train, y_train)

print(f"Deep tree:    train={deep.score(X_train, y_train):.3f}  test={deep.score(X_test, y_test):.3f}")
# Deep tree:    train=1.000  test=0.956   ← memorized, then stumbled
print(f"Shallow tree: train={shallow.score(X_train, y_train):.3f}  test={shallow.score(X_test, y_test):.3f}")
# Shallow tree: train=0.962  test=0.956   ← honest, generalizes

The deep tree's perfect training score is a lie — it learned the noise. The shallow tree scores lower on training but the same on unseen data. The test score is the only score that counts.

train_test_split done properly

  • random_state — fixes the shuffle so your results are reproducible. Always set it.
  • test_size=0.2–0.3 — hold out 20–30% for the final verdict.
  • stratify=y for classification — keeps class proportions identical in both splits, critical for small or imbalanced datasets.
  • Never touch the test set until the final evaluation — no tuning on it, no peeking.
X_train, X_test, y_train, y_test = train_test_split(
    iris.data, iris.target, test_size=0.2, random_state=42, stratify=iris.target)

Cross-validation: one split is noisy

A single train/test split can be lucky or unlucky. Cross-validation rotates the test fold — with cv=5, the data is split into 5 parts, the model trains on 4 and tests on the 5th, five times:

from sklearn.model_selection import cross_val_score

model = DecisionTreeClassifier(random_state=1)
scores = cross_val_score(model, iris.data, iris.target, cv=5)
print("Fold scores:", [f"{s:.3f}" for s in scores])
# Fold scores: ['0.967', '0.967', '0.900', '1.000', '1.000']
print(f"Mean: {scores.mean():.4f}")   # Mean: 0.9667

Report the mean ± std across folds — "96.7% ± 3.6%" is an honest claim; a single 100% fold is not. Use cross-validation to compare models and tune settings, then confirm the winner once on the held-out test set.

Picking the right metric for the job

  • Regression: MAE when errors cost linearly, RMSE when big errors are expensive, R² for a quick "variance explained" summary.
  • Balanced classification: accuracy is fine.
  • Imbalanced classification (fraud, disease): accuracy lies — optimize precision/recall/F1 or ROC-AUC instead.
  • Ranking/recommendation: precision@k, NDCG — how good are the top results?

Rule of thumb: the metric is the business goal translated into math. If false alarms cost support tickets, optimize precision. If misses cost lives, optimize recall. Say the tradeoff out loud before you train.

Key takeaways

  • Overfitting = memorizing training noise; the train/test gap exposes it.
  • Split with random_state for reproducibility and stratify for classification.
  • Cross-validation (cv=5) gives an honest mean ± std — use it to compare models.
  • Never tune on the test set; it gets one final look, at the end.
  • Choose the metric from the business cost of errors, not habit.

Next in this series: Feature Engineering Essentials for Machine Learning.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number