End-to-End ML Project: From Raw Data to Predictions

Part 7 of the Python for AI/ML track. Last updated: September 2026.

Six posts of theory end here. A real ML project is a pipeline: raw messy data in, a saved model file out, predictions on demand. This post builds one completely — used-car price prediction — reusing everything from Parts 1–6.

Step 1: Load and inspect

1,000 used-car listings: brand, year, mileage → price. Real-world mess included — missing values:

import numpy as np
import pandas as pd

rng = np.random.default_rng(7)
n = 1000
brands = rng.choice(["Toyota", "Honda", "Ford", "BMW"], n)
years = rng.integers(2005, 2024, n)
mileage = rng.uniform(5_000, 200_000, n)
price = (30_000 - (2024 - years) * 1_200 - mileage * 0.08
         + np.where(brands == "BMW", 8_000, 0) + rng.normal(0, 1500, n))

df = pd.DataFrame({"brand": brands, "year": years,
                   "mileage": mileage, "price": price.round(2)})
df.loc[rng.choice(n, 40, replace=False), "mileage"] = np.nan   # missing data happens
df.loc[rng.choice(n, 25, replace=False), "brand"] = None

print(df.shape)            # (1000, 4)
print(df.isna().sum())     # brand 25 missing, mileage 40 missing
print(df.head(3))

Step 2: Split first, then preprocess

Part 4's rule: split before any fitting, so nothing leaks from test to train:

from sklearn.model_selection import train_test_split

X = df.drop(columns="price")
y = df["price"]
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)   # (800, 3) (200, 3)

Step 3: Build the preprocessing pipeline

Part 5's pattern — median-impute + scale the numbers, most-frequent-impute + one-hot the brand, all inside a ColumnTransformer so cross-validation stays honest:

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestRegressor

numeric = ["year", "mileage"]
categorical = ["brand"]

preprocess = ColumnTransformer([
    ("num", Pipeline([
        ("impute", SimpleImputer(strategy="median")),
        ("scale", StandardScaler()),
    ]), numeric),
    ("cat", Pipeline([
        ("impute", SimpleImputer(strategy="most_frequent")),
        ("encode", OneHotEncoder(handle_unknown="ignore")),
    ]), categorical),
])

model = Pipeline([
    ("preprocess", preprocess),
    ("regressor", RandomForestRegressor(n_estimators=100, random_state=42)),
])

Step 4: Train and evaluate

from sklearn.metrics import mean_absolute_error, r2_score

model.fit(X_train, y_train)
pred = model.predict(X_test)

print(f"MAE: ${mean_absolute_error(y_test, pred):,.2f}")  # MAE: $1,500.12
print(f"R²:  {r2_score(y_test, pred):.4f}")               # R²:  0.9469

Off by ~$1,500 on average, explaining 95% of price variation — solid for a first pass. (A next iteration might tune n_estimators / max_depth with cross-validation from Part 4.)

Step 5: Save the model, predict on new data

A model you cannot reuse is a demo, not a product. joblib serializes the whole pipeline — preprocessing included — so prediction needs zero manual feature work:

import joblib

joblib.dump(model, "car_price_model.pkl")   # ship this file

# Later — or in another service entirely:
loaded = joblib.load("car_price_model.pkl")
new_cars = pd.DataFrame({
    "brand": ["Toyota", "BMW"],
    "year": [2020, 2018],
    "mileage": [45_000, 60_000],
})
print(loaded.predict(new_cars))   # [20433.89  ...] — raw inputs, clean outputs

Because the pipeline holds the imputer, scaler, and encoder, the loaded model accepts raw listings — missing values and all — and returns prices. That is the whole game: raw data in, predictions out, one file.

The full project checklist

  • Load & inspect — shapes, missing values, distributions (Parts 1–2 of the DS track).
  • Split first — train/test before any fitting; stratify for classification (Part 4).
  • Preprocess in a pipeline — impute, scale, encode per column type (Part 5).
  • Train a strong baseline — random forest / gradient boosting for tables (Parts 2–3).
  • Evaluate honestly — MAE/RMSE/R² or precision/recall on the held-out set (Part 4).
  • Save with joblib — the pipeline is the deployable artifact.

Key takeaways

  • An ML project is a pipeline: load → split → preprocess → train → evaluate → save.
  • Splitting before preprocessing and wrapping everything in a Pipeline eliminates leakage by construction.
  • joblib persists the entire pipeline — preprocessing travels with the model.
  • Start with a strong baseline (random forest), then iterate with cross-validation.
  • You now own the complete classical-ML workflow, plus a working neural net from Part 6.

You've completed the Python for AI/ML track! Browse every tutorial on the Python topic page. Next on the roadmap: Python for Web & APIs.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number