End-to-End ML Project: From Raw Data to Predictions
Part 7 of the Python for AI/ML track. Last updated: September 2026.
Six posts of theory end here. A real ML project is a pipeline: raw messy data in, a saved model file out, predictions on demand. This post builds one completely — used-car price prediction — reusing everything from Parts 1–6.
Step 1: Load and inspect
1,000 used-car listings: brand, year, mileage → price. Real-world mess included — missing values:
import numpy as np
import pandas as pd
rng = np.random.default_rng(7)
n = 1000
brands = rng.choice(["Toyota", "Honda", "Ford", "BMW"], n)
years = rng.integers(2005, 2024, n)
mileage = rng.uniform(5_000, 200_000, n)
price = (30_000 - (2024 - years) * 1_200 - mileage * 0.08
+ np.where(brands == "BMW", 8_000, 0) + rng.normal(0, 1500, n))
df = pd.DataFrame({"brand": brands, "year": years,
"mileage": mileage, "price": price.round(2)})
df.loc[rng.choice(n, 40, replace=False), "mileage"] = np.nan # missing data happens
df.loc[rng.choice(n, 25, replace=False), "brand"] = None
print(df.shape) # (1000, 4)
print(df.isna().sum()) # brand 25 missing, mileage 40 missing
print(df.head(3))
Step 2: Split first, then preprocess
Part 4's rule: split before any fitting, so nothing leaks from test to train:
from sklearn.model_selection import train_test_split
X = df.drop(columns="price")
y = df["price"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape) # (800, 3) (200, 3)
Step 3: Build the preprocessing pipeline
Part 5's pattern — median-impute + scale the numbers, most-frequent-impute + one-hot the brand, all inside a ColumnTransformer so cross-validation stays honest:
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestRegressor
numeric = ["year", "mileage"]
categorical = ["brand"]
preprocess = ColumnTransformer([
("num", Pipeline([
("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler()),
]), numeric),
("cat", Pipeline([
("impute", SimpleImputer(strategy="most_frequent")),
("encode", OneHotEncoder(handle_unknown="ignore")),
]), categorical),
])
model = Pipeline([
("preprocess", preprocess),
("regressor", RandomForestRegressor(n_estimators=100, random_state=42)),
])
Step 4: Train and evaluate
from sklearn.metrics import mean_absolute_error, r2_score
model.fit(X_train, y_train)
pred = model.predict(X_test)
print(f"MAE: ${mean_absolute_error(y_test, pred):,.2f}") # MAE: $1,500.12
print(f"R²: {r2_score(y_test, pred):.4f}") # R²: 0.9469
Off by ~$1,500 on average, explaining 95% of price variation — solid for a first pass. (A next iteration might tune n_estimators / max_depth with cross-validation from Part 4.)
Step 5: Save the model, predict on new data
A model you cannot reuse is a demo, not a product. joblib serializes the whole pipeline — preprocessing included — so prediction needs zero manual feature work:
import joblib
joblib.dump(model, "car_price_model.pkl") # ship this file
# Later — or in another service entirely:
loaded = joblib.load("car_price_model.pkl")
new_cars = pd.DataFrame({
"brand": ["Toyota", "BMW"],
"year": [2020, 2018],
"mileage": [45_000, 60_000],
})
print(loaded.predict(new_cars)) # [20433.89 ...] — raw inputs, clean outputs
Because the pipeline holds the imputer, scaler, and encoder, the loaded model accepts raw listings — missing values and all — and returns prices. That is the whole game: raw data in, predictions out, one file.
The full project checklist
- Load & inspect — shapes, missing values, distributions (Parts 1–2 of the DS track).
- Split first — train/test before any fitting; stratify for classification (Part 4).
- Preprocess in a pipeline — impute, scale, encode per column type (Part 5).
- Train a strong baseline — random forest / gradient boosting for tables (Parts 2–3).
- Evaluate honestly — MAE/RMSE/R² or precision/recall on the held-out set (Part 4).
- Save with joblib — the pipeline is the deployable artifact.
Key takeaways
- An ML project is a pipeline: load → split → preprocess → train → evaluate → save.
- Splitting before preprocessing and wrapping everything in a Pipeline eliminates leakage by construction.
- joblib persists the entire pipeline — preprocessing travels with the model.
- Start with a strong baseline (random forest), then iterate with cross-validation.
- You now own the complete classical-ML workflow, plus a working neural net from Part 6.
You've completed the Python for AI/ML track! Browse every tutorial on the Python topic page. Next on the roadmap: Python for Web & APIs.
Comments
Post a Comment