Linear Regression with scikit-learn: Predicting Numbers

Part 2 of the Python for AI/ML track. Last updated: September 2026.

Regression means predicting a number: tomorrow's temperature, a stock's close, a house's price. The simplest and most useful regression model is also the oldest — linear regression, the line of best fit. Do not let the simplicity fool you: it is the baseline every fancier model must beat.

The intuition: drawing the best line

Given points on a scatter plot, linear regression finds the straight line that minimizes the squared distances to the points. The model learns one weight per feature plus an intercept:

price = w × sqft + b

Training = finding the w and b that make the line hug the data. Prediction = plugging a new sqft into the equation.

A worked example: house prices

Synthetic but realistic data — price grows ~$300 per square foot with noise:

import numpy as np
from sklearn.linear_model import LinearRegression

rng = np.random.default_rng(42)
sqft = rng.uniform(500, 3500, 200)
price = 50_000 + 300 * sqft + rng.normal(0, 25_000, 200)

X = sqft.reshape(-1, 1)      # sklearn wants a 2D array: (n_samples, n_features)
model = LinearRegression()
model.fit(X, price)

print(f"Learned: price = {model.coef_[0]:.2f} × sqft + {model.intercept_:.2f}")
# Learned: price = 305.49 × sqft + 39524.33  ← recovered the true formula!

print(f"2000 sqft → ${model.predict([[2000]])[0]:,.2f}")
# 2000 sqft → $650,511.15

The model recovered the true relationship (300 × sqft + 50,000) from noisy data alone. That is learning.

Evaluating regression: MAE, RMSE, R²

A model is only as good as its errors. Three metrics cover nearly everything:

  • MAE (Mean Absolute Error) — average |error| in original units. "Off by $19,878 on average." Easy to explain.
  • RMSE (Root Mean Squared Error) — like MAE but punishes big errors harder. Use it when large mistakes are expensive.
  • R² (R-squared) — fraction of variance explained, 0 to 1. 0.99 means the model captures 99% of price variation.
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score

pred = model.predict(X)
print(f"MAE:  ${mean_absolute_error(price, pred):,.2f}")   # MAE:  $19,878.61
print(f"RMSE: ${mean_squared_error(price, pred) ** 0.5:,.2f}")  # RMSE: $24,947.55
print(f"R²:   {r2_score(price, pred):.4f}")                # R²:   0.9907

Multiple features

Real houses need more than square footage. Adding features is trivial — one column per feature:

bedrooms = rng.integers(1, 6, 200)
price2 = 40_000 + 280 * sqft + 15_000 * bedrooms + rng.normal(0, 25_000, 200)

X2 = np.column_stack([sqft, bedrooms])
m2 = LinearRegression().fit(X2, price2)
print(m2.coef_)        # [  280.x  15000.x] — one weight per feature
print(f"R²: {m2.score(X2, price2):.4f}")

Each coefficient now reads as "holding other features fixed, one more bedroom adds ~$15k." That interpretability is linear regression's superpower — and why it survives in finance, medicine, and policy.

Key takeaways

  • Regression predicts numbers; linear regression fits the best line (or hyperplane).
  • sklearn needs 2D feature arrays — reshape single features with .reshape(-1, 1).
  • MAE = average error in real units; RMSE punishes big misses; R² = fraction of variance explained.
  • Coefficients are interpretable: each weight is the feature's price tag, all else equal.
  • Always start with linear regression — it is the baseline every complex model must beat.

Next in this series: Classification with Python: Logistic Regression and Decision Trees.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number