Feature Engineering Essentials for Machine Learning

Part 5 of the Python for AI/ML track. Last updated: September 2026.

Ask a Kaggle grandmaster what wins competitions and the answer is rarely the fanciest model — it is better features. Feature engineering is how you feed the algorithm data it can actually digest: scaled numbers, encoded categories, no leaks. This post covers the three operations that handle 95% of tabular data.

1. Scaling: put features on the same ruler

Many algorithms (KNN, SVM, neural nets, regularized regression) measure distances or take gradient steps. If income is in dollars (20,000–200,000) and age in years (18–70), income dominates everything. StandardScaler fixes it: subtract the mean, divide by the standard deviation:

import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier

rng = np.random.default_rng(0)
income = rng.uniform(20_000, 200_000, 300)
age = rng.uniform(18, 70, 300)
y = ((income / 1000) * 0.02 + age * 0.5 + rng.normal(0, 3, 300) > 25).astype(int)
X = np.column_stack([income, age])
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=1)

print("KNN, no scaling:      ", round(KNeighborsClassifier().fit(X_train, y_train).score(X_test, y_test), 4))
# KNN, no scaling:       0.4778  ← worse than guessing

scaler = StandardScaler().fit(X_train)          # learn mean/std from TRAIN only
X_train_s = scaler.transform(X_train)
X_test_s = scaler.transform(X_test)            # apply the same transform to test
print("KNN, after scaling:   ", round(KNeighborsClassifier().fit(X_train_s, y_train).score(X_test_s, y_test), 4))
# KNN, after scaling:    0.9222

Same model, same data — 0.48 to 0.92 just from scaling. (Tree-based models like random forests do not need scaling, but it never hurts.)

2. Encoding categoricals: numbers only

Models eat numbers, not strings. For a nominal category (brand: Toyota/Honda/Ford), one-hot encoding creates one binary column per value. Never use 1/2/3 labels for nominal data — the model will think Ford (3) is "more than" Toyota (1).

import pandas as pd
from sklearn.preprocessing import OneHotEncoder

df = pd.DataFrame({"brand": ["Toyota", "Honda", "Ford", "Toyota"]})
enc = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
print(enc.fit_transform(df[["brand"]]))
# [[0. 0. 1.]   ← Toyota
#  [0. 1. 0.]   ← Honda
#  [1. 0. 0.]   ← Ford
#  [0. 0. 1.]]
print(enc.get_feature_names_out())  # ['brand_Ford' 'brand_Honda' 'brand_Toyota']

3. The leakage trap: fit on train ONLY

The single most common beginner bug: calling .fit() on the whole dataset before splitting. The scaler then "sees" the test set's mean and std — information leaks from test into training, and your evaluation is quietly inflated. The rule is absolute:

  • fit (learn parameters) — training data only.
  • transform (apply) — train, test, and all future data.

The professional pattern: ColumnTransformer + Pipeline

Real data mixes numeric and categorical columns, each needing different treatment. ColumnTransformer applies the right steps per column; Pipeline chains preprocessing + model into one object that cannot leak — cross-validation then handles the fit/transform split correctly for you:

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestClassifier

numeric = ["income", "age"]
categorical = ["brand"]

preprocess = ColumnTransformer([
    ("num", Pipeline([
        ("impute", SimpleImputer(strategy="median")),
        ("scale", StandardScaler()),
    ]), numeric),
    ("cat", Pipeline([
        ("impute", SimpleImputer(strategy="most_frequent")),
        ("encode", OneHotEncoder(handle_unknown="ignore")),
    ]), categorical),
])

model = Pipeline([
    ("preprocess", preprocess),
    ("classifier", RandomForestClassifier(random_state=42)),
])

model.fit(X_train_df, y_train)       # one fit: imputes, scales, encodes, trains
print(model.score(X_test_df, y_test))  # honest — no leakage possible

Bonus: the pipeline also fixes missing values (SimpleImputer) in the same pass. From here on, every model you build should live inside a pipeline.

Key takeaways

  • Scale features for distance/gradient-based models — StandardScaler is the default.
  • One-hot encode nominal categories; never use 1/2/3 labels for unordered values.
  • Fit on train, transform everywhere — fitting preprocessing on all data is leakage.
  • ColumnTransformer + Pipeline is the professional pattern: mixed column types, imputation, and zero leakage in one object.
  • Feature quality beats model choice — invest your time here first.

Next in this series: Neural Networks Crash Course with PyTorch.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number