Posts

End-to-End ML Project: From Raw Data to Predictions

Part 7 of the Python for AI/ML track. Last updated: September 2026. Six posts of theory end here. A real ML project is a pipeline : raw messy data in, a saved model file out, predictions on demand. This post builds one completely — used-car price prediction — reusing everything from Parts 1–6. Step 1: Load and inspect 1,000 used-car listings: brand, year, mileage → price. Real-world mess included — missing values: import numpy as np import pandas as pd rng = np.random.default_rng(7) n = 1000 brands = rng.choice(["Toyota", "Honda", "Ford", "BMW"], n) years = rng.integers(2005, 2024, n) mileage = rng.uniform(5_000, 200_000, n) price = (30_000 - (2024 - years) * 1_200 - mileage * 0.08 + np.where(brands == "BMW", 8_000, 0) + rng.normal(0, 1500, n)) df = pd.DataFrame({"brand": brands, "year": years, "mileage": mileage, "price": price.round(2)}) df.loc[rng.choice(n, 40,...

Neural Networks Crash Course with PyTorch

Part 6 of the Python for AI/ML track. Last updated: September 2026. Everything so far was classical ML — and for tabular data it is often all you need. Neural networks earn their keep elsewhere: images, text, audio, and problems with mountains of data where they learn their own features. This crash course gives you the mental model plus a working PyTorch network. When neural nets beat classical ML Unstructured data — images, speech, raw text. A random forest cannot look at pixels; a CNN can. Huge datasets — classical models plateau; deep networks keep improving with more data. Feature learning — the network invents its own features instead of you engineering them. When not to use them: small tabular datasets (a few thousand rows). There, gradient boosting usually wins with 10x less fuss. Neural nets are a power tool, not a default. Setup and tensors pip install torch --index-url https://download.pytorch.org/whl/cpu A tensor is PyTorch's array — like a NumPy array ...

Feature Engineering Essentials for Machine Learning

Part 5 of the Python for AI/ML track. Last updated: September 2026. Ask a Kaggle grandmaster what wins competitions and the answer is rarely the fanciest model — it is better features . Feature engineering is how you feed the algorithm data it can actually digest: scaled numbers, encoded categories, no leaks. This post covers the three operations that handle 95% of tabular data. 1. Scaling: put features on the same ruler Many algorithms (KNN, SVM, neural nets, regularized regression) measure distances or take gradient steps . If income is in dollars (20,000–200,000) and age in years (18–70), income dominates everything. StandardScaler fixes it: subtract the mean, divide by the standard deviation: import numpy as np from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from sklearn.neighbors import KNeighborsClassifier rng = np.random.default_rng(0) income = rng.uniform(20_000, 200_000, 300) age = rng.uniform(18, 70, 300) y = ((inco...

Model Evaluation Done Right: Train-Test Split and Cross-Validation

Part 4 of the Python for AI/ML track. Last updated: September 2026. A model that scores 100% on its training data can still be worthless. Evaluation — measuring performance honestly on data the model never saw — is what separates ML from numerology. Get this wrong and you ship a model that fails on day one. The trap: overfitting An unconstrained decision tree memorizes the training set — every quirk, every noise point. Watch the gap between train and test scores: from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeClassifier iris = load_iris() X_train, X_test, y_train, y_test = train_test_split( iris.data, iris.target, test_size=0.3, random_state=1) deep = DecisionTreeClassifier(random_state=1).fit(X_train, y_train) # no limits shallow = DecisionTreeClassifier(max_depth=2, random_state=1).fit(X_train, y_train) print(f"Deep tree: train={deep.score(X_train, y_train):.3f} test={deep...

Classification with Python: Logistic Regression and Decision Trees

Part 3 of the Python for AI/ML track. Last updated: September 2026. Regression predicts numbers; classification predicts categories : spam or not, churned or retained, which species this flower is. It is the most commercially used branch of ML, and two algorithms cover most real cases — logistic regression and decision trees . The intuition A classifier draws boundaries between categories. Logistic regression (despite the name, it is a classifier) draws one straight boundary and outputs probabilities — "87% chance this is versicolor." A decision tree asks a series of yes/no questions — "petal length > 2.5? → …" — carving the data into boxes. Straight boundary vs boxy boundaries; both are useful. Head-to-head on the Iris dataset from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import accuracy_sc...

Linear Regression with scikit-learn: Predicting Numbers

Part 2 of the Python for AI/ML track. Last updated: September 2026. Regression means predicting a number: tomorrow's temperature, a stock's close, a house's price. The simplest and most useful regression model is also the oldest — linear regression , the line of best fit. Do not let the simplicity fool you: it is the baseline every fancier model must beat. The intuition: drawing the best line Given points on a scatter plot, linear regression finds the straight line that minimizes the squared distances to the points. The model learns one weight per feature plus an intercept: price = w × sqft + b Training = finding the w and b that make the line hug the data. Prediction = plugging a new sqft into the equation. A worked example: house prices Synthetic but realistic data — price grows ~$300 per square foot with noise: import numpy as np from sklearn.linear_model import LinearRegression rng = np.random.default_rng(42) sqft = rng.uniform(500, 3500, 200) price = 50_000 +...

Machine Learning with Python: Roadmap and scikit-learn Setup

Part 1 of the Python for AI/ML track. Last updated: September 2026. You finished the Python core track and the Data Science track — you can wrangle data with pandas and summarize it with statistics. Now the fun part: teaching the machine to learn patterns from data by itself . That is machine learning, and Python is its home turf. What machine learning actually is Traditional programming: you write the rules, data goes in, answers come out. Machine learning flips it: data and answers go in, rules come out. You show the algorithm thousands of examples ("this house sold for $650k"), and it figures out the pattern (price grows ~$300 per square foot). Then it applies that pattern to houses it has never seen. Every ML project follows the same workflow: Collect data — rows of examples with known outcomes. Pick features — the columns the model gets to look at. Train — the algorithm adjusts itself to fit the examples. Evaluate — test it on examples it never saw during t...