Neural Networks Crash Course with PyTorch
Part 6 of the Python for AI/ML track. Last updated: September 2026.
Everything so far was classical ML — and for tabular data it is often all you need. Neural networks earn their keep elsewhere: images, text, audio, and problems with mountains of data where they learn their own features. This crash course gives you the mental model plus a working PyTorch network.
When neural nets beat classical ML
- Unstructured data — images, speech, raw text. A random forest cannot look at pixels; a CNN can.
- Huge datasets — classical models plateau; deep networks keep improving with more data.
- Feature learning — the network invents its own features instead of you engineering them.
When not to use them: small tabular datasets (a few thousand rows). There, gradient boosting usually wins with 10x less fuss. Neural nets are a power tool, not a default.
Setup and tensors
pip install torch --index-url https://download.pytorch.org/whl/cpu
A tensor is PyTorch's array — like a NumPy array with two superpowers: it can live on a GPU, and it remembers how it was computed so gradients flow backwards automatically (autograd). That second trick is what makes training possible:
import torch x = torch.tensor([2.0], requires_grad=True) # track operations on x y = x ** 2 + 3 * x # y = x² + 3x y.backward() # dy/dx = 2x + 3 = 7 at x=2 print(x.grad) # tensor([7.])
One line of calculus, done for any function you can write. Backpropagation is just this, scaled to millions of parameters.
A tiny neural network: MLPs
A multi-layer perceptron stacks layers of neurons: each layer = linear transform + non-linearity (ReLU). We will classify the two-moons dataset — two interleaved crescents no straight line can separate:
import torch.nn as nn
from sklearn.datasets import make_moons
from sklearn.model_selection import train_test_split
X, y = make_moons(n_samples=1000, noise=0.2, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
Xtr = torch.tensor(X_train, dtype=torch.float32)
ytr = torch.tensor(y_train, dtype=torch.float32).unsqueeze(1)
Xte = torch.tensor(X_test, dtype=torch.float32)
model = nn.Sequential(
nn.Linear(2, 16), # input (x, y) → 16 neurons
nn.ReLU(), # non-linearity: the bend that curves boundaries
nn.Linear(16, 16),
nn.ReLU(),
nn.Linear(16, 1), # → single logit
nn.Sigmoid(), # squash to a probability
)
print(model)
The training loop
Training = repeat three steps: predict, measure error (loss), nudge weights against the gradient. The optimizer (Adam) handles the nudging:
criterion = nn.BCELoss() # binary cross-entropy loss
optimizer = torch.optim.Adam(model.parameters(), lr=0.05)
for epoch in range(1, 301):
model.train()
optimizer.zero_grad() # clear old gradients
loss = criterion(model(Xtr), ytr) # 1. forward pass → loss
loss.backward() # 2. backpropagate gradients
optimizer.step() # 3. update weights
if epoch % 100 == 0:
print(f"Epoch {epoch}: loss = {loss.item():.4f}")
# Epoch 100: loss = 0.2814
# Epoch 200: loss = 0.2133
# Epoch 300: loss = 0.1966 ← falling loss = learning
model.eval()
with torch.no_grad(): # no gradients needed for evaluation
preds = (model(Xte) > 0.5).float()
acc = (preds.squeeze() == torch.tensor(y_test, dtype=torch.float32)).float().mean()
print(f"Test accuracy: {acc:.4f}") # Test accuracy: ~0.96
The network learned a curved boundary around the moons — something logistic regression fundamentally cannot do. That curved boundary is the whole point of hidden layers.
Key takeaways
- Use neural nets for unstructured data and huge datasets; classical ML usually wins on small tables.
- Tensors = NumPy arrays + GPU + autograd (automatic gradients).
- An MLP is Linear → ReLU → Linear → …; hidden layers learn curved boundaries.
- The training loop is always: forward pass → loss → backward → optimizer step.
- Evaluate with torch.no_grad() — gradients are for training only.
Next in this series: End-to-End ML Project: From Raw Data to Predictions.
Comments
Post a Comment