Neural Networks Crash Course with PyTorch

Part 6 of the Python for AI/ML track. Last updated: September 2026.

Everything so far was classical ML — and for tabular data it is often all you need. Neural networks earn their keep elsewhere: images, text, audio, and problems with mountains of data where they learn their own features. This crash course gives you the mental model plus a working PyTorch network.

When neural nets beat classical ML

  • Unstructured data — images, speech, raw text. A random forest cannot look at pixels; a CNN can.
  • Huge datasets — classical models plateau; deep networks keep improving with more data.
  • Feature learning — the network invents its own features instead of you engineering them.

When not to use them: small tabular datasets (a few thousand rows). There, gradient boosting usually wins with 10x less fuss. Neural nets are a power tool, not a default.

Setup and tensors

pip install torch --index-url https://download.pytorch.org/whl/cpu

A tensor is PyTorch's array — like a NumPy array with two superpowers: it can live on a GPU, and it remembers how it was computed so gradients flow backwards automatically (autograd). That second trick is what makes training possible:

import torch

x = torch.tensor([2.0], requires_grad=True)  # track operations on x
y = x ** 2 + 3 * x                            # y = x² + 3x
y.backward()                                  # dy/dx = 2x + 3 = 7 at x=2
print(x.grad)                                 # tensor([7.])

One line of calculus, done for any function you can write. Backpropagation is just this, scaled to millions of parameters.

A tiny neural network: MLPs

A multi-layer perceptron stacks layers of neurons: each layer = linear transform + non-linearity (ReLU). We will classify the two-moons dataset — two interleaved crescents no straight line can separate:

import torch.nn as nn
from sklearn.datasets import make_moons
from sklearn.model_selection import train_test_split

X, y = make_moons(n_samples=1000, noise=0.2, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

Xtr = torch.tensor(X_train, dtype=torch.float32)
ytr = torch.tensor(y_train, dtype=torch.float32).unsqueeze(1)
Xte = torch.tensor(X_test, dtype=torch.float32)

model = nn.Sequential(
    nn.Linear(2, 16),   # input (x, y) → 16 neurons
    nn.ReLU(),          # non-linearity: the bend that curves boundaries
    nn.Linear(16, 16),
    nn.ReLU(),
    nn.Linear(16, 1),   # → single logit
    nn.Sigmoid(),       # squash to a probability
)
print(model)

The training loop

Training = repeat three steps: predict, measure error (loss), nudge weights against the gradient. The optimizer (Adam) handles the nudging:

criterion = nn.BCELoss()                        # binary cross-entropy loss
optimizer = torch.optim.Adam(model.parameters(), lr=0.05)

for epoch in range(1, 301):
    model.train()
    optimizer.zero_grad()                       # clear old gradients
    loss = criterion(model(Xtr), ytr)           # 1. forward pass → loss
    loss.backward()                             # 2. backpropagate gradients
    optimizer.step()                            # 3. update weights
    if epoch % 100 == 0:
        print(f"Epoch {epoch}: loss = {loss.item():.4f}")
# Epoch 100: loss = 0.2814
# Epoch 200: loss = 0.2133
# Epoch 300: loss = 0.1966   ← falling loss = learning

model.eval()
with torch.no_grad():                           # no gradients needed for evaluation
    preds = (model(Xte) > 0.5).float()
    acc = (preds.squeeze() == torch.tensor(y_test, dtype=torch.float32)).float().mean()
print(f"Test accuracy: {acc:.4f}")              # Test accuracy: ~0.96

The network learned a curved boundary around the moons — something logistic regression fundamentally cannot do. That curved boundary is the whole point of hidden layers.

Key takeaways

  • Use neural nets for unstructured data and huge datasets; classical ML usually wins on small tables.
  • Tensors = NumPy arrays + GPU + autograd (automatic gradients).
  • An MLP is Linear → ReLU → Linear → …; hidden layers learn curved boundaries.
  • The training loop is always: forward pass → loss → backward → optimizer step.
  • Evaluate with torch.no_grad() — gradients are for training only.

Next in this series: End-to-End ML Project: From Raw Data to Predictions.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number