NumPy Crash Course: Arrays, Vectorization and Broadcasting
Part 1 of the Python for Data Science track. Last updated: September 2026.
Almost every data-science library — pandas, scikit-learn, TensorFlow — is built on top of NumPy. It gives Python fast, memory-efficient arrays and the math to operate on them. If the core track taught you Python, this post teaches you Python at data scale.
Install it inside your virtual environment (see Part 1 of the core track) with:
pip install numpy
Why NumPy? Speed
Python lists store pointers to objects scattered in memory. A NumPy array stores raw numbers in one contiguous C block — so numeric loops run in compiled C instead of interpreted Python:
import numpy as np
data = np.arange(1_000_000)
# Pure-Python loop: slow
total = 0
for x in data:
total += x
# Vectorized: the loop happens in C — typically 50-100x faster
total = data.sum()
doubled = data * 2 # every element at once, no loop
print(doubled[:5]) # [0 2 4 6 8]
The rule of data science in Python: if you are writing a loop over numbers, you are probably doing it wrong.
Creating arrays
a = np.array([1, 2, 3, 4]) # from a list z = np.zeros(5) # [0. 0. 0. 0. 0.] o = np.ones((2, 3)) # 2 rows x 3 columns of 1.0 r = np.arange(0, 10, 2) # like range(): [0 2 4 6 8] l = np.linspace(0, 1, 5) # 5 evenly spaced: [0. 0.25 0.5 0.75 1.] rnd = np.random.rand(3) # 3 random floats in [0, 1) print(a.dtype, o.shape) # e.g. int64 (2, 3) — dtype and shape matter
Indexing and slicing
m = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]])
print(m[0, 2]) # 3 — row 0, column 2
print(m[1]) # [4 5 6] — whole row
print(m[:, 0]) # [1 4 7] — whole column (: means "all")
print(m[0:2, 1:3]) # 2x2 block → [[2 3] [5 6]]
# Boolean masks: filter without loops
big = m[m > 5] # [6 7 8 9]
Broadcasting: arithmetic on mismatched shapes
When shapes differ, NumPy stretches the smaller one — aligning dimensions from the right; each dimension must match or be 1:
a = np.array([[1, 2, 3],
[4, 5, 6]])
print(a + 10) # scalar → every element
col = np.array([[100], [200]]) # shape (2, 1)
print(a + col) # column stretches across → [[101 102 103] [204 205 206]]
# Normalize columns: subtract each column's mean — one line, no loops
centered = a - a.mean(axis=0)
Essential functions
x = np.array([3, 7, 2, 9, 4]) print(x.mean(), x.sum(), x.std()) # 5.0 25 2.607... print(np.sort(x)) # [2 3 4 7 9] m = np.arange(12).reshape(3, 4) # reshape: 12 elements → 3x4 print(m.shape) # (3, 4) print(m.T.shape) # transpose → (4, 3) print(m.sum(axis=0)) # sum DOWN the rows → one value per column
Key takeaways
- NumPy arrays are fast because numeric loops run in C, not Python.
- Vectorize — operate on whole arrays instead of writing loops.
- Broadcasting stretches smaller shapes; dimensions must match or be 1.
- Learn mean / sum / std / reshape cold — you will use them in every analysis.
Next in this series: pandas Part 1: DataFrames, Selection and Data Cleaning.
Comments
Post a Comment