NumPy Crash Course: Arrays, Vectorization and Broadcasting

Part 1 of the Python for Data Science track. Last updated: September 2026.

Almost every data-science library — pandas, scikit-learn, TensorFlow — is built on top of NumPy. It gives Python fast, memory-efficient arrays and the math to operate on them. If the core track taught you Python, this post teaches you Python at data scale.

Install it inside your virtual environment (see Part 1 of the core track) with:

pip install numpy

Why NumPy? Speed

Python lists store pointers to objects scattered in memory. A NumPy array stores raw numbers in one contiguous C block — so numeric loops run in compiled C instead of interpreted Python:

import numpy as np

data = np.arange(1_000_000)

# Pure-Python loop: slow
total = 0
for x in data:
    total += x

# Vectorized: the loop happens in C — typically 50-100x faster
total = data.sum()
doubled = data * 2        # every element at once, no loop
print(doubled[:5])        # [0 2 4 6 8]

The rule of data science in Python: if you are writing a loop over numbers, you are probably doing it wrong.

Creating arrays

a = np.array([1, 2, 3, 4])      # from a list
z = np.zeros(5)                 # [0. 0. 0. 0. 0.]
o = np.ones((2, 3))             # 2 rows x 3 columns of 1.0
r = np.arange(0, 10, 2)         # like range(): [0 2 4 6 8]
l = np.linspace(0, 1, 5)        # 5 evenly spaced: [0. 0.25 0.5 0.75 1.]
rnd = np.random.rand(3)         # 3 random floats in [0, 1)

print(a.dtype, o.shape)         # e.g. int64 (2, 3) — dtype and shape matter

Indexing and slicing

m = np.array([[1, 2, 3],
              [4, 5, 6],
              [7, 8, 9]])

print(m[0, 2])       # 3 — row 0, column 2
print(m[1])          # [4 5 6] — whole row
print(m[:, 0])       # [1 4 7] — whole column (: means "all")
print(m[0:2, 1:3])   # 2x2 block → [[2 3] [5 6]]

# Boolean masks: filter without loops
big = m[m > 5]       # [6 7 8 9]

Broadcasting: arithmetic on mismatched shapes

When shapes differ, NumPy stretches the smaller one — aligning dimensions from the right; each dimension must match or be 1:

a = np.array([[1, 2, 3],
              [4, 5, 6]])

print(a + 10)                    # scalar → every element
col = np.array([[100], [200]])   # shape (2, 1)
print(a + col)                   # column stretches across → [[101 102 103] [204 205 206]]

# Normalize columns: subtract each column's mean — one line, no loops
centered = a - a.mean(axis=0)

Essential functions

x = np.array([3, 7, 2, 9, 4])
print(x.mean(), x.sum(), x.std())   # 5.0 25 2.607...
print(np.sort(x))                   # [2 3 4 7 9]

m = np.arange(12).reshape(3, 4)     # reshape: 12 elements → 3x4
print(m.shape)                      # (3, 4)
print(m.T.shape)                    # transpose → (4, 3)
print(m.sum(axis=0))                # sum DOWN the rows → one value per column

Key takeaways

  • NumPy arrays are fast because numeric loops run in C, not Python.
  • Vectorize — operate on whole arrays instead of writing loops.
  • Broadcasting stretches smaller shapes; dimensions must match or be 1.
  • Learn mean / sum / std / reshape cold — you will use them in every analysis.

Next in this series: pandas Part 1: DataFrames, Selection and Data Cleaning.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number