Statistics for Data Science with Python
Part 7 of the Python for Data Science track. Last updated: September 2026.
Statistics is the language data speaks. You do not need a math degree — you need five concepts and the Python to compute them. Each one below comes with its data-science use case.
Mean, median, mode: three "averages"
import statistics import numpy as np scores = [78, 85, 92, 65, 88, 95, 72, 81, 90, 250] # 250: a data-entry error? print(statistics.mean(scores)) # dragged up by the outlier print(statistics.median(scores)) # 86.5 — the robust center print(np.median(scores)) # same, NumPy style
Use case: reporting "average" revenue or latency — median when outliers lurk, mean otherwise. Choosing wrong here misleads every stakeholder.
Variance and standard deviation: how spread out?
print(statistics.stdev(scores)) # sample standard deviation print(np.std(scores, ddof=1)) # identical — ddof=1 means "sample", not population # Use case: two delivery services both average 30 min — # the one with std 5 is reliable, the one with std 20 is a gamble.
The normal distribution
Heights, measurement errors, test scores — countless real quantities pile up in the bell curve. The 68-95-99.7 rule: about 68% of values fall within 1 std of the mean, about 95% within 2:
import matplotlib.pyplot as plt
heights = np.random.normal(loc=170, scale=10, size=1000) # mean 170 cm, std 10
plt.hist(heights, bins=30)
plt.title("Simulated Heights — the bell curve appears")
plt.show()
inside = np.mean(np.abs(heights - 170) < 10) # fraction within 1 std?
print(f"{inside:.1%} within one std of the mean") # about 68%
Correlation: do two things move together?
import pandas as pd
df = pd.DataFrame({
"ad_spend": [10, 20, 30, 40, 50],
"sales": [15, 28, 44, 58, 72]})
print(df["ad_spend"].corr(df["sales"])) # ≈ 0.999 — strong positive link
# df.corr() gives the whole matrix — pair it with a seaborn heatmap (Part 4)
Use case: feature selection — correlated features are candidates for your model. And the mantra: correlation is not causation, but it tells you where to look.
Sampling: learning from a slice
import random population = list(range(100_000)) sample = random.sample(population, 1000) # unbiased random sample print(statistics.mean(sample)) # ≈ 50,000 — close to the true mean # In pandas: df.sample(n=1000, random_state=42) — random_state makes it reproducible
Use case: you cannot plot 100M rows — sample 10k and your charts look the same. Train/test splits are sampling too (much more in the AI/ML track).
Key takeaways
- Median resists outliers; mean does not — pick deliberately.
- Standard deviation measures reliability and risk, not just spread.
- The normal distribution and the 68-95-99.7 rule describe countless real datasets.
- Correlation finds relationships (not causes); sampling makes big data workable.
You've completed the Python for Data Science track! Next on the roadmap: Python for AI/ML.
Comments
Post a Comment