Statistics for Data Science with Python

Part 7 of the Python for Data Science track. Last updated: September 2026.

Statistics is the language data speaks. You do not need a math degree — you need five concepts and the Python to compute them. Each one below comes with its data-science use case.

Mean, median, mode: three "averages"

import statistics
import numpy as np

scores = [78, 85, 92, 65, 88, 95, 72, 81, 90, 250]  # 250: a data-entry error?

print(statistics.mean(scores))    # dragged up by the outlier
print(statistics.median(scores))  # 86.5 — the robust center
print(np.median(scores))         # same, NumPy style

Use case: reporting "average" revenue or latency — median when outliers lurk, mean otherwise. Choosing wrong here misleads every stakeholder.

Variance and standard deviation: how spread out?

print(statistics.stdev(scores))  # sample standard deviation
print(np.std(scores, ddof=1))    # identical — ddof=1 means "sample", not population

# Use case: two delivery services both average 30 min —
# the one with std 5 is reliable, the one with std 20 is a gamble.

The normal distribution

Heights, measurement errors, test scores — countless real quantities pile up in the bell curve. The 68-95-99.7 rule: about 68% of values fall within 1 std of the mean, about 95% within 2:

import matplotlib.pyplot as plt

heights = np.random.normal(loc=170, scale=10, size=1000)  # mean 170 cm, std 10
plt.hist(heights, bins=30)
plt.title("Simulated Heights — the bell curve appears")
plt.show()

inside = np.mean(np.abs(heights - 170) < 10)   # fraction within 1 std?
print(f"{inside:.1%} within one std of the mean")  # about 68%

Correlation: do two things move together?

import pandas as pd

df = pd.DataFrame({
    "ad_spend": [10, 20, 30, 40, 50],
    "sales":    [15, 28, 44, 58, 72]})
print(df["ad_spend"].corr(df["sales"]))  # ≈ 0.999 — strong positive link

# df.corr() gives the whole matrix — pair it with a seaborn heatmap (Part 4)

Use case: feature selection — correlated features are candidates for your model. And the mantra: correlation is not causation, but it tells you where to look.

Sampling: learning from a slice

import random

population = list(range(100_000))
sample = random.sample(population, 1000)      # unbiased random sample
print(statistics.mean(sample))               # ≈ 50,000 — close to the true mean

# In pandas: df.sample(n=1000, random_state=42) — random_state makes it reproducible

Use case: you cannot plot 100M rows — sample 10k and your charts look the same. Train/test splits are sampling too (much more in the AI/ML track).

Key takeaways

  • Median resists outliers; mean does not — pick deliberately.
  • Standard deviation measures reliability and risk, not just spread.
  • The normal distribution and the 68-95-99.7 rule describe countless real datasets.
  • Correlation finds relationships (not causes); sampling makes big data workable.

You've completed the Python for Data Science track! Next on the roadmap: Python for AI/ML.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number