Statistics for Data Science with Python
Part 7 of the Python for Data Science track. Last updated: September 2026. Statistics is the language data speaks. You do not need a math degree — you need five concepts and the Python to compute them. Each one below comes with its data-science use case. Mean, median, mode: three "averages" import statistics import numpy as np scores = [78, 85, 92, 65, 88, 95, 72, 81, 90, 250] # 250: a data-entry error? print(statistics.mean(scores)) # dragged up by the outlier print(statistics.median(scores)) # 86.5 — the robust center print(np.median(scores)) # same, NumPy style Use case: reporting "average" revenue or latency — median when outliers lurk, mean otherwise. Choosing wrong here misleads every stakeholder. Variance and standard deviation: how spread out? print(statistics.stdev(scores)) # sample standard deviation print(np.std(scores, ddof=1)) # identical — ddof=1 means "sample", not population # Use case: two delivery services both...