Posts

Statistics for Data Science with Python

Part 7 of the Python for Data Science track. Last updated: September 2026. Statistics is the language data speaks. You do not need a math degree — you need five concepts and the Python to compute them. Each one below comes with its data-science use case. Mean, median, mode: three "averages" import statistics import numpy as np scores = [78, 85, 92, 65, 88, 95, 72, 81, 90, 250] # 250: a data-entry error? print(statistics.mean(scores)) # dragged up by the outlier print(statistics.median(scores)) # 86.5 — the robust center print(np.median(scores)) # same, NumPy style Use case: reporting "average" revenue or latency — median when outliers lurk, mean otherwise. Choosing wrong here misleads every stakeholder. Variance and standard deviation: how spread out? print(statistics.stdev(scores)) # sample standard deviation print(np.std(scores, ddof=1)) # identical — ddof=1 means "sample", not population # Use case: two delivery services both...

Exploratory Data Analysis: A Complete Walkthrough

Part 6 of the Python for Data Science track. Last updated: September 2026. Exploratory Data Analysis (EDA) is the detective work before modeling: load the data, interrogate it, clean it, and let it tell you what is going on. This post does one complete EDA end-to-end on a small sales dataset — follow the thinking , not just the code. 1. Load import pandas as pd from io import StringIO csv_data = """date,product,region,units,price 2026-09-01,Widget A,North,10,25.0 2026-09-01,Widget B,South,5,40.0 2026-09-02,Widget A,North,8,25.0 2026-09-02,Widget C,East,12,15.0 2026-09-03,Widget B,South,,40.0 2026-09-03,Widget A,West,15,25.0 2026-09-04,Widget C,East,9,15.0 2026-09-04,Widget B,North,7,40.0 2026-09-05,Widget A,South,11,25.0 2026-09-05,Widget C,West,6, 2026-09-06,Widget A,North,10,25.0 2026-09-06,Widget A,North,10,25.0""" df = pd.read_csv(StringIO(csv_data)) 2. Inspect: interrogate before trusting print(df.head()) print(df.info()) # 12 rows...

SQL with Python: sqlite3, pandas and SQLAlchemy Basics

Part 5 of the Python for Data Science track. Last updated: September 2026. Most data lives in databases, not CSVs. Python talks to SQL databases natively via sqlite3 (built in, zero setup — perfect for learning) and scales up with pandas and SQLAlchemy . Connecting and creating a table import sqlite3 conn = sqlite3.connect("shop.db") # creates the file if it doesn't exist cur = conn.cursor() cur.execute("""CREATE TABLE IF NOT EXISTS customers ( id INTEGER PRIMARY KEY, name TEXT, city TEXT)""") conn.commit() Inserting safely: parameterized queries # ? placeholders — values travel SEPARATELY from the SQL text cur.execute("INSERT INTO customers (name, city) VALUES (?, ?)", ("Ada", "London")) cur.execute("INSERT INTO customers (name, city) VALUES (?, ?)", ("Grace", "New York")) conn.commit() # NEVER build SQL with f-strings: f"SELECT ... WHERE city = '{city}'...

Data Visualization with Matplotlib and Seaborn

Part 4 of the Python for Data Science track. Last updated: September 2026. A table of numbers convinces nobody; a chart does. Matplotlib is the foundation — full control, verbose. Seaborn sits on top — prettier defaults, statistical plots in one line. Learn both. Install them in your virtual environment with: pip install matplotlib seaborn Your first plot: line chart import matplotlib.pyplot as plt months = ["Jan", "Feb", "Mar", "Apr"] revenue = [120, 150, 130, 180] plt.plot(months, revenue, marker="o", label="2026") plt.title("Monthly Revenue") plt.xlabel("Month") plt.ylabel("Revenue ($k)") plt.legend() plt.show() Every chart needs a title, axis labels, and a legend — unlabeled charts are a rookie tell. Bar charts products = ["A", "B", "C"] units = [45, 78, 62] plt.bar(products, units, color="green") plt.title("Units Sold by Product"...

pandas Part 2: GroupBy, Merging and Pivot Tables

Part 3 of the Python for Data Science track. Last updated: September 2026. Part 2 covered single-table basics. Real analysis means combining tables and summarizing groups — the pandas equivalent of SQL's GROUP BY and JOIN. That is this post. import pandas as pd sales = pd.DataFrame({ "region": ["North", "North", "South", "South", "East"], "rep": ["Ana", "Ben", "Cara", "Dan", "Eli"], "amount": [120, 200, 150, 90, 300] }) GroupBy: split-apply-combine # Total per region print(sales.groupby("region")["amount"].sum()) # Multiple stats at once with agg report = sales.groupby("region").agg( total=("amount", "sum"), average=("amount", "mean"), deals=("amount", "count")) print(report) groupby splits rows into groups, applies a function, and ...

pandas Part 1: DataFrames, Selection and Data Cleaning

Part 2 of the Python for Data Science track. Last updated: September 2026. pandas is the spreadsheet of Python: labeled rows and columns you can filter, clean, and summarize in one-liners. It is built on NumPy, so everything from Part 1 applies underneath. Install it in your virtual environment with: pip install pandas Series and DataFrames A Series is one labeled column; a DataFrame is a table of them: import pandas as pd s = pd.Series([10, 20, 30], index=["a", "b", "c"]) df = pd.DataFrame({ "name": ["Ada", "Grace", "Katherine"], "age": [36, 85, 103], "city": ["London", "New York", "Virginia"] }) print(df) Reading data and first inspection df = pd.read_csv("employees.csv") # CSV file, URL, or read_excel() for spreadsheets df.head() # first 5 rows — always LOOK at your data first df.info() # row count, column dtypes, ...

NumPy Crash Course: Arrays, Vectorization and Broadcasting

Part 1 of the Python for Data Science track. Last updated: September 2026. Almost every data-science library — pandas, scikit-learn, TensorFlow — is built on top of NumPy . It gives Python fast, memory-efficient arrays and the math to operate on them. If the core track taught you Python, this post teaches you Python at data scale . Install it inside your virtual environment (see Part 1 of the core track) with: pip install numpy Why NumPy? Speed Python lists store pointers to objects scattered in memory. A NumPy array stores raw numbers in one contiguous C block — so numeric loops run in compiled C instead of interpreted Python: import numpy as np data = np.arange(1_000_000) # Pure-Python loop: slow total = 0 for x in data: total += x # Vectorized: the loop happens in C — typically 50-100x faster total = data.sum() doubled = data * 2 # every element at once, no loop print(doubled[:5]) # [0 2 4 6 8] The rule of data science in Python: if you are writing a loo...