Exploratory Data Analysis: A Complete Walkthrough
Part 6 of the Python for Data Science track. Last updated: September 2026.
Exploratory Data Analysis (EDA) is the detective work before modeling: load the data, interrogate it, clean it, and let it tell you what is going on. This post does one complete EDA end-to-end on a small sales dataset — follow the thinking, not just the code.
1. Load
import pandas as pd from io import StringIO csv_data = """date,product,region,units,price 2026-09-01,Widget A,North,10,25.0 2026-09-01,Widget B,South,5,40.0 2026-09-02,Widget A,North,8,25.0 2026-09-02,Widget C,East,12,15.0 2026-09-03,Widget B,South,,40.0 2026-09-03,Widget A,West,15,25.0 2026-09-04,Widget C,East,9,15.0 2026-09-04,Widget B,North,7,40.0 2026-09-05,Widget A,South,11,25.0 2026-09-05,Widget C,West,6, 2026-09-06,Widget A,North,10,25.0 2026-09-06,Widget A,North,10,25.0""" df = pd.read_csv(StringIO(csv_data))
2. Inspect: interrogate before trusting
print(df.head()) print(df.info()) # 12 rows — but are all columns complete? print(df.describe()) # any wild values in units/price? print(df.isna().sum()) # units: 1 missing, price: 1 missing print(df.duplicated().sum()) # 1 duplicate row — the Sep 6 Widget A line repeats
Findings so far: two missing values, one duplicated row. Never analyze before inspecting — totals computed on dirty data are fiction.
3. Clean
df = df.drop_duplicates() # 12 → 11 rows df["units"] = df["units"].fillna(df["units"].median()) # median: robust to outliers df["price"] = df["price"].fillna(df["price"].median()) df["date"] = pd.to_datetime(df["date"]) df["revenue"] = df["units"] * df["price"] # derived column: the real metric print(df.isna().sum().sum()) # 0 — clean
4. Describe: aggregate
print(df.groupby("product")["revenue"].sum().sort_values(ascending=False))
print(df.groupby("region")["revenue"].sum().sort_values(ascending=False))
print("Total revenue:", df["revenue"].sum())
5. Visualize
import matplotlib.pyplot as plt
df.groupby("date")["revenue"].sum().plot(kind="line", marker="o")
plt.title("Daily Revenue")
plt.ylabel("Revenue ($)")
plt.show()
df.groupby("product")["revenue"].sum().plot(kind="bar", color="green")
plt.title("Revenue by Product")
plt.ylabel("Revenue ($)")
plt.show()
6. Insights: what the data says
- Widget A drives revenue through volume — mid-priced, but sells the most units. Protect its supply.
- North is the strongest region; West the weakest — is West under-served or just small? Worth investigating, not assuming.
- Revenue dips mid-week (Sep 3–4) — check whether that was inventory, marketing, or noise before acting.
- The data had 3 defects (2 missing values, 1 duplicate) — found in minutes by inspecting first. Dirty data is the norm, not the exception.
Notice the pattern: load → inspect → clean → describe → visualize → conclude. Every EDA follows it; the questions change, the workflow does not.
Key takeaways
- EDA workflow: load, inspect, clean, describe, visualize, conclude.
- Inspect with head / info / describe / isna / duplicated before trusting any number.
- Create derived columns (like revenue) that match the business question.
- End with written insights, not just charts — analysis without conclusions is trivia.
Next in this series: Statistics for Data Science with Python.
Comments
Post a Comment