Exploratory Data Analysis: A Complete Walkthrough

Part 6 of the Python for Data Science track. Last updated: September 2026.

Exploratory Data Analysis (EDA) is the detective work before modeling: load the data, interrogate it, clean it, and let it tell you what is going on. This post does one complete EDA end-to-end on a small sales dataset — follow the thinking, not just the code.

1. Load

import pandas as pd
from io import StringIO

csv_data = """date,product,region,units,price
2026-09-01,Widget A,North,10,25.0
2026-09-01,Widget B,South,5,40.0
2026-09-02,Widget A,North,8,25.0
2026-09-02,Widget C,East,12,15.0
2026-09-03,Widget B,South,,40.0
2026-09-03,Widget A,West,15,25.0
2026-09-04,Widget C,East,9,15.0
2026-09-04,Widget B,North,7,40.0
2026-09-05,Widget A,South,11,25.0
2026-09-05,Widget C,West,6,
2026-09-06,Widget A,North,10,25.0
2026-09-06,Widget A,North,10,25.0"""

df = pd.read_csv(StringIO(csv_data))

2. Inspect: interrogate before trusting

print(df.head())
print(df.info())             # 12 rows — but are all columns complete?
print(df.describe())         # any wild values in units/price?
print(df.isna().sum())       # units: 1 missing, price: 1 missing
print(df.duplicated().sum()) # 1 duplicate row — the Sep 6 Widget A line repeats

Findings so far: two missing values, one duplicated row. Never analyze before inspecting — totals computed on dirty data are fiction.

3. Clean

df = df.drop_duplicates()                              # 12 → 11 rows
df["units"] = df["units"].fillna(df["units"].median())  # median: robust to outliers
df["price"] = df["price"].fillna(df["price"].median())
df["date"] = pd.to_datetime(df["date"])
df["revenue"] = df["units"] * df["price"]               # derived column: the real metric
print(df.isna().sum().sum())  # 0 — clean

4. Describe: aggregate

print(df.groupby("product")["revenue"].sum().sort_values(ascending=False))
print(df.groupby("region")["revenue"].sum().sort_values(ascending=False))
print("Total revenue:", df["revenue"].sum())

5. Visualize

import matplotlib.pyplot as plt

df.groupby("date")["revenue"].sum().plot(kind="line", marker="o")
plt.title("Daily Revenue")
plt.ylabel("Revenue ($)")
plt.show()

df.groupby("product")["revenue"].sum().plot(kind="bar", color="green")
plt.title("Revenue by Product")
plt.ylabel("Revenue ($)")
plt.show()

6. Insights: what the data says

  • Widget A drives revenue through volume — mid-priced, but sells the most units. Protect its supply.
  • North is the strongest region; West the weakest — is West under-served or just small? Worth investigating, not assuming.
  • Revenue dips mid-week (Sep 3–4) — check whether that was inventory, marketing, or noise before acting.
  • The data had 3 defects (2 missing values, 1 duplicate) — found in minutes by inspecting first. Dirty data is the norm, not the exception.

Notice the pattern: load → inspect → clean → describe → visualize → conclude. Every EDA follows it; the questions change, the workflow does not.

Key takeaways

  • EDA workflow: load, inspect, clean, describe, visualize, conclude.
  • Inspect with head / info / describe / isna / duplicated before trusting any number.
  • Create derived columns (like revenue) that match the business question.
  • End with written insights, not just charts — analysis without conclusions is trivia.

Next in this series: Statistics for Data Science with Python.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number