pandas Part 1: DataFrames, Selection and Data Cleaning

Part 2 of the Python for Data Science track. Last updated: September 2026.

pandas is the spreadsheet of Python: labeled rows and columns you can filter, clean, and summarize in one-liners. It is built on NumPy, so everything from Part 1 applies underneath.

Install it in your virtual environment with:

pip install pandas

Series and DataFrames

A Series is one labeled column; a DataFrame is a table of them:

import pandas as pd

s = pd.Series([10, 20, 30], index=["a", "b", "c"])
df = pd.DataFrame({
    "name": ["Ada", "Grace", "Katherine"],
    "age": [36, 85, 103],
    "city": ["London", "New York", "Virginia"]
})
print(df)

Reading data and first inspection

df = pd.read_csv("employees.csv")   # CSV file, URL, or read_excel() for spreadsheets

df.head()        # first 5 rows — always LOOK at your data first
df.info()        # row count, column dtypes, non-null counts
df.describe()    # count/mean/std/min/max for numeric columns
df.shape         # (rows, columns)

head → info → describe is the reflex for every new dataset. Burn it in.

Selecting data: [], loc, iloc

df["name"]            # one column → Series
df[["name", "age"]]   # several columns → DataFrame
df.loc[0, "city"]     # by LABEL: row 0, column "city"
df.iloc[0:2, 1:3]     # by POSITION: rows 0-1, columns 1-2
df[df["age"] > 50]    # boolean filter → rows where age > 50

Use loc for labels, iloc for positions. Mixing them up is the classic pandas bug.

Handling missing values

df.isna().sum()                                # missing-value count per column

df = df.dropna()                               # drop rows with ANY missing value
# df = df.dropna(subset=["age"])               # or only when THESE columns are missing

df["city"] = df["city"].fillna("Unknown")              # fill with a placeholder
df["age"] = df["age"].fillna(df["age"].median())       # or a statistic

Never ignore missing data — isna().sum() first, then decide: drop or fill.

Renaming columns and fixing dtypes

df = df.rename(columns={"name": "full_name", "city": "location"})
df["age"] = df["age"].astype(float)     # fix a wrong dtype
print(df.dtypes)                        # verify: object = text, int64/float64 = numbers

Key takeaways

  • DataFrame = labeled table; Series = single column.
  • New dataset? Run head → info → describe before anything else.
  • loc = by label, iloc = by position; boolean filters select rows.
  • Find missing values with isna().sum(), then dropna() or fillna().

Next in this series: pandas Part 2: GroupBy, Merging and Pivot Tables.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number