pandas Part 1: DataFrames, Selection and Data Cleaning
Part 2 of the Python for Data Science track. Last updated: September 2026.
pandas is the spreadsheet of Python: labeled rows and columns you can filter, clean, and summarize in one-liners. It is built on NumPy, so everything from Part 1 applies underneath.
Install it in your virtual environment with:
pip install pandas
Series and DataFrames
A Series is one labeled column; a DataFrame is a table of them:
import pandas as pd
s = pd.Series([10, 20, 30], index=["a", "b", "c"])
df = pd.DataFrame({
"name": ["Ada", "Grace", "Katherine"],
"age": [36, 85, 103],
"city": ["London", "New York", "Virginia"]
})
print(df)
Reading data and first inspection
df = pd.read_csv("employees.csv") # CSV file, URL, or read_excel() for spreadsheets
df.head() # first 5 rows — always LOOK at your data first
df.info() # row count, column dtypes, non-null counts
df.describe() # count/mean/std/min/max for numeric columns
df.shape # (rows, columns)
head → info → describe is the reflex for every new dataset. Burn it in.
Selecting data: [], loc, iloc
df["name"] # one column → Series df[["name", "age"]] # several columns → DataFrame df.loc[0, "city"] # by LABEL: row 0, column "city" df.iloc[0:2, 1:3] # by POSITION: rows 0-1, columns 1-2 df[df["age"] > 50] # boolean filter → rows where age > 50
Use loc for labels, iloc for positions. Mixing them up is the classic pandas bug.
Handling missing values
df.isna().sum() # missing-value count per column
df = df.dropna() # drop rows with ANY missing value
# df = df.dropna(subset=["age"]) # or only when THESE columns are missing
df["city"] = df["city"].fillna("Unknown") # fill with a placeholder
df["age"] = df["age"].fillna(df["age"].median()) # or a statistic
Never ignore missing data — isna().sum() first, then decide: drop or fill.
Renaming columns and fixing dtypes
df = df.rename(columns={"name": "full_name", "city": "location"})
df["age"] = df["age"].astype(float) # fix a wrong dtype
print(df.dtypes) # verify: object = text, int64/float64 = numbers
Key takeaways
- DataFrame = labeled table; Series = single column.
- New dataset? Run head → info → describe before anything else.
- loc = by label, iloc = by position; boolean filters select rows.
- Find missing values with isna().sum(), then dropna() or fillna().
Next in this series: pandas Part 2: GroupBy, Merging and Pivot Tables.
Comments
Post a Comment