Web Scraping with Python: requests and BeautifulSoup
Part 2 of the Python for Automation track. Last updated: September 2026.
Prices change, new listings appear, competitors update their pages — and checking all of it by hand does not scale. Web scraping is automation's eyes: fetch a page, pull out the structured data hiding in the HTML, and hand it to the rest of your pipeline. This post covers the whole flow with requests and BeautifulSoup, plus how to scrape without getting blocked — or being a bad citizen.
Fetching a page with requests
requests is the standard HTTP library (pip install requests). Three habits from the start: a real User-Agent header (many sites block the default), a timeout (never hang forever), and checking the status code:
import requests
headers = {"User-Agent": "Mozilla/5.0 (price-tracker; contact: you@example.com)"}
try:
r = requests.get("https://example.com/deals", headers=headers, timeout=10)
r.raise_for_status() # raises on 404, 500, ...
except requests.RequestException as e:
print("Fetch failed:", e)
else:
print(r.status_code) # 200
html = r.text # the page source as a string
Naming yourself in the User-Agent with a contact address is basic scraping etiquette — site operators can reach you instead of just blocking you.
Parsing HTML with BeautifulSoup
BeautifulSoup (pip install beautifulsoup4) turns HTML soup into a searchable tree. The two methods you will use constantly are select (all matches) and select_one (first match), both taking CSS selectors:
from bs4 import BeautifulSoup html = """Wireless Mouse
$24.99""" soup = BeautifulSoup(html, "html.parser") for card in soup.select("div.product"): # every product card name = card.select_one("h2.name").get_text(strip=True) price = float(card.select_one("span.price").get_text(strip=True).replace("$", "")) print(name, "->", price) # Wireless Mouse -> 24.99 # USB-C Hub -> 39.5USB-C Hub
$39.50
CSS selectors you already know from web development transfer directly: div.product (class), #main (id), table tr td (descendants), a[href] (attribute). Open the page in your browser's dev tools, right-click an element, "Copy selector" — then simplify it.
Extracting structured data
The goal is never "some HTML" — it is a list of dicts, one per item, ready for a CSV, a database, or the next script. Clean as you extract:
import csv
items = []
for card in soup.select("div.product"):
items.append({
"name": card.select_one("h2.name").get_text(strip=True),
"price": float(card.select_one("span.price").get_text(strip=True).replace("$", "")),
"rating": float(card.select_one("span.rating").get_text(strip=True)),
})
with open("products.csv", "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=["name", "price", "rating"])
writer.writeheader()
writer.writerows(items)
print(f"Wrote {len(items)} products to products.csv")
# Wrote 2 products to products.csv
Scraping responsibly
Scraping is a privilege the site owner grants implicitly. Four rules keep you out of trouble:
- Check robots.txt. It states which paths are off-limits. Python can read it for you:
import urllib.robotparser
rp = urllib.robotparser.RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()
print(rp.can_fetch("my-price-tracker", "https://example.com/deals")) # True = allowed
- Rate-limit yourself. A human reads a page every few seconds; your loop should not fire 100 requests per second. time.sleep(2) between requests is polite and avoids IP bans.
- Identify yourself. The User-Agent with contact info from the first example.
- Cache while developing. Save pages to disk and parse the saved copy — re-fetching the same page 50 times while debugging is how you get blocked.
import time
from pathlib import Path
def fetch_cached(url, cache_file="page.html"):
p = Path(cache_file)
if p.exists(): # develop against the saved copy
return p.read_text()
time.sleep(2) # be polite on the real request
r = requests.get(url, headers=headers, timeout=10)
r.raise_for_status()
p.write_text(r.text)
return r.text
Key takeaways
- requests fetches pages: always set a User-Agent, a timeout, and call raise_for_status().
- BeautifulSoup parses HTML: select / select_one with CSS selectors, get_text(strip=True) for clean text.
- Scrape into a list of dicts — structured data is the whole point; dump it to CSV with csv.DictWriter.
- Scrape responsibly: check robots.txt, sleep between requests, identify yourself, and cache pages while developing.
Next in this series: Scheduling Python Scripts: cron, Task Scheduler, and Email Alerts.
Comments
Post a Comment