Web Scraping with Python: requests and BeautifulSoup

Part 2 of the Python for Automation track. Last updated: September 2026.

Prices change, new listings appear, competitors update their pages — and checking all of it by hand does not scale. Web scraping is automation's eyes: fetch a page, pull out the structured data hiding in the HTML, and hand it to the rest of your pipeline. This post covers the whole flow with requests and BeautifulSoup, plus how to scrape without getting blocked — or being a bad citizen.

Fetching a page with requests

requests is the standard HTTP library (pip install requests). Three habits from the start: a real User-Agent header (many sites block the default), a timeout (never hang forever), and checking the status code:

import requests

headers = {"User-Agent": "Mozilla/5.0 (price-tracker; contact: you@example.com)"}

try:
    r = requests.get("https://example.com/deals", headers=headers, timeout=10)
    r.raise_for_status()          # raises on 404, 500, ...
except requests.RequestException as e:
    print("Fetch failed:", e)
else:
    print(r.status_code)          # 200
    html = r.text                 # the page source as a string

Naming yourself in the User-Agent with a contact address is basic scraping etiquette — site operators can reach you instead of just blocking you.

Parsing HTML with BeautifulSoup

BeautifulSoup (pip install beautifulsoup4) turns HTML soup into a searchable tree. The two methods you will use constantly are select (all matches) and select_one (first match), both taking CSS selectors:

from bs4 import BeautifulSoup

html = """

Wireless Mouse

$24.99 4.5

USB-C Hub

$39.50 4.8
""" soup = BeautifulSoup(html, "html.parser") for card in soup.select("div.product"): # every product card name = card.select_one("h2.name").get_text(strip=True) price = float(card.select_one("span.price").get_text(strip=True).replace("$", "")) print(name, "->", price) # Wireless Mouse -> 24.99 # USB-C Hub -> 39.5

CSS selectors you already know from web development transfer directly: div.product (class), #main (id), table tr td (descendants), a[href] (attribute). Open the page in your browser's dev tools, right-click an element, "Copy selector" — then simplify it.

Extracting structured data

The goal is never "some HTML" — it is a list of dicts, one per item, ready for a CSV, a database, or the next script. Clean as you extract:

import csv

items = []
for card in soup.select("div.product"):
    items.append({
        "name": card.select_one("h2.name").get_text(strip=True),
        "price": float(card.select_one("span.price").get_text(strip=True).replace("$", "")),
        "rating": float(card.select_one("span.rating").get_text(strip=True)),
    })

with open("products.csv", "w", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "price", "rating"])
    writer.writeheader()
    writer.writerows(items)
print(f"Wrote {len(items)} products to products.csv")
# Wrote 2 products to products.csv

Scraping responsibly

Scraping is a privilege the site owner grants implicitly. Four rules keep you out of trouble:

  1. Check robots.txt. It states which paths are off-limits. Python can read it for you:
import urllib.robotparser

rp = urllib.robotparser.RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()
print(rp.can_fetch("my-price-tracker", "https://example.com/deals"))  # True = allowed
  1. Rate-limit yourself. A human reads a page every few seconds; your loop should not fire 100 requests per second. time.sleep(2) between requests is polite and avoids IP bans.
  2. Identify yourself. The User-Agent with contact info from the first example.
  3. Cache while developing. Save pages to disk and parse the saved copy — re-fetching the same page 50 times while debugging is how you get blocked.
import time
from pathlib import Path

def fetch_cached(url, cache_file="page.html"):
    p = Path(cache_file)
    if p.exists():                      # develop against the saved copy
        return p.read_text()
    time.sleep(2)                       # be polite on the real request
    r = requests.get(url, headers=headers, timeout=10)
    r.raise_for_status()
    p.write_text(r.text)
    return r.text

Key takeaways

  • requests fetches pages: always set a User-Agent, a timeout, and call raise_for_status().
  • BeautifulSoup parses HTML: select / select_one with CSS selectors, get_text(strip=True) for clean text.
  • Scrape into a list of dicts — structured data is the whole point; dump it to CSV with csv.DictWriter.
  • Scrape responsibly: check robots.txt, sleep between requests, identify yourself, and cache pages while developing.

Next in this series: Scheduling Python Scripts: cron, Task Scheduler, and Email Alerts.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number