Automate the Boring Stuff: Files, Folders, and Bulk Renaming with Python
Part 1 of the Python for Automation track. Last updated: September 2026.
Every developer has a downloads folder that looks like a crime scene: installers next to PDFs next to 400 screenshots named Screenshot (47).png. Renaming and sorting those files by hand is exactly the kind of repetitive work computers exist to do. This post builds your first automation: a script that organizes any messy folder for you, plus the file-handling toolkit every later post in this track will use.
pathlib: the modern way to handle paths
Forget os.path.join and string concatenation. pathlib treats paths as objects, and the / operator joins them:
from pathlib import Path home = Path.home() # /home/you (or C:\Users\you) downloads = home / "Downloads" # the / operator joins paths print(downloads.exists()) # True print(downloads.is_dir()) # True print(downloads.suffix) # '' — it's a folder, not a file f = downloads / "report.pdf" print(f.name) # 'report.pdf' print(f.stem) # 'report' print(f.suffix) # '.pdf' print(f.parent) # .../Downloads
Paths are objects with useful properties — .name, .stem, .suffix, .parent — so you stop parsing filenames with string splits.
Finding files with glob patterns
glob matches files by pattern; rglob searches recursively through subfolders:
from pathlib import Path
downloads = Path.home() / "Downloads"
pdfs = list(downloads.glob("*.pdf")) # PDFs directly in Downloads
print(len(pdfs))
all_images = list(downloads.rglob("*.jpg")) # every .jpg, any depth
all_images += list(downloads.rglob("*.png"))
big_files = [f for f in downloads.iterdir() # files over 100 MB
if f.is_file() and f.stat().st_size > 100_000_000]
for f in big_files:
print(f.name, f"{f.stat().st_size / 1e6:.0f} MB")
Project: organize the downloads folder
The classic automation. Sort every file into a subfolder by type, and never overwrite an existing file — rename to report-1.pdf, report-2.pdf instead:
from pathlib import Path
EXT_MAP = {
".pdf": "Documents", ".txt": "Documents", ".docx": "Documents",
".jpg": "Images", ".jpeg": "Images", ".png": "Images", ".gif": "Images",
".mp3": "Audio", ".wav": "Audio",
".zip": "Archives", ".tar": "Archives", ".gz": "Archives",
".py": "Code", ".js": "Code",
}
def organize(folder: Path) -> int:
moved = 0
for f in folder.iterdir():
if not f.is_file():
continue
dest_dir = folder / EXT_MAP.get(f.suffix.lower(), "Other")
dest_dir.mkdir(exist_ok=True)
target = dest_dir / f.name
i = 1
while target.exists(): # avoid overwriting
target = dest_dir / f"{f.stem}-{i}{f.suffix}"
i += 1
f.rename(target)
moved += 1
return moved
if __name__ == "__main__":
n = organize(Path.home() / "Downloads")
print(f"Organized {n} files.")
# Organized 7 files.
Bulk renaming
Camera dumps like IMG_20260101_120000.jpg are unreadable. Rename a whole set in one loop — sorted() keeps the numbering in shooting order:
from pathlib import Path
photos = Path("/tmp/autotest/rename_demo")
for i, p in enumerate(sorted(photos.glob("IMG_*.jpg")), start=1):
new = p.with_name(f"vacation-2026-{i:02d}.jpg")
print(f"{p.name} -> {new.name}")
p.rename(new)
# IMG_20260101_120000.jpg -> vacation-2026-01.jpg
# IMG_20260102_130000.jpg -> vacation-2026-02.jpg
# IMG_20260103_140000.jpg -> vacation-2026-03.jpg
Moving, copying, and deleting safely
shutil handles the heavy operations — it works across drives, where Path.rename can fail:
import shutil
from pathlib import Path
src = Path("draft.docx")
shutil.copy2(src, Path("backup") / src.name) # copy2 preserves timestamps
shutil.move("old-report.pdf", Path("Archives") / "old-report.pdf")
One safety rule: never use os.remove in an automation script until you are certain — a bug deletes real files with no undo. During development, install send2trash (pip install send2trash) so deletes go to the recycle bin instead:
# pip install send2trash
from send2trash import send2trash
send2trash("temp-notes.txt") # goes to trash — recoverable
The reusable script: argparse
A script you run by hand should take arguments instead of hardcoding paths. argparse gives you a real command-line interface in ten lines:
import argparse
from pathlib import Path
from organize import organize # the function from above, saved as organize.py
parser = argparse.ArgumentParser(description="Sort a messy folder by file type.")
parser.add_argument("folder", type=Path, help="folder to organize")
parser.add_argument("--dry-run", action="store_true",
help="show what would move, without moving anything")
args = parser.parse_args()
print(f"Would organize: {args.folder}" if args.dry_run else f"Organizing {args.folder}...")
# python organizer.py ~/Downloads --dry-run
--dry-run is the professional habit: every destructive automation should have a mode that only reports what it would do.
Key takeaways
- pathlib beats os.path: paths are objects, / joins them, .suffix / .stem parse them.
- glob("*.pdf") finds files by pattern; rglob searches subfolders too.
- Bulk operations are loops over iterdir() + rename; shutil.move/copy2 works across drives.
- Never overwrite silently — rename to name-1.ext instead; never hard-delete in a script, use send2trash.
- Give scripts argparse arguments and a --dry-run mode before you trust them.
Next in this series: Web Scraping with Python: requests and BeautifulSoup.
Comments
Post a Comment