Posts

Streaming: Spark Structured Streaming

The nightly revenue report from the first post in this track now finishes in twenty minutes — a win. But product's "dashboard that updates every minute" request is still open, and yesterday's incident made it urgent: a pricing bug overcharged customers for three hours before anyone noticed, because the batch job only runs at midnight. "We need to see what's happening now ," the on-call engineer says. "Not what happened yesterday." That's the streaming question — "what's happening?" — and this post answers it with the engine you already know. Spark Structured Streaming takes the DataFrame API from posts 3–4 and points it at data that never ends: instead of learning a separate streaming API, you write what looks like a batch query, and the engine executes it incrementally as new rows arrive. By the end you'll understand the mental model, the one interview concept that matters (event time vs processing time), watermarks,...

File Formats & the Lakehouse

Your nightly revenue export is a CSV. It was fine when it was 40 MB. Now it's 900 GB, it takes an hour to download, and the analyst's "quick question" — sum revenue by region — reads the entire file to touch one column. Someone suggests Parquet. Someone else says "just put it in the data lake." A third person says "Delta Lake," which sounds like a place you'd go fishing. This post settles all three. You'll learn how Parquet actually stores data (column chunks, pages, encodings — not hand-waving), measure real compression differences between SNAPPY, GZIP, and no compression, watch dictionary encoding kick in on a low-cardinality column and fall back on a unique one, build a hive-partitioned directory tree, and then run a real Delta Lake table — create, append, schema evolution, time travel — with plain Java and no Spark engine. By the end you'll know what a lakehouse is made of, and why "the file format" turned out to be on...

Spark SQL Deep Dive

Back in post 1, the nightly revenue report scanned a 2 TB orders table and took four hours. Here's the part nobody said in that meeting: the report selects three columns and filters on last week's dates. The query asks for a thimbleful; the engine drinks the ocean. Most of those four hours isn't computing revenue — it's reading bytes the answer never needed. This post is about not reading the ocean. Spark SQL's speed on analytical queries comes less from clever computation than from aggressive avoidance : an optimizer (Catalyst) that rewrites your query to skip data, a file format (Parquet) that tells the reader exactly which chunks can be skipped, and a partitioning model that decides how the remaining work splits across machines. You'll learn the four stages every query passes through, what predicate pushdown and column pruning actually do and why they matter, how to read EXPLAIN output like a plan instead of a wall of text, and why partition keys and ske...