Spark Core: RDDs, DataFrames & the Java API
Post 1 drew the map and gave you a four-line Spark smoke test. Now it's time to earn those four lines. That nightly revenue report — the one that takes four hours and times out — is a batch job: read a few hundred gigabytes of orders, filter, group, sum, write. This post is the machinery underneath it: the RDD, the unit Spark thinks in; laziness, the reason your program does nothing until the last line; partitions, the unit it scales in; and caching, the one performance lever that is entirely your responsibility. Every claim with a transcript below ran in this track's lab — Spark 3.5.3, local[2] , Temurin JDK 17 — and the two sections that couldn't run here are labeled honestly instead of faked. Lab honesty, up front. Everything in the RDD sections ran for real in spark-shell --master "local[2]" on this track's lab machine: the toDebugString lineage, the laziness demonstration, the partition counts with their stage-tracker lines, and the caching experim...