Posts

Showing posts with the label Java-DataEng

Spark Core: RDDs, DataFrames & the Java API

Post 1 drew the map and gave you a four-line Spark smoke test. Now it's time to earn those four lines. That nightly revenue report — the one that takes four hours and times out — is a batch job: read a few hundred gigabytes of orders, filter, group, sum, write. This post is the machinery underneath it: the RDD, the unit Spark thinks in; laziness, the reason your program does nothing until the last line; partitions, the unit it scales in; and caching, the one performance lever that is entirely your responsibility. Every claim with a transcript below ran in this track's lab — Spark 3.5.3, local[2] , Temurin JDK 17 — and the two sections that couldn't run here are labeled honestly instead of faked. Lab honesty, up front. Everything in the RDD sections ran for real in spark-shell --master "local[2]" on this track's lab machine: the toDebugString lineage, the laziness demonstration, the partition counts with their stage-tracker lines, and the caching experim...

HDFS & MapReduce: The Foundations

Picture the nightly revenue report from post 1 — one big SQL query that scans 2 TB and times out. Now imagine it doesn't live in Postgres at all. The orders live in files on a hundred cheap machines, and instead of one database grinding through them, a hundred small programs each read their own slice and combine answers. That's the idea this post is built on — and the first thing you need to understand about it is where the bytes actually sit . Everything in data engineering starts with storage, and the storage this whole industry grew out of is HDFS. HDFS and MapReduce are old technology — HDFS from 2006, MapReduce from the 2004 Google paper. You will likely never write a production MapReduce job. But here's why this post exists: every tool you'll meet later leaks this foundation. Spark's shuffle is MapReduce's shuffle. Partitioning schemes only make sense once you've watched a hot key stall a reducer. And when a Spark job mysteriously reads slowly fro...

Data Engineering on the JVM: The Landscape

Your Spring Boot service from the Data & Messaging track is doing well. Too well: the Postgres holding its orders just crossed 2 TB, the nightly revenue report — one big SQL query — now takes four hours and times out half the time, and product just asked for "a dashboard that updates every minute." Someone in the meeting says "we need a data platform." Everyone nods. Nobody can draw what that means. This post draws the map. Data engineering on the JVM is a crowded territory — Hadoop, Spark, Flink, Cassandra, HBase, Redis, ClickHouse, Kafka — and every tool's marketing claims it does everything. It doesn't. Each one was built for one shape of work, and choosing wrong is the most expensive mistake in this space: a batch engine forced into real-time, or a serving store forced into analytics, fails slowly and expensively. By the end of this post you'll know what each piece does, why the JVM ended up running nearly all of it, and which tool fits which...