Posts

Capstone: End-to-End Pipeline

Ten posts ago you couldn't draw the data platform. Now you're going to run one. This capstone wires every row of the platform diagram from post 1 into a single working pipeline: a stream of e-commerce order events flows through ingest → validate → transform → store → serve , and every stage runs for real on this lab machine — the event generator, the validator with its dead-letter queue, a Spark job with a genuine shuffle, a Parquet file with inspectable internals, and two serving systems answering the two questions they're built for. By the end you'll have watched the same numbers come out of three different systems and agree with each other. That agreement is the whole game. Lab honesty, up front. Everything shown as a run below actually ran: the generator (5,000 events), the validator (113 rejects), the Spark RDD transform in local mode (real reduceByKey shuffle), the Parquet write/read with footer inspection, the Redis serving demo (Jedis), and the ClickHouse ...

Redis & ClickHouse: Storage & Serving

Post 1 opened with a scene: a 2 TB Postgres, a four-hour revenue report, and a product manager asking for "a dashboard that updates every minute." The track since then has built the pipeline that feeds that dashboard — Kafka carrying events, Spark crunching them. But a pipeline that nobody can read from at speed is a write-only achievement. The dashboard has two questions it must answer fast: "what are this user's features right now ?" (milliseconds, millions of times a day) and "how did revenue trend by country this quarter?" (billions of rows, seconds). One system cannot answer both well. This post gives you the two that can: Redis for the millisecond question, ClickHouse for the billion-row question. The Data & Messaging track's "Caching in Java" post already covered Redis-as-cache — cache-aside, eviction, TTLs, thundering herds. We won't repeat any of that. This post goes deeper into Redis-as- serving-layer : its data str...

HBase: Strong Consistency on Hadoop

Your team already runs Hadoop — HDFS holds the lake, Spark crunches it nightly (posts 2–5), and Cassandra from the last post absorbs the write firehose. Then product drops the requirement that breaks the plan: the user-profile service needs millisecond point reads on two billion profiles, and every read must reflect the latest write. Not eventually. Now. Cassandra can be tuned toward strong consistency, but the team asks a sharper question: is there a store that gives strong row-level consistency by construction , on the HDFS cluster we already operate? That's HBase — Google's BigTable design, rebuilt on Hadoop. And it comes with a catch that defines the entire post: in HBase, the row key is the schema, and a bad one doesn't just slow you down — it funnels your whole cluster's writes through a single machine. Lab honesty, up front. The labs in this post are marked [NOT RUN] : this sandbox blocks raw TCP from Java processes, so HBase's JVM daemons can't com...

Cassandra: Wide-Column at Scale

Your clickstream pipeline is humming: Kafka carries the events, your streaming job counts them, and product just asked for the obvious next feature — "recent activity" on the profile page. The query is simple: give me the last 50 events for user X , served in milliseconds, while the write firehose never stops. Your Postgres can answer it today with an index on (user_id, event_time) — and at ten times the write volume, with a second data center in the plan, you start pricing out what that index costs you on every single insert. This post introduces the store built for exactly that shape of work: Apache Cassandra , the wide-column database that trades ad-hoc querying for write throughput and scale. But Cassandra's real lesson isn't a new query language — it's a new design habit. In a relational database you model the entities and query however you like afterward. In Cassandra you do the opposite: you model the query first, and the table follows. Get that backw...