System Design Interviews: A Practical Primer (With a 4-Step Framework)

System design rounds decide senior offers — and most candidates fail them without writing a single wrong answer. They fail on structure. This primer gives you a 4-step framework that works for any design question, the back-of-envelope math to memorize, and the building blocks interviewers expect you to reach for.

What interviewers actually score

There is no "correct" architecture. Interviewers score four things:

  • Structured thinking — do you break a fuzzy problem into pieces in a sensible order, or jump straight to "use Kubernetes"?
  • Trade-off reasoning — every choice has a cost. "We pick X because Y, accepting Z" is the sentence that earns points.
  • Technical breadth — do you know the standard building blocks (below) and when each one earns its place?
  • Communication — think out loud. A silent genius and a stuck candidate look identical for 45 minutes.

Notice what is not on the list: naming a specific cloud product, or drawing the "perfect" diagram.

The 4-step framework (for a 45-minute round)

Step 1 — Clarify requirements (5 minutes)

Never start designing. Ask:

  • Functional: what must the system do? (e.g. "shorten URLs and redirect" — but do we need custom aliases? analytics? expiry?")
  • Non-functional: scale (DAU? reads vs writes?), latency expectations, availability target.
  • Scope: "Should I cover the analytics dashboard, or focus on the core shorten/redirect path?" Interviewers love this question — it shows judgment.

Step 2 — Back-of-envelope estimation (5 minutes)

Turn vague scale into numbers. You only need a few conversions (memorize the table below):

  • Requests/sec from daily active users.
  • Storage from (objects × average size).
  • Bandwidth from (requests × payload size).

Example: "10 million short URLs created per day ≈ 116 writes/sec average, maybe 300/sec at peak. At 500 bytes of metadata each, that's ~5 GB/day, ~1.8 TB/year." You just proved you think in numbers, not adjectives.

Step 3 — High-level design (15 minutes)

Draw boxes and arrows: client → load balancer → app servers → database, plus cache and CDN where they obviously help. Sketch the core API:

POST /api/shorten   { "longUrl": "...", "customAlias": "optional" }
                     -> { "shortUrl": "jvmk.us/aB3x9" }
GET  /{shortKey}     -> 302 redirect to the long URL

Keep it simple. You will refine it in step 4 — that is the point.

Step 4 — Deep dive (15 minutes)

Pick the 2–3 hardest parts and go deep: the bottleneck, the scaling story, the failure modes. "What breaks first when traffic 10x's? The single database — so we add read replicas, then shard by short-key hash." This is where senior candidates separate from the pack.

Back-of-envelope numbers to memorize

ConversionRule of thumb
Requests/day → per second1M/day ≈ 12/sec average; assume 2–3x at peak
Storage1M objects × 1 KB = 1 GB; × 1 MB = 1 TB
Memory reference~100 nanoseconds
SSD random read~150 microseconds
Disk seek~10 milliseconds
Same-datacenter network round trip~0.5 milliseconds
Cross-region round trip~100–200 milliseconds

The pattern that matters: memory is ~1,000x faster than SSD, which is ~100x faster than disk. That single insight justifies every cache you will ever propose.

The building blocks cheat sheet

You do not need twenty tools. Know these eight and their one-line "use when":

  • Load balancer — spread traffic across app servers; no single point of failure.
  • Cache (Redis/Memcached) — serve hot reads in milliseconds; put it in front of anything read-heavy.
  • CDN — static assets and cacheable content, served from near the user.
  • Relational DB + read replicas — structured data with joins; replicas absorb read scale.
  • NoSQL (key-value / document) — massive scale, simple access patterns, flexible schema.
  • Message queue (Kafka/RabbitMQ) — decouple slow work (emails, analytics) from the request path.
  • Object storage (S3-style) — images, videos, backups; effectively infinite.
  • Search index (Elasticsearch-style) — full-text search the database cannot do well.

The 5 concepts behind 80% of follow-up questions

  1. Caching — where do you cache, and how do you invalidate? (TTL, write-through vs cache-aside.)
  2. Database scaling — replication (copies for reads) vs sharding (splitting writes). Know which solves which.
  3. Consistency — CAP in one line: when the network splits, you choose availability or consistency, not both. Most large systems pick availability and accept eventual consistency.
  4. Async processing — anything not needed for the response goes on a queue.
  5. Rate limiting — protect every public API; token bucket is the algorithm to name.

How to talk through a trade-off (template)

Use this sentence shape every time: "I'd pick X here because reason, accepting cost." Example: "I'd pick a NoSQL store for the URL mappings because the access pattern is a pure key lookup at high write volume, accepting that we lose ad-hoc joins — which this system never needs."

Practice checklist

  • Can you do the QPS/storage/bandwidth math for any DAU number in under 2 minutes?
  • Can you draw the 8 building blocks from memory and state each one's "use when"?
  • Can you explain CAP without jargon, in 30 seconds?
  • Have you talked through one full design out loud, timed at 45 minutes?

If yes to all four, you are ready for the worked examples.

In this series

  1. System Design Interviews: A Practical Primer (this post) — the 4-step framework.
  2. System Design: URL Shortener and Rate Limiter, End to End — the classic warm-up questions.
  3. System Design: How to Design a Social Media Feed — fan-out, ranking, and the celebrity problem.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number