Top 40 System Design Interview Questions and Answers (2026 Edition)
System design interviews test judgment, not memorization: can you scope a vague problem, size it, draw a sane architecture, and defend your trade-offs — including what happens when things fail? These are the 40 questions that come up most often — each with a crisp, interview-ready answer. Every answer links to a full worked design in this series when you want the end-to-end walkthrough.
How to use this: read the question, say your answer out loud, then check. Interviewers score decision rules and trade-offs, not definitions — "I'd pick X because reason, accepting cost." And the meta-rule for every answer below: don't just name a technology. Say what property the requirement needs, which mechanism provides it, at what cost — and how the system behaves when that mechanism fails.
Framework & Estimation
1. What's the 4-step framework for a 45-minute system design round?
Answer: Clarify requirements (5 min) — functional vs non-functional, nail the scope. Back-of-envelope estimation (5 min) — QPS, storage, bandwidth. High-level design (15 min) — and this is where strong candidates explicitly identify the API and data model: what data exists, what operations consumers need, how they interact — before drawing boxes. Deep dive (15 min) — the one hard part the interviewer probes. The trap: jumping to components before clarifying scope — that's how you design the wrong system confidently. Deep dive →
2. How do you do back-of-envelope estimation?
Answer: convert DAU to QPS with the rule of thumb: 1M requests/day ≈ 12/sec average. Then estimate peak — but state your assumption out loud instead of memorising a multiple: "I'll assume peak is 3x average for this consumer product; for this B2B tool I'd assume 1.5x." Peak/average ratio is workload-dependent. Also split read QPS from write QPS early — that split is what drives most architecture decisions. Storage: 1M objects × 1KB = 1GB; bandwidth = QPS × object size. The point isn't precision — it's showing you size the problem before designing. Deep dive →
3. Why do interviewers insist you clarify requirements first?
Answer: real systems fail on misunderstood requirements, not missing components. Ask what the core features are (functional) and what scale, latency, and consistency the system needs (non-functional). The trap: designing for a billion users when the ask was an MVP — clarification sets the scale you're actually designing for. Deep dive →
4. How do you talk through a trade-off without rambling?
Answer: use one sentence shape every time: "I'd pick X here because [reason], accepting [cost]." Example: "I'd pick a NoSQL store for the URL mappings because the access pattern is pure key lookup at high write volume, accepting the loss of ad-hoc joins — which this system never needs." Interviewers score the explicit cost acceptance, not the choice itself. Deep dive →
5. What latency numbers should you know, and how should you use them?
Answer: don't memorise magic constants — hardware and topology vary too much. Think in orders of magnitude: CPU/cache → memory → local SSD → network/service call → cross-region communication, each step roughly 10–1000x slower than the last. The interview lesson: every network boundary and every storage boundary costs latency — know where your hot path crosses them. That single insight justifies every cache, every co-location decision, and every "why is this slow?" diagnosis you'll ever give. Deep dive →
Caching & Scalability
6. Where do you put caches in a system, and how do you invalidate them?
Answer: cache at every layer that repeats work: a CDN for static assets, then an application cache (Redis-style) in front of read-heavy data. Invalidate with TTLs where some staleness is acceptable, and explicit invalidation on write where data must be fresh. The trap: proposing a cache with no invalidation story — stale-data bugs are the classic follow-up. Deep dive →
7. Cache-aside vs read-through vs write-through — which, when?
Answer: cache-aside (the app checks the cache, falls back to the DB, and populates the cache) is the most common — simple, and only hot data ever gets cached. Read-through moves the fill logic into the cache layer itself. Write-through updates the cache as part of the write path, which reduces stale-cache windows — but don't claim it buys system-wide read-after-write consistency; that depends on cache topology, replicas, concurrency, and failure semantics. It costs write latency and more complicated failure handling. Default to cache-aside unless you genuinely need the tighter staleness bound. Deep dive →
8. How do you handle a cache stampede (thundering herd)?
Answer: when a hot key expires, thousands of requests hit the database at once. Mitigate with request coalescing (only one request rebuilds the key), probabilistic early expiry (refresh before the TTL under load), or stale-while-revalidate. The interview signal: you know that caches create their own failure modes. Deep dive →
9. "How do you scale to 1M requests per second?" — what's the senior first sentence?
Answer: "1M requests per second of what?" — 1M cached GETs, 1M payment writes, and 1M video segments are completely different systems. Clarify first: read/write mix → payload size → latency SLO → geographic distribution → consistency requirements. Then architect: stateless app servers behind a load balancer, read replicas and caching for reads, sharding for writes, CDN for static content, async queues to buffer spikes. Say which layer handles which order of magnitude — that mapping is the senior answer. Deep dive →
10. What happens when your cache goes down?
Answer: decide up front: fail open (serve from the DB, alert loudly) if availability matters more, or fail closed (reject traffic) if the database can't survive the load and protection matters more. There is no universally right answer — the interviewer wants the explicit trade-off. The follow-up they love: "your DB melts under the stampede — now what?" Deep dive →
Databases & Consistency
11. SQL vs NoSQL — how do you decide?
Answer: don't start from "SQL vs NoSQL" — start from the workload: access patterns, transaction boundaries, consistency requirements, query flexibility, data volume, and scale. Then the technology becomes the consequence: complex joins and multi-row transactions point relational; simple key lookups at extreme write volume point elsewhere. Modern distributed SQL and cloud databases blur the old binary anyway. The trap: picking NoSQL for "scale" when the workload is 10k writes a day with heavy joins — scale you don't have doesn't justify the trade-off. Deep dive →
12. Replication vs sharding — which solves which problem?
Answer: replication copies data to more nodes — primarily for availability, durability, and redundancy, and depending on the architecture it can also scale reads. Partitioning/sharding divides ownership of the dataset across nodes — that's what you reach for when one node can't handle the storage or write workload, at the cost of painful cross-shard queries. Know which solves which; mixing them up is the classic wrong answer. Deep dive →
13. How do you pick a shard key — and what is consistent hashing actually for?
Answer: these are two separate decisions. The shard key decides logical data ownership and locality — pick one that distributes writes evenly and matches your query pattern (e.g., sharding by customer_id keeps one customer's data together). The trap is the hot shard: one huge tenant funnels all writes to a single shard. Consistent hashing is a different technique: a way to map keys onto a changing set of nodes while minimising remapping when nodes join or leave. A sharded relational database doesn't necessarily use it — know which problem each solves. Deep dive →
14. Explain the CAP theorem the way you'd actually use it in an interview.
Answer: CAP matters when a network partition occurs — and you apply it per operation, not per application. For the operation you're designing, decide: is it safer to reject/delay requests, or to continue and risk stale or divergent state? A like count being seconds stale is acceptable; an account balance going stale or two writers conflicting on inventory may not be — it depends on oversell tolerance. Don't brand an entire application "AP". Useful later: PACELC — even without a partition, distributed systems trade latency against stronger consistency. Deep dive →
15. How do you store images and videos at scale?
Answer: for large media objects at scale, prefer object storage (S3-style) for the bytes — effectively infinite, cheap, durable — with a CDN in front for delivery, and keep only metadata and URLs in the database. The standard follow-up deep dive is the media pipeline: upload → transcode → thumbnails → CDN. Deep dive →
16. Two users try to buy the final item simultaneously — how do you prevent overselling?
Answer: this forces you past generic "use strong consistency." Name the business invariant (inventory can't go negative), then the mechanism that protects it: an atomic conditional update ("decrement where stock > 0"), a transaction with proper isolation, or serialisation of the contended operation. Say what breaks if you get it wrong — and what the mechanism costs in throughput. Interviewers ask this to see whether you reason about invariants, not just recite consistency levels. Deep dive →
Messaging, Retries & Async
17. When do you put a message queue in a design?
Answer: use asynchronous messaging when work can happen after the request — especially when you need decoupling, buffering, retries, or independent consumers (notifications, analytics, search indexing). Not every background operation warrants a queue; "anything not needed for the response" is too broad. And teach the cost honestly: eventual consistency + operational complexity + duplicate or reordered delivery. The real interview question is "why a queue rather than a synchronous call?" — answer that, don't just draw one. Deep dive →
18. How do you guarantee exactly-once processing?
Answer: the honest answer: distributed systems realistically offer at-least-once delivery, so you design idempotent consumers — dedupe on a stable event ID so replays are harmless. Two patterns senior candidates name: the idempotency key / stable event ID on every operation, and the transactional outbox — write the state change and the event to the database in one transaction, then have a relay publish the event, so a DB commit can never succeed while its event is lost. The sentence that scores: "I don't rely on the broker alone for business exactly-once; I make the operation idempotent and make state-change plus event publication reliable." Deep dive →
19. How do you make a write API safe to retry?
Answer: the client sends an idempotency key (a stable operation ID) with the request; the server stores the key with the result of the first execution and returns the stored result on replay instead of re-executing. Dedupe on the key, expire keys after a sensible window. This is more valuable than another cache question — payment and order APIs live or die on it. The trap: retrying a non-idempotent write and charging the customer twice. Deep dive →
20. How do you prevent a database update succeeding while its event fails to publish?
Answer: the transactional outbox pattern: in the same database transaction that applies the state change, insert the event into an outbox table. A separate relay process reads the outbox and publishes to the message broker, marking events sent. If the relay crashes, it replays from the outbox — the event can never be lost while the commit succeeded. This is the bridge between database state and message publication that makes Q18's guarantees real. Deep dive →
21. How do you design a rate limiter?
Answer: the algorithm to name is token bucket: each client has a bucket refilled at a fixed rate, and each request costs a token. Put the limiter at the API gateway or as middleware. For the state, you need a shared, coordinated home across API instances — a Redis-style shared store is one common implementation, not architecture law. Then say what happens when that store is unavailable (fail open? fail closed?) — that's the follow-up that separates vocabulary from judgment. At real scale, shard the limiter state by client key, because the limiter itself becomes the bottleneck. Deep dive →
22. Fixed window vs sliding window vs token bucket vs leaky bucket — how do they differ?
Answer: know all four and their trade-offs: fixed window — simplest, but suffers boundary bursts (double the limit across a window edge). Sliding window — smoother and more accurate, but keeps more state. Token bucket — allows controlled bursts while enforcing the average rate; a common choice when bursts are legitimate. Leaky bucket — smooths output to a steady rate. Don't universally "recommend" one — match the algorithm to the traffic shape: bursty-but-bounded → token bucket; must-never-exceed → leaky bucket or sliding window. Deep dive →
Reliability & Operations
23. Traffic spikes to 10x normal — how does your system survive?
Answer: layer the defenses: autoscale the stateless tier, let queues buffer what can't be served immediately, shed low-priority load and degrade gracefully (serve cached or stale content), and rate-limit abusive clients. The interview signal: you think in graceful degradation, not just "add more servers." Deep dive →
24. What happens when a downstream service gets slow — not down, just slow?
Answer: slow is worse than down, because callers pile up waiting. Defend in layers: timeouts on every outbound call, bounded retries with exponential backoff + jitter (jitter prevents retry storms from synchronising), then a circuit breaker that stops calling a failing dependency, plus load shedding to protect yourself. The senior insight: retries can amplify an outage — unbounded retries turn one slow dependency into a cascading failure. Always bound them. Deep dive →
25. What is backpressure, and where do you apply it?
Answer: backpressure is the system pushing back on producers when consumers can't keep up — instead of accepting unbounded work and dying. Apply it at every handoff: bounded queues (reject or block when full), bounded concurrency (fixed worker pools), throttling or rejecting producers (HTTP 429, client-side rate limits), and watch consumer lag as the signal. The alternative to backpressure is an OOM kill at 3 AM — say that, and the interviewer knows you've operated something. Deep dive →
26. How do you design for a regional outage?
Answer: first name the model: active/passive (one region serves, another stands by — simpler, but failover takes time and the standby must actually work) vs active/active (multiple regions serve — better availability, but now you own cross-region replication, conflict handling, and routing). Then the mechanics: replicated data with a defined RTO/RPO, health-checked DNS or traffic failover, and an honest statement of consistency consequences during failover. The trap: claiming "multi-region" without saying what happens to in-flight writes when you flip. Deep dive →
27. RTO vs RPO — what do they change in your design?
Answer: RTO (recovery time objective) = how long you can be unavailable; RPO (recovery point objective) = how much data you can afford to lose. They turn "design for DR" from hand-waving into engineering: RPO of zero means synchronous replication (and its latency cost); RPO of minutes allows async replication; RTO of minutes demands automated failover and practiced runbooks, while RTO of hours tolerates manual steps. Ask for them — or state your assumptions — before designing anything "highly available." Deep dive →
28. How do you know the system is healthy?
Answer: most candidates design systems and never discuss operating them — this question fixes that. Define SLIs (what you measure: latency, errors) and SLOs (the targets), track the four golden signals — latency, traffic, errors, saturation — plus queue lag and dependency health. Build dashboards around user-visible symptoms, alert on SLO burn (not on every CPU spike), and make sure every alert links to a runbook. The senior sentence: "If nobody would do anything different at 3 AM, it's not a page." Deep dive →
Worked Design Prompts
29. How do you design a URL shortener?
Answer: API: POST /shorten returns a short key; GET /{key} redirects — and here's the classic follow-up: 301 vs 302 is a trade-off, not a memorised answer. A permanent (301) redirect can be cached, reducing server traffic, but makes changing the destination and collecting analytics harder; a temporary (302) redirect keeps control at the cost of more hits. Discuss it; don't hardcode it. Generate keys from a counter encoded in base62 (or a pre-generated key pool), store the mapping in a NoSQL store keyed by short URL, cache hot mappings. Keep analytics off the synchronous redirect path — click events go to a queue. Deep dive →
30. How do you generate unique short keys at scale?
Answer: the modern options: random/base62 IDs (simple, unpredictable), distributed ID generation (Snowflake-style: timestamp + worker ID + sequence), or allocated ranges handed to app servers — plus hash-with-collision-handling where appropriate. A single counter plus base62 is simplest and collision-free, but the counter is a bottleneck. State the real trade-off: coordination cost vs collision probability vs predictability/security (sequential IDs are guessable — fine for some products, not for private links). Deep dive →
31. How do you design a social media news feed?
Answer: fan-out on write: when a user posts, push the post ID into each follower's precomputed feed, stored in cache or the database. Reads then avoid the database on the hot path — they page through the precomputed feed. The hard part isn't the happy path; it's the celebrity problem. Deep dive →
32. Fan-out on write vs fan-out on read — when do you use each?
Answer: fan-out on write gives fast reads but explodes on celebrities; fan-out on read computes the feed at request time — it handles celebrities naturally but makes every read slower. The experienced answer is hybrid, tuned by follower count, activity, and cost — there's no single threshold where one flips to the other. Normal users fan-out on write; high-follower accounts fan-out on read; the boundary moves with your traffic and budget. Deep dive →
33. A celebrity posts during the Super Bowl — what breaks, and how do you fix it?
Answer: the fan-out write storm breaks — conceptually, one post to 100M followers means up to 100M feed writes (treat that as illustrative worst-case arithmetic; production systems batch, reference, rank, and distribute it). Fix it by routing celebrity accounts to fan-out on read, and queue the push with backpressure so the write spike is absorbed instead of crashing the system. This is the single most-asked social-feed follow-up. Deep dive →
34. How do you keep feed reads fast at very high read volume?
Answer: precomputation moves the expensive fan-out and ranking work off the request path, so reads are served largely from prepared, cacheable state — paginated, without the database on the hot path. (Real feed assembly still involves feed IDs, post hydration, ranking, privacy filtering, and pagination — "a single cache lookup" is too magical.) If the interviewer pushes ("cache is cold?"), fall back to fan-out on read with tight limits — degraded but alive. Deep dive →
35. What consistency does a social feed actually need?
Answer: different interactions need different guarantees — that's the real lesson, not "social feed = eventual consistency." General feed browsing: eventual consistency is usually fine. Your own new post: read-your-writes — users absolutely notice when their post doesn't appear. Like counts: approximate/stale is usually acceptable. Privacy changes: strong, fast propagation — a deleted post must disappear now. Match the guarantee to what users can perceive and what the business requires. Deep dive →
36. How do you add click analytics to a URL shortener without slowing redirects?
Answer: keep analytics off the synchronous critical path — the principle is that redirect availability must never depend on the analytics backend. Emit a click event to a message queue on each redirect and aggregate from the stream (streaming aggregation is fine; it doesn't have to be batch). The redirect stays fast; analytics lands shortly after, which is fine for dashboards. Deep dive →
37. Malicious users are creating millions of short links — what do you do?
Answer: that's what the rate limiter is for — the two designs compose. Apply quotas and rate limits by authenticated principal first; IP is one additional signal, not identity (NATs and proxies mean many users share IPs). Require auth or a CAPTCHA for anonymous creation, and set quotas. The interview signal: you connect the two systems instead of treating abuse as a separate problem. Deep dive →
38. How do you design a notification system?
Answer: a great modern design prompt because it exercises everything: ingest events → check user preferences (channels, quiet hours) → fan-out across push/email/SMS → retry with backoff when providers fail → dedupe so a retry storm doesn't spam users → schedule and rate-limit sends. Decide per channel: push can be best-effort, but an OTP SMS failing needs escalation. And name the failure modes: provider outage, duplicate delivery, preference-store staleness. Deep dive →
39. How do you design real-time chat?
Answer: the shape: long-lived connections (WebSockets) from clients to gateway servers, presence tracking who's online, messages routed to the recipient's connection or stored for offline users, ordering per conversation (sequence numbers), and delivery/read state. Partition conversations across servers; replicate for availability. The follow-ups interviewers love: "two devices, one user — where does the message go?" and "how do you show 'typing…' without melting the servers?" Deep dive →
40. How do you evolve an API or schema without breaking old clients?
Answer: backward compatibility as a discipline: only additive changes (new optional fields, new endpoints), never rename or remove in place. For breaking changes: expand → migrate → contract — ship the new shape alongside the old, move writers, then move readers, then delete. Version explicitly where the contract demands it, and use dual reads/writes during migrations where necessary. This is daily FDE work: production systems are never greenfield. Deep dive →
In this series
- Top 40 System Design Interview Questions and Answers (2026 Edition) (this post) — start here.
- System Design Interviews: A Practical Primer (With a 4-Step Framework) — the 4-step framework.
- System Design: URL Shortener and Rate Limiter, End to End — the classic warm-up questions.
- System Design: How to Design a Social Media Feed — fan-out, ranking, and the celebrity problem.
Related: System Design Interviews: A Practical Primer — the 4-step framework these questions plug into.
Comments
Post a Comment