Resilience: Timeouts, Retries, Circuit Breakers & Bulkheads
Two incidents, one night, same on-call engineer.
Incident one, 2:40 AM. The payments downstream hiccuped for forty seconds — a bad deploy on their side, already rolling back. Our checkout service did exactly what its config told it to do: every failed call was retried, five times, immediately, no delay. Four hundred in-flight checkouts times five instant retries turned a forty-second blip into a self-inflicted DDoS. When the downstream came back up, it was greeted by two thousand requests arriving in the same millisecond — the retry backlog plus fresh traffic — and fell over again. This time for twenty minutes. The downstream team fixed their deploy in under a minute; our retry policy kept the outage alive for twenty.
Incident two, 4:15 AM. While the dust settled, finance found 214 customers billed twice for a single order. Same root shape: a charge request timed out after the downstream had committed the payment, the response never made it back, and the retry — doing its job — charged the card a second time. No bug in the retry code. No bug in the charge code. The bug was the combination: a retry policy wrapped around an operation that must never run twice.
Both incidents are one subject wearing two masks: what your service does when a call fails. This post builds the production answer with Resilience4j (version 2.4.0, the current release on Maven Central): timeouts that bound every wait, retries with exponential backoff and jitter, circuit breakers that stop calling what's already down, bulkheads that cap the blast radius, and rate limiters that protect the downstream from you. Every snippet below was compiled with javac and run on OpenJDK 21.0.3, and every output block is the real output — including two findings that surprised me and that I'll flag honestly when we get there.
The one rule this whole post orbits: never retry blindly — first ask whether the operation is idempotent. We'll prove why with a ledger that shows a double charge in real output, then fix it with idempotency keys.
Lab setup: Resilience4j without Maven
No Maven on this machine, so the jars come straight from Maven Central with curl. Resilience4j's modules are separate artifacts — pull the ones this post uses, plus an SLF4J API and a simple binding so the library's logging stays quiet:
$ mkdir -p libs && cd libs
$ V=2.4.0
$ for a in resilience4j-core resilience4j-retry resilience4j-circuitbreaker \
resilience4j-bulkhead resilience4j-ratelimiter resilience4j-timelimiter; do
curl -s -f -O "https://repo1.maven.org/maven2/io/github/resilience4j/$a/$V/$a-$V.jar"
done
$ curl -s -f -O "https://repo1.maven.org/maven2/org/slf4j/slf4j-api/2.0.13/slf4j-api-2.0.13.jar"
$ curl -s -f -O "https://repo1.maven.org/maven2/org/slf4j/slf4j-simple/2.0.13/slf4j-simple-2.0.13.jar"
$ CP=$(ls *.jar | tr '\n' ':')
$ javac -cp "$CP" -d classes src/*.java && java -cp "classes:$CP" RetryDemo
That's the whole build system for this post: curl, javac -cp, java -cp. Keep that $CP variable handy — every demo runs the same way.
The failure you should design for
Not every failure deserves the same response. Before reaching for a pattern, classify what went wrong:
- Fail-fast (bad request, auth denied, validation error): retrying is pointless — the next attempt fails identically. Return the error.
- Fail-transient (timeout, 503, connection reset): the operation might succeed on retry. This is the only category retries are for.
- Fail-slow (downstream hangs, thread pool exhausted, GC storm): nothing failed yet, but everything is late. Timeouts, bulkheads, and circuit breakers are the answer — a retry just adds load to a drowning system.
The five patterns in this post map onto those categories: timeouts bound the wait, retries handle the transient, circuit breakers and bulkheads contain the slow, and rate limiters keep you from becoming someone else's slow. Order matters — we'll stack them at the end.
1. Timeouts: every call gets a deadline
A call with no timeout is a thread you've donated to the downstream forever. If the downstream hangs, your thread hangs; if enough threads hang, your pool is exhausted and your service is down even though your code is fine. The timeout is the cheapest resilience pattern there is: decide, up front, the longest you'll wait.
Resilience4j's TimeLimiter wraps a CompletableFuture-producing call and throws TimeoutException when the budget expires. Here the downstream sleeps 2 seconds; the caller gives up after 500ms:
import io.github.resilience4j.timelimiter.TimeLimiter;
import io.github.resilience4j.timelimiter.TimeLimiterConfig;
import java.time.Duration;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
public class TimeoutDemo {
public static void main(String[] args) throws Exception {
TimeLimiter limiter = TimeLimiter.of(TimeLimiterConfig.custom()
.timeoutDuration(Duration.ofMillis(500)) // give up after 500ms
.cancelRunningFuture(true) // ask the stuck call to stop
.build());
var exec = Executors.newSingleThreadExecutor();
java.util.concurrent.atomic.AtomicLong taskEnd = new java.util.concurrent.atomic.AtomicLong(-1);
long t0 = System.currentTimeMillis();
try {
limiter.executeFutureSupplier(() ->
CompletableFuture.supplyAsync(() -> {
try { Thread.sleep(2000); } // downstream hangs 2s
catch (InterruptedException e) { taskEnd.set(System.currentTimeMillis() - t0); return "interrupted"; }
taskEnd.set(System.currentTimeMillis() - t0);
return "slow-but-eventually-OK";
}, exec));
} catch (java.util.concurrent.TimeoutException e) {
System.out.println("caller: TimeoutException at t=" + (System.currentTimeMillis() - t0) + "ms");
}
Thread.sleep(2500); // wait past the downstream's 2s sleep, then ask when the worker really stopped
System.out.println("worker task actually finished at t=" + taskEnd.get() + "ms");
exec.shutdownNow();
}
}
The caller behaved exactly as configured — TimeoutException at 515ms. But the instrumented worker tells the real story:
caller: TimeoutException at t=515ms
worker task actually finished at t=2007ms
Honest finding #1: cancelRunningFuture(true) did not stop the stuck task. The caller got its TimeoutException at 515ms, but the worker thread slept the full 2 seconds and finished at 2007ms. I verified this at the JDK level too: CompletableFuture.cancel(true) on a supplyAsync task marks the future cancelled but does not interrupt the runner thread (the task only saw an interrupt when the executor itself was shut down). I tested three executor types — same result on all of them.
The lesson is the point of this section: a timeout frees the caller, not the resource. Your request thread moves on at 500ms, but the downstream call — and whatever thread, connection, or socket it holds — keeps running to its own conclusion. Under load, timeouts alone don't bound resource consumption. That's what bulkheads are for (section 4), and it's why the patterns compose.
One more timeout rule before we move on: a timeout must come from the workload's latency budget, never from a default that felt fine. 500ms is right here because the demo's downstream is local; against a real payment gateway with a p99 of 800ms, a 500ms timeout would amputate healthy calls. Measure the downstream's latency distribution first, then set the timeout above its p99 with room to spare — and make it configuration, not a constant.
2. Retries: backoff, jitter, and the thundering herd
A retry without a delay is just the same failure, faster — that's what caused incident one. The fix has two parts: backoff (wait longer between each attempt, so a struggling downstream gets breathing room) and jitter (randomize the wait, so a hundred clients don't all retry in the same millisecond).
The demo below wraps a scripted flaky downstream — it throws on the first three calls, succeeds on the fourth — in a Resilience4j Retry with exponential backoff (base 100ms, ×2 each attempt) and ±50% random jitter:
import io.github.resilience4j.core.IntervalFunction;
import io.github.resilience4j.retry.Retry;
import io.github.resilience4j.retry.RetryConfig;
import java.util.concurrent.atomic.AtomicInteger;
public class RetryDemo {
static final AtomicInteger ATTEMPTS = new AtomicInteger(0);
static long start;
// Scripted flaky downstream: fails the first 3 calls, succeeds on the 4th.
static String flakyCall() {
int n = ATTEMPTS.incrementAndGet();
long at = System.currentTimeMillis() - start;
System.out.printf(" attempt %d at t=%4d ms -> ", n, at);
if (n < 4) { System.out.println("THROWS (503 from flaky-service)"); throw new IllegalStateException("503"); }
System.out.println("succeeds");
return "OK";
}
public static void main(String[] args) {
Retry retry = Retry.of("flaky", RetryConfig.custom()
.maxAttempts(6)
.intervalFunction(IntervalFunction.ofExponentialRandomBackoff(100, 2.0, 0.5))
.retryExceptions(IllegalStateException.class)
.build());
System.out.println("Backoff plan: base=100ms, multiplier=2, jitter=+/-50% (100 -> 200 -> 400 ...)");
start = System.currentTimeMillis();
String r = retry.executeSupplier(RetryDemo::flakyCall);
System.out.println("result: " + r + " in " + ATTEMPTS.get() + " attempts, "
+ (System.currentTimeMillis() - start) + " ms wall clock");
}
}
Two runs of the same program. Watch the waits between attempts — the plan is identical, but the jitter makes every run different:
$ java -cp "classes:$CP" RetryDemo
Backoff plan: base=100ms, multiplier=2, jitter=+/-50% (100 -> 200 -> 400 ...)
attempt 1 at t= 3 ms -> THROWS (503 from flaky-service)
attempt 2 at t= 150 ms -> THROWS (503 from flaky-service)
attempt 3 at t= 338 ms -> THROWS (503 from flaky-service)
attempt 4 at t= 571 ms -> succeeds
result: OK in 4 attempts, 580 ms wall clock
$ java -cp "classes:$CP" RetryDemo # same code, second run
Backoff plan: base=100ms, multiplier=2, jitter=+/-50% (100 -> 200 -> 400 ...)
attempt 1 at t= 3 ms -> THROWS (503 from flaky-service)
attempt 2 at t= 179 ms -> THROWS (503 from flaky-service)
attempt 3 at t= 383 ms -> THROWS (503 from flaky-service)
attempt 4 at t= 712 ms -> succeeds
result: OK in 4 attempts, 716 ms wall clock
The waits grew exponentially in both runs (~147ms, ~188ms, ~233ms the first time; ~176ms, ~204ms, ~329ms the second) — but never identically. That's the jitter doing its job. Here's the first run on a timeline:
Why jitter: the thundering herd
Backoff alone has a flaw that only shows up at scale. Suppose the downstream dies for 30 seconds and 400 clients are retrying with a clean exponential backoff of 1s, 2s, 4s, 8s… They all failed at roughly the same moment, so their retry clocks are synchronized: at t=30s every one of them fires its next attempt within the same few milliseconds. The downstream — just recovered, caches cold, connections re-establishing — is hit by 400 simultaneous requests. That's incident one. The recovery causes the second outage.
Principle: backoff gives the downstream time; jitter stops your clients from acting as one. Resilience4j's IntervalFunction.ofExponentialRandomBackoff(initialMillis, multiplier, randomizationFactor) bakes both in — use it instead of hand-rolling Thread.sleep loops.
What not to retry (and the retry budget)
Retries are for transient failures. Retrying anything else is either wasteful or harmful:
| Failure | Retry? | Why |
|---|---|---|
| HTTP 429 / 503, connection reset, socket timeout | Yes, with backoff + jitter | Transient — the next attempt may genuinely succeed. |
| HTTP 400, 401, 403, 404, 422 | No | Deterministic — the request itself is wrong; retrying burns quota and logs. |
| Non-idempotent write (charge, book, send) | Not without an idempotency key | The retry may double-apply. Section "BREAK IT" proves it. |
| Downstream confirmed down (breaker open) | No | The circuit breaker already decided. Retrying is arguing with the verdict. |
In Resilience4j this is the retryExceptions / ignoreExceptions config: retry on the transient exception types, and let everything else propagate immediately. And cap the total damage with a retry budget — maxAttempts bounds the attempts per call, but also think in totals: if every request retries 5 times with 400ms waits, one slow downstream turns each of your threads into a 2-second hostage. The budget is per-call and per-system: max attempts, plus bulkheads (section 4) so retries can't consume every thread you own.
3. Circuit breakers: stop calling what's down
Retries assume the downstream might recover. A circuit breaker handles the case where it clearly hasn't: after enough failures, stop calling it at all for a while, fail fast instead, and probe cautiously before trusting it again. It's the electrical breaker in your fuse box — when the circuit is faulting, you don't keep flipping the switch; you open the breaker, wait, then test.
Three states, and the transitions are the whole pattern:
The demo runs 8 calls, one second apart, against a downstream that fails its first 4 calls and then recovers. Breaker config: trip at ≥50% failure rate over a 4-call window, wait 2 seconds in OPEN, allow 2 probe calls in HALF_OPEN. State transitions are printed by the breaker's event publisher — this is the real output:
import io.github.resilience4j.circuitbreaker.CircuitBreaker;
import io.github.resilience4j.circuitbreaker.CircuitBreakerConfig;
import java.time.Duration;
import java.util.concurrent.atomic.AtomicInteger;
public class CircuitDemo {
static final AtomicInteger CALLS = new AtomicInteger(0);
static long t0;
// Downstream that fails for the first 4 calls, then recovers.
static String downstream() {
int n = CALLS.incrementAndGet();
if (n <= 4) throw new IllegalStateException("downstream is DOWN");
return "OK";
}
public static void main(String[] args) throws Exception {
CircuitBreaker cb = CircuitBreaker.of("payment", CircuitBreakerConfig.custom()
.failureRateThreshold(50)
.minimumNumberOfCalls(4)
.slidingWindowSize(4)
.waitDurationInOpenState(Duration.ofSeconds(2))
.permittedNumberOfCallsInHalfOpenState(2)
.build());
cb.getEventPublisher().onStateTransition(e ->
System.out.printf(" t=%4dms *** STATE: %s -> %s%n", System.currentTimeMillis() - t0,
e.getStateTransition().getFromState(), e.getStateTransition().getToState()));
t0 = System.currentTimeMillis();
for (int i = 1; i <= 8; i++) {
String r;
try { r = cb.executeSupplier(CircuitDemo::downstream); }
catch (io.github.resilience4j.circuitbreaker.CallNotPermittedException e) {
r = "REJECTED (breaker open)";
} catch (Exception e) { r = "failed: " + e.getMessage(); }
System.out.printf("t=%4dms call %d: %-22s [breaker=%s, downstreamCalls=%d]%n",
System.currentTimeMillis() - t0, i, r, cb.getState(), CALLS.get());
Thread.sleep(1000);
}
}
}
$ java -cp "classes:$CP" CircuitDemo
initial state: CLOSED
t= 82ms call 1: failed: downstream is DOWN [breaker=CLOSED, downstreamCalls=1]
t=1130ms call 2: failed: downstream is DOWN [breaker=CLOSED, downstreamCalls=2]
t=2135ms call 3: failed: downstream is DOWN [breaker=CLOSED, downstreamCalls=3]
t=3146ms *** STATE: CLOSED -> OPEN
t=3146ms call 4: failed: downstream is DOWN [breaker=OPEN, downstreamCalls=4]
t=4148ms call 5: REJECTED (breaker open) [breaker=OPEN, downstreamCalls=4]
t=5150ms *** STATE: OPEN -> HALF_OPEN
t=5152ms call 6: OK [breaker=HALF_OPEN, downstreamCalls=5]
t=6155ms *** STATE: HALF_OPEN -> CLOSED
t=6156ms call 7: OK [breaker=CLOSED, downstreamCalls=6]
t=7158ms call 8: OK [breaker=CLOSED, downstreamCalls=7]
Read it as a story: calls 1–3 fail but the breaker stays CLOSED — it needs a minimum of 4 calls before it trusts its own statistics. Call 4 completes the window at 4/4 failures (100% ≥ 50%) and the breaker trips on that call's failure. Call 5 never reaches the downstream at all — rejected in microseconds, and note downstreamCalls stays at 4. After the 2-second wait, call 6 is admitted as a half-open probe and succeeds; call 7's success completes the probe window with a 0% failure rate, and the breaker closes. The downstream got exactly one probe call while it was recovering — not a retry storm.
Honest finding #2: the half-open verdict uses ≥, and I confirmed it in the source. While tuning this demo I ran a variant where the first probe failed and the second succeeded: 1 failure + 1 success = 50% failure rate, and the breaker went back to OPEN — because Resilience4j's CircuitBreakerMetrics re-opens when failureRate >= failureRateThreshold (I read the 2.4.0 source to confirm). With a 50% threshold and 2 probe calls, one failed probe is enough to re-open. If that surprises you, it should: set the threshold and the probe count together, deliberately, or the half-open state will behave in ways you didn't intend.
Principle: the breaker protects the downstream from you, and you from the downstream. While OPEN, your service fails fast (and can serve a fallback — a cached price, a "try again" page) instead of queueing threads behind a dead dependency.
4. Bulkheads: cap the blast radius
Named after the watertight compartments in a ship's hull: one flooded compartment shouldn't sink the vessel. A bulkhead limits how many concurrent calls may enter a downstream — everyone else is rejected immediately instead of queueing. Remember honest finding #1: timeouts free the caller but not the worker. The bulkhead is what actually bounds the workers.
Here the "downstream" is a 2-connection pool (maxConcurrentCalls(2)), and callers refuse to wait (maxWaitDuration(ZERO)). Two tasks grab both slots and hold them; the third arrives while the pool is full:
import io.github.resilience4j.bulkhead.Bulkhead;
import io.github.resilience4j.bulkhead.BulkheadConfig;
import java.time.Duration;
import java.util.concurrent.*;
public class BulkheadDemo {
public static void main(String[] args) throws Exception {
Bulkhead bulkhead = Bulkhead.of("db-pool", BulkheadConfig.custom()
.maxConcurrentCalls(2) // only 2 slots, like a 2-connection pool
.maxWaitDuration(Duration.ZERO) // don't queue: reject immediately
.build());
CountDownLatch bothInside = new CountDownLatch(2);
CountDownLatch release = new CountDownLatch(1);
ExecutorService pool = Executors.newFixedThreadPool(3);
Runnable slowTask = () -> {
try {
String r = bulkhead.executeSupplier(() -> {
bothInside.countDown();
try { release.await(5, TimeUnit.SECONDS); }
catch (InterruptedException e) { Thread.currentThread().interrupt(); }
return "db result";
});
System.out.println(Thread.currentThread().getName() + ": " + r);
} catch (io.github.resilience4j.bulkhead.BulkheadFullException e) {
System.out.println(Thread.currentThread().getName() + ": REJECTED -> " + e.getMessage());
}
};
pool.submit(slowTask);
pool.submit(slowTask);
bothInside.await(); // wait until 2 calls hold both slots
pool.submit(slowTask); // 3rd call: no slot, no waiting -> rejected
Thread.sleep(500);
release.countDown();
pool.shutdown();
pool.awaitTermination(5, TimeUnit.SECONDS);
System.out.println("bulkhead metrics: availableConcurrentCalls="
+ bulkhead.getMetrics().getAvailableConcurrentCalls());
}
}
$ java -cp "classes:$CP" BulkheadDemo
pool-1-thread-3: REJECTED -> Bulkhead 'db-pool' is full and does not permit further calls
pool-1-thread-2: db result
pool-1-thread-1: db result
bulkhead metrics: availableConcurrentCalls=2
(Thread names and print order vary run to run — the point is stable: two calls served, one rejected with BulkheadFullException.) The third caller fails fast instead of waiting behind two stuck calls. In production this is what keeps a slow database from eating every request thread: the bulkhead around the DB pool rejects excess load, and those rejections become clean 503s your load balancer can route elsewhere — instead of a thread pool that fills up and takes the whole service down with it.
Two bulkhead flavors exist: the SemaphoreBulkhead used here (limits concurrent calls) and a ThreadPoolBulkhead (runs calls on a dedicated thread pool, isolating them from your main pool entirely). Start with the semaphore version; reach for the thread-pool version when one downstream's slowness must not be able to starve your request threads at all.
5. Rate limiters: protect the downstream from yourself
Bulkheads bound concurrency; rate limiters bound throughput over time. Different problem: a downstream that handles 100 requests/second fine but falls over at 500. Your service, retrying happily, can be the thing that pushes it over. The rate limiter is a polite "not right now" — with timeoutDuration(ZERO), an immediate rejection instead of a queue:
import io.github.resilience4j.ratelimiter.RateLimiter;
import io.github.resilience4j.ratelimiter.RateLimiterConfig;
import io.github.resilience4j.ratelimiter.RequestNotPermitted;
import java.time.Duration;
public class RateLimitDemo {
public static void main(String[] args) throws Exception {
RateLimiter limiter = RateLimiter.of("api", RateLimiterConfig.custom()
.limitForPeriod(2) // 2 calls ...
.limitRefreshPeriod(Duration.ofSeconds(1)) // ... per second
.timeoutDuration(Duration.ZERO) // don't wait for a permit: reject
.build());
for (int i = 1; i <= 4; i++) {
try {
String r = limiter.executeSupplier(() -> "served");
System.out.println("request " + i + ": " + r
+ " (available permits now: " + limiter.getMetrics().getAvailablePermissions() + ")");
} catch (RequestNotPermitted e) {
System.out.println("request " + i + ": REJECTED -> " + e.getMessage());
}
}
Thread.sleep(1100); // new period starts
System.out.println("request 5 (after 1s refresh): " + limiter.executeSupplier(() -> "served"));
}
}
$ java -cp "classes:$CP" RateLimitDemo
request 1: served (available permits now: 1)
request 2: served (available permits now: 0)
request 3: REJECTED -> RateLimiter 'api' does not permit further calls
request 4: REJECTED -> RateLimiter 'api' does not permit further calls
request 5 (after 1s refresh): served
Two permits per second, third and fourth requests rejected, permits refresh after the period. In production, put the rate limiter on the client side of calls to a quota'd downstream (a payment gateway's "100 req/s" plan, a partner API) — it turns "we blew past their quota and got banned" into "we shed our own excess load first." Note the asymmetry with retries: a rejected call here should generally not be retried immediately — the permit refresh is the backoff.
BREAK IT: the retry that billed twice
Everything so far assumed retries are safe. They aren't — not by default. This is incident two from the opening, reproduced exactly: a charge that succeeds on attempt 1, but the response is lost on the wire, so the caller only sees a timeout. The retry policy, doing precisely what it was configured to do, charges the card again. Watch the ledger:
import io.github.resilience4j.retry.Retry;
import io.github.resilience4j.retry.RetryConfig;
import java.time.Duration;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
public class DoubleChargeDemo {
// The "bank ledger": every successful charge appends here.
static final List<String> LEDGER = new ArrayList<>();
static final AtomicInteger CALLS = new AtomicInteger(0);
/**
* Charge the card. Scripted disaster: attempt #1 DOES charge the card,
* then the network drops the response, so the caller only sees a timeout.
* Attempt #2+ works normally.
*/
static String charge(String card, int cents) {
int n = CALLS.incrementAndGet();
if (n == 1) {
LEDGER.add("charged " + card + " $" + cents / 100.0); // money moved!
throw new RuntimeException("timeout: response lost after commit");
}
LEDGER.add("charged " + card + " $" + cents / 100.0);
return "receipt#" + n;
}
public static void main(String[] args) {
Retry retry = Retry.of("payment", RetryConfig.custom()
.maxAttempts(3)
.waitDuration(Duration.ofMillis(200))
.retryExceptions(RuntimeException.class)
.build());
System.out.println("Customer clicks Pay once: $49.99");
try {
String receipt = retry.executeSupplier(() -> charge("card-4242", 4999));
System.out.println("caller sees: " + receipt);
} catch (Exception e) {
System.out.println("caller sees failure: " + e.getMessage());
}
System.out.println("--- bank ledger: " + LEDGER.size()
+ (LEDGER.size() == 1 ? " charge" : " charges") + " ---";
LEDGER.forEach(e -> System.out.println(" " + e));
System.out.println(LEDGER.size() == 1
? "OK: charged exactly once"
: "DOUBLE CHARGE: the customer was billed " + LEDGER.size() + " times for one click");
}
}
$ java -cp "classes:$CP" DoubleChargeDemo
Customer clicks Pay once: $49.99
caller sees: receipt#2
--- bank ledger: 2 charges ---
charged card-4242 $49.99
charged card-4242 $49.99
DOUBLE CHARGE: the customer was billed 2 times for one click
The caller sees one clean receipt. The ledger shows two charges. Nothing threw, nothing logged a warning — the retry worked perfectly and the customer was still billed twice. This is the scenario the rule exists for: a timeout tells you the response was lost, never that the work didn't happen. "No response" and "not executed" are different facts, and a retry treats them as the same.
The core rule, stated as a gate: before wrapping any operation in a retry, answer one question — is this operation idempotent? An operation is idempotent if performing it N times has the same effect as performing it once. Reads are idempotent. PUT /orders/123 {"status":"shipped"} is idempotent. POST /charges is not — each call creates a new charge. If the answer is no, you have exactly two choices: make it idempotent (below), or don't retry it.
The fix: idempotency keys
The industry-standard fix: the client generates a unique key once per logical operation and sends it with every attempt, including retries. The server stores the key with the first execution's result; a repeat key returns the stored result instead of re-executing. Same disaster script, now with the server deduplicating:
import io.github.resilience4j.retry.Retry;
import io.github.resilience4j.retry.RetryConfig;
import java.time.Duration;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.atomic.AtomicInteger;
public class IdempotentChargeDemo {
static final List<String> LEDGER = new ArrayList<>();
static final Map<String, String> SEEN_KEYS = new HashMap<>(); // idempotency store
static final AtomicInteger CALLS = new AtomicInteger(0);
/** Same disaster script, but the server dedupes on idempotencyKey. */
static String charge(String idempotencyKey, String card, int cents) {
int n = CALLS.incrementAndGet();
if (SEEN_KEYS.containsKey(idempotencyKey)) {
return SEEN_KEYS.get(idempotencyKey) + " (duplicate suppressed)";
}
String receipt = "receipt#" + n;
SEEN_KEYS.put(idempotencyKey, receipt);
LEDGER.add("charged " + card + " $" + cents / 100.0);
if (n == 1) throw new RuntimeException("timeout: response lost after commit");
return receipt;
}
public static void main(String[] args) {
Retry retry = Retry.of("payment", RetryConfig.custom()
.maxAttempts(3)
.waitDuration(Duration.ofMillis(200))
.retryExceptions(RuntimeException.class)
.build());
String key = "order-8842-attempt"; // generated ONCE by the client, reused on retry
System.out.println("Customer clicks Pay once: $49.99 (idempotency-key=" + key + ")");
String receipt = retry.executeSupplier(() -> charge(key, "card-4242", 4999));
System.out.println("caller sees: " + receipt);
System.out.println("--- bank ledger: " + LEDGER.size()
+ (LEDGER.size() == 1 ? " charge" : " charges") + " ---");
LEDGER.forEach(e -> System.out.println(" " + e));
System.out.println(LEDGER.size() == 1
? "OK: charged exactly once — retry was safe"
: "STILL BROKEN: " + LEDGER.size() + " charges");
}
}
$ java -cp "classes:$CP" IdempotentChargeDemo
Customer clicks Pay once: $49.99 (idempotency-key=order-8842-attempt)
caller sees: receipt#1 (duplicate suppressed)
--- bank ledger: 1 charge ---
charged card-4242 $49.99
OK: charged exactly once — retry was safe
Attempt 1 charged the card and stored the key before the response was lost. Attempt 2 arrived with the same key, the server recognized it, and returned the original receipt without touching the ledger. One click, one charge, one receipt — and the retry policy didn't have to change at all.
The details that make idempotency keys work in production:
- Generate the key once per user intent, not per attempt. The classic bug is generating a fresh UUID inside the retry loop — every attempt then looks like a new operation and the dedupe never fires. Generate before the first attempt; reuse across all of them.
- Scope the key correctly.
order-8842dedupes the order's payment — good. A key scoped too broadly ("all charges today") blocks legitimate distinct charges; too narrowly and duplicates slip through. - Store key → result, with a TTL. The store (Redis, a DB table) must survive the server restarting between attempt 1 and attempt 2, and keys must expire — a payment key kept forever is a slow storage leak and a replay hazard.
- Design the HTTP layer for it. This is why REST distinguishes
PUT(idempotent by definition — same URL, same body, same effect) fromPOST(not idempotent). When you must usePOSTfor a state-changing call, accept anIdempotency-Keyheader — Stripe's API is the canonical example.
Principle: idempotency is what makes retries morally permissible, not just technically configured. Backoff and jitter decide when to retry; idempotency decides whether you're allowed to.
Putting it together: the stack order
In a real service these patterns compose — one call passes through several of them. Resilience4j modules each expose decorateSupplier, and you nest them innermost-first: the last decoration applied is the outermost layer, the first to see the call. The sane order:
Supplier<String> downstream = () -> { /* the real call */ };
// innermost: the bulkhead guards the resource closest to the metal
Supplier<String> guarded = Retry.decorateSupplier(retry,
CircuitBreaker.decorateSupplier(circuitBreaker,
Bulkhead.decorateSupplier(bulkhead, downstream)));
// outermost (applied last, sees the call first): the retry,
// and a TimeLimiter around the whole thing for the hard deadline
So a call flows retry → circuit breaker → bulkhead → downstream, with the time limiter bounding the total. Why this order? The retry is outermost because a rejection from the breaker or the bulkhead usually shouldn't be retried — those are decisions, not transient failures. The breaker sits outside the bulkhead so it counts the downstream's real health, not your own rejections. The bulkhead is innermost because it's the last gate before the scarce resource.
The demo wires all four around the flaky downstream (fails once, then succeeds):
$ java -cp "classes:$CP" StackedDemo
stacked result: ok-on-try-2
downstream calls: 2 | breaker state: CLOSED | calls that succeeded after a retry: 1
First attempt failed, the outer retry fired once, the second attempt passed through a CLOSED breaker and an available bulkhead slot, and the whole thing completed inside the 2-second time limit. Four patterns, one call, each doing exactly its job.
Honest finding #3, minor: Resilience4j's docs show a Decorators fluent builder for this composition, but that class is not shipped in the 2.4.0 module jars on Maven Central (I checked every artifact). The nested decorateSupplier calls above are the way that compiles against the real 2.4.0 jars — and frankly the nesting makes the layer order more explicit than the builder hid it.
Cheat sheet: interview one-liners
- "How do you make a service resilient to downstream failures?" — Timeouts on every call, retries with exponential backoff + jitter for transient failures only, circuit breakers to fail fast when the downstream is down, bulkheads to cap concurrent load, rate limiters to respect quotas. Never retry without asking if the operation is idempotent.
- "Why add jitter to backoff?" — Without it, all clients' retry clocks stay synchronized and the recovery moment becomes a thundering herd that re-crushes the downstream. Jitter spreads retries across the window; the average wait doesn't change.
- "What does a circuit breaker do that retries don't?" — Retries assume recovery is likely; the breaker handles confirmed outage: after a failure-rate threshold it rejects calls instantly for a cooldown, then admits a few half-open probes before trusting the downstream again.
- "Your timeout fired but the downstream still did the work — explain." — A timeout frees the caller, not the resource: the worker thread/connection keeps running. (Measured in this post: caller freed at 515ms, worker finished at 2007ms.) Bound the resource separately with a bulkhead.
- "How do you safely retry a payment?" — Idempotency keys: client generates one key per user intent, server dedupes on it and returns the stored result for repeats. A timeout means the response was lost, never that the work didn't happen.
- "What shouldn't you retry?" — 4xx (deterministic), non-idempotent writes without keys, and anything the circuit breaker already rejected.
What's next
You now have the failure-handling half of production Java: every call has a deadline, transient failures get patient retries, confirmed outages get a breaker, scarce resources get a bulkhead, quotas get a rate limiter — and no retry ever runs without passing the idempotency gate. But resilience only matters if the service stays up while you ship it: the next post, Production Deployments: Health Checks, Graceful Shutdown & Safe Rollouts, covers the deploy half — health checks that tell the orchestrator the truth, shutdown hooks that drain in-flight work instead of killing it, database migrations that don't lock the table, API versioning that doesn't break old clients, and rollouts you can reverse.
Field check before you move on: (1) Take RetryDemo and change the jitter factor from 0.5 to 0.0 — run it three times and confirm the waits are now identical across runs (147→188→233 becomes exactly 100→200→400); then set it to 1.0 and observe the spread. (2) In CircuitDemo, change permittedNumberOfCallsInHalfOpenState to 1 and make the downstream fail its 5th call too — verify the breaker re-opens on the single failed probe and stays open. (3) Write a 30-line variant of DoubleChargeDemo where the idempotency key is (incorrectly) generated inside the retried supplier — confirm the ledger shows 2 charges again, proving the key must be generated once per intent, not per attempt. (4) In BulkheadDemo, change maxWaitDuration from ZERO to 2 seconds and re-run — the third call should now wait and succeed instead of being rejected; decide which behavior your own DB pool wants and write down why.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment