Production Deployments: Health Checks, Graceful Shutdown & Safe Rollouts
The deploy went out at 4 PM on a Friday. Four instances behind a load balancer, one after another, the pipeline reported green. Then the support channel lit up: "checkout fails with a connection error, but only sometimes." Retry and it worked. Wait five minutes and it worked. The team stared at the green pipeline, the healthy dashboards, the passing tests — and at customers getting 502 Bad Gateway in a neat one-in-four pattern. The rolling deploy had done exactly what it was told. The service, however, had done something nobody told it not to do: every instance accepted traffic the moment its JVM started, and dropped every in-flight request the moment SIGTERM arrived.
This post is the missing half of "it works on my machine": health checks, graceful shutdown, and safe rollouts. You'll build a service that tells the load balancer the truth about itself (/health/live vs /health/ready), a JVM that refuses to drop a single in-flight request on shutdown, and the deployment patterns that make Friday deploys boring: rolling deploys, feature flags, expand/contract migrations, versioned APIs, and a rollback plan you write before you need it. Every demo was compiled with javac and run on OpenJDK 21.0.3, with real output pasted below.
Demo 1: liveness vs readiness — two questions, two endpoints
A load balancer asks your instance one question — "can I send you traffic?" — but that question is really two:
- Liveness: is this process alive? If not, restart it. A deadlocked-but-breathing process fails liveness; a JVM that just started passes it.
- Readiness: is this instance fit to take traffic right now? A freshly started JVM whose caches are cold, whose pool has zero connections, whose schema check hasn't run — alive, but not ready. An instance mid-shutdown — alive, but not ready.
Conflate them and you get the Friday incident: the balancer saw "alive" and routed traffic to an instance that hadn't finished warming up, then kept routing to instances that were already draining. Here they are as separate endpoints. Readiness starts false (warm-up), flips true, then flips back to false when the shutdown drain begins — while liveness stays 200 the whole time:
import com.sun.net.httpserver.HttpServer;
import com.sun.net.httpserver.HttpExchange;
import java.io.IOException;
import java.io.OutputStream;
import java.net.InetSocketAddress;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicBoolean;
public class HealthServer {
static final AtomicBoolean READY = new AtomicBoolean(false);
static HttpServer server;
public static void main(String[] args) throws Exception {
int port = Integer.parseInt(System.getenv().getOrDefault("PORT", "18080"));
server = HttpServer.create(new InetSocketAddress(port), 0);
server.setExecutor(Executors.newFixedThreadPool(8));
// LIVENESS: is this process breathing?
server.createContext("/health/live", ex -> reply(ex, 200, "{\"status\":\"alive\"}"));
// READINESS: is this instance fit to take traffic right now?
server.createContext("/health/ready", ex -> {
if (READY.get()) reply(ex, 200, "{\"status\":\"ready\"}");
else reply(ex, 503, "{\"status\":\"not-ready\"}");
});
server.createContext("/hello", ex ->
reply(ex, 200, "{\"greeting\":\"hello from instance\"}"));
// Simulate warm-up: readiness flips true only after startup work completes.
server.start();
System.out.println("server started on port " + port + " (ready=false, warming up)");
Thread.sleep(3_000); // e.g. caches, connection pools, schema checks
READY.set(true);
System.out.println("warm-up done: ready=true");
// After 9 more seconds, begin a graceful shutdown (Demo 2 reuses this pattern).
Thread.sleep(9_000);
System.out.println("SIGTERM received: entering drain mode (ready=false, no new work)");
READY.set(false);
Thread.sleep(4_000);
System.out.println("in-flight requests drained: stopping server");
server.stop(0);
System.out.println("server stopped");
}
static void reply(HttpExchange ex, int code, String body) throws IOException {
byte[] bytes = body.getBytes(StandardCharsets.UTF_8);
ex.getResponseHeaders().add("Content-Type", "application/json");
ex.sendResponseHeaders(code, bytes.length);
try (OutputStream os = ex.getResponseBody()) { os.write(bytes); }
}
}
Now curl it at three moments — during warm-up, after warm-up, and during the drain:
$ javac HealthServer.java && java HealthServer &
## t ≈ 1s — still warming up:
$ curl -s -w "\nHTTP %{http_code}\n" http://localhost:18080/health/live
{"status":"alive"}
HTTP 200
$ curl -s -w "\nHTTP %{http_code}\n" http://localhost:18080/health/ready
{"status":"not-ready"}
HTTP 503
## t ≈ 5s — warm-up done:
$ curl -s -w "\nHTTP %{http_code}\n" http://localhost:18080/health/ready
{"status":"ready"}
HTTP 200
## t ≈ 14s — shutdown drain in progress:
$ curl -s -w "\nHTTP %{http_code}\n" http://localhost:18080/health/live
{"status":"alive"}
HTTP 200
$ curl -s -w "\nHTTP %{http_code}\n" http://localhost:18080/health/ready
{"status":"not-ready"}
HTTP 503
The server log tells the same story from the inside:
server started on port 18080 (ready=false, warming up)
warm-up done: ready=true
SIGTERM received: entering drain mode (ready=false, no new work)
in-flight requests drained: stopping server
server stopped
503 but keeps you alive; only liveness failure triggers a restart. During warm-up and during drain, the truthful answer is "alive, not ready" — and "alive" alone is a lie that costs you requests.The state machine your instance actually moves through — this is the contract every deployment strategy in this post depends on:
One gotcha the demo makes visible: readiness must flip before the listener stops accepting connections, and the balancer needs time to notice. Flipping the flag to false and immediately killing the socket is the same lie in reverse — the balancer may still route to you for one more probe interval. Production shutdown sequences look like: readiness → false, sleep for one probe interval (so the balancer drains you), then stop the listener, then drain in-flight, then exit.
Demo 2: graceful shutdown — SIGTERM, mid-request, zero dropped connections
Readiness keeps new traffic away. But what about the request that arrived 200ms before SIGTERM and needs 3 more seconds? The default JVM behavior on kill is to die mid-request. A graceful shutdown instead runs this sequence:
- SIGTERM arrives →
Runtime.addShutdownHookfires. - Stop accepting new connections (the listener closes).
- Wait for in-flight requests to finish (bounded — a timeout is the backstop, not the plan).
- Exit. Only then does the orchestrator's kill timer (
SIGKILL) matter — and it should never need to fire.
The demo below does this for real: a /slow endpoint that takes 3 seconds, an AtomicInteger counting in-flight requests, and a shutdown hook. The shell script fires a client request, then sends a real kill -TERM while the request is mid-flight:
import com.sun.net.httpserver.HttpServer;
import com.sun.net.httpserver.HttpExchange;
import java.io.IOException;
import java.io.OutputStream;
import java.net.InetSocketAddress;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.Executors;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.atomic.AtomicInteger;
public class GracefulShutdown {
static final AtomicInteger IN_FLIGHT = new AtomicInteger(0);
static HttpServer server;
static ExecutorService pool;
public static void main(String[] args) throws Exception {
int port = 18081;
server = HttpServer.create(new InetSocketAddress(port), 0);
pool = Executors.newFixedThreadPool(4);
server.setExecutor(pool);
server.createContext("/slow", ex -> {
IN_FLIGHT.incrementAndGet();
System.out.println(" handler: started slow request, in-flight=" + IN_FLIGHT.get());
try {
try { Thread.sleep(3_000); } // a real request takes time
catch (InterruptedException ie) { Thread.currentThread().interrupt(); }
reply(ex, 200, "{\"result\":\"slow-request-completed\"}");
System.out.println(" handler: finished slow request");
} finally {
IN_FLIGHT.decrementAndGet();
}
});
server.start();
System.out.println("server started on port " + port + " (pid visible in shell)");
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
System.out.println("SHUTDOWN HOOK: SIGTERM received, in-flight=" + IN_FLIGHT.get());
System.out.println("SHUTDOWN HOOK: stop accepting new connections; wait up to 10s for drain...");
server.stop(10); // listener stops now; in-flight exchanges get up to 10s to finish
System.out.println("SHUTDOWN HOOK: listener stopped, in-flight=" + IN_FLIGHT.get()
+ " — exiting cleanly");
}));
Thread.currentThread().join(); // stay alive until the signal arrives
}
static void reply(HttpExchange ex, int code, String body) throws IOException {
byte[] bytes = body.getBytes(StandardCharsets.UTF_8);
ex.getResponseHeaders().add("Content-Type", "application/json");
ex.sendResponseHeaders(code, bytes.length);
try (OutputStream os = ex.getResponseBody()) { os.write(bytes); }
}
}
And the run — the client request starts, SIGTERM lands 1.5 seconds later while it's still running, and the client gets its full response:
$ bash run-demoB.sh
server pid=39248
[shell] SIGTERM -> java 39248 while the request is mid-flight
[shell] client finished; waiting for server process to exit
[shell] probing new connection AFTER shutdown (must be refused):
HTTP 000
curl exit=7 (connection refused - new connections rejected)
[shell] server exited; final logs:
=== client response:
{"result":"slow-request-completed"}
=== server log:
server started on port 18081 (pid visible in shell)
handler: started slow request, in-flight=1
SHUTDOWN HOOK: SIGTERM received, in-flight=1
SHUTDOWN HOOK: stop accepting new connections; wait up to 10s for drain...
handler: finished slow request
SHUTDOWN HOOK: listener stopped, in-flight=0 — exiting cleanly
Read that log as the contract it is: the hook saw one in-flight request, refused everything new (curl exit 7, connection refused, HTTP 000), let the running handler finish, and exited with in-flight=0. The client — the thing your customers actually experience — never knew a shutdown happened.
server.stop(10) and not server.stop(0)? On the JDK's built-in HttpServer, stop(0) force-terminates immediately — in-flight exchanges get dropped, which is the ungraceful behavior we're fixing. The delay argument is the drain budget: listener closes now, in-flight work gets up to that many seconds. I verified this the hard way: the first version of this demo used stop(0) plus manual pool draining and the mid-flight client got an empty response. The graceful version above is what ships in the post.Rolling deploys: why the two demos above are one mechanism
A rolling deploy replaces instances one (or a few) at a time: take instance 1 out of the pool, shut it down gracefully, start the new version, wait for its readiness to report 200, add it back, move to instance 2. Every step is one of the two demos:
Note what makes this safe: the only synchronization between the pipeline and the instance is the two health endpoints. The pipeline can't see inside the JVM, so the instance must be honest — Demo 1's readiness flag and Demo 2's drain are the entire protocol. Two more strategies reuse the same protocol with different traffic shapes:
- Blue-green: run the whole new fleet (green) alongside the old (blue), wait for green's readiness, then flip the balancer 100% at once. Fast rollback (flip back), but you pay for double capacity and the flip is a single big event.
- Canary: route a small percentage (1–5%) of traffic to the new version and watch error rates and latency before widening. Best for catching "works in staging, dies on real traffic" — which is precisely what the opening incident was.
which docker kubectl returns nothing), so nothing containerized below was run here. The Dockerfile and manifest are shown so you can run them yourself — they are marked, and no output is invented for them.# Dockerfile — NOT run here (no Docker on this machine).
# Shown so you can build it yourself: the SIGTERM story matters here.
FROM eclipse-temurin:21-jre
WORKDIR /app
COPY GracefulShutdown.jar app.jar
# Exec form (JSON array) is load-bearing: the JVM becomes PID 1 and
# receives SIGTERM directly. Shell form wraps it in /bin/sh, which
# swallows the signal and your graceful shutdown never fires.
ENTRYPOINT ["java", "-jar", "app.jar"]
# k8s manifest fragment — NOT run here (no cluster on this machine).
livenessProbe:
httpGet: { path: /health/live, port: 8080 }
periodSeconds: 10
readinessProbe:
httpGet: { path: /health/ready, port: 8080 }
periodSeconds: 5 # balancer notices "not ready" within ~5s
terminationGracePeriodSeconds: 30 # SIGKILL backstop; drain must finish inside this
# preStop hook (optional belt-and-braces): sleep one probe interval after
# flipping readiness off, so the balancer drains you BEFORE the listener closes.
That ENTRYPOINT comment is the one that bites real teams: with the shell form, PID 1 is /bin/sh, which does not forward SIGTERM to your JVM — your shutdown hook never runs, terminationGracePeriodSeconds expires, and you get SIGKILL mid-request. Exec form makes the JVM PID 1 so the signal lands where your hook is.
Demo 3: feature flags — deploy the code, release the behavior
Deploys answer "which code runs." Feature flags answer "which behavior is on" — and they change the answer at runtime, without a deploy. The use cases stack up fast: roll a risky pricing change to 5% of users, kill a misbehaving feature in seconds instead of rolling back a whole deploy, let product toggle a UI without an engineer. The demo seeds the flag from an env var and flips it live through an admin endpoint (in production that endpoint is a flag service — LaunchDarkly, Unleash, a config push — not a hand-rolled HTTP switch, but the semantics are identical):
import com.sun.net.httpserver.HttpServer;
import com.sun.net.httpserver.HttpExchange;
import java.io.IOException;
import java.io.OutputStream;
import java.net.InetSocketAddress;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicBoolean;
public class FeatureFlag {
// seeded from environment; mutable at runtime
static final AtomicBoolean NEW_PRICING =
new AtomicBoolean(Boolean.parseBoolean(
System.getenv().getOrDefault("NEW_PRICING", "false")));
public static void main(String[] args) throws Exception {
int port = 18082;
HttpServer server = HttpServer.create(new InetSocketAddress(port), 0);
server.setExecutor(Executors.newFixedThreadPool(4));
server.createContext("/price", ex -> {
String body;
if (NEW_PRICING.get()) {
// new code path: tiered pricing, behind the flag
body = "{\"price\":90,\"engine\":\"v2-tiered\",\"note\":\"loyalty discount applied\"}";
} else {
// old code path: flat pricing
body = "{\"price\":100,\"engine\":\"v1-flat\"}";
}
reply(ex, 200, body);
});
// admin switch — stands in for a flag service
server.createContext("/admin/flag", ex -> {
String q = ex.getRequestURI().getQuery(); // enabled=true|false
boolean on = q != null && q.contains("enabled=true");
NEW_PRICING.set(on);
reply(ex, 200, "{\"flag\":\"NEW_PRICING\",\"enabled\":" + on + "}");
});
server.start();
System.out.println("started on port " + port
+ ", NEW_PRICING=" + NEW_PRICING.get() + " (from env)");
Thread.sleep(30_000);
server.stop(0);
}
// same reply() helper as HealthServer
static void reply(HttpExchange ex, int code, String body) throws IOException {
byte[] bytes = body.getBytes(StandardCharsets.UTF_8);
ex.getResponseHeaders().add("Content-Type", "application/json");
ex.sendResponseHeaders(code, bytes.length);
try (OutputStream os = ex.getResponseBody()) { os.write(bytes); }
}
}
$ NEW_PRICING=false java FeatureFlag &
started on port 18082, NEW_PRICING=false (from env)
$ curl -s http://localhost:18082/price ## flag off
{"price":100,"engine":"v1-flat"}
$ curl -s "http://localhost:18082/admin/flag?enabled=true" ## flip it live
{"flag":"NEW_PRICING","enabled":true}
$ curl -s http://localhost:18082/price ## same process, new behavior
{"price":90,"engine":"v2-tiered","note":"loyalty discount applied"}
$ curl -s "http://localhost:18082/admin/flag?enabled=false" ## instant rollback
{"flag":"NEW_PRICING","enabled":false}
$ curl -s http://localhost:18082/price
{"price":100,"engine":"v1-flat"}
No rebuild, no restart, no deploy pipeline — the behavior changed in milliseconds, and "rolling back" the pricing change was one HTTP call. That's the point: flags decouple deployment from release. The code ships on Tuesday's quiet deploy; the behavior turns on Thursday when the team is watching — and turns off in seconds if the metrics move the wrong way.
else branch is tech debt the moment the flag is permanent; (3) never nest flags more than one deep in a single code path. A codebase with 200 stale flags is worse than no flags at all, because nobody knows which behavior is actually live.Database migrations: the expand/contract pattern
Here's the deploy problem flags can't solve: the schema. During a rolling deploy, v1 and v2 code run against the same database at the same time. If your migration renames full_name to first_name/last_name in one step, v1 instances crash the moment the migration runs (column gone) and v2 instances crash before it runs (columns missing). The fix is to never make a breaking change in one step — expand, then contract:
The dual-write code (conceptual sketch — the pattern is what's being taught, not this exact code against a live DB):
// Phase 2 — deployable alongside v1: write BOTH representations, read the OLD one.
void saveUser(Connection c, String first, String last) throws SQLException {
String full = first + " " + last;
try (var ps = c.prepareStatement(
"INSERT INTO users(full_name, first_name, last_name) VALUES (?,?,?)")) {
ps.setString(1, full); // keeps v1 working
ps.setString(2, first); // feeds v2
ps.setString(3, last);
ps.executeUpdate();
}
}
// Phase 3 — after backfill: read the NEW columns. full_name still written,
// so rolling back to v1 remains safe.
User loadUser(Connection c, long id) throws SQLException {
try (var ps = c.prepareStatement(
"SELECT first_name, last_name FROM users WHERE id=?")) {
...
}
}
// Phase 4 — only when zero v1 instances remain: DROP COLUMN full_name.
API versioning and backward compatibility
The database has a twin problem one layer up: clients of your API upgrade on their own schedule. Your v2 deploy doesn't update the mobile app on a million phones. The rules that keep old clients working:
- Additive changes only, without a version bump. Adding a field to a response, adding an optional request parameter, adding a new endpoint — old clients ignore what they don't know. Renaming a field, changing its type, removing it, or changing its meaning are breaking changes, full stop.
- Tolerant readers. Clients should ignore unknown fields (this is why JSON libraries default to lenient deserialization — and why turning on "fail on unknown properties" in a client is a deploy hazard). Servers should ignore unknown request fields too, so a new client can talk to an old server during the rollout window.
- Version explicitly when you must break.
/v1/orderskeeps serving the old contract while/v2/ordersserves the new one. Version in the path (most common for public APIs) or in a header — but pick one and document it. And versioning is a promise: v1 keeps working until its announced sunset date, with a migration guide, not until someone gets impatient. - Deprecation is a process, not a header. Mark it deprecated, announce the sunset date, measure v1 traffic until it's near zero, then remove. "Nobody uses v1 anymore" is a metric you check, not a feeling you have.
// Tolerant reader: new field 'loyaltyTier' added server-side.
// Old clients (built before it existed) keep working — they just don't see it.
{"orderId":"ord-9917","total":149.99,"loyaltyTier":"gold"} <- v2 response
// v1 client deserializes orderId + total, ignores loyaltyTier. No crash, no deploy.
// Breaking change done right: new version, old one untouched.
GET /v1/orders/ord-9917 -> {"orderId":"ord-9917","total":149.99}
GET /v2/orders/ord-9917 -> {"id":"ord-9917","total":{"amount":149.99,"currency":"USD"}}
The rollback plan you write before you need it
Every deploy plan has a rollback section. Most are fiction — "roll back if issues arise" — written by someone who has never tried to roll back at 2 AM. A real one names all four:
- The trigger: which metric, which threshold, who decides. "Error rate on
/checkoutabove 1% for 5 minutes → on-call rolls back, no meeting required." A rollback that needs a committee is not a rollback plan. - The mechanism: for code, that's the previous artifact (blue-green flips are the fastest: one balancer change). For flags, it's the flag flip from Demo 3 — seconds, not minutes. For migrations, it's expand/contract: because phase 3 never dropped the old column, v1 code still runs.
- The data story: can the new version's writes be read by the old version? If v2 wrote rows v1 can't parse, rolling back the code doesn't roll back the data. This is why dual-write phases exist and why "we'll figure out the data" is the most expensive sentence in a postmortem.
- The drill: run it in staging quarterly. A rollback path you've never executed is a hope, and hope is not an engineering strategy.
Cheat sheet: the interview one-liners
- "Liveness vs readiness?" — Liveness: is the process alive (restart me if not). Readiness: am I fit for traffic right now (warm-up, drain). Conflating them routes traffic to instances that can't serve it.
- "What does your shutdown hook do?" — On SIGTERM: flip readiness off, wait one probe interval, stop accepting connections, drain in-flight with a bounded timeout, then exit. Never drop a request the client already sent.
- "Why did the deploy 502?" — Usual suspects: no readiness gate (traffic during warm-up), no drain (SIGTERM mid-request), shell-form Docker ENTRYPOINT swallowing SIGTERM so the hook never fired.
- "Rolling vs blue-green vs canary?" — Rolling: one-by-one, constant capacity, slow. Blue-green: full parallel fleet, instant flip and instant rollback, double cost. Canary: percentage-based, best for catching real-traffic-only failures.
- "How do you change a column safely?" — Expand/contract: add the new column, dual-write, backfill, switch reads, drop the old column only after all old instances are gone. Four boring phases, zero breaking changes.
- "What's a safe API change?" — Additive: new fields, new optional params, new endpoints. Breaking changes get a new version (
/v2), a sunset date, and measured v1 traffic before removal. - "Deploy vs release?" — Deploy puts code on servers; release exposes behavior to users. Feature flags decouple them — ship code Tuesday, flip behavior Thursday, roll back in seconds.
Field check
Reproduce the lab on your own machine — everything here needs only a JDK:
- Run
HealthServerand curl/health/liveand/health/readyduring warm-up, after warm-up, and during the drain. Confirm you see200/503 → 200/200 → 200/503. - Run
GracefulShutdown, start acurlagainst/slow, thenkill -TERMthe Java PID mid-request. Confirm the client receives the full JSON body and a subsequent connection is refused. - Break it on purpose: change
server.stop(10)toserver.stop(0)and rerun step 2. Observe the dropped client — this is the default behavior you're protecting against. - Run
FeatureFlag, flipNEW_PRICINGvia the admin endpoint, and confirm/pricechanges without a restart. Then write down your flag's expiry date and who owns its removal — that's the other half of the feature. - Take one schema change from your own project and write it as four expand/contract phases. If any phase can't be deployed while the previous code version is still running, the plan is wrong — redo it.
What's next
You can now ship without fearing Friday: instances that tell the truth about their health, JVMs that drain instead of dropping, deploys that roll forward one honest instance at a time, flags that separate shipping code from releasing behavior, migrations that never break the running version, and a rollback plan with a trigger, a mechanism, a data story, and a drill.
But "safe" isn't "fast." The next post proves your service can take the load it's about to get: Load Testing & Performance Budgets — generating realistic traffic, reading percentiles without lying to yourself, and setting budgets that fail the build before users feel the pain.
Follow the full track: Java Learning Roadmap 2026
Comments
Post a Comment