Capstone: Production-Ready Spring Boot Service
It's your first deploy night. The order service passed every test, the demo went perfectly, and at 9 PM you shipped it. By 2 AM you have four separate fires and they are all the same fire:
- Kubernetes killed a pod mid-request during the rollout — there was no graceful shutdown, so in-flight checkouts just died.
- The one error you can see in the logs has no request id, so you can't tell which of the 40,000 log lines belong to the failing checkout.
- The pricing service had a 30-second wobble. Your service retried every failed call instantly, with no backoff — turning their wobble into your retry storm, which turned into their outage.
- And the database password? It's in
application.yml, which is in git, which the new contractor cloned yesterday.
None of these are coding bugs. The code was correct. What's missing is everything around the code — the production checklist this whole track has been building: observability, configuration, resilience, safe deployments, and the load test that proves the budget. This capstone puts all of it into one service, running for real.
What "running for real" means here — stated up front. This machine has no PostgreSQL, no Kafka, no Redis, no Docker, and no Maven. So the lab service is built on the pure JDK (com.sun.net.httpserver, virtual threads) with real Resilience4j jars from Maven Central — every log line, metric, and status code below is real output from a real run. PostgreSQL, Kafka, Redis, and Docker appear as the target architecture and as a complete Spring Boot version of the same service, clearly marked not run in this lab. I show no Spring output because there is none to show.
The whole post in one diagram — bright boxes ran in this lab, dashed boxes are the production target:
1. The production-readiness checklist
Before the code: the checklist. Every item is one thing that bites you on deploy night, each taught in its own post earlier in this track, each implemented in the lab below. If you take one artifact from this post, take this table — run it against your own service:
| # | Checklist item | In the lab | Taught in |
|---|---|---|---|
| 1 | Structured logs with trace IDs | One JSON line per event; X-Request-Id on every response | Post 1 · Observability |
| 2 | RED metrics: rate, errors, duration | Hand-rolled registry, Prometheus-style /metrics with p50/p95/p99 | Post 1 · Observability |
| 3 | Config from the environment; secrets never in code or logs | APP_PORT precedence demo; DB_PASSWORD logged as <redacted> | Post 2 · Configuration & Secrets |
| 4 | Timeout on every downstream call | 2s deadline around the pricing call → real 504 | Post 3 · Resilience |
| 5 | Retry with backoff — only for idempotent, retryable failures | 3 attempts, 200ms backoff; fail-fast signals are not retried | Post 3 · Resilience |
| 6 | Circuit breaker — fail fast when the downstream is down | CLOSED → OPEN → HALF_OPEN → CLOSED, all in the logs; 503 while open | Post 3 · Resilience |
| 7 | Bulkhead — bounded concurrency per downstream | Semaphore of 8; saturation → 429, not a pile-up | Post 3 · Resilience |
| 8 | Liveness vs readiness probes | /health/live vs /health/ready (ready follows the breaker) | Post 4 · Deployments |
| 9 | Graceful shutdown with drain | SIGTERM → stop sending traffic → up to 5s drain → exit 0 | Post 4 · Deployments |
| 10 | Load test + latency budget | 100-request run feeding the p50/p95 percentiles | Post 5 · Load Testing |
| 11 | Automated tests: unit + contract | JUnit 5, 12 tests, all green | Tooling & Testing track |
| 12 | CI/CD pipeline | GitHub Actions workflow (Spring section — reviewed, not run) | This capstone |
2. The lab service: what it is and isn't
The lab is a small order service — POST /api/orders prices an item through a downstream pricing service, validates the request, stores the order, and exposes health and metrics endpoints. It is one Java program (ordersvc package, ~700 lines), compiled with javac on OpenJDK 21.0.3, using the real Resilience4j 2.1.0 jars (resilience4j-core, resilience4j-retry, resilience4j-circuitbreaker) downloaded from Maven Central. The HTTP layer is com.sun.net.httpserver on a virtual-thread-per-task executor — the same virtual threads from the Concurrency track, doing what they're for: one thread per request, blocking cheaply.
ConcurrentHashMap. Both are honest stand-ins with a defined seam: the pricing client is an interface, the store is a repository class, and the Spring Boot section shows what plugs into those seams in production. Everything else — the HTTP handling, validation, logging, metrics, retry/breaker behavior, config, probes, shutdown — is the real thing, running.3. Configuration: the environment wins, secrets stay secret
Twelve-factor config with a fixed precedence — environment variable beats system property beats default — and the resolution is logged with its source, so "why is it listening on that port?" is always answerable:
// Config.java — precedence: env > system property > default
public static Entry resolve(String envKey, String sysKey, String def) {
String env = System.getenv(envKey);
if (env != null && !env.isBlank()) return new Entry(env, "env:" + envKey);
String sys = System.getProperty(sysKey);
if (sys != null && !sys.isBlank()) return new Entry(sys, "sysprop:" + sysKey);
return new Entry(def, "default");
}
Started with both APP_PORT=18080 in the environment and -Dapp.port=18081 on the command line, the real startup log reads (timestamps trimmed for readability):
{"level":"INFO","logger":"ordersvc","trace":"-","event":"config.loaded",
"port":18080,"port.source":"env:APP_PORT",
"pricing.failures":4,"pricing.failures.source":"env:PRICING_FAILURES",
"pricing.sleepMs":0,"db.password":"<redacted>"}
Port 18080 — the environment won, and the log proves which source each value came from. And DB_PASSWORD was set to a real secret for this run; it appears in the log only as <redacted>. The value exists in the process environment; it never appears in code, in logs, or in any committed file. That's the whole of Post 2's secrets rule in one line: a secret in a log is a secret already shared.
4. REST + validation at the boundary
Validation happens before anything downstream is touched — no pricing call, no store write, no wasted work for a request that was never valid:
// OrderService.java — the boundary contract, unit-tested
static List<String> validate(Map<String, String> body) {
List<String> errs = new ArrayList<>();
if (body.getOrDefault("customerId", "").trim().isEmpty()) errs.add("customerId is required");
if (body.getOrDefault("item", "").trim().isEmpty()) errs.add("item is required");
String q = body.get("qty");
try {
int qty = Integer.parseInt(q == null ? "" : q.trim());
if (qty < 1 || qty > 1000) errs.add("qty must be between 1 and 1000");
} catch (NumberFormatException e) {
errs.add("qty must be an integer");
}
return errs;
}
Real responses from the running service:
$ curl -s -w "\nHTTP %{http_code}\n" -X POST localhost:18080/api/orders \
-H 'Content-Type: application/json' -d '{"customerId":"","item":"widget","qty":0}'
{"error":"validation_failed","details":"customerId is required; qty must be between 1 and 1000"}
HTTP 400
$ curl -s -w "\nHTTP %{http_code}\n" localhost:18080/api/orders/999
{"error":"not_found","id":999}
HTTP 404
The 400 names every problem in one response — a client shouldn't have to play whack-a-mole with your validator. The rejected request also increments orders_rejected_total{reason="validation"}, because rejected traffic is a metric worth watching: a spike in 400s is usually a broken client deploy, and you want to see it before the client team tells you.
5. Structured logs with trace IDs
Every request gets a trace id at the edge, returned in the X-Request-Id header and attached to every log line the request produces. One JSON object per line — greppable, shippable to Loki/ELK, joinable with traces later:
$ curl -s -i -X POST localhost:18080/api/orders -H 'Content-Type: application/json' \
-d '{"customerId":"c-42","item":"widget","qty":3}' | head -8
HTTP/1.1 502 Bad Gateway
Date: Tue, 06 Oct 2026 06:36:14 GMT
Content-type: application/json
X-request-id: 5e131f89
Content-length: 74
{"error":"pricing_failed","detail":"pricing service failed after retries"}
(This first order hit the simulated pricing outage — the resilience story starts in the next section. Watch the trace id 5e131f89:)
$ grep 5e131f89 server.log
{"level":"WARN","logger":"ordersvc","trace":"5e131f89","event":"pricing.failed",
"error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 3)"}
{"level":"INFO","logger":"ordersvc","trace":"5e131f89","event":"request.done",
"method":"POST","route":"orders","status":502,"durationMs":507}
Header and logs agree: one id ties the client's failed request to exactly its server-side story. That's the 2 AM difference between "grep 40,000 lines" and "grep one id".
"trace":"-" instead of the request's id — the trace id lives in a ThreadLocal, and the resilience work runs on a timeout-pool thread the ThreadLocal doesn't reach. In production, OpenTelemetry's context propagation carries the trace across thread hops automatically; a hand-rolled ThreadLocal does not. I'm showing the gap rather than hiding it, because "my trace ids vanish inside async code" is a classic production surprise.6. Metrics: the RED numbers
Rate, Errors, Duration — the three numbers that tell you whether the service is healthy, on one /metrics endpoint in Prometheus text format. Counters carry labels (method, route, status); latencies feed p50/p95/p99. After the demo sequence plus a 100-request latency run, the real exposition:
$ curl -s localhost:18080/metrics
http_requests_total{method="GET",route="orders",status="200"} 100
http_requests_total{method="GET",route="orders",status="404"} 1
http_requests_total{method="GET",route="ready",status="200"} 1
http_requests_total{method="GET",route="ready",status="503"} 1
http_requests_total{method="POST",route="orders",status="201"} 1
http_requests_total{method="POST",route="orders",status="400"} 1
http_requests_total{method="POST",route="orders",status="502"} 1
http_requests_total{method="POST",route="orders",status="503"} 2
orders_created_total 1
orders_rejected_total{reason="validation"} 1
pricing_rejected_total 2
http_request_ms_count 108
http_request_ms_sum 775
http_request_ms_p50 0
http_request_ms_p95 2
http_request_ms_p99 231
circuitbreaker_state{breaker="pricing"} 0
circuitbreaker_failed_calls{breaker="pricing"} 0
circuitbreaker_not_permitted_calls{breaker="pricing"} 0
Read it like an operator: 108 requests, one 502, two fast 503s, one validation rejection, and the latency percentiles say the service itself is fast (p95 2ms) — the pain is all downstream, which is exactly what the next section fixes. The breaker gauges read 0 because the circuit is currently CLOSED; they're sliding-window counters, so they describe the current window, not history. In production this endpoint is scraped by Prometheus and the RED dashboard is the first thing you open when the pager goes off — Post 1's whole point.
One request's journey through every layer — the pipeline the rest of this post walks through:
7. Resilience: retry, circuit breaker, bulkhead, timeout
This is Post 3's toolkit, wired as one stack around the pricing call — bulkhead outermost, then the timeout deadline, then retry, then the circuit breaker closest to the call:
// Pricing.java — the resilience stack (real Resilience4j 2.1.0)
Retry retry = Retry.of("pricing", RetryConfig.custom()
.maxAttempts(3)
.waitDuration(Duration.ofMillis(200)) // backoff between attempts
.retryExceptions(RuntimeException.class)
.ignoreExceptions(CallNotPermittedException.class, // fail-fast: never retry
BulkheadFullException.class) // saturated: never retry
.build());
CircuitBreaker cb = CircuitBreaker.of("pricing", CircuitBreakerConfig.custom()
.failureRateThreshold(50) // ≥50% failures…
.minimumNumberOfCalls(4) // …over at least 4 calls…
.slidingWindowSize(4) // …in the last 4 calls → OPEN
.waitDurationInOpenState(Duration.ofSeconds(5)) // then one half-open probe
.build());
// Retry outermost, breaker innermost: Retry(CircuitBreaker(call))
Supplier<Double> guarded = Retry.decorateSupplier(retry,
CircuitBreaker.decorateSupplier(cb, () -> callDelegate(item)));
The core rule from Post 3, enforced in code: never retry blindly — is the failure retryable? A downstream 500 might clear; a circuit-open rejection or a saturated bulkhead will not, so those are excluded from retry. Retrying a fail-fast signal is pure load — the retry storm from the opening incident, rebuilt as a feature.
Now watch it work. The pricing simulator is configured to fail its first 4 calls — a deterministic downstream outage:
Order 1 — retry fights, retry loses. Three attempts, 200ms apart, all fail. The log shows the backoff working exactly as configured, and the request still fails — correctly — after 507ms:
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.retry",
"retryAttempt":1,"error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 1)"}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.retry",
"retryAttempt":2,"error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 2)"}
{"level":"WARN","logger":"ordersvc","trace":"5e131f89","event":"pricing.failed",
"error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 3)"}
{"level":"INFO","logger":"ordersvc","trace":"5e131f89","event":"request.done",
"method":"POST","route":"orders","status":502,"durationMs":507}
Order 2 — the breaker opens mid-request. The fourth failed call trips the threshold (4 calls, 100% failure ≥ 50%), and the breaker opens while the request is still being retried:
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.circuit",
"from":"CLOSED","to":"OPEN"}
$ curl -s -X POST localhost:18080/api/orders -H 'Content-Type: application/json' \
-d '{"customerId":"c-43","item":"widget","qty":1}'
{"error":"pricing_unavailable","detail":"circuit breaker is OPEN — failing fast","retryable":true}
Order 3 — fail fast. With the breaker open, the request never touches the downstream at all:
$ time curl -s -o /dev/null -w "HTTP %{http_code}\n" -X POST localhost:18080/api/orders \
-H 'Content-Type: application/json' -d '{"customerId":"c-44","item":"widget","qty":1}'
HTTP 503
real 0m0.011s
11 milliseconds. Compare: the first order burned 507ms plus three downstream calls to discover the outage; this one spends 11ms and zero downstream calls to report it. That gap — between "discover the outage per request" and "remember the outage across requests" — is the entire economic argument for a circuit breaker.
After the 5-second open window — the half-open probe. One request is let through; it succeeds; the breaker closes:
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.circuit",
"from":"OPEN","to":"HALF_OPEN"}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.circuit",
"from":"HALF_OPEN","to":"CLOSED"}
$ curl -s -X POST localhost:18080/api/orders -H 'Content-Type: application/json' \
-d '{"customerId":"c-45","item":"gadget","qty":2}'
{"id":1,"customerId":"c-45","item":"gadget","qty":2,"price":19.99,"status":"CREATED"}
The state machine, as configured:
The timeout story, on a second instance. With the downstream configured to sleep 5 seconds per call and the deadline at 2 seconds, one order returns 504 — and note the elapsed time:
$ time curl -s -w "\nHTTP %{http_code}\n" -X POST localhost:18081/api/orders \
-H 'Content-Type: application/json' -d '{"customerId":"c-9","item":"widget","qty":1}'
{"error":"pricing_timeout","detail":"pricing timed out after 2000ms"}
HTTP 504
real 0m2.180s
{"level":"WARN","logger":"ordersvc","trace":"3eda351c","event":"resilience.timeout",
"item":"widget","timeoutMs":2000}
{"level":"INFO","logger":"ordersvc","trace":"3eda351c","event":"request.done",
"method":"POST","route":"orders","status":504,"durationMs":2099}
2.1 seconds, not 5 — the deadline did its job. And here's the design detail worth keeping: the 2s deadline wraps the whole guarded call — retry included — so a hung downstream costs this request ~2s total, not attempts × timeout. That's the overall-deadline pattern: per-attempt backoff inside, one deadline outside. Your worst case is a number you chose, not a number you discover.
pricing.call attempt firing after the timeout already returned 504 to the client. Cancelling the future interrupted the sleeping attempt, but the retry living inside the cancelled task fired its next attempt anyway — the orphan kept a thread busy for its full 5-second sleep while nobody waited for the answer. Timeouts protect your latency; they don't stop work already in flight downstream. The bulkhead (8 permits) is what bounds how many such orphans can pile up. Defense in depth means no single mechanism has to be perfect.One more Post 3 rule, visible in the 400s from section 4: the pricing call is an idempotent read, so retrying it is safe. If this were POST /payments/charge, blind retry would double-charge — that's why the retry config retries only failures, never fail-fast signals, and why non-idempotent operations get idempotency keys instead of retries. The rule travels with you; the config stays here.
8. Liveness vs readiness: two different questions
Post 4's distinction, in one handler: liveness asks "is the process alive?" (if not, restart it); readiness asks "should this instance receive traffic right now?" (if not, pull it from the load balancer but leave it running). Conflating them is how you get Kubernetes restarting a pod that was merely waiting on a downstream — turning a wobble into a restart loop:
// OrderService.java — readiness reflects the world, liveness reflects the process
private int handleReady(HttpExchange ex) throws IOException {
String pricingState = pricing.breaker().getState().name();
boolean pricingOk = !"OPEN".equals(pricingState);
boolean up = ready && pricingOk; // ready flips false the moment shutdown begins
String body = "{\"status\":\"" + (up ? "UP" : "DOWN") + "\",\"checks\":{"
+ "\"store\":\"UP\","
+ "\"pricing\":\"" + (pricingOk ? "UP" : "circuit OPEN") + "\"}}";
return send(ex, up ? 200 : 503, body);
}
While the breaker was open in the demo above, readiness told the truth:
$ curl -s -w "\nHTTP %{http_code}\n" localhost:18080/health/ready
{"status":"DOWN","checks":{"store":"UP","pricing":"circuit OPEN"}}
HTTP 503
…and after the half-open probe closed the circuit:
$ curl -s -w "\nHTTP %{http_code}\n" localhost:18080/health/ready
{"status":"UP","checks":{"store":"UP","pricing":"UP"}}
HTTP 200
The pod stays alive the whole time — no restart, no lost warmup — it just stops receiving traffic while its downstream is down, then rejoins when the probe succeeds. In production the checks object grows a real database ping and a Kafka producer check; the shape doesn't change.
9. Graceful shutdown: the drain
The opening incident's first fire — Kubernetes killing a pod mid-request. The fix is a shutdown hook with an order of operations: fail readiness first (so the load balancer stops sending new traffic), then wait for in-flight requests to finish, with a bounded grace period:
// OrderService.java — main()
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
Log.warn("shutdown.signal",
"detail", "SIGTERM received — refusing new traffic, draining in-flight requests");
svc.ready = false; // 1. readiness fails: load balancer looks away
server.stop(5); // 2. up to 5s for in-flight requests to finish
Log.warn("shutdown.drained", "inFlight", svc.inFlight.get());
Log.warn("shutdown.complete", "detail", "graceful shutdown finished");
}, "shutdown-hook"));
Real run: a 3-second request in flight, SIGTERM delivered half a second in:
$ curl -s "localhost:18080/api/orders?sleepMs=3000" > /dev/null &
$ sleep 0.5; kill -TERM $PID; wait $PID
{"level":"WARN","logger":"ordersvc","trace":"-","event":"shutdown.signal",
"detail":"SIGTERM received — refusing new traffic, draining in-flight requests"}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"shutdown.drained","inFlight":0}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"shutdown.complete",
"detail":"graceful shutdown finished"}
The slow request completed inside the grace period (inFlight:0 — nothing was abandoned), the process exited 0, and the whole sequence took ~2.6 seconds. Note the ordering subtlety the hook gets right: if you stop the server before failing readiness, there's a window where the load balancer still routes to a dying pod. Readiness first, always.
10. Tests: 12 green, and what each group proves
JUnit 5 via the platform console launcher (no Maven on this machine — the standalone jar runs anywhere). The full suite:
$ java -jar junit-platform-console-standalone-1.10.3.jar \
--class-path "out:lib/..." --select-package ordersvc
[ 12 tests found ]
[ 12 tests started ]
[ 12 tests successful ]
[ 0 tests failed ]
What the 12 cover, by group:
- Validation contract (5 tests) — blank ids, zero/negative/non-numeric quantities, and the 1000/1001 boundary. The boundary behavior is the specification; the test is where it's written down.
- Config precedence (2 tests) — system property beats default; default applies when nothing is set. (Environment-variable precedence is proven by the live startup log in section 3 — you can't set env vars from inside the JVM under test, so the honest proof lives in the run, not the suite.)
- Transactional outbox (1 test) — creating an order writes the row and the
OrderCreatedevent atomically. The test drains the outbox and asserts the event carries the right order id — the contract the Kafka relay will depend on. - Metrics math (2 tests) — percentiles computed from known samples (p50=30, p95=50 on [10..50]); counters increment per label-set and render in the exposition format.
- Resilience4j behavior (2 tests) — retry really attempts 3 times before succeeding; the breaker really opens after the failure threshold and the guarded supplier body never runs again while open (asserted with a call counter — the fail-fast property, pinned down).
HttpClient against the running server) cannot run here — it fails with a sandbox policy error, not an application error. The HTTP round-trips are instead proven by the curl transcript throughout this post: real client, real server, real status codes. Twelve of twelve runnable tests green; the thirteenth proof is the demo above.11. The Spring Boot target — the same service, production-shaped
./mvnw spring-boot:run and docker compose up; the lab outputs above are the behavior to compare against.The lab's hand-rolled pieces each have a Spring Boot equivalent. The mapping, first — then the code:
| Lab piece (ran above) | Spring Boot equivalent (below, not run) |
|---|---|
com.sun.net.httpserver + virtual threads | Spring MVC on virtual threads (spring.threads.virtual.enabled=true) |
Hand-rolled validate() | Jakarta Bean Validation (@Valid, @NotBlank, @Min/@Max) |
Log JSON lines + ThreadLocal trace | Logback JSON encoder + Micrometer Tracing (trace id automatic, cross-thread) |
Hand-rolled Metrics + /metrics | Micrometer + Actuator /actuator/prometheus |
| Resilience4j builders in code | Resilience4j Spring Boot starter, config in application.yml |
/health/live, /health/ready | Actuator /actuator/health/liveness, /actuator/health/readiness |
Shutdown hook + server.stop(5) | server.shutdown=graceful + spring.lifecycle.timeout-per-shutdown-phase |
Config.resolve() precedence | Spring's property-source precedence (env > system properties > application.yml) |
| In-memory store + outbox list | Spring Data JPA (PostgreSQL) + transactional outbox table, relay to Kafka |
curl transcript | CI pipeline running the same checks on every push |
The controller — compare with createOrder in the lab:
// OrderController.java (Spring Boot 3.x — NOT RUN IN THIS LAB)
@RestController
@RequestMapping("/api/orders")
public class OrderController {
private final OrderService orders; // @Transactional service: JPA + outbox in one tx
@PostMapping
public ResponseEntity<Order> create(@Valid @RequestBody OrderRequest req) {
// @Valid runs the Jakarta constraints before this line executes —
// the lab's validate() method, declarative instead of imperative.
// Resilience4j annotations wrap the pricing call in the service layer.
return ResponseEntity.status(HttpStatus.CREATED).body(orders.create(req));
}
@GetMapping("/{id}")
public Order get(@PathVariable long id) {
return orders.get(id).orElseThrow(() -> new OrderNotFoundException(id));
}
}
// OrderRequest.java
public record OrderRequest(
@NotBlank String customerId,
@NotBlank String item,
@Min(1) @Max(1000) int qty) {}
The configuration — one file replacing the lab's Config, Log, Metrics, and resilience builders:
# application.yml (NOT RUN IN THIS LAB)
server:
port: ${APP_PORT:8080} # env wins, default 8080 — lab's Config.resolve
shutdown: graceful
spring:
threads.virtual.enabled: true # virtual threads, like the lab executor
datasource:
url: ${DB_URL} # secrets from env — never committed
username: ${DB_USER}
password: ${DB_PASSWORD}
jpa.properties.hibernate.jdbc.time_zone: UTC
lifecycle.timeout-per-shutdown-phase: 30s
kafka.producer.bootstrap-servers: ${KAFKA_BOOTSTRAP_SERVERS}
management:
endpoints.web.exposure.include: health,metrics,prometheus
endpoint.health.probes.enabled: true # /actuator/health/liveness + /readiness
metrics.tags.application: ordersvc
resilience4j:
retry:
instances.pricing:
max-attempts: 3
wait-duration: 200ms
ignore-exceptions: # fail-fast signals are not retried
- io.github.resilience4j.circuitbreaker.CallNotPermittedException
circuitbreaker:
instances.pricing:
failure-rate-threshold: 50
minimum-number-of-calls: 4
sliding-window-size: 4
wait-duration-in-open-state: 5s
timelimiter:
instances.pricing:
timeout-duration: 2s
logging:
pattern.console: '{"ts":"%d{ISO8601}","level":"%p","trace":"%X{traceId}",...}%n'
The container and the pipeline:
# Dockerfile (NOT RUN IN THIS LAB — multi-stage, no JDK in the runtime image)
FROM eclipse-temurin:21-jdk AS build
WORKDIR /app
COPY . .
RUN ./mvnw -q -DskipTests package
FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --from=build /app/target/ordersvc-*.jar app.jar
EXPOSE 8080
ENTRYPOINT ["java","-jar","app.jar"]
# .github/workflows/ci.yml (NOT RUN IN THIS LAB)
name: ci
on: [push]
jobs:
build:
runs-on: ubuntu-latest
services:
postgres: { image: postgres:16, env: { POSTGRES_PASSWORD: test } }
kafka: { image: bitnami/kafka:3.7 }
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { java-version: '21', distribution: 'temurin' }
- run: ./mvnw -q verify # unit + contract tests, Postgres + Kafka up
- run: docker build -t ordersvc:${{ github.sha }} .
Notice what the Spring version doesn't change: the checklist. Validation still at the boundary, retries still only for retryable failures, the breaker still fails fast, readiness still gates traffic, shutdown still drains. Frameworks change the spelling; the production properties stay the same. That's why the lab was worth building on the bare JDK — if you understand each mechanism without the framework, the framework's version is configuration, not magic.
Field check: earn the checklist yourself
Reading this post gives you the checklist; running these gives you the scars. Do all four on your own machine (JDK 21, the Resilience4j jars, no other dependencies):
- Predict, then run. Set
PRICING_FAILURES=7and write down the exact status-code sequence you expect for five consecutivePOST /api/orderscalls — then run it and compare. If your prediction was wrong, find the exact call where the breaker opened and explain why. (Hint: count circuit-breaker calls, not requests — retry multiplies them.) - Close the trace-id gap. The lab's retry logs carry
"trace":"-"because theThreadLocaldoesn't cross into the timeout-pool thread. Fix it: capture the trace id before submitting to the pool and restore it inside the task (try/finally to clear it). Re-run order 1 and show oneresilience.retryline carrying the request's trace id. - Drill the drain. Start the service, fire
curl "localhost:8080/api/orders?sleepMs=5000"in the background,kill -TERMthe server after half a second, and verify: the slow request still returns 200, the log showsshutdown.drainedwithinFlight:0, and the process exits 0. Then shorten the grace period below the sleep and watch what changes. - Audit your own service. Take the 12-row checklist from section 1 to the nearest thing you've deployed — a side project counts — and mark each row present, missing, or unknown. The three most common "missing" answers across readers will be 3 (secrets), 6 (breaker), and 9 (drain). Fix one this week.
What's next
This was the destination. The whole Production Java track in six lines:
- Observability: Logs, Metrics & Traces — structured JSON logs, trace IDs that follow a request, RED metrics with real percentiles; the difference between "grep 40,000 lines" and "grep one id."
- Configuration & Secrets: 12-Factor on the JVM — env beats system property beats default, every value logged with its source, secrets in the environment and never in code or logs.
- Resilience: Timeouts, Retries, Circuit Breakers & Bulkheads — the core rule: never retry blindly, only retry what retry can fix; deadlines bound your worst case; breakers turn per-request discovery into remembered knowledge.
- Production Deployments: Health Checks, Graceful Shutdown & Safe Rollouts — liveness vs readiness as two different questions, readiness-first shutdown ordering, rollouts that don't kill in-flight work.
- Load Testing & Performance Budgets — budgets before bottlenecks: k6/Gatling basics, p99 as the number users feel, testing the deployed shape not the laptop shape.
- Capstone: Production-Ready Spring Boot Service (this post) — all of it in one running service: REST, validation, structured logs, metrics, Resilience4j, health probes, graceful shutdown, tests, and the Spring Boot target it maps to.
And zoom out once more, because this post was never just the end of one track. Every track in this roadmap was a supply line to this service: Core Java gave you the language; Spring Boot & APIs gave you the web layer; Tooling & Testing gave you the build and the test suite; Concurrency & JVM gave you the virtual threads under the hood and the diagnosis toolkit for when it misbehaves; Data & Messaging gave you the persistence and event patterns the outbox is reaching for; Production Java gave you everything around the code that keeps the pager quiet. The capstone is where they all land.
From here, the roadmap hub is your map and your reference shelf — every post, every track, in order:
Continue: Java Learning Roadmap 2026
Go back through it with the checklist from section 1 in hand. Pick the service you care about most — the side project, the work project, the one that paged you last month — and run the twelve rows against it. The roadmap got you here; the checklist keeps you from getting paged at 2 AM. That was the whole point.
Comments
Post a Comment