Capstone: Production-Ready Spring Boot Service

It's your first deploy night. The order service passed every test, the demo went perfectly, and at 9 PM you shipped it. By 2 AM you have four separate fires and they are all the same fire:

  • Kubernetes killed a pod mid-request during the rollout — there was no graceful shutdown, so in-flight checkouts just died.
  • The one error you can see in the logs has no request id, so you can't tell which of the 40,000 log lines belong to the failing checkout.
  • The pricing service had a 30-second wobble. Your service retried every failed call instantly, with no backoff — turning their wobble into your retry storm, which turned into their outage.
  • And the database password? It's in application.yml, which is in git, which the new contractor cloned yesterday.

None of these are coding bugs. The code was correct. What's missing is everything around the code — the production checklist this whole track has been building: observability, configuration, resilience, safe deployments, and the load test that proves the budget. This capstone puts all of it into one service, running for real.

What "running for real" means here — stated up front. This machine has no PostgreSQL, no Kafka, no Redis, no Docker, and no Maven. So the lab service is built on the pure JDK (com.sun.net.httpserver, virtual threads) with real Resilience4j jars from Maven Central — every log line, metric, and status code below is real output from a real run. PostgreSQL, Kafka, Redis, and Docker appear as the target architecture and as a complete Spring Boot version of the same service, clearly marked not run in this lab. I show no Spring output because there is none to show.

The whole post in one diagram — bright boxes ran in this lab, dashed boxes are the production target:

Target architecture — bright ran in this lab, dashed is the production target Client curl · k6 · browser ordersvc — THIS LAB (pure JDK) HTTP server · virtual-thread executor Validation → 400s at the boundary Trace IDs · structured JSON logs Metrics registry → /metrics Resilience: bulkhead · timeout · retry · breaker flaky pricing simulator stands in for a real HTTP downstream in-memory store + outbox stands in for PostgreSQL PRODUCTION TARGET PostgreSQL (JPA) Kafka (outbox relay) Redis (cache) Prometheus + Grafana Loki (logs) · Tempo K8s → /health/* probes Docker image · CI/CD pipeline ran in this lab — real output below production target — Spring Boot section, not run here

1. The production-readiness checklist

Before the code: the checklist. Every item is one thing that bites you on deploy night, each taught in its own post earlier in this track, each implemented in the lab below. If you take one artifact from this post, take this table — run it against your own service:

#Checklist itemIn the labTaught in
1Structured logs with trace IDsOne JSON line per event; X-Request-Id on every responsePost 1 · Observability
2RED metrics: rate, errors, durationHand-rolled registry, Prometheus-style /metrics with p50/p95/p99Post 1 · Observability
3Config from the environment; secrets never in code or logsAPP_PORT precedence demo; DB_PASSWORD logged as <redacted>Post 2 · Configuration & Secrets
4Timeout on every downstream call2s deadline around the pricing call → real 504Post 3 · Resilience
5Retry with backoff — only for idempotent, retryable failures3 attempts, 200ms backoff; fail-fast signals are not retriedPost 3 · Resilience
6Circuit breaker — fail fast when the downstream is downCLOSED → OPEN → HALF_OPEN → CLOSED, all in the logs; 503 while openPost 3 · Resilience
7Bulkhead — bounded concurrency per downstreamSemaphore of 8; saturation → 429, not a pile-upPost 3 · Resilience
8Liveness vs readiness probes/health/live vs /health/ready (ready follows the breaker)Post 4 · Deployments
9Graceful shutdown with drainSIGTERM → stop sending traffic → up to 5s drain → exit 0Post 4 · Deployments
10Load test + latency budget100-request run feeding the p50/p95 percentilesPost 5 · Load Testing
11Automated tests: unit + contractJUnit 5, 12 tests, all greenTooling & Testing track
12CI/CD pipelineGitHub Actions workflow (Spring section — reviewed, not run)This capstone

2. The lab service: what it is and isn't

The lab is a small order service — POST /api/orders prices an item through a downstream pricing service, validates the request, stores the order, and exposes health and metrics endpoints. It is one Java program (ordersvc package, ~700 lines), compiled with javac on OpenJDK 21.0.3, using the real Resilience4j 2.1.0 jars (resilience4j-core, resilience4j-retry, resilience4j-circuitbreaker) downloaded from Maven Central. The HTTP layer is com.sun.net.httpserver on a virtual-thread-per-task executor — the same virtual threads from the Concurrency track, doing what they're for: one thread per request, blocking cheaply.

Honesty note, stated plainly: the "pricing service" is a simulator that fails its first N calls on purpose (deterministic, so the resilience demo is reproducible), and the "database" is a ConcurrentHashMap. Both are honest stand-ins with a defined seam: the pricing client is an interface, the store is a repository class, and the Spring Boot section shows what plugs into those seams in production. Everything else — the HTTP handling, validation, logging, metrics, retry/breaker behavior, config, probes, shutdown — is the real thing, running.

3. Configuration: the environment wins, secrets stay secret

Twelve-factor config with a fixed precedence — environment variable beats system property beats default — and the resolution is logged with its source, so "why is it listening on that port?" is always answerable:

// Config.java — precedence: env > system property > default
public static Entry resolve(String envKey, String sysKey, String def) {
    String env = System.getenv(envKey);
    if (env != null && !env.isBlank()) return new Entry(env, "env:" + envKey);
    String sys = System.getProperty(sysKey);
    if (sys != null && !sys.isBlank()) return new Entry(sys, "sysprop:" + sysKey);
    return new Entry(def, "default");
}

Started with both APP_PORT=18080 in the environment and -Dapp.port=18081 on the command line, the real startup log reads (timestamps trimmed for readability):

{"level":"INFO","logger":"ordersvc","trace":"-","event":"config.loaded",
 "port":18080,"port.source":"env:APP_PORT",
 "pricing.failures":4,"pricing.failures.source":"env:PRICING_FAILURES",
 "pricing.sleepMs":0,"db.password":"<redacted>"}

Port 18080 — the environment won, and the log proves which source each value came from. And DB_PASSWORD was set to a real secret for this run; it appears in the log only as <redacted>. The value exists in the process environment; it never appears in code, in logs, or in any committed file. That's the whole of Post 2's secrets rule in one line: a secret in a log is a secret already shared.

4. REST + validation at the boundary

Validation happens before anything downstream is touched — no pricing call, no store write, no wasted work for a request that was never valid:

// OrderService.java — the boundary contract, unit-tested
static List<String> validate(Map<String, String> body) {
    List<String> errs = new ArrayList<>();
    if (body.getOrDefault("customerId", "").trim().isEmpty()) errs.add("customerId is required");
    if (body.getOrDefault("item", "").trim().isEmpty()) errs.add("item is required");
    String q = body.get("qty");
    try {
        int qty = Integer.parseInt(q == null ? "" : q.trim());
        if (qty < 1 || qty > 1000) errs.add("qty must be between 1 and 1000");
    } catch (NumberFormatException e) {
        errs.add("qty must be an integer");
    }
    return errs;
}

Real responses from the running service:

$ curl -s -w "\nHTTP %{http_code}\n" -X POST localhost:18080/api/orders \
    -H 'Content-Type: application/json' -d '{"customerId":"","item":"widget","qty":0}'
{"error":"validation_failed","details":"customerId is required; qty must be between 1 and 1000"}
HTTP 400

$ curl -s -w "\nHTTP %{http_code}\n" localhost:18080/api/orders/999
{"error":"not_found","id":999}
HTTP 404

The 400 names every problem in one response — a client shouldn't have to play whack-a-mole with your validator. The rejected request also increments orders_rejected_total{reason="validation"}, because rejected traffic is a metric worth watching: a spike in 400s is usually a broken client deploy, and you want to see it before the client team tells you.

5. Structured logs with trace IDs

Every request gets a trace id at the edge, returned in the X-Request-Id header and attached to every log line the request produces. One JSON object per line — greppable, shippable to Loki/ELK, joinable with traces later:

$ curl -s -i -X POST localhost:18080/api/orders -H 'Content-Type: application/json' \
    -d '{"customerId":"c-42","item":"widget","qty":3}' | head -8
HTTP/1.1 502 Bad Gateway
Date: Tue, 06 Oct 2026 06:36:14 GMT
Content-type: application/json
X-request-id: 5e131f89
Content-length: 74

{"error":"pricing_failed","detail":"pricing service failed after retries"}

(This first order hit the simulated pricing outage — the resilience story starts in the next section. Watch the trace id 5e131f89:)

$ grep 5e131f89 server.log
{"level":"WARN","logger":"ordersvc","trace":"5e131f89","event":"pricing.failed",
 "error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 3)"}
{"level":"INFO","logger":"ordersvc","trace":"5e131f89","event":"request.done",
 "method":"POST","route":"orders","status":502,"durationMs":507}

Header and logs agree: one id ties the client's failed request to exactly its server-side story. That's the 2 AM difference between "grep 40,000 lines" and "grep one id".

One real imperfection, kept in: the retry-attempt log lines below carry "trace":"-" instead of the request's id — the trace id lives in a ThreadLocal, and the resilience work runs on a timeout-pool thread the ThreadLocal doesn't reach. In production, OpenTelemetry's context propagation carries the trace across thread hops automatically; a hand-rolled ThreadLocal does not. I'm showing the gap rather than hiding it, because "my trace ids vanish inside async code" is a classic production surprise.

6. Metrics: the RED numbers

Rate, Errors, Duration — the three numbers that tell you whether the service is healthy, on one /metrics endpoint in Prometheus text format. Counters carry labels (method, route, status); latencies feed p50/p95/p99. After the demo sequence plus a 100-request latency run, the real exposition:

$ curl -s localhost:18080/metrics
http_requests_total{method="GET",route="orders",status="200"} 100
http_requests_total{method="GET",route="orders",status="404"} 1
http_requests_total{method="GET",route="ready",status="200"} 1
http_requests_total{method="GET",route="ready",status="503"} 1
http_requests_total{method="POST",route="orders",status="201"} 1
http_requests_total{method="POST",route="orders",status="400"} 1
http_requests_total{method="POST",route="orders",status="502"} 1
http_requests_total{method="POST",route="orders",status="503"} 2
orders_created_total 1
orders_rejected_total{reason="validation"} 1
pricing_rejected_total 2
http_request_ms_count 108
http_request_ms_sum 775
http_request_ms_p50 0
http_request_ms_p95 2
http_request_ms_p99 231
circuitbreaker_state{breaker="pricing"} 0
circuitbreaker_failed_calls{breaker="pricing"} 0
circuitbreaker_not_permitted_calls{breaker="pricing"} 0

Read it like an operator: 108 requests, one 502, two fast 503s, one validation rejection, and the latency percentiles say the service itself is fast (p95 2ms) — the pain is all downstream, which is exactly what the next section fixes. The breaker gauges read 0 because the circuit is currently CLOSED; they're sliding-window counters, so they describe the current window, not history. In production this endpoint is scraped by Prometheus and the RED dashboard is the first thing you open when the pager goes off — Post 1's whole point.

One request's journey through every layer — the pipeline the rest of this post walks through:

1 · Trace ID assigned X-Request-Id header · ThreadLocal for the request's lifetime 2 · Validation bad input → 400 now; downstream never sees it 3 · Metrics + timer start http_requests_total · latency sample on completion 4 · Bulkhead (8 permits) saturated → 429 immediately, never a queue pile-up 5 · Timeout (2s deadline) one deadline around the whole pricing call → 504 6 · Retry ×3 + Circuit breaker 200ms backoff · breaker open → fail fast with 503 7 · Store + outbox, atomically one lock: the row and the OrderCreated event commit together 8 · Response + structured log 201 · one JSON line carrying the trace id Amber stages are the resilience stack — Post 3 of this track, running below.

7. Resilience: retry, circuit breaker, bulkhead, timeout

This is Post 3's toolkit, wired as one stack around the pricing call — bulkhead outermost, then the timeout deadline, then retry, then the circuit breaker closest to the call:

// Pricing.java — the resilience stack (real Resilience4j 2.1.0)
Retry retry = Retry.of("pricing", RetryConfig.custom()
        .maxAttempts(3)
        .waitDuration(Duration.ofMillis(200))          // backoff between attempts
        .retryExceptions(RuntimeException.class)
        .ignoreExceptions(CallNotPermittedException.class, // fail-fast: never retry
                          BulkheadFullException.class)     // saturated: never retry
        .build());

CircuitBreaker cb = CircuitBreaker.of("pricing", CircuitBreakerConfig.custom()
        .failureRateThreshold(50)          // ≥50% failures…
        .minimumNumberOfCalls(4)           // …over at least 4 calls…
        .slidingWindowSize(4)              // …in the last 4 calls → OPEN
        .waitDurationInOpenState(Duration.ofSeconds(5)) // then one half-open probe
        .build());

// Retry outermost, breaker innermost: Retry(CircuitBreaker(call))
Supplier<Double> guarded = Retry.decorateSupplier(retry,
        CircuitBreaker.decorateSupplier(cb, () -> callDelegate(item)));

The core rule from Post 3, enforced in code: never retry blindly — is the failure retryable? A downstream 500 might clear; a circuit-open rejection or a saturated bulkhead will not, so those are excluded from retry. Retrying a fail-fast signal is pure load — the retry storm from the opening incident, rebuilt as a feature.

Now watch it work. The pricing simulator is configured to fail its first 4 calls — a deterministic downstream outage:

Order 1 — retry fights, retry loses. Three attempts, 200ms apart, all fail. The log shows the backoff working exactly as configured, and the request still fails — correctly — after 507ms:

{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.retry",
 "retryAttempt":1,"error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 1)"}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.retry",
 "retryAttempt":2,"error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 2)"}
{"level":"WARN","logger":"ordersvc","trace":"5e131f89","event":"pricing.failed",
 "error":"java.lang.RuntimeException: pricing service unavailable (simulated failure 3)"}
{"level":"INFO","logger":"ordersvc","trace":"5e131f89","event":"request.done",
 "method":"POST","route":"orders","status":502,"durationMs":507}

Order 2 — the breaker opens mid-request. The fourth failed call trips the threshold (4 calls, 100% failure ≥ 50%), and the breaker opens while the request is still being retried:

{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.circuit",
 "from":"CLOSED","to":"OPEN"}
$ curl -s -X POST localhost:18080/api/orders -H 'Content-Type: application/json' \
    -d '{"customerId":"c-43","item":"widget","qty":1}'
{"error":"pricing_unavailable","detail":"circuit breaker is OPEN — failing fast","retryable":true}

Order 3 — fail fast. With the breaker open, the request never touches the downstream at all:

$ time curl -s -o /dev/null -w "HTTP %{http_code}\n" -X POST localhost:18080/api/orders \
    -H 'Content-Type: application/json' -d '{"customerId":"c-44","item":"widget","qty":1}'
HTTP 503

real    0m0.011s

11 milliseconds. Compare: the first order burned 507ms plus three downstream calls to discover the outage; this one spends 11ms and zero downstream calls to report it. That gap — between "discover the outage per request" and "remember the outage across requests" — is the entire economic argument for a circuit breaker.

After the 5-second open window — the half-open probe. One request is let through; it succeeds; the breaker closes:

{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.circuit",
 "from":"OPEN","to":"HALF_OPEN"}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"resilience.circuit",
 "from":"HALF_OPEN","to":"CLOSED"}
$ curl -s -X POST localhost:18080/api/orders -H 'Content-Type: application/json' \
    -d '{"customerId":"c-45","item":"gadget","qty":2}'
{"id":1,"customerId":"c-45","item":"gadget","qty":2,"price":19.99,"status":"CREATED"}

The state machine, as configured:

CLOSED calls flow through OPEN fail fast · 503 HALF_OPEN one probe call ≥50% fail / 4 calls after 5s probe succeeds → CLOSED · probe fails → back to OPEN

The timeout story, on a second instance. With the downstream configured to sleep 5 seconds per call and the deadline at 2 seconds, one order returns 504 — and note the elapsed time:

$ time curl -s -w "\nHTTP %{http_code}\n" -X POST localhost:18081/api/orders \
    -H 'Content-Type: application/json' -d '{"customerId":"c-9","item":"widget","qty":1}'
{"error":"pricing_timeout","detail":"pricing timed out after 2000ms"}
HTTP 504

real    0m2.180s
{"level":"WARN","logger":"ordersvc","trace":"3eda351c","event":"resilience.timeout",
 "item":"widget","timeoutMs":2000}
{"level":"INFO","logger":"ordersvc","trace":"3eda351c","event":"request.done",
 "method":"POST","route":"orders","status":504,"durationMs":2099}

2.1 seconds, not 5 — the deadline did its job. And here's the design detail worth keeping: the 2s deadline wraps the whole guarded call — retry included — so a hung downstream costs this request ~2s total, not attempts × timeout. That's the overall-deadline pattern: per-attempt backoff inside, one deadline outside. Your worst case is a number you chose, not a number you discover.

What the timeout does not do — observed, not theorized: the log shows a second pricing.call attempt firing after the timeout already returned 504 to the client. Cancelling the future interrupted the sleeping attempt, but the retry living inside the cancelled task fired its next attempt anyway — the orphan kept a thread busy for its full 5-second sleep while nobody waited for the answer. Timeouts protect your latency; they don't stop work already in flight downstream. The bulkhead (8 permits) is what bounds how many such orphans can pile up. Defense in depth means no single mechanism has to be perfect.

One more Post 3 rule, visible in the 400s from section 4: the pricing call is an idempotent read, so retrying it is safe. If this were POST /payments/charge, blind retry would double-charge — that's why the retry config retries only failures, never fail-fast signals, and why non-idempotent operations get idempotency keys instead of retries. The rule travels with you; the config stays here.

8. Liveness vs readiness: two different questions

Post 4's distinction, in one handler: liveness asks "is the process alive?" (if not, restart it); readiness asks "should this instance receive traffic right now?" (if not, pull it from the load balancer but leave it running). Conflating them is how you get Kubernetes restarting a pod that was merely waiting on a downstream — turning a wobble into a restart loop:

// OrderService.java — readiness reflects the world, liveness reflects the process
private int handleReady(HttpExchange ex) throws IOException {
    String pricingState = pricing.breaker().getState().name();
    boolean pricingOk = !"OPEN".equals(pricingState);
    boolean up = ready && pricingOk;   // ready flips false the moment shutdown begins
    String body = "{\"status\":\"" + (up ? "UP" : "DOWN") + "\",\"checks\":{"
            + "\"store\":\"UP\","
            + "\"pricing\":\"" + (pricingOk ? "UP" : "circuit OPEN") + "\"}}";
    return send(ex, up ? 200 : 503, body);
}

While the breaker was open in the demo above, readiness told the truth:

$ curl -s -w "\nHTTP %{http_code}\n" localhost:18080/health/ready
{"status":"DOWN","checks":{"store":"UP","pricing":"circuit OPEN"}}
HTTP 503

…and after the half-open probe closed the circuit:

$ curl -s -w "\nHTTP %{http_code}\n" localhost:18080/health/ready
{"status":"UP","checks":{"store":"UP","pricing":"UP"}}
HTTP 200

The pod stays alive the whole time — no restart, no lost warmup — it just stops receiving traffic while its downstream is down, then rejoins when the probe succeeds. In production the checks object grows a real database ping and a Kafka producer check; the shape doesn't change.

9. Graceful shutdown: the drain

The opening incident's first fire — Kubernetes killing a pod mid-request. The fix is a shutdown hook with an order of operations: fail readiness first (so the load balancer stops sending new traffic), then wait for in-flight requests to finish, with a bounded grace period:

// OrderService.java — main()
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
    Log.warn("shutdown.signal",
        "detail", "SIGTERM received — refusing new traffic, draining in-flight requests");
    svc.ready = false;          // 1. readiness fails: load balancer looks away
    server.stop(5);             // 2. up to 5s for in-flight requests to finish
    Log.warn("shutdown.drained", "inFlight", svc.inFlight.get());
    Log.warn("shutdown.complete", "detail", "graceful shutdown finished");
}, "shutdown-hook"));

Real run: a 3-second request in flight, SIGTERM delivered half a second in:

$ curl -s "localhost:18080/api/orders?sleepMs=3000" > /dev/null &
$ sleep 0.5; kill -TERM $PID; wait $PID

{"level":"WARN","logger":"ordersvc","trace":"-","event":"shutdown.signal",
 "detail":"SIGTERM received — refusing new traffic, draining in-flight requests"}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"shutdown.drained","inFlight":0}
{"level":"WARN","logger":"ordersvc","trace":"-","event":"shutdown.complete",
 "detail":"graceful shutdown finished"}

The slow request completed inside the grace period (inFlight:0 — nothing was abandoned), the process exited 0, and the whole sequence took ~2.6 seconds. Note the ordering subtlety the hook gets right: if you stop the server before failing readiness, there's a window where the load balancer still routes to a dying pod. Readiness first, always.

10. Tests: 12 green, and what each group proves

JUnit 5 via the platform console launcher (no Maven on this machine — the standalone jar runs anywhere). The full suite:

$ java -jar junit-platform-console-standalone-1.10.3.jar \
    --class-path "out:lib/..." --select-package ordersvc
[        12 tests found           ]
[        12 tests started         ]
[        12 tests successful      ]
[         0 tests failed          ]

What the 12 cover, by group:

  • Validation contract (5 tests) — blank ids, zero/negative/non-numeric quantities, and the 1000/1001 boundary. The boundary behavior is the specification; the test is where it's written down.
  • Config precedence (2 tests) — system property beats default; default applies when nothing is set. (Environment-variable precedence is proven by the live startup log in section 3 — you can't set env vars from inside the JVM under test, so the honest proof lives in the run, not the suite.)
  • Transactional outbox (1 test) — creating an order writes the row and the OrderCreated event atomically. The test drains the outbox and asserts the event carries the right order id — the contract the Kafka relay will depend on.
  • Metrics math (2 tests) — percentiles computed from known samples (p50=30, p95=50 on [10..50]); counters increment per label-set and render in the exposition format.
  • Resilience4j behavior (2 tests) — retry really attempts 3 times before succeeding; the breaker really opens after the failure threshold and the guarded supplier body never runs again while open (asserted with a call counter — the fail-fast property, pinned down).
One sandbox limitation, stated plainly: this machine blocks raw TCP from the JVM, so an in-JVM HTTP end-to-end test (Java HttpClient against the running server) cannot run here — it fails with a sandbox policy error, not an application error. The HTTP round-trips are instead proven by the curl transcript throughout this post: real client, real server, real status codes. Twelve of twelve runnable tests green; the thirteenth proof is the demo above.

11. The Spring Boot target — the same service, production-shaped

Honesty note, stated plainly and once: everything in this section was not run in this lab — no Maven, no Docker, no PostgreSQL, no Kafka, no Redis on this machine. The code is written as the production target the lab maps to: reviewed carefully, but unexecuted. There is no Spring output below because none exists. Run it yourself with ./mvnw spring-boot:run and docker compose up; the lab outputs above are the behavior to compare against.

The lab's hand-rolled pieces each have a Spring Boot equivalent. The mapping, first — then the code:

Lab piece (ran above)Spring Boot equivalent (below, not run)
com.sun.net.httpserver + virtual threadsSpring MVC on virtual threads (spring.threads.virtual.enabled=true)
Hand-rolled validate()Jakarta Bean Validation (@Valid, @NotBlank, @Min/@Max)
Log JSON lines + ThreadLocal traceLogback JSON encoder + Micrometer Tracing (trace id automatic, cross-thread)
Hand-rolled Metrics + /metricsMicrometer + Actuator /actuator/prometheus
Resilience4j builders in codeResilience4j Spring Boot starter, config in application.yml
/health/live, /health/readyActuator /actuator/health/liveness, /actuator/health/readiness
Shutdown hook + server.stop(5)server.shutdown=graceful + spring.lifecycle.timeout-per-shutdown-phase
Config.resolve() precedenceSpring's property-source precedence (env > system properties > application.yml)
In-memory store + outbox listSpring Data JPA (PostgreSQL) + transactional outbox table, relay to Kafka
curl transcriptCI pipeline running the same checks on every push

The controller — compare with createOrder in the lab:

// OrderController.java (Spring Boot 3.x — NOT RUN IN THIS LAB)
@RestController
@RequestMapping("/api/orders")
public class OrderController {

    private final OrderService orders;   // @Transactional service: JPA + outbox in one tx

    @PostMapping
    public ResponseEntity<Order> create(@Valid @RequestBody OrderRequest req) {
        // @Valid runs the Jakarta constraints before this line executes —
        // the lab's validate() method, declarative instead of imperative.
        // Resilience4j annotations wrap the pricing call in the service layer.
        return ResponseEntity.status(HttpStatus.CREATED).body(orders.create(req));
    }

    @GetMapping("/{id}")
    public Order get(@PathVariable long id) {
        return orders.get(id).orElseThrow(() -> new OrderNotFoundException(id));
    }
}

// OrderRequest.java
public record OrderRequest(
        @NotBlank String customerId,
        @NotBlank String item,
        @Min(1) @Max(1000) int qty) {}

The configuration — one file replacing the lab's Config, Log, Metrics, and resilience builders:

# application.yml (NOT RUN IN THIS LAB)
server:
  port: ${APP_PORT:8080}                       # env wins, default 8080 — lab's Config.resolve
  shutdown: graceful
spring:
  threads.virtual.enabled: true                # virtual threads, like the lab executor
  datasource:
    url: ${DB_URL}                             # secrets from env — never committed
    username: ${DB_USER}
    password: ${DB_PASSWORD}
  jpa.properties.hibernate.jdbc.time_zone: UTC
  lifecycle.timeout-per-shutdown-phase: 30s
  kafka.producer.bootstrap-servers: ${KAFKA_BOOTSTRAP_SERVERS}
management:
  endpoints.web.exposure.include: health,metrics,prometheus
  endpoint.health.probes.enabled: true          # /actuator/health/liveness + /readiness
  metrics.tags.application: ordersvc
resilience4j:
  retry:
    instances.pricing:
      max-attempts: 3
      wait-duration: 200ms
      ignore-exceptions:                        # fail-fast signals are not retried
        - io.github.resilience4j.circuitbreaker.CallNotPermittedException
  circuitbreaker:
    instances.pricing:
      failure-rate-threshold: 50
      minimum-number-of-calls: 4
      sliding-window-size: 4
      wait-duration-in-open-state: 5s
  timelimiter:
    instances.pricing:
      timeout-duration: 2s
logging:
  pattern.console: '{"ts":"%d{ISO8601}","level":"%p","trace":"%X{traceId}",...}%n'

The container and the pipeline:

# Dockerfile (NOT RUN IN THIS LAB — multi-stage, no JDK in the runtime image)
FROM eclipse-temurin:21-jdk AS build
WORKDIR /app
COPY . .
RUN ./mvnw -q -DskipTests package

FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --from=build /app/target/ordersvc-*.jar app.jar
EXPOSE 8080
ENTRYPOINT ["java","-jar","app.jar"]
# .github/workflows/ci.yml (NOT RUN IN THIS LAB)
name: ci
on: [push]
jobs:
  build:
    runs-on: ubuntu-latest
    services:
      postgres: { image: postgres:16, env: { POSTGRES_PASSWORD: test } }
      kafka:    { image: bitnami/kafka:3.7 }
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { java-version: '21', distribution: 'temurin' }
      - run: ./mvnw -q verify          # unit + contract tests, Postgres + Kafka up
      - run: docker build -t ordersvc:${{ github.sha }} .

Notice what the Spring version doesn't change: the checklist. Validation still at the boundary, retries still only for retryable failures, the breaker still fails fast, readiness still gates traffic, shutdown still drains. Frameworks change the spelling; the production properties stay the same. That's why the lab was worth building on the bare JDK — if you understand each mechanism without the framework, the framework's version is configuration, not magic.

Capstone principle: production-ready is a checklist, not a framework. Every item on it is a deploy-night scar someone else already earned. Run the table from section 1 against your own service once a quarter; the one you skip is the one that pages you.

Field check: earn the checklist yourself

Reading this post gives you the checklist; running these gives you the scars. Do all four on your own machine (JDK 21, the Resilience4j jars, no other dependencies):

  1. Predict, then run. Set PRICING_FAILURES=7 and write down the exact status-code sequence you expect for five consecutive POST /api/orders calls — then run it and compare. If your prediction was wrong, find the exact call where the breaker opened and explain why. (Hint: count circuit-breaker calls, not requests — retry multiplies them.)
  2. Close the trace-id gap. The lab's retry logs carry "trace":"-" because the ThreadLocal doesn't cross into the timeout-pool thread. Fix it: capture the trace id before submitting to the pool and restore it inside the task (try/finally to clear it). Re-run order 1 and show one resilience.retry line carrying the request's trace id.
  3. Drill the drain. Start the service, fire curl "localhost:8080/api/orders?sleepMs=5000" in the background, kill -TERM the server after half a second, and verify: the slow request still returns 200, the log shows shutdown.drained with inFlight:0, and the process exits 0. Then shorten the grace period below the sleep and watch what changes.
  4. Audit your own service. Take the 12-row checklist from section 1 to the nearest thing you've deployed — a side project counts — and mark each row present, missing, or unknown. The three most common "missing" answers across readers will be 3 (secrets), 6 (breaker), and 9 (drain). Fix one this week.

What's next

This was the destination. The whole Production Java track in six lines:

  • Observability: Logs, Metrics & Traces — structured JSON logs, trace IDs that follow a request, RED metrics with real percentiles; the difference between "grep 40,000 lines" and "grep one id."
  • Configuration & Secrets: 12-Factor on the JVM — env beats system property beats default, every value logged with its source, secrets in the environment and never in code or logs.
  • Resilience: Timeouts, Retries, Circuit Breakers & Bulkheads — the core rule: never retry blindly, only retry what retry can fix; deadlines bound your worst case; breakers turn per-request discovery into remembered knowledge.
  • Production Deployments: Health Checks, Graceful Shutdown & Safe Rollouts — liveness vs readiness as two different questions, readiness-first shutdown ordering, rollouts that don't kill in-flight work.
  • Load Testing & Performance Budgets — budgets before bottlenecks: k6/Gatling basics, p99 as the number users feel, testing the deployed shape not the laptop shape.
  • Capstone: Production-Ready Spring Boot Service (this post) — all of it in one running service: REST, validation, structured logs, metrics, Resilience4j, health probes, graceful shutdown, tests, and the Spring Boot target it maps to.

And zoom out once more, because this post was never just the end of one track. Every track in this roadmap was a supply line to this service: Core Java gave you the language; Spring Boot & APIs gave you the web layer; Tooling & Testing gave you the build and the test suite; Concurrency & JVM gave you the virtual threads under the hood and the diagnosis toolkit for when it misbehaves; Data & Messaging gave you the persistence and event patterns the outbox is reaching for; Production Java gave you everything around the code that keeps the pager quiet. The capstone is where they all land.

From here, the roadmap hub is your map and your reference shelf — every post, every track, in order:

Continue: Java Learning Roadmap 2026

Go back through it with the checklist from section 1 in hand. Pick the service you care about most — the side project, the work project, the one that paged you last month — and run the twelve rows against it. The roadmap got you here; the checklist keeps you from getting paged at 2 AM. That was the whole point.

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number