Posts

Capstone: Production-Ready Spring Boot Service

It's your first deploy night. The order service passed every test, the demo went perfectly, and at 9 PM you shipped it. By 2 AM you have four separate fires and they are all the same fire: Kubernetes killed a pod mid-request during the rollout — there was no graceful shutdown, so in-flight checkouts just died. The one error you can see in the logs has no request id, so you can't tell which of the 40,000 log lines belong to the failing checkout. The pricing service had a 30-second wobble. Your service retried every failed call instantly, with no backoff — turning their wobble into your retry storm, which turned into their outage. And the database password? It's in application.yml , which is in git, which the new contractor cloned yesterday. None of these are coding bugs. The code was correct. What's missing is everything around the code — the production checklist this whole track has been building: observability, configuration, resilience, safe deployments...

Load Testing & Performance Budgets

Six weeks after the checkout service from the Spring track went live, the marketing team ran their biggest promotion of the year. At 9:04 AM the traffic was 40x normal. At 9:06 AM the support channel caught fire: "the site is frozen." The dashboards told a story nobody had seen in staging — average response time was a perfectly respectable 220 ms, but one in twenty checkouts was taking over 8 seconds , and the payment page was timing out for the unluckiest customers. The team restarted the service, traffic fell, and the incident "resolved itself." Nobody could say what the service could actually handle, because nobody had ever asked it. That question — what can this service actually handle, and how does it behave when it can't? — is load testing. Not "does it work" (functional tests answer that) but "does it work under load , and where exactly does it stop working." This post builds that discipline for Java APIs: what a performance budget ...

Production Deployments: Health Checks, Graceful Shutdown & Safe Rollouts

The deploy went out at 4 PM on a Friday. Four instances behind a load balancer, one after another, the pipeline reported green. Then the support channel lit up: "checkout fails with a connection error, but only sometimes." Retry and it worked. Wait five minutes and it worked. The team stared at the green pipeline, the healthy dashboards, the passing tests — and at customers getting 502 Bad Gateway in a neat one-in-four pattern. The rolling deploy had done exactly what it was told. The service, however, had done something nobody told it not to do: every instance accepted traffic the moment its JVM started, and dropped every in-flight request the moment SIGTERM arrived. This post is the missing half of "it works on my machine": health checks, graceful shutdown, and safe rollouts . You'll build a service that tells the load balancer the truth about itself ( /health/live vs /health/ready ), a JVM that refuses to drop a single in-flight request on shutdown, a...