Production Diagnosis: Thread Dumps, Heap Dumps, JFR & async-profiler
It's 2 AM. The API's p99 went from 150ms to 3 seconds, CPU is pinned at 95%, and it works fine on your laptop. Restarting "fixed" it last time — for four hours. This post is the toolkit for the time restarting doesn't fix it: four tools, each answering one question, mapped to the symptom in front of you. 1. The scenario we'll diagnose Two failure shapes cover most production mysteries. Learn to tell them apart first, because each one reaches for a different tool: High CPU, slow responses — threads are doing something expensive (hot loop, lock contention, pathological regex). The threads are guilty; find what they're executing. Growing memory, eventual OOM or long GC pauses — objects are accumulating . The heap is guilty; find what's being retained. Requests hang forever, CPU idle — threads are waiting on each other. Nobody's guilty yet; find the deadlock. Here's a lab specimen containing all three sins. Run it, then diagnose it li...