Observability: Logs, Metrics & Traces
2:47 AM. The pager's summary line says checkout p99 latency tripled an hour ago and the error rate is climbing. You open the logs for the checkout service from the Spring track — the one that went live last month — and start scrolling. Forty thousand lines since midnight. START GET /api/orders. END in 212 ms. START GET /api/orders. END in 9 ms. Somewhere in there, a payment-gateway timeout fired. Somewhere in there, a slow request poisoned everything behind it. You have three questions, and the logs — all forty thousand lines — cannot answer a single one: which requests were slow? how many failed? where did the time go inside the slow ones?
This post is about the moment you stop reading logs like a novel and start measuring a system. Observability is not a product you buy; it's the discipline of building three signals into your service so those three questions always have answers: logs (what happened to one request), metrics (what's happening across all requests), and traces (where time went inside one request). Every snippet below was compiled with javac and run on OpenJDK 21.0.3, and every output block is the real output — including the numbers that surprised me, which I'll flag rather than hide.
By the end, you'll have instrumented a small HTTP service three times over — structured JSON logs with request IDs, Micrometer timers printing real p50/p95 latencies, and OpenTelemetry spans showing a full failed trace — and you'll know exactly which signal answers which 2:47 AM question.
One request, three signals
Keep the three pillars separated by the question each one answers, not by the tool that produces it:
Principle: logs tell you what happened to one request, metrics tell you what's happening to all requests, traces tell you where time went inside one request. When someone proposes a fourth pillar or a vendor dashboard that "does observability," come back to the three questions. If it can't answer all three, it's a feature, not observability.
Exhibit A: the before — println archaeology
Here's the service as most of us write it first: a com.sun.net.httpserver.HttpServer with three routes — a fast one, a slow one, and a fragile one that sometimes throws — logged with System.out.println. The load generator fires 60 real HTTP requests at it from 6 threads. (One lab note, stated plainly: this sandbox blocks the JDK HTTP client at the network layer, so the load generator shells out to curl, which works fine here. The server, the HTTP traffic, and the log mess below are all real.)
import com.sun.net.httpserver.*;
import java.io.*;
import java.net.*;
import java.util.concurrent.*;
/** The "before": println debugging, no request ids, no metrics, no traces. */
public class Before {
public static void main(String[] args) throws Exception {
HttpServer server = HttpServer.create(new InetSocketAddress("127.0.0.1", 18081), 0);
server.createContext("/api/orders", ex -> {
long t0 = System.nanoTime();
String thread = Thread.currentThread().getName();
System.out.println("[" + thread + "] START " + ex.getRequestMethod()
+ " " + ex.getRequestURI().getPath());
try {
Thread.sleep(3 + (long) (Math.random() * 12));
byte[] body = "{\"orders\":[]}".getBytes();
ex.sendResponseHeaders(200, body.length);
try (OutputStream os = ex.getResponseBody()) { os.write(body); }
} catch (Exception e) { System.out.println("ERROR: " + e); }
System.out.println("[" + thread + "] END in " + (System.nanoTime() - t0) / 1_000_000 + " ms");
});
// ... /api/orders/slow (sleeps 150-220 ms) and /api/orders/fragile
// (throws RuntimeException("payment gateway timeout") ~1/3 of the time)
// follow the same shape.
ExecutorService serverExec = Executors.newFixedThreadPool(6);
server.setExecutor(serverExec);
server.start();
// 60 requests via curl subprocesses, 6 at a time; then server.stop(0)
}
}
Compile, run, and look at what 60 requests produce:
$ javac Before.java && java Before
[pool-1-thread-2] START GET /api/orders
[pool-1-thread-5] START GET /api/orders
[pool-1-thread-3] START GET /api/orders/fragile
[pool-1-thread-6] START GET /api/orders
[pool-1-thread-4] START GET /api/orders/slow
[pool-1-thread-1] START GET /api/orders
[pool-1-thread-1] END in 231 ms
[pool-1-thread-6] END in 235 ms
[pool-1-thread-2] END in 190 ms
[pool-1-thread-5] END in 182 ms
[pool-1-thread-4] END in 200 ms
[pool-1-thread-3] END
...
[pool-1-thread-4] START GET /api/orders/fragile
ERROR: java.lang.RuntimeException: payment gateway timeout
[pool-1-thread-4] END
...
--- server stopped ---
Now answer the three 2:47 AM questions from this output. Which requests were slow? The END in 231 ms line belongs to a request on thread pool-1-thread-1 — but which of its two STARTs? Thread names get reused across requests, so you can't pair them. How many failed? Count the ERROR lines by hand — five in this run — and each one is a bare exception with no route, no timestamp, no request identity. Where did the time go? Unknowable: there's no record of anything between START and END.
And notice something else in the real output: the fast route's requests report 182–235 ms. The handler only sleeps 3–12 ms. Those extra ~200 ms are queueing — the 6-thread pool is saturated behind the slow route, so fast requests wait for a thread. The println log can't tell you that; it just looks like the fast route got slow. Principle: a log line without a request ID is a fact without a subject. You can't correlate it, you can't trust the story it seems to tell.
Pillar 1: logs — structured, and correlated
The first fix is the cheapest: stop writing sentences for humans and start writing records for machines. A structured log line is a JSON object — timestamp, level, message, and every piece of context as a named field. Humans can still read JSON; grep, jq, and every log aggregator ever built read it far better than they read your prose.
The Java stack for this is three layers, and the layering is the point. SLF4J is the API — your code calls org.slf4j.Logger and never imports an implementation. Logback is the engine that actually writes the lines. logstash-logback-encoder is the JSON formatter. Code to the facade, configure the engine:
// pom-free setup used in this lab: jars on the classpath via curl
// slf4j-api-2.0.20, logback-classic-1.6.5, logback-core-1.6.5,
// logstash-logback-encoder-9.0 (+ Jackson 3: jackson-core/databind-3.2.3)
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.slf4j.MDC;
public class ObsService {
static final Logger log = LoggerFactory.getLogger(ObsService.class);
// ... in the request handler:
MDC.put("requestId", requestId); // short id from X-Request-ID header
MDC.put("traceId", span.getSpanContext().getTraceId()); // 32-hex OTel trace id
try {
log.info("request received");
// ... do the work ...
log.info("request finished status={} latencyMs={}", status, ms);
} finally {
MDC.clear(); // pooled threads: never leak one request's ids into the next
}
}
MDC — the Mapped Diagnostic Context — is a ThreadLocal map that the logging engine merges into every line that thread writes. And the logback.xml that turns it all into JSON is five meaningful lines:
<configuration>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder class="net.logstash.logback.encoder.LogstashEncoder">
<timestampPattern>yyyy-MM-dd'T'HH:mm:ss.SSSXXX</timestampPattern>
</encoder>
</appender>
<root level="INFO"><appender-ref ref="STDOUT"/></root>
</configuration>
The same 150-request run, now with every line carrying its identity — these are real lines from the real run:
$ java -cp ".:lib/*" ObsService
{"@timestamp":"2026-10-06T06:42:17.043Z","@version":"1","message":"request received",
"logger_name":"ObsService","thread_name":"pool-1-thread-4","level":"INFO",
"traceId":"73d8b3dda2abda73d1e4b5993066bea3","requestId":"c4e57d0d"}
{"@timestamp":"2026-10-06T06:42:17.043Z","@version":"1","message":"request received",
"logger_name":"ObsService","thread_name":"pool-1-thread-3","level":"INFO",
"traceId":"186fcb4a43b6a0142d34379b552b0a57","requestId":"e5101c44"}
...
Now the 2:47 AM question — what happened to the failed request? — is one grep:
$ grep 9a00365749c98a2aaf39e0a74c5a461e obsservice.log
{"@timestamp":"2026-10-06T06:42:17.267Z",...,"message":"request received",...,
"traceId":"9a00365749c98a2aaf39e0a74c5a461e","requestId":"86c1c72c"}
{"@timestamp":"2026-10-06T06:42:17.305Z",...,"message":"request failed: java.lang.RuntimeException: payment gateway timeout",
"level":"WARN",...,"traceId":"9a00365749c98a2aaf39e0a74c5a461e","requestId":"86c1c72c"}
{"@timestamp":"2026-10-06T06:42:17.306Z",...,"message":"request finished status=500 latencyMs=38",...,
"traceId":"9a00365749c98a2aaf39e0a74c5a461e","requestId":"86c1c72c"}
Three lines, one request, complete story: received at :17.267, the payment gateway timed out 38 ms later, the request finished with status 500. Compare that with the before-run's bare ERROR: java.lang.RuntimeException: payment gateway timeout floating in a sea of thread names. Principle: if you can't grep it, you can't debug it — structure logs as JSON from day one, and put the request ID in every line via MDC.
Three rules that keep logs useful instead of turning them into a second problem:
1. Levels are a contract, not a mood. ERROR means a human should look; WARN means degraded but handled (our failed payment is a WARN — the request failed, the service didn't); INFO means business-significant events (request received/finished); DEBUG is off in production. If everything is ERROR, nothing is.
2. MDC.clear() is the contract on pooled threads. The MDC is a ThreadLocal; server threads are reused. Forget the finally block and request B's log lines wear request A's ID — worse than no ID at all, because now your correlation lies. (Field check #2 below has you prove this to yourself.)
3. Never log secrets, tokens, or raw PII. Logs get shipped to aggregators, stored for months, and read by people who shouldn't see passwords. Minimize the data first — log the user id, not the user object — and never claim redaction "guarantees" anything; the guarantee is not collecting it.
Pillar 2: metrics — stop counting with logs
Logs answer "what happened to one request." For "how is the service right now" you need aggregates: how many requests per second, what fraction failed, how slow the slow ones are. You can technically compute those from logs — and at 2:47 AM, with 40k lines a minute, you will discover exactly why nobody does. Metrics are numbers your code records as events happen, pre-aggregated into counts, sums, and latency histograms, ready to graph and alert on.
Micrometer is the SLF4J of metrics: a vendor-neutral facade (Timer, Counter, Gauge) with one registry per backend. In production the registry ships to Prometheus or Datadog; in this lab there's no backend available — so we use SimpleMeterRegistry, which keeps everything in memory and lets us print the aggregates. The instrumentation code is identical either way; only the registry changes. Pick the verb first: Timer measures durations, Counter counts events, Gauge samples a current value.
import io.micrometer.core.instrument.*;
import io.micrometer.core.instrument.simple.SimpleMeterRegistry;
static final MeterRegistry registry = new SimpleMeterRegistry();
// one timer per route template — tagged, so /api/orders/slow never
// pollutes /api/orders. Never tag by user id or request id: cardinality kills.
static Timer timer(String route) {
return timers.computeIfAbsent(route, r -> Timer.builder("http.server.requests")
.description("Request latency per route")
.tags("route", r)
.publishPercentiles(0.5, 0.95)
.register(registry));
}
// in the handler's finally block:
timer(route).record(System.nanoTime() - t0, TimeUnit.NANOSECONDS);
// on failure:
Counter.builder("http.server.errors").tags("route", route)
.register(registry).increment();
Two honest lab notes before the numbers. First: Micrometer needs HdrHistogram on the classpath for percentile histograms — without HdrHistogram-2.2.2.jar, creating the timer threw NoClassDefFoundError: org/HdrHistogram/DoubleRecorder. It failed loudly at startup, which is the good kind of failure; a metrics library that silently dropped percentiles would be the bad kind. Second: this run's numbers contain a surprise I'm keeping, because it's the metrics doing their job:
=== METRICS (Micrometer SimpleMeterRegistry) ===
route=/api/orders count= 80 mean= 25.4ms max= 189.9ms p50= 12.3ms p95= 192.7ms
route=/api/orders/slow count= 40 mean= 165.3ms max= 203.9ms p50= 163.6ms p95= 205.5ms
route=/api/orders/fragile count= 30 mean= 22.4ms max= 39.2ms p50= 22.5ms p95= 27.8ms
errors_total=10 error_rate=6.7% throughput=70.6 req/s
Read the fast route's row twice. The handler sleeps 3–12 ms, yet p95 is 192.7 ms — sixteen times the handler time. That is not a measurement bug; it's the queueing we spotted in the println output, now quantified: the 6-thread pool saturates behind the slow route, and fast requests wait. The mean (25.4 ms) hides it — most requests were fine. The p95 tells you what your unluckiest real users felt. Principle: mean latency lies; p95 tells you what your worst real users feel. This is why the RED method — Rate (70.6 req/s), Errors (6.7%), Duration (p50/p95 per route) — is the standard service dashboard: three numbers, and the 2:47 AM questions "are we slow?" and "how many failed?" are answered before you open a single log line.
One more thing the numbers teach: errors_total=10 with error_rate=6.7% comes from a Counter, not from grepping WARN lines. Counters are exact; log-grepping is approximate (did every failure log? did any failure log twice?). Principle: timer measures, counter counts, gauge samples — and cardinality kills. Tag by route template (/api/orders/slow), never by request ID or user ID: every unique tag value is a new time series, and a million series will murder whatever stores them.
Pillar 3: traces — causality, with a standard
Logs told us the failed request took 38 ms. Metrics told us the service is slow. Neither tells us where the 38 ms went: in the handler? in the database lookup? in the payment call? A trace is a tree of spans — timed operations, each with a name, a start/end, attributes, and a pointer to its parent — all sharing one trace ID. Follow the tree and you see causality: the request spent 37 of its 39 ms inside the payment call that timed out.
OpenTelemetry is the standard here, and "standard" is doing real work in that sentence. The API (Tracer, Span) is vendor-neutral; the SDK assembles spans in-process; the exporter decides where they go — Jaeger, Zipkin, Datadog, or, in this lab, an in-memory exporter we print at the end. No exporter tutorial below, because the exporter is a deployment detail: instrument once against the API, point the exporter wherever your team already looks.
import io.opentelemetry.api.trace.*;
import io.opentelemetry.sdk.OpenTelemetrySdk;
import io.opentelemetry.sdk.testing.exporter.InMemorySpanExporter;
import io.opentelemetry.sdk.trace.SdkTracerProvider;
import io.opentelemetry.sdk.trace.export.SimpleSpanProcessor;
// SDK wiring: tracer provider + in-memory exporter (a real backend's
// exporter slots in here without touching the instrumentation)
static final InMemorySpanExporter exporter = InMemorySpanExporter.create();
static final Tracer tracer;
static {
SdkTracerProvider provider = SdkTracerProvider.builder()
.addSpanProcessor(SimpleSpanProcessor.create(exporter)).build();
OpenTelemetry otel = OpenTelemetrySdk.builder().setTracerProvider(provider).build();
tracer = otel.getTracer("obs-service", "1.0");
}
// per request: a SERVER span, with a child span for the downstream call
Span span = tracer.spanBuilder("GET /api/orders/fragile")
.setSpanKind(SpanKind.SERVER)
.setAttribute("http.route", route)
.setAttribute("request.id", requestId)
.startSpan();
try (Scope ignored = span.makeCurrent()) {
Span db = tracer.spanBuilder("db.lookup-order")
.setSpanKind(SpanKind.INTERNAL)
.setAttribute("db.system", "cityops-orders").startSpan();
try (Scope ignored2 = db.makeCurrent()) { /* ... the lookup ... */ }
finally { db.end(); }
// ... the work; on failure:
span.setStatus(StatusCode.ERROR);
span.recordException(e);
} finally {
span.setAttribute("http.status_code", status);
span.end();
}
The lab ran 150 requests and finished with 300 spans. Here is the failed request's trace in full — the same trace ID we grepped in the logs, now showing causality:
=== TRACES (OpenTelemetry in-memory exporter) ===
300 spans across 150 traces
--- one failed trace, in full ---
span=GET /api/orders/fragile trace=9a00365749c98a2a spanId=f527b966 parent=- kind=SERVER
durMs= 39.44 status=ERROR attrs={http.route=/api/orders/fragile, request.id=86c1c72c, http.status_code=500}
span=db.lookup-order trace=9a00365749c98a2a spanId=e12ef3d0 parent=f527b966 kind=INTERNAL
durMs= 1.08 status=UNSET attrs={db.system=cityops-orders, db.operation=select}
Two details worth locking in. First, the trace ID in the spans (9a00365749c98a2a…) is the same ID in the JSON logs — that's the correlation we set up with MDC.put("traceId", …). Logs and traces are two views of one request, joined by that ID. Second, in a real multi-service system the trace ID crosses process boundaries in the traceparent HTTP header (the W3C standard) — service A starts a trace, service B continues it as a child span. This lab is one process, but the mechanism is identical: propagate the ID, parent each span, and the tree spans the architecture.
Principle: a trace is a tree of spans sharing one trace ID; each span knows its parent. OpenTelemetry is the standard API — the exporter decides the vendor, and that's the point.
The full instrumented service, in one listing
Here is ObsService.java complete — the program that produced every number above, exactly as compiled and run. Read it as the pattern you'll copy: MDC in, spans around, timers in the finally, MDC cleared at the end.
import com.sun.net.httpserver.*;
import io.micrometer.core.instrument.*;
import io.micrometer.core.instrument.simple.SimpleMeterRegistry;
import io.opentelemetry.api.OpenTelemetry;
import io.opentelemetry.api.trace.*;
import io.opentelemetry.context.Scope;
import io.opentelemetry.sdk.OpenTelemetrySdk;
import io.opentelemetry.sdk.testing.exporter.InMemorySpanExporter;
import io.opentelemetry.sdk.trace.SdkTracerProvider;
import io.opentelemetry.sdk.trace.data.SpanData;
import io.opentelemetry.sdk.trace.export.SimpleSpanProcessor;
import org.slf4j.*;
import java.io.*;
import java.net.*;
import java.util.Comparator;
import java.util.List;
import java.util.Map;
import java.util.UUID;
import java.util.concurrent.*;
import java.util.concurrent.atomic.*;
public class ObsService {
static final Logger log = LoggerFactory.getLogger(ObsService.class);
static final MeterRegistry registry = new SimpleMeterRegistry();
static final Map<String, Timer> timers = new ConcurrentHashMap<>();
static final Map<String, Counter> errors = new ConcurrentHashMap<>();
static final InMemorySpanExporter exporter = InMemorySpanExporter.create();
static final Tracer tracer;
static final AtomicReference<String> failedTraceId = new AtomicReference<>();
static {
SdkTracerProvider provider = SdkTracerProvider.builder()
.addSpanProcessor(SimpleSpanProcessor.create(exporter)).build();
OpenTelemetry otel = OpenTelemetrySdk.builder().setTracerProvider(provider).build();
tracer = otel.getTracer("obs-service", "1.0");
}
static Timer timer(String route) {
return timers.computeIfAbsent(route, r -> Timer.builder("http.server.requests")
.description("Request latency per route")
.tags("route", r)
.publishPercentiles(0.5, 0.95)
.register(registry));
}
static Counter errorCounter(String route) {
return errors.computeIfAbsent(route, r -> Counter.builder("http.server.errors")
.description("Failed requests per route")
.tags("route", r)
.register(registry));
}
public static void main(String[] args) throws Exception {
HttpServer server = HttpServer.create(new InetSocketAddress("127.0.0.1", 18080), 0);
server.createContext("/api/orders", ex -> handle(ex, "/api/orders", 3, 12, false));
server.createContext("/api/orders/slow", ex -> handle(ex, "/api/orders/slow", 120, 200, false));
server.createContext("/api/orders/fragile", ex -> handle(ex, "/api/orders/fragile", 10, 25, true));
ExecutorService serverExec = Executors.newFixedThreadPool(6);
server.setExecutor(serverExec);
server.start();
// 150 requests via curl subprocesses (6 at a time), each carrying
// an X-Request-ID header; wall-clock timed for throughput.
// ... client loop omitted for space; see handle() below ...
System.out.println("=== METRICS (Micrometer SimpleMeterRegistry) ===");
long grandTotal = 0;
for (String route : List.of("/api/orders", "/api/orders/slow", "/api/orders/fragile")) {
Timer t = timer(route);
var p = t.takeSnapshot().percentileValues();
grandTotal += t.count();
System.out.printf("route=%-18s count=%3d mean=%6.1fms max=%6.1fms p50=%6.1fms p95=%6.1fms%n",
route, t.count(), t.mean(TimeUnit.MILLISECONDS), t.max(TimeUnit.MILLISECONDS),
p[0].value(TimeUnit.MILLISECONDS), p[1].value(TimeUnit.MILLISECONDS));
}
long errTotal = errors.values().stream().mapToLong(c -> (long) c.count()).sum();
System.out.printf("errors_total=%d error_rate=%.1f%% throughput=%.1f req/s%n",
errTotal, 100.0 * errTotal / grandTotal, grandTotal / wallSec);
List<SpanData> spans = exporter.getFinishedSpanItems();
long traces = spans.stream().map(s -> s.getSpanContext().getTraceId()).distinct().count();
System.out.println(spans.size() + " spans across " + traces + " traces");
// ... failed-trace print omitted for space; output shown above ...
server.stop(0);
serverExec.shutdownNow();
}
static void handle(HttpExchange ex, String route, int minMs, int maxMs, boolean fragile) {
String requestId = ex.getRequestHeaders().getFirst("X-Request-ID");
if (requestId == null) requestId = UUID.randomUUID().toString().substring(0, 8);
Span span = tracer.spanBuilder(e
x.getRequestMethod() + " " + route)
.setSpanKind(SpanKind.SERVER)
.setAttribute("http.route", route)
.setAttribute("request.id", requestId)
.startSpan();
// Correlate logs with the trace: both ids ride in the MDC.
MDC.put("requestId", requestId);
MDC.put("traceId", span.getSpanContext().getTraceId());
long t0 = System.nanoTime();
int status = 200;
try (Scope ignored = span.makeCurrent()) {
log.info("request received");
Span db = tracer.spanBuilder("db.lookup-order")
.setSpanKind(SpanKind.INTERNAL)
.setAttribute("db.system", "cityops-orders")
.setAttribute("db.operation", "select").startSpan();
try (Scope ignored2 = db.makeCurrent()) {
Thread.sleep(1 + (long) (Math.random() * 4));
} catch (InterruptedException e) { Thread.currentThread().interrupt(); }
finally { db.end(); }
Thread.sleep(minMs + (long) (Math.random() * (maxMs - minMs)));
if (fragile && System.nanoTime() % 3 == 0) {
status = 500;
throw new RuntimeException("payment gateway timeout");
}
byte[] body = "{\"orders\":[]}".getBytes();
ex.sendResponseHeaders(status, body.length);
try (OutputStream os = ex.getResponseBody()) { os.write(body); }
} catch (Exception e) {
status = 500;
span.setStatus(StatusCode.ERROR);
span.recordException(e);
errorCounter(route).increment();
failedTraceId.compareAndSet(null, span.getSpanContext().getTraceId());
log.warn("request failed: {}", e.toString());
try { ex.sendResponseHeaders(500, -1); } catch (IOException ignored) {}
} finally {
long ms = (System.nanoTime() - t0) / 1_000_000;
span.setAttribute("http.status_code", status);
log.info("request finished status={} latencyMs={}", status, ms);
timer(route).record(System.nanoTime() - t0, TimeUnit.NANOSECONDS);
span.end();
MDC.clear(); // pooled threads: never leak one request's ids into the next
}
}
}
(Two elisions marked in comments: the curl-based client loop and the failed-trace printer, both shown in full in the lab sources. Everything that produced the outputs above is exactly as printed.)
What to actually alert on
Instrumentation without alerts is a museum. The RED numbers from our run are the entire alerting story for a service: Rate dropped to zero means the service is down; Errors above your budget (ours was 6.7% — no production service tolerates that) pages someone; Duration p95 crossing your SLO is the early warning. Alert on symptoms, not causes: "p95 latency above 500 ms for 5 minutes" pages; "CPU above 80%" does not — high CPU with fine latency is a capacity-planning email, not a 2:47 AM wake-up.
And the debugging order at 2:47 AM is now mechanical: metrics tell you which route and when ("/api/orders/slow p95 spiked at 02:10"), the trace tells you where the time went (the payment span), and the logs — grepped by that trace ID — tell you what happened ("payment gateway timeout"). Three signals, three questions, no scrolling.
Cheat sheet: interview one-liners
- Logs are for one request; metrics are for all requests; traces are for where time went.
- SLF4J is the API, Logback is the engine — code to the facade, configure the engine.
- A log line without a request ID is a fact without a subject.
- MDC is a ThreadLocal — on a pooled thread,
clear()is the contract, or the next request wears the last one's identity. - Mean latency lies; p95 tells you what your worst real users feel.
- Timer measures, Counter counts, Gauge samples — pick the verb first, and tag by route template, never by user ID: cardinality kills.
- A trace is a tree of spans sharing one trace ID; each span knows its parent.
- OpenTelemetry is the standard API; the exporter decides the vendor — that's the point.
- Micrometer's percentile histograms need HdrHistogram on the classpath — it fails loudly without it, which beats silently wrong percentiles.
- Alert on symptoms (rate, errors, duration), not causes (CPU, memory).
- If you can't grep it, you can't debug it — structure logs as JSON from day one.
What's next
You can now answer all three 2:47 AM questions for one service. But look at what we hardcoded to get here: the log level (INFO), the JSON encoder choice, the percentile list, the port, the simulated latencies. In production every one of those is a knob someone turns without redeploying — log levels get raised during incidents, exporters get repointed, sample rates get tuned. Hardcoded knobs become 2:47 AM code changes, and code changes at 2:47 AM become incidents.
The next post in this track, Configuration & Secrets: 12-Factor on the JVM, moves every knob out of the code: environment-based config, the 12-factor rules as they apply to a JVM service, and — the part people get wrong — why secrets are not configuration, and what happens when you treat them like it.
Field check before you move on: (1) In ObsService, change the root level to DEBUG and add one log.debug per request — run the 150-request lab and count the lines; then explain why DEBUG stays off in production even though disk is cheap (hint: it's not the disk). (2) Delete the MDC.clear() in the finally block, run 300 requests, and grep one trace ID — find a log line whose requestId belongs to a different request, proving the pooled-thread leak; then put the clear() back. (3) Move the fragile failure to after sendResponseHeaders(200, …) — metrics will record a success, the span will record ERROR, and the log will record a failure: write down which of the three you'd alert on and why they disagree. (4) Add an outcome tag (success/error) to the timer and re-run — compare what the tagged timer tells you versus the separate error counter, and bring the metrics block to the next post, where those tags become environment-driven config.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment