JMH: Microbenchmarking Done Right
A teammate once rewrote the hottest string-building loop in our checkout service and "proved" the new version was 40% faster. The proof was a loop with System.nanoTime() around it, run once, printed twice. We shipped it on the strength of that number. P99 latency got worse. The rollback was the easy part — the hard part was admitting the benchmark had never measured the code at all. It had measured the JIT compiler’s leftovers.
Microbenchmarking on the JVM is adversarial. Your opponent is the smartest optimizer in the room, and it rewrites your code while you time it. JMH — the Java Microbenchmark Harness — is the OpenJDK project’s own harness, built because JVM engineers got tired of watching everyone lose this fight. This post shows you how to stop losing it.
The four ways a hand-rolled benchmark lies
The classic naive benchmark looks like this. You have written it, or you will:
public class NaiveTiming {
public static void main(String[] args) {
long start = System.nanoTime();
for (int i = 0; i < 10_000_000; i++) {
Math.sqrt(i); // result thrown away
}
long elapsed = System.nanoTime() - start;
System.out.println("sqrt loop: " + (elapsed / 1_000_000) + " ms");
}
}
Typical output:
sqrt loop: 4 ms
Ten million square roots in 4 milliseconds? A real sqrt costs on the order of 10–20 nanoseconds — ten million of them should take well over 100 ms. This number is confidently, spectacularly wrong, for four independent reasons:
- No warmup — you timed cold code. The JVM starts by interpreting bytecode. After roughly a thousand invocations the C1 compiler produces quick machine code; after heavier profiling, C2 recompiles with aggressive optimization. Your single timing captures interpreted or C1 code — the exact code that will never run in a warmed-up production JVM.
- Dead-code elimination — the JIT deleted your work. The result of
Math.sqrt(i)is never used. The C2 compiler can prove that and delete the call — very likely the entire loop. You didn’t time ten million square roots. You timed an empty loop. - Constant folding — constants get precomputed. If your inputs are literals or effectively-final constants, the JIT can compute the answer once at compile time and replace the loop body with the result. Your "benchmark" then measures how fast the JVM loads a constant.
- GC pauses and OS noise — one sample is no sample. A single timed run includes whatever garbage-collection pause, thread-scheduling hiccup, or turbo-boost clock ramp happened to occur during those milliseconds. One number carries no error bars; it tells you nothing about the distribution.
Decision rule: if your timing harness doesn’t actively defeat the JIT, you aren’t measuring your code — you’re measuring whatever the JIT happened to leave behind.
JMH: a harness that does the boring parts right
JMH doesn’t make your code faster. It makes your numbers honest by handling the four problems above on every run:
- Warmup iterations — runs the benchmark several times before measuring, letting C1/C2 compilation settle into steady state. Warmup results are discarded, never reported.
- Measurement iterations — measures repeatedly and reports a statistical score with error bars, not a single number.
- Forks — spawns a fresh JVM per fork, isolating each run from the previous one’s compilation profile and GC history.
- Blackhole — an opaque result sink the JIT cannot see through. Consuming results through it defeats dead-code elimination: the work has an observable effect, so it can’t be deleted.
The centerpiece: a complete JMH benchmark
Let’s measure a real question from the Core track’s Strings post: in a loop, is + concatenation really slower than StringBuilder — and by how much? That post proved String is immutable; this measures what the immutability costs when you ignore the advice. Every annotation below is doing real work, which we unpack after the code:
package com.javamakeuse.bench;
import org.openjdk.jmh.annotations.Benchmark;
import org.openjdk.jmh.annotations.BenchmarkMode;
import org.openjdk.jmh.annotations.Fork;
import org.openjdk.jmh.annotations.Measurement;
import org.openjdk.jmh.annotations.Mode;
import org.openjdk.jmh.annotations.OutputTimeUnit;
import org.openjdk.jmh.annotations.Scope;
import org.openjdk.jmh.annotations.Setup;
import org.openjdk.jmh.annotations.State;
import org.openjdk.jmh.annotations.Warmup;
import org.openjdk.jmh.infra.Blackhole;
import java.util.concurrent.TimeUnit;
@BenchmarkMode(Mode.Throughput)
@OutputTimeUnit(TimeUnit.SECONDS)
@Warmup(iterations = 3)
@Measurement(iterations = 5)
@Fork(1)
@State(Scope.Benchmark)
public class StringConcatBenchmark {
private String[] words;
@Setup
public void setup() {
words = new String[200];
for (int i = 0; i < words.length; i++) {
words[i] = "word" + i;
}
}
@Benchmark
public void concatWithPlus(Blackhole bh) {
String result = "";
for (String w : words) {
result = result + w;
}
bh.consume(result);
}
@Benchmark
public void concatWithStringBuilder(Blackhole bh) {
StringBuilder sb = new StringBuilder();
for (String w : words) {
sb.append(w);
}
bh.consume(sb.toString());
}
}
What each piece is for:
@State(Scope.Benchmark)— the benchmark instance itself holds the input data.Scope.Benchmarkmeans one shared instance (useScope.Threadwhen threads must not share state).@Setup— builds the 200-word input once, before timing starts, so setup cost never pollutes the measurement.@Benchmark— marks the methods to time. JMH injects theBlackholeparameter automatically.bh.consume(result)— makes the result “observable” so dead-code elimination can’t delete the work. Returning the value instead would also work;Blackholeis the explicit form.@BenchmarkMode(Mode.Throughput)— measure operations per second (alternatives:AverageTime,SampleTime,SingleShotTime).@OutputTimeUnit(TimeUnit.SECONDS)— report throughput per second.@Warmup(iterations = 3)/@Measurement(iterations = 5)— three warmup iterations discarded, five measured.@Fork(1)— one fork, one fresh JVM. This keeps the example fast; for numbers you’d defend in a design review, use 3–5 forks.
Running it: Maven
JMH ships as two artifacts: jmh-core (the runtime) and jmh-generator-annprocess (the annotation processor that generates the runner code from your @Benchmark methods). You need both — without the generator, your benchmarks silently don’t run. Package everything into a self-contained jar with the shade plugin:
<dependencies>
<dependency>
<groupId>org.openjdk.jmh</groupId>
<artifactId>jmh-core</artifactId>
<version>1.37</version>
</dependency>
<dependency>
<groupId>org.openjdk.jmh</groupId>
<artifactId>jmh-generator-annprocess</artifactId>
<version>1.37</version>
<scope>provided</scope>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-shade-plugin</artifactId>
<version>3.5.1</version>
<executions>
<execution>
<phase>package</phase>
<goals><goal>shade</goal></goals>
<configuration>
<finalName>benchmarks</finalName>
<transformers>
<transformer implementation="org.apache.maven.plugins.shade.resource.ManifestResourceTransformer">
<mainClass>org.openjdk.jmh.Main</mainClass>
</transformer>
</transformers>
</configuration>
</execution>
</executions>
</plugin>
</plugins>
</build>
(Check for newer artifact versions than the ones pinned above — the coordinates are the stable part.) Then build and run:
mvn clean package
java -jar target/benchmarks.jar StringConcat
The trailing StringConcat filters which benchmarks run — useful when the jar holds many. Run without it to execute everything.
Gradle users: add implementation 'org.openjdk.jmh:jmh-core:1.37' and annotationProcessor 'org.openjdk.jmh:jmh-generator-annprocess:1.37' to your dependencies, apply the community me.champeau.jmh plugin (check for its latest version), and run ./gradlew jmh. The plugin wires up the same annotation processing and fat-jar assembly the Maven setup above does by hand.
How to read the output
A few minutes later, JMH prints one line per @Benchmark method:
Benchmark Mode Cnt Score Error Units
StringConcatBenchmark.concatWithPlus thrpt 5 268441.203 ± 17482.510 ops/s
StringConcatBenchmark.concatWithStringBuilder thrpt 5 3120938.776 ± 91024.337 ops/s
Each column:
- Benchmark — the method that ran. Two lines: the two approaches, side by side.
- Mode —
thrptis throughput: higher is better. (InAverageTimemode this column would sayavgtand smaller would be better.) - Cnt — the number of measurement iterations aggregated: 1 fork × 5 iterations = 5.
- Score — the mean: ~268k ops/s for
+, ~3.1M ops/s forStringBuilder. About an 11× difference — real this time, because the harness defeated the JIT instead of being defeated by it. - Error — the ± is the half-width of the 99.9% confidence interval. The true score is very likely inside Score ± Error. This is the column hand-rolled benchmarks don’t have.
- Units —
ops/s, because we asked for throughput per second. Average-time mode would reportns/opinstead.
Golden rule: never claim a win when the error bars overlap. If method A scores 100 ± 30 and method B scores 110 ± 30, the intervals intersect — the data cannot tell them apart, and “B is 10% faster” is fiction. Shrink the error (more iterations, more forks, quieter machine) or admit the tie.
Micro vs macro: where JMH ends and Gatling begins
JMH answers “how fast is this method?” — one method, one JVM, nanosecond precision, allocation behavior under a microscope. It cannot tell you what happens when 500 users call it at once. That’s a different question with a different tool:
- JMH (micro): a single method in isolation. Finds the slow method and proves the fix with honest numbers.
- Gatling (macro): the whole system under concurrent load — p99 latency, throughput ceilings, connection-pool exhaustion, GC pauses that only appear at sustained allocation rates. Proves the system survives real traffic. (This track covers it in the post “Load Testing with Gatling: Prove Your API Survives Traffic”.)
Decision rule: use JMH to find the slow method; use Gatling to prove the system survives traffic. Micro tells you what to fix. Macro tells you whether the fix mattered.
The four traps
Trap 1 — Benchmarking without warmup
Deleting @Warmup “to save time” measures interpreted code and calls it a result. Symptom: the first iteration’s numbers look nothing like the rest. The warmup iterations are the cheapest insurance in the whole harness — keep them, and be suspicious of any benchmark whose warmup and measurement scores are identical (that usually means the workload is too trivial to compile differently, or something is very wrong).
Trap 2 — Forgetting to consume the result
Every benchmark method must do something observable with its result: return it, or feed it to a Blackhole. A method that computes and discards is a method the JIT is allowed to delete. This is the single most common JMH bug in the wild, and its signature is unmistakable: suspiciously perfect numbers — like ten million square roots in 4 milliseconds.
Trap 3 — Benchmarking on a noisy machine
Laptop with Slack, a browser, and a video call running; turbo boost ramping the clock up and down; thermal throttling kicking in halfway through the run. Your numbers will wobble between runs and the error bars will tell you so — that’s what they’re for. Close everything, plug in, and benchmark on a machine as close to production as you can get. Allocation-heavy code adds GC noise on top of OS noise; the Core track’s JVM & GC post explains what a pause does to your percentiles.
Trap 4 — Micro-optimizing what doesn’t matter
A 2× win on code that consumes 0.1% of total runtime is a rounding error — and it usually costs readability to get there. Profile the hot path first, then benchmark only the code the profiler says matters. JMH measures precisely; precision applied to the wrong method is still waste. This is the trap the developer in the opening story fell into in reverse: they had a precise number for the wrong conclusion.
What’s next
You now have the full tooling loop: build it (Maven/Gradle), prove it works (JUnit, Mockito, Testcontainers), ship it (Docker, CI/CD), prove it survives (Gatling), keep it clean (static analysis), and measure it honestly (JMH). The next track — Spring Boot & APIs — puts all of this to work building real services, starting with your first REST API.
Field check before you move on: take the StringConcatBenchmark above, run it, then change @Measurement(iterations = 5) to iterations = 1 and run it again. Watch what happens to the Error column. That difference — between one sample and a measured distribution — is the entire point of this post.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment