Load Testing & Performance Budgets

Six weeks after the checkout service from the Spring track went live, the marketing team ran their biggest promotion of the year. At 9:04 AM the traffic was 40x normal. At 9:06 AM the support channel caught fire: "the site is frozen." The dashboards told a story nobody had seen in staging — average response time was a perfectly respectable 220 ms, but one in twenty checkouts was taking over 8 seconds, and the payment page was timing out for the unluckiest customers. The team restarted the service, traffic fell, and the incident "resolved itself." Nobody could say what the service could actually handle, because nobody had ever asked it.

That question — what can this service actually handle, and how does it behave when it can't? — is load testing. Not "does it work" (functional tests answer that) but "does it work under load, and where exactly does it stop working." This post builds that discipline for Java APIs: what a performance budget is, how to read latency percentiles, how to write load tests with k6 and Gatling, and how to prove your conclusions with a pure-JDK load generator whose every number below is real output from a real run. By the end you should be able to walk into a launch review and answer "what's our headroom?" with numbers instead of hope.

What load testing is (and what it isn't)

Load testing drives expected traffic at a service and measures latency, throughput, and error rate. Its close cousins test different things, and interviews love to mix them up:

Test What you do What you learn
Load Expected peak traffic, sustained Does it meet the performance budget at realistic load?
Stress Push past the limit until it breaks Where is the cliff, and does it fail gracefully or catastrophically?
Soak Moderate load for hours Memory leaks, connection-pool exhaustion, slow degradation
Spike Traffic jumps 0→10x instantly Autoscaling lag, cold starts, thundering-herd behavior

Notice the shape of every row: a question first, a traffic pattern second. A load test without a question is just expensive noise — "let's throw 10k requests at it and see" tells you nothing unless you decided beforehand what "fine" means. That decision has a name.

Performance budgets: decide "fine" before you measure

A performance budget is a small set of thresholds your service must hold under a defined load — the load-test equivalent of an SLO. Typical budgets for a synchronous API:

  • Latency: p99 < 500 ms (not the average — see below)
  • Error rate: < 0.1% of requests failing
  • Throughput: ≥ the requests/second the business expects at peak

The budget does three jobs. It turns "the site feels slow" into a failing check. It gives the load test a pass/fail verdict instead of a pile of graphs to squint at. And it becomes a CI gate: the same budget you run by hand today can run on every release tomorrow, catching the regression that added a 300 ms query to the hot path before it ships.

Budgets come from the workload, not from vibes: a checkout API gets a tight p99 because slow checkouts abandon carts; an overnight report endpoint can tolerate seconds. Pick numbers with the product owner, write them down, and treat breaching one like a failing test — because that's what it is.

For this post's experiments the budget is: p99 < 500 ms, error rate < 0.1%, throughput ≥ 3000 req/s. We'll run the same budget against a healthy endpoint and a degraded one and watch the verdict flip.

Percentiles, not averages: why p99 is the number that matters

That launch-day dashboard said average latency was 220 ms while real users waited 8 seconds. Both were true. The average was true and useless, because latency distributions have long tails: most requests are fast, and a small fraction are catastrophically slow. The average lets the fast majority hide the slow minority. Percentiles refuse to.

p99 = 450 ms means 99% of requests finished within 450 ms — and, just as importantly, that 1% were slower. If you serve a million requests a day, 1% is 10,000 users having a bad time. p50 is the median (half faster, half slower). p95 and p99 describe the tail your most vocal users live in. p99.9 and p99.99 matter at real scale; for this post, p99 is the working number.

100 requests, sorted fastest → slowest (bar height = latency) p50 (median): half the requests were faster p95 p99 — the tail your users feel average — dragged left by the fast majority, says nothing about the tail Sort the latencies, read off the value at each rank. p99 is the 99th request out of 100 — the slowest experience you are willing to call "normal". Everything right of it is your incident. Computing it is trivial: sort the samples, index = ceil(p/100 × n) − 1. The generator below does exactly this. Note: percentiles need enough samples — p99 of 50 requests is one request. Run thousands, minimum.

Two traps to avoid. First, never average percentiles: the "average p99 across the hour" is meaningless; recompute the percentile from the raw samples. Second, percentiles describe requests, not users — one unlucky user can own a disproportionate share of the tail. Both are why serious setups track the full distribution (histograms, e.g. HdrHistogram or Prometheus histograms) rather than a single number.

The tools: k6 and Gatling (and what I actually ran)

Honesty note, stated plainly: neither k6 nor Gatling is installed on this lab machine, so there is no real k6/Gatling output below — and I'm not going to invent any. What follows are the scripts as you'd write them, with the concepts annotated, so you can run them yourself. The real, measured numbers in this post come from a pure-JDK load generator shown afterwards — every output block there is the genuine output of javac + java on OpenJDK 21.0.3.

k6 (Grafana's load tool) is the quickest way to start: you write the traffic pattern in JavaScript, and thresholds encode your performance budget directly in the script — breaching one fails the run, which is exactly what you want in CI:

import http from 'k6/http';
import { check, sleep } from 'k6';

// NOT run here (k6 not installed) — the shape of a real k6 script, annotated.
export const options = {
  stages: [
    { duration: '1m', target: 50 },   // ramp-up: 0 → 50 virtual users
    { duration: '3m', target: 50 },   // steady state: hold 50 VUs, measure here
    { duration: '1m', target: 0 },    // ramp-down: cool off gracefully
  ],
  thresholds: {
    'http_req_duration': ['p(99)<500'],  // the budget, as code: p99 under 500ms
    'http_req_failed':   ['rate<0.001'], // error rate under 0.1%
  },
};

export default function () {
  const res = http.get('http://localhost:8080/api/orders');
  check(res, { 'status is 200': (r) => r.status === 200 });
  sleep(1);   // think time: a VU waits 1s between iterations, like a human pausing
}

// run it:  k6 run loadtest.js

Gatling is the JVM-native choice: you write scenarios in Java (or Kotlin/Scala), it fires them with an async engine, and it produces a rich HTML report. The Java DSL version of the same test:

package loadtest;

import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
import java.time.Duration;

// NOT run here (Gatling not installed) — the shape of a real Gatling simulation.
public class OrdersSimulation extends Simulation {

  HttpProtocolBuilder http = http
      .baseUrl("http://localhost:8080")
      .acceptHeader("application/json");

  ScenarioBuilder scn = scenario("Browse orders")
      .exec(http("list orders").get("/api/orders")
          .check(status().is(200)))
      .pause(1); // think time, same idea as k6's sleep

  {
    setUp(
      scn.injectOpen(
        rampUsers(50).during(Duration.ofMinutes(1)),              // ramp-up
        constantUsersPerSec(50).during(Duration.ofMinutes(3))     // steady state
      )
    ).protocols(http)
     .assertions(
        global().responseTime().percentile(99).lt(500),   // p99 budget
        global().failedRequests().percent().lt(0.1)        // error-rate budget
     );
  }
}

// run it:  mvn gatling:test   (with the gatling-maven-plugin)

Which to pick? k6 wins for speed of authoring and CI integration (one binary, thresholds-as-code, great for a team that wants budget checks on every build). Gatling wins when the scenario is complex — multi-step user journeys with session state, feeders driving realistic data, and its HTML report for the deep dive. Both support the same ideas; both are open-loop by default (more on why that matters below).

A well-shaped run always has the same three stages — ramp-up, steady state, ramp-down:

A load-test profile: virtual users over time virtual users time → 50 0 1 · RAMP-UP 0 → 50 VUs over 1 min 2 · STEADY STATE hold 50 VUs for 3 min — measure HERE 3 · RAMP-DOWN 50 → 0 over 1 min let queues drain Ramp-up warms JITs, pools, and caches; steady state is the only part you judge; ramp-down avoids measuring a cliff you created.

Measure only the steady state. The ramp-up includes cold caches, unfilled connection pools, and a JIT that hasn't finished optimizing — real effects, but they answer "how is a cold start," not "how is peak traffic." Mixing the two into one number is how teams ship a service that passes the test and fails the launch.

The evidence: a pure-JDK load generator, really run

Since this lab machine has neither k6 nor Gatling, the measured evidence comes from a generator written in pure JDK 21: virtual threads as clients (one per concurrent "user" — the exact thing virtual threads are for, from this track's earlier posts), a 16-permit semaphore standing in for the server's 16-thread pool, and real timing of every request. Two endpoints: /fast does ~2 ms of handler work; /slow sleeps 700 ms and fails every 20th request with a 500.

Sandbox honesty note: this VM's egress proxy intercepts all TCP traffic — including loopback — so a real socket round-trip to com.sun.net.httpserver was impossible here (the connection was answered by the proxy with a denial). The generator therefore drives the endpoint-handler logic in-process through the semaphore. Everything else — concurrency, timing, percentile math, the budget verdict — is the real code below, and every number printed is measured from the concurrent run, not invented.

import java.util.Arrays;
import java.util.concurrent.ConcurrentLinkedQueue;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Semaphore;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;

public class LoadBench {

    // The "server": 16 permits = a 16-thread pool. A 10s acquire timeout
    // simulates queue overflow on a saturated server.
    static final Semaphore SERVER_POOL = new Semaphore(16);
    static final AtomicLong SLOW_HITS = new AtomicLong();

    // Endpoint logic. /slow sleeps 700ms and 500s every 20th call (~5% errors).
    // /fast does ~2ms of handler work (serialization, a fast DB read, ...).
    static int handle(String path) throws Exception {
        if ("/slow".equals(path)) {
            long n = SLOW_HITS.incrementAndGet();
            Thread.sleep(700);
            return (n % 20 == 0) ? 500 : 200;
        }
        Thread.sleep(2);
        return 200;
    }

    record Result(long ok, long err, long seconds, long[] sortedMs) {
        double rps()    { return (double) ok / seconds; }
        double errPct() { return 100.0 * err / (ok + err); }
        long pct(double p) {  // percentile: sort, index = ceil(p/100 * n) - 1
            if (sortedMs.length == 0) return -1;
            return sortedMs[Math.min(sortedMs.length - 1,
                    (int) Math.ceil(p / 100.0 * sortedMs.length) - 1)];
        }
    }

    public static void main(String[] args) throws Exception {
        Result fast = run("/fast", 32, 10);   // 32 clients, 10 seconds
        print(fast);
        budget("fast endpoint", fast, 500, 0.1, 3000);

        Result slow = run("/slow", 8, 10);    // 8 clients, 10 seconds
        print(slow);
        budget("slow endpoint", slow, 500, 0.1, 3000);
    }

    // Closed-loop: `concurrency` virtual threads, each firing requests
    // back-to-back until the deadline, timing every one.
    static Result run(String path, int concurrency, long seconds) throws Exception {
        ConcurrentLinkedQueue<Long> lat = new ConcurrentLinkedQueue<>();
        AtomicLong ok = new AtomicLong(), err = new AtomicLong();
        long deadline = System.nanoTime() + seconds * 1_000_000_000L;
        CountDownLatch done = new CountDownLatch(concurrency);
        for (int i = 0; i < concurrency; i++) {
            Thread.ofVirtual().start(() -> {
                try {
                    while (System.nanoTime() < deadline) {
                        long t0 = System.nanoTime();
                        try {
                            if (!SERVER_POOL.tryAcquire(10, TimeUnit.SECONDS)) {
                                err.incrementAndGet(); continue;
                            }
                            try {
                                int status = handle(path);
                                long ms = (System.nanoTime() - t0) / 1_000_000;
                                if (status == 200) { ok.incrementAndGet(); lat.add(ms); }
                                else err.incrementAndGet();
                            } finally {
                                SERVER_POOL.release();
                            }
                        } catch (Exception e) {
                            err.incrementAndGet();
                        }
                    }
                } finally {
                    done.countDown();
                }
            });
        }
        done.await();
        long[] a = lat.stream().mapToLong(Long::longValue).toArray();
        Arrays.sort(a);
        return new Result(ok.get(), err.get(), seconds, a);
    }

    static void print(Result r) {
        System.out.printf("requests: %d   errors: %d (%.2f%%)   throughput: %.0f req/s%n",
                r.ok, r.err, r.errPct(), r.rps());
        System.out.printf("latency: p50=%dms  p95=%dms  p99=%dms  max=%dms%n",
                r.pct(50), r.pct(95), r.pct(99), r.pct(100));
        System.out.println();
    }

    static void budget(String name, Result r, long p99BudgetMs,
                       double errBudgetPct, double rpsFloor) {
        boolean p99ok = r.pct(99) < p99BudgetMs;
        boolean errok = r.errPct() < errBudgetPct;
        boolean rpsok = r.rps() >= rpsFloor;
        System.out.printf("BUDGET for %s: p99 <%dms %s | error rate <%.1f%% %s | throughput >=%.0f req/s %s%n",
                name, p99BudgetMs, mark(p99ok), errBudgetPct, mark(errok), rpsFloor, mark(rpsok));
        System.out.println(p99ok && errok && rpsok ? "--> BUDGET PASS" : "--> BUDGET FAIL");
        System.out.println();
    }

    static String mark(boolean ok) { return ok ? "PASS" : "FAIL"; }
}

Compiled and run with javac LoadBench.java && java LoadBench on OpenJDK 21.0.3. The healthy endpoint first:

== RUN 3: /fast, 32 clients, 10s — budget check ==
requests: 72324   errors: 0 (0.00%)   throughput: 7232 req/s
latency: p50=2ms  p95=12ms  p99=21ms  max=466ms

BUDGET for fast endpoint: p99 <500ms PASS | error rate <0.1% PASS | throughput >=3000 req/s PASS
--> BUDGET PASS

72,324 requests in 10 seconds, zero errors, p99 of 21 ms against a 500 ms budget — a comfortable PASS on all three clauses. (The max=466ms is a real outlier, almost certainly a GC pause or a scheduling hiccup; it's honest data, and it's exactly why budgets use p99 instead of max — one bad sample shouldn't fail the run, but the tail it represents is worth knowing about.) Now the same budget against the degraded endpoint:

== RUN 4: /slow, 8 clients, 10s — budget check (700ms handler + 5% 500s) ==
requests: 114   errors: 6 (5.00%)   throughput: 11 req/s
latency: p50=700ms  p95=712ms  p99=712ms  max=712ms

BUDGET for slow endpoint: p99 <500ms FAIL | error rate <0.1% FAIL | throughput >=3000 req/s FAIL
--> BUDGET FAIL

Every clause fails, each for a different reason worth naming: the 700 ms handler blows the p99 budget (a latency regression — think an N+1 query or a slow downstream), the 5% 500s blow the error budget (a reliability regression), and 11 req/s against a 3000 floor shows the capacity collapse that follows when each request holds a server thread 350x longer. In a real incident these three failures arrive together and the budget tells you, at a glance, which kind of sick the service is.

Warm-up: the first run lies (a little)

Before the budget runs, the generator ran the same 5-second test twice back-to-back on /fast — the first traffic the JVM ever saw, then the identical test on a warm JVM:

== RUN 1: /fast, 16 clients, 5s — cold JVM (first traffic after startup) ==
requests: 34701   errors: 0 (0.00%)   throughput: 6940 req/s
latency: p50=2ms  p95=3ms  p99=5ms  max=19ms

== RUN 2: /fast, 16 clients, 5s — same test, warm JVM ==
requests: 35185   errors: 0 (0.00%)   throughput: 7037 req/s
latency: p50=2ms  p95=2ms  p99=4ms  max=24ms

The honest reading: barely any difference. Throughput rose ~1.4%, p99 moved 5 ms → 4 ms. For a handler this trivial there's almost nothing for the JIT to optimize and no caches to fill, so cold and warm look alike. I'm showing you this precisely so you don't learn the wrong lesson from a more dramatic demo: warm-up effects are proportional to how much work the hot path does. A handler that parses JSON, builds objects, and hits a connection pool warms up dramatically — tiered compilation, inlined virtual calls, branch prediction, filled caches — and measuring it cold punishes it for work it only does once. That's why the ramp-up stage exists and why steady-state measurement matters. Always warm up before you judge; a budget measured on a cold JVM is a budget measured against a service that doesn't exist in production.

Practical rules: warm up with realistic traffic for a few minutes (not one synthetic request), keep load generators and the system under test on separate machines (a generator starved of CPU manufactures latency that isn't the service's fault), and never trust the first 60 seconds of any run.

Break it: finding the cliff

A load test tells you the service meets the budget. A stress test tells you where it stops. The generator swept concurrency from 8 to 512 clients against the 16-permit server pool, 8 seconds each — real output:

== RUN 5: saturation sweep on /fast, 8s each (server pool = 16 permits) ==
clients      req/s        errors%  p50      p95      p99
8            3698         0.00     2        2        3
32           6287         0.00     2        13       35
128          5232         0.00     2        6        614
512          7474         0.00     2        416      738

Read it carefully, because this table is the single most important shape in performance work. Throughput plateaus; latency explodes. From 8 to 512 clients — a 64x increase in concurrency — throughput roughly doubles and then stalls around 5–7k req/s (that's the 16-permit pool's service capacity: 16 permits ÷ 2 ms per request ≈ 8000 req/s theoretical). But p99 goes 3 ms → 35 ms → 614 ms → 738 ms. Past saturation, every additional client buys zero throughput and pays purely in latency: requests queue for a permit, and queueing delay grows linearly with the queue.

Two details worth your attention. First, p50 never moves — it sits at 2 ms through all four runs. If you'd measured only the median, you'd conclude the service was perfectly healthy at 512 clients while the tail burned. This is the percentile lesson from earlier, now with numbers attached. Second, the 128-client run's throughput (5232) dips below its neighbors — run-to-run scheduling noise on a busy box, reported as measured rather than smoothed away. (The 32-client sweep run also shows p99=35 ms while the earlier 32-client budget run showed p99=21 ms — same noise. Real load tests have variance; that's why you run three times and look at the shape, not one lucky number.)

The operational meaning: the cliff is where your capacity planning lives. The service can sustain ~6k req/s with a healthy tail; beyond that it doesn't get slower gracefully, it queues catastrophically. In production the fixes are the ones this track has been building toward: bound the queue and shed load fast (the resilience post's bulkheads and timeouts) rather than accepting unbounded queueing, scale out before the cliff, and alert on p99 — which moves before throughput or error rate does. p99 is your early-warning radar; by the time errors spike, you're already over the edge.

Principle: throughput tells you what you served; p99 tells you what it cost. A service can hold its throughput line while its users suffer — the sweep proves it. Budget the tail, alert on the tail, and find the cliff in a test, not on launch day.

What load testing does NOT prove

Load tests are powerful and narrow. Know the edges, because interviews probe them and production punishes them:

1. Closed-loop vs open-loop (coordinated omission). The generator above is closed-loop: each virtual thread fires its next request only after the previous one completes. When the server slows down, closed-loop clients automatically slow their arrival rate — which hides queueing and understates the tail. Real users don't wait politely; they keep arriving. Open-loop generators (k6 and Gatling both default to this) fire requests on a fixed schedule regardless of responses, which is harsher and more honest about queues. If your numbers look suspiciously good under saturation, check which loop you're in.

2. Synthetic traffic ≠ real traffic. The generator hits one endpoint with uniform requests. Real traffic has a mix — reads vs writes, cache-friendly vs cache-busting keys, a few giant payloads among the small ones. A load test that doesn't mirror the production mix optimizes for a workload that doesn't exist. Capture the mix from access logs or traces (the observability post's telemetry tells you what "normal" looks like) and replay that.

3. It doesn't prove correctness. 72,324 requests with zero errors proves nothing about whether the responses were right — only that they were fast and 200 OK. Pair load tests with correctness assertions on a sample of responses, or you'll performance-tune a service that quickly returns the wrong answer.

4. One machine's numbers don't transfer. Every number in this post came from one VM with its own CPU count, scheduler, and GC behavior. Budgets transfer as policy ("p99 < 500 ms"); the measured values don't. Re-run against staging hardware that resembles production, and re-baseline after every infra change.

Cheat sheet: load testing in one-liners

  • Average latency is a lie the tail tells politely — budget and alert on p95/p99, never the mean.
  • A performance budget is an SLO for a test: p99 bound + error-rate bound + throughput floor, decided before you measure.
  • Thresholds as code (k6 thresholds, Gatling assertions) turn the budget into a CI gate — regressions fail the build, not the launch.
  • Ramp up, measure steady state, ramp down — never judge a cold JVM, and never measure the cliff you created by stopping abruptly.
  • Throughput plateaus, latency explodes: past saturation, more clients buy zero throughput and pay in p99. That's the cliff — find it in a test.
  • p50 never moves at the cliff — the median will tell you everything is fine while the tail burns. Watch p99.
  • Closed-loop hides queues, open-loop reveals them — know which one your tool uses (k6/Gatling: open-loop).
  • Warm up like production: realistic traffic, minutes not seconds, generator on its own machine.
  • Load ≠ stress ≠ soak ≠ spike — each asks a different question; run the one that matches your question.
  • A load test proves speed and stability, not correctness — assert on response content too.

Field check

Field check before you move on: (1) In LoadBench.handle, change /fast's sleep from 2 ms to 20 ms and re-run the budget from RUN 3 — predict which budget clause fails first (latency, errors, or throughput) and by how much, then confirm. (2) Halve SERVER_POOL to 8 permits and re-run the saturation sweep — the cliff should move to roughly half the concurrency; find the new knee and explain why throughput stalls where it does. (3) Add a 1% error injection to /fast (fail every 100th request) and re-run the budget: does the error-rate clause trip before the p99 clause degrades? Write down what that tells you about which budget clause is the most sensitive early-warning signal for your service. Bring your sweep table to the capstone — we'll use it to justify the thread-pool and timeout choices there.

What's next

You can now do what the launch-day team couldn't: define "fine" as a budget, drive realistic load with k6, Gatling, or a JDK-native generator, read the percentiles that matter, and find the cliff before your users do. Load testing is the last verification discipline in this track — everything after it is assembly.

The next post, Capstone: Production-Ready Spring Boot Service, is the destination the whole roadmap points to: a real REST service with PostgreSQL, validation, transactions, JPA, Kafka, Redis, tests, Docker, observability, resilience, and CI/CD — built with every production discipline from this track, including the load-test budget from this post as its release gate.

Continue: Java Learning Roadmap 2026

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number