Load Testing with Gatling: Prove Your API Survives Traffic

It's launch week. The checkout API passes all 412 unit tests. Integration tests are green. The demo on the founder's laptop is flawless. At 9:03 on Monday morning, 500 real users hit "Buy" at once — and the API falls over. Latency climbs from 80 milliseconds to 40 seconds, then requests start dying outright. The team rolls back, but there is nothing to roll back to: the code is correct. The logic was never the problem.

The problem was the assumptions nobody tested: a web-server thread pool of 20 trying to serve 500 concurrent users, a database connection pool of 10 that every request fights over, and garbage-collection pauses that only show up when the allocation rate is sustained. Unit tests answer "does it work?" Load tests answer "does it survive?" This post is about the second question — with a tool that speaks your language.

What actually breaks at 500 concurrent users

When an API dies under load, the cause is almost always one of three capacity failures, not a logic bug:

  • Thread pool exhaustion. Your server (Tomcat, Jetty, Netty's event loop aside) has a fixed number of worker threads. Five hundred concurrent requests arrive; twenty threads can do work. The rest queue. Queues grow, timeouts fire, clients retry — and retries make the queue worse. Death spiral.
  • Database connection pool too small. Each request borrows a connection, does its queries, returns it. With a pool of 10 and 500 borrowers, most of your request latency is spent waiting for a connection, not running SQL. Your APM tool will show the queries as fast and the endpoints as slow — a classic misdirection.
  • GC pauses under sustained allocation. On a laptop, a young-generation collection takes a few milliseconds and nobody notices. Under load, the allocation rate is 50× higher, collections take longer, and the worst pauses land exactly when traffic peaks. The p99 latency graph develops a sawtooth pattern that has nothing to do with your code.

Decision rule: every capacity number in your config — thread pool size, connection pool size, heap size — is a guess until you've measured it under load. Load testing turns those guesses into measurements.

Why Gatling, for a Java developer

Gatling is a load-testing tool written in Scala that runs on the JVM — and since version 3.7 it ships a first-class Java DSL, so you write your load scenarios in plain Java. That matters for three reasons:

  • It's code, not clicks. Scenarios are Java classes: version-controlled, diffable, reviewable in a pull request. When someone asks "what exactly did we test?", you point at a file, not at a recorded script from a GUI session.
  • It's a real programming language. Loops, feeders, dynamic data from CSVs, conditional checks — all the things that turn a toy test into a realistic traffic model, without a visual-scripting tax.
  • The HTML report is the best in the business. One self-contained index.html with response-time distributions, percentiles, throughput graphs, and per-request breakdowns. You'll see exactly how to read it below.

The alternatives: JMeter is the old guard — mature, GUI-driven, but heavier and awkward to version-control. k6 is excellent and scriptable, but it's JavaScript and outside the JVM; if your team lives in Java, keeping the toolchain in one language keeps the barrier to writing tests low. Decision rule: pick the load tool your team can read in code review. For a Java team, that's Gatling.

Setup: the Gatling Maven plugin

Gatling ships as a Maven plugin plus two test-scoped dependencies. Add this to your pom.xml (the 3.13.x line is current as of this writing — any 3.13.x works):

<properties>
    <gatling.version>3.13.0</gatling.version>
    <maven.compiler.release>21</maven.compiler.release>
</properties>

<dependencies>
    <dependency>
        <groupId>io.gatling.highcharts</groupId>
        <artifactId>gatling-charts-highcharts</artifactId>
        <version>${gatling.version}</version>
        <scope>test</scope>
    </dependency>
    <dependency>
        <groupId>io.gatling</groupId>
        <artifactId>gatling-java</artifactId>
        <version>${gatling.version}</version>
        <scope>test</scope>
    </dependency>
</dependencies>

<build>
    <plugins>
        <plugin>
            <groupId>io.gatling</groupId>
            <artifactId>gatling-maven-plugin</artifactId>
            <version>4.13.0</version>
        </plugin>
    </plugins>
</build>

Two things to note: gatling-charts-highcharts is what generates the HTML report, and gatling-java is the Java DSL itself. Both are test-scoped — they never ship in your production artifact. (Gradle users: the io.gatling.gradle plugin is the equivalent — apply it, put simulations under src/gatling/java, run ./gradlew gatlingRun. The build-tool mechanics are covered in this track's Maven & Gradle post; here we only need the plugin on the classpath.)

The simulation: a complete, runnable test

This is the centerpiece. One Java class that hammers a product API with a realistic mix: browsing products (reads) and creating products (writes), with per-user think time, a CSV feeder so virtual users don't all request the same product, and assertions that fail the run if the API misbehaves. Save it as src/test/java/com/acme/perf/ProductApiSimulation.java:

package com.acme.perf;

import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;

import java.time.Duration;

import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;

public class ProductApiSimulation extends Simulation {

    // Shared defaults every request in this simulation inherits.
    HttpProtocolBuilder httpProtocol = http
        .baseUrl("http://localhost:8080")
        .acceptHeader("application/json")
        .contentTypeHeader("application/json");

    // One product id per virtual user, picked at random from the CSV.
    // File lives at src/test/resources/product-ids.csv, header row: productId
    FeederBuilder<String> productIds = csv("product-ids.csv").random();

    ChainBuilder browse = exec(
        http("GET product by id")
            .get("/api/products/#{productId}")
            .check(
                status().is(200),
                jsonPath("$.id").exists()));

    ChainBuilder create = exec(
        http("POST create product")
            .post("/api/products")
            .body(StringBody("{\"name\":\"load-test-#{productId}\",\"price\":99.99}"))
            .asJson()
            .check(
                status().is(201),
                jsonPath("$.id").exists()));

    ScenarioBuilder scn = scenario("Product API browse + create")
        .feed(productIds)
        .exec(browse)
        .pause(Duration.ofSeconds(1), Duration.ofSeconds(3))
        .exec(create);

    {
        setUp(
            scn.injectOpen(
                rampUsers(50).during(Duration.ofSeconds(60)),
                constantUsersPerSec(5).during(Duration.ofMinutes(2)))
        )
        .protocols(httpProtocol)
        .assertions(
            global().failedRequests().percent().lt(1.0),
            global().responseTime().percentile3().lt(800),
            details("GET product by id").responseTime().percentile4().lt(1200));
    }
}

The matching feeder file, src/test/resources/product-ids.csv:

productId
P-1001
P-1002
P-1003
P-1004
P-1005

Walking through the important pieces:

  • Protocol (http.baseUrl(...)): base URL and default headers, inherited by every request. Change one line to point the same scenario at staging or production.
  • Feeder (csv(...).random()): each virtual user draws a random productId, referenced as #{productId} in URLs and request bodies. Without this, every user requests the same product and your cache hit rate is a fantasy — trap #3 below.
  • Checks (status().is(200), jsonPath("$.id").exists()): per-request correctness. A request that returns HTTP 500 is marked KO (failed) — it's load that surfaced the bug, and the check is what catches it.
  • Pause (pause(1s, 3s)): random think time between 1 and 3 seconds. Real users don't fire requests back-to-back; without pauses you model a DDoS, not traffic.
  • Injection profile: 50 users ramp up over 60 seconds, then a steady 5 new users per second for 2 minutes. The ramp is what this diagram shows — load rising gradually, then holding:
Virtual users over time 50 25 0 0s 60s 180s time → users → rampUsers(50) .during(60s) constantUsersPerSec(5) — steady load 50 concurrent users, new arrivals replacing finished ones wind-down
  • Assertions: the verdict for the whole run. Less than 1% failed requests, p95 response time under 800ms globally, p99 under 1200ms for the GET. If any assertion fails, the build fails — which is exactly what you want when this runs in CI.

Decision rule: write the assertions first, from your SLA. "p95 under 800ms" isn't a number Gatling invented — it's your product requirement, encoded as a build gate. A load test without assertions is a fireworks show: pretty, meaningless.

Running it: what to expect

From the project root, with your API running on localhost:8080:

mvn gatling:test

(To pick a specific simulation when you have several: mvn gatling:test -Dgatling.simulationClass=com.acme.perf.ProductApiSimulation.) Gatling prints progress as the injection runs, then a console summary like this:

================================================================================
---- Global Information --------------------------------------------------------
> request count                                  15230 (OK=15215  KO=15   )
> min response time                                  3 (OK=3      KO=4100 )
> max response time                               5200 (OK=1900   KO=5200 )
> mean response time                               214 (OK=208    KO=4300 )
> 50th percentile                                  120 (OK=119    KO=4300 )
> 95th percentile                                  640 (OK=620    KO=4600 )
> 99th percentile                                 1500 (OK=1400   KO=5000 )
---- Assertions ----------------------------------------------------------------
> Global: percentage of failed requests is less than 1.0   : true
> Global: 95th percentile of response time is less than 800: true
================================================================================
Reports generated in 0s.
Please open the following file: /home/you/shop/target/gatling/productapisimulation-20261004120000/index.html

The real gold is that index.html — open it in a browser. Gatling writes every run to target/gatling/<simulation-name>-<timestamp>/, so runs never overwrite each other. Decision rule: run the same profile against a staging environment sized like production, or the numbers mean nothing. A green run on your laptop proves your laptop is fast.

Reading the HTML report: the numbers that matter

The report is dense, so here's the annotated tour — the four numbers to read first, and what each one is telling you:

Global Information — ProductApiSimulation Requests 15,230 OK 15,215 KO (failed) 15 (0.1%) Mean response time 214 ms Throughput 181 req/s Response time distribution p50 120ms p95 640ms p99 1.5s min 3ms max 5.2s slow tail → investigate the requests at the far right KO% > 0 is a failed test — investigate, don't round down p95 is the number your SLA cares about — not the mean (hides the tail) flat throughput ceiling = downstream bottleneck (DB, external API)

What each number means, and what to do about it:

  • Requests/sec (throughput). Completed requests per second. Healthy runs show throughput rising with the ramp, then holding. If throughput plateaus while you keep adding users, you've found a ceiling — something downstream (database, external API, thread pool) can't go faster. The fix is never "add more load".
  • p50 / p95 / p99. Percentiles: p50 is the median, p95 means 95% of requests were faster than this, p99 the same for 99%. SLAs are written against p95/p99, not the mean — the mean hides the tail. A p50 of 120ms with a p99 of 1.5s means most users are happy and the unlucky 1% are furious.
  • Mean response time. Useful for trends, dangerous alone. Never report the mean as "our API takes X ms" — percentiles tell the real story.
  • KO% (failed requests). Any failed request under load is a signal. Even 0.1% at 15,000 requests is 15 failures — probably timeouts or refused connections, and they'll be your users' failures on launch day too.

A short sample interpretation of the numbers above: p95 640ms and KO 0.1% — the assertions pass, but p99 is 1.5s and throughput flattened at 181 req/s from 60 users onward. That's a queueing signature: requests are waiting, not computing. First suspects: the web-server thread pool and the database connection pool. Second suspect, if the p99 spikes arrive on a regular cadence: GC pauses — the JVM memory & GC post is your next stop for reading those.

Wiring it into CI: a nightly gate

Load tests don't belong in every pull-request build — they take minutes, need a real environment, and flaky network conditions cause false alarms. They belong on a schedule: nightly against staging, or as a merge gate on main. A minimal GitHub Actions sketch (the Actions mechanics get their own post later in this series):

name: nightly-load-test
on:
  schedule:
    - cron: "0 3 * * *"   # 3am UTC, when nobody is deploying
jobs:
  load-test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21" }
      - run: mvn -q -B gatling:test
      - uses: actions/upload-artifact@v4
        with:
          name: gatling-report
          path: target/gatling/*/

If an assertion fails, the Maven build fails, and the job goes red — your team wakes up to a report showing exactly when the p95 regressed. Keep the assertions slightly looser in CI than in your SLA (environments wobble); the trend across nightly runs is the real signal.

When it breaks: 4 Gatling traps

Trap 1 — The load generator is the bottleneck

Symptom: throughput flatlines, but the API's CPU and database look bored. The injector machine itself — often the same laptop running the build — can't generate requests fast enough: not enough threads, not enough ephemeral ports, GC pauses in the test JVM. You're measuring your laptop, not your API. For serious loads, run injectors on dedicated machines (or several) close to the target, and watch the injector's own CPU first when numbers look suspicious.

Trap 2 — Checks vs assertions confusion

check(status().is(200)) validates one response; assertions(...) judges the whole run. Beginners put SLA logic in checks ("if this request is slow, mark it failed") and then wonder why the run is green with a terrible p99. Keep the separation: checks verify correctness of individual responses, assertions encode the performance verdict for the build.

Trap 3 — No feeders, fantasy traffic

Without a feeder, every virtual user requests /api/products/P-1001 — the same product, the same cache line, the same database page. Your cache hit rate is 100% and your database is asleep. Real traffic spreads across keys. Feeder-less tests don't just under-test; they actively mislead. If you see suspiciously great numbers, ask whether the traffic was real before celebrating.

Trap 4 — Ramping too fast: the thundering herd

rampUsers(500).during(10) dumps 500 users in 10 seconds — no real launch looks like that, and the result is a synchronized burst that trips rate limiters and connection queues in ways gradual traffic never would. You'll "find" bottlenecks that don't exist in production and miss the slow-burn ones (connection leaks, memory growth) that do. Ramp like reality: minutes, not seconds. If production diagnosis under load ever points at thread exhaustion, the production diagnosis post shows how to confirm it with thread dumps.

What's next

You now have the full loop: a scenario in code, a realistic injection profile, assertions wired to your SLA, a report you can actually read, and a nightly job that keeps the whole thing honest. The habit to build is running the profile before every capacity-sensitive change — new query, new dependency, new cache — and comparing against last week's numbers, not against your hopes.

Field check before you move on: take the simulation above, point it at a real endpoint you own (even a local one), and run mvn gatling:test. Open the report, find the p95 and the KO%, and write down one sentence: "At N users, our p95 is X and the first thing to saturate would be Y." If you can't name Y yet, run it again with more users until the graph tells you.

Continue: Java Learning Roadmap 2026

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number