Mutation Testing with PIT: Does Your Test Suite Actually Catch Bugs?
Monday, 8 AM. Refunds are overpaying — customers who should get $90 back are getting $110. The diff, when someone finds it, is one character: price * (1 - discountRate) became price * (1 + discountRate) in a late-night refactor. The test suite is green. Line coverage: 100%. Branch coverage: 100%. The coverage gate passed with room to spare. Every number the previous post taught you to trust is green, and the math is backwards.
This is the gap the last post left open on purpose. Coverage measures whether your tests run the code. It cannot tell you whether they would notice if the code were wrong. The test that "covered" the refund math had no assertion — it called the method, ignored the result, and passed. Coverage loved it. The bug shipped anyway.
Mutation testing closes that gap by fighting back. It makes tiny, deliberate changes to your code — flips a - to a +, turns a > into a >=, deletes a method call — then runs your test suite against each mutated version. If a test fails, the mutant is killed: your suite caught the bug. If every test still passes, the mutant survived — and you have found a hole in your tests, not your code. This post runs PIT, the standard mutation testing tool for the JVM, against the same DiscountCalculator from the coverage post.
Mutation testing in one paragraph
A mutant is your program with one small syntactic change applied — the kind of change a tired developer makes at 1 AM. A mutation operator is the rule that creates mutants: negate a conditional, replace an arithmetic operator, remove a method call, return a default value instead of the computed one. PIT applies these operators across your code, runs your full test suite once per mutant, and reports the mutation score: killed mutants ÷ total mutants. A score of 100% means every seeded bug was caught. A surviving mutant is a concrete, reproducible demonstration that a specific class of bug can slip past your suite today.
The premise is deliberately adversarial, in the same spirit as this track's JMH post: don't trust the harness, make the harness prove itself. Coverage asks "did you run it?" Mutation testing asks "would you have noticed?"
The example, again
The same class from the coverage post — unchanged, 100% line and branch coverage with the full test suite from that post:
package com.javamakeuse.billing;
public class DiscountCalculator {
public double apply(double price, double discountRate) {
if (price < 0) {
throw new IllegalArgumentException("price must be >= 0");
}
if (discountRate < 0 || discountRate > 0.5) {
throw new IllegalArgumentException("discountRate must be in [0, 0.5]");
}
return price * (1 - discountRate);
}
}
Now add the test from the end of the coverage post — the one with no assertion. It exists in more codebases than anyone admits:
@Test
void discountDoesNotThrow() {
calc.apply(100.0, 0.10); // passes — and proves nothing
}
Suppose the suite is exactly this: the happy-path assertion test plus this assert-less one. Coverage is 100% / 100%. The gate is green. Watch what PIT does to it.
A surviving mutant, step by step
PIT's math mutator takes the return statement and applies the smallest possible change:
--- DiscountCalculator.java (original)
+++ DiscountCalculator.java (mutant 1: MATH — replaced - with +)
@@
- return price * (1 - discountRate);
+ return price * (1 + discountRate);
PIT compiles this mutant and runs the whole suite against it:
tenPercentOffHundredasserts90.0. The mutant returns110.0. The assertion fails — mutant killed. This is the assertion doing its job.discountDoesNotThrowcalls the method and checks nothing.110.0comes back, nobody looks at it, the test passes — mutant survived against this test.
Now the punchline: remove tenPercentOffHundred from the suite — or imagine a module where the assert-less style is the norm — and mutant 1 survives completely. The report flags it: SURVIVED, with the exact line and the exact change. That row in the report is a written confession: "if someone flips this operator, nothing in your suite will notice."
Contrast with a mutant the suite does catch. The conditional-boundary mutator changes the price guard:
--- DiscountCalculator.java (original)
+++ DiscountCalculator.java (mutant 2: CONDITIONAL BOUNDARY — replaced < with <=)
@@
- if (price < 0) {
+ if (price <= 0) {
Coverage can't see this change — both versions execute identically on every test that never passes exactly 0.0. But add one boundary test:
@Test
void zeroPriceIsFree() {
assertEquals(0.0, calc.apply(0.0, 0.10), 0.0001);
}
The mutant throws on 0.0; the test fails — mutant killed. Notice what just happened: mutation testing told us which test to write. The surviving-mutant list is a to-do list for your test suite, prioritized by real bug classes.
Adding PIT to Maven
PIT is a Maven plugin invoked with the mutationCoverage goal. For JUnit 5 suites it needs its JUnit 5 plugin as a plugin dependency — without it, PIT silently finds no tests:
<plugin>
<groupId>org.pitest</groupId>
<artifactId>pitest-maven</artifactId>
<version>0.26</version>
<dependencies>
<dependency>
<groupId>org.pitest</groupId>
<artifactId>pitest-junit5-plugin</artifactId>
<version>1.2.3</version>
</dependency>
</dependencies>
<configuration>
<targetClasses>
<param>com.javamakeuse.billing.*</param>
</targetClasses>
<targetTests>
<param>com.javamakeuse.billing.*</param>
</targetTests>
<mutators>
<mutator>DEFAULTS</mutator>
</mutators>
<outputFormats>
<value>HTML</value>
</outputFormats>
<timeoutFactor>1.25</timeoutFactor>
</configuration>
</plugin>
(Check for newer artifact versions than the ones pinned above — the coordinates are the stable part.) The configuration knobs that matter:
targetClasses/targetTests— scope the run. Never point PIT at your whole codebase on day one; start with one package, the way you'd start coverage with one module.mutators—DEFAULTSis the well-tested set (conditionals, math, increments, return-value and void-call mutators). Stronger sets exist; start here.timeoutFactor— mutants that loop forever are killed by timeout, not by your tests. 1.25× the normal test time is a sane default; raise it if you see timeout-kills on slow integration tests.
Run it after compiling the tests:
mvn test-compile org.pitest:pitest-maven:mutationCoverage
The HTML report lands in target/pit-reports/. Gradle users: the community info.solidsoft.pitest plugin (check for its latest version) wires the same goal to ./gradlew pitest with pitest { targetClasses = ['com.javamakeuse.billing.*'] } in the build script.
Reading the mutation report
Open target/pit-reports/index.html. The headline is the mutation score per package — killed ÷ total. Below it, every mutant gets a row: the file and line, the operator applied, and one of three verdicts:
- KILLED — at least one test failed on this mutant. Your suite catches this bug class. (Mutant 2 above, once
zeroPriceIsFreeexists.) - SURVIVED — the full suite passed on broken code. This is the actionable row: it names the exact line and the exact change your tests are blind to. (Mutant 1, against the assert-less test.)
- NO_COVERAGE — no test even executed the mutated line. This is the coverage post's red line wearing a different hat — fix it by writing a test that reaches the code, then check the mutant dies.
The report also shows test strength — killed ÷ (killed + survived), i.e. the score ignoring uncovered code. Test strength answers "of the code my tests reach, how much do they actually verify?" — which is precisely the question the assert-less test was gaming. When line coverage is 100% but test strength is 60%, your suite runs everything and verifies little.
Read it the way you read the coverage report: start with the SURVIVED rows, write the test each one is asking for, re-run, watch the row flip to KILLED. The list shrinks as your suite gets honest.
Mutation testing in practice
Mutation testing is the most expensive tool in this track, and the guidance has to be honest about that:
- It's slow. PIT runs your suite once per mutant. Hundreds of mutants × a multi-minute suite = hours on a real codebase. This is a nightly or scheduled CI job, not a per-PR gate — the same way the Gatling load tests from this track run on a schedule, not on every commit. The GitHub Actions post's workflow gets a
schedule: cronjob for this, separate from the PR pipeline. - Scope it. Start with the package that handles money, auth, or data integrity — the code where a wrong-but-covered line costs you the most. Expand as the score stabilizes.
- Equivalent mutants exist. Occasionally a mutant is semantically identical to the original (e.g. reordering two independent operations) — no test can kill it. Adjudicate these by hand, exclude them, and move on. A handful of equivalents is normal; a flood of them means your mutator set is too aggressive for the code.
- Don't chase 100%. The score is a trend metric, like coverage: watch it climb as the suite improves, investigate when it drops. A team that mandates 100% mutation score gets a team that writes mutants-aware tests instead of good tests.
- Incremental runs help. PIT supports history files that skip re-running mutants in code that hasn't changed — check the current PIT docs for the exact configuration. On a large codebase this is the difference between "runs nightly" and "never runs."
Decision rules: run mutation testing where bugs are expensive and coverage is already high — that's exactly where its signal is strongest. Treat every SURVIVED row as a test-writing task, not a code bug. And keep it off the PR critical path: a tool that takes hours belongs in the nightly build, where a red result is investigated in the morning, not a merge blocker at 6 PM.
The full tooling loop
Ten posts, one working loop. Build it, prove it works, prove it honestly, ship it, prove it survives:
- Maven vs Gradle: Builds Demystified — the build itself.
- JUnit 5 & Mockito: Testing That Catches Bugs — the test suite.
- Testcontainers: Real Databases in Tests — integration tests against real infrastructure.
- Docker for Java Devs: Layered Jars & Jib — packaging for production.
- CI/CD with GitHub Actions for Java Projects — the pipeline everything runs in.
- Load Testing with Gatling: Prove Your API Survives Traffic — the system under real load.
- Static Analysis for Java: SpotBugs, Checkstyle & SonarQube — the code kept clean.
- JMH: Microbenchmarking Done Right — the numbers kept honest.
- JaCoCo Code Coverage: Prove Your Tests Actually Run the Code — the suite kept complete.
- Mutation Testing with PIT: Does Your Test Suite Actually Catch Bugs? — the suite kept honest.
The through-line of the whole track: every tool here exists because a simpler measurement lied to someone. Hand-rolled benchmarks lied about speed; green builds lied about shipping; 100% coverage lied about verification. Each post added one tool that refuses the previous lie. Mutation testing is the last one — after it, the only thing left to distrust is the code review, and that's a people problem no plugin can fix.
The next track — Spring Boot & APIs — puts this entire loop to work building real services, starting with your first REST API.
Field check before you move on: take the DiscountCalculator and the assert-less test, run PIT, and confirm mutant 1 shows SURVIVED. Now add the assertion from tenPercentOffHundred and re-run — watch the row flip to KILLED. Then write one more test for any remaining SURVIVED row. That loop — survive, write, kill — is mutation testing as a practice, not just a tool.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment