From Sanskriti Khandelwal | Product & Market Analysis
AI-Generated Unit Tests: Judge Them by Bugs Caught, Not Lines Covered
On this page
An AI model wrote tests for one Java function that reached 100% line and branch coverage, then caught 4% of the deliberately injected bugs. That gap is the whole problem with how AI testing tools are sold. Coverage tells you which lines ran. It does not tell you whether AI generated unit tests would fail when the code is wrong.
Key takeaways
- Coverage and bug detection come apart for AI-written tests. A 2025 study found Llama-3.3 tests for one HumanEval-Java subject at 100% line and branch coverage and a 4% mutation score.
- Mutation-guided generation is where the measured gains are. Meta's ACH system wrote 571 tests, and 277 of them killed injected faults without adding any line coverage, so a coverage gate would have thrown them away.
- Most AI tests encode what the code does today, bugs included. A University of Luxembourg study found LLM test generation mainly captures actual program behaviour, which makes catching existing bugs difficult.
- Your team can measure this without vendor help. Run a mutation tool such as PIT on one module before and after the AI tool, and count the extra mutants its tests kill.
The direct answer. AI generated unit tests reliably raise code coverage, but coverage does not show whether they catch bugs. Mutation testing does: it injects small faults and counts how many the tests detect. Published studies show AI tests ranging from a 4% to an 89% mutation score depending on method. Measure that score on your own code before you buy.
This post is written for the QA lead or head of delivery deciding whether an AI testing tool earns its seat. That reader converts on a measurement method, not a model comparison.
Why coverage is the wrong scoreboard for AI generated unit tests
Line coverage is the share of your code that runs at least once while the test suite runs. It is cheap to compute, every CI system reports it, and it rises fast when a model writes tests. That is exactly why vendors lead with it.
The trouble is that a line can run without anything checking its result. A test that calls a function and asserts nothing about the output still counts that function as covered. So does a test that asserts something trivially true. Coverage measures execution. It does not measure verification.
Mutation testing measures verification directly. A mutation tool makes a small change to your code, such as swapping a greater-than for a greater-than-or-equal, and reruns the tests. Each changed copy is a mutant. If at least one test fails, the mutant is killed. The mutation score is the share of mutants your suite kills.
This is not a new objection invented for AI. In 2014, Laura Inozemtseva and Reid Holmes generated 31,000 test suites for five Java systems of up to 724,000 lines and measured each one. Once suite size was held constant, they found only a low to moderate correlation between coverage and fault detection. Stronger coverage types, such as branch or modified condition coverage, gave no extra insight. Their conclusion was that coverage helps find under-tested code but should not be used as a quality target.
The same year, René Just and colleagues tested whether mutants stand in for real bugs. Using 357 real faults from five open-source applications, they found a statistically significant correlation between mutant detection and real fault detection, independent of code coverage. That result is why mutation score is the accepted proxy for bug-finding power in the research literature.
My position is blunt. Coverage should be retired as the headline metric for evaluating AI testing tools. It is still useful as a map of untested code. It is a poor scoreboard for a tool whose main talent is producing a lot of executed lines quickly.
What the published studies say about AI testing and real bugs
The evidence base is small, recent and mostly academic. It is still enough to show a pattern. When a tool is optimised for coverage, it delivers coverage. When it is optimised for killing mutants, it kills mutants.
A coverage-first tool in production at Meta
Meta's TestGen-LLM, published at FSE 2024, is the clearest industrial example of the coverage-first design. It extends existing Kotlin test classes and keeps a new test only if it builds, passes, and adds line coverage.
The filter is strict for a reason. In Meta's evaluation on Instagram components, 75% of test classes had a new test that built, 57% had one that passed reliably, and 25% had one that raised coverage. Engineers accepted 73% of the improvements it submitted for production. That is a real result. It is also measured entirely in coverage. The paper does not report a mutation score or a count of bugs found.
TestGen-LLM also throws away any test that fails on its first run. It does this because, without a specification, it cannot tell a real bug from a wrong assertion. So by design it only writes regression tests that pin down current behaviour. That choice is sensible engineering. It also means the tool cannot catch a bug that is already in the code.
A mutation-first tool from the same company
A year later, Meta published ACH, which flips the order. It first generates mutants that represent a specific class of fault, in this case privacy regressions. Then it asks a model to write tests that kill those mutants.
ACH ran across 10,795 Android Kotlin classes on 7 Meta platforms, producing 9,095 usable mutants and 571 tests. Engineers in Messenger and WhatsApp test-a-thons accepted 73% of the tests reviewed. The striking number is the overlap with coverage. Of the 571 tests, 277 would have been discarded by a line-coverage-only rule, even though each caught a fault no existing test caught.
The paper's own comparison makes the point sharper. TestGen-LLM produced tests for 32% of classes and killed 2.4% of mutants. ACH produced tests for 5.3% of classes and killed 15%. Touching more of the codebase did not mean catching more faults.
Open benchmarks outside Meta
Academic studies on public code tell the same story with more variance. The MutGen paper from Lero, the University of Ottawa and Huawei compared plain LLM prompting, mutation-guided prompting and EvoSuite, a search-based generator, on two Java benchmarks. Mutation-guided generation reached 89.5% and 89.1% mutation scores. Plain prompting reached 77.9% and 69.9%. All three approaches sat above 92% line coverage.
A larger study from the University of Luxembourg generated 216,300 tests across 690 Java classes with GPT-3.5, GPT-4, Mistral 7B and Mixtral 8x7B. It reported compilation failure rates of up to 86%. Where LLM tests did pass, they often targeted interfaces, abstract classes or trivial methods, which produced a 0% mutation score for those tests.
The TestPilot study from GitHub Next and university collaborators reported a median 70.2% statement coverage on 25 npm packages with gpt-3.5-turbo. It also found that a median 61.4% of generated tests contained non-trivial assertions. That second number is the one to notice. Roughly 4 in 10 tests in a typical package checked nothing meaningful, and coverage counted them all.
The oracle problem: AI tests that confirm the bug
An oracle is the part of a test that decides whether the output is correct. It is usually the assertion. Writing the code that calls a function is easy. Knowing what the right answer should be is the hard part of testing, and it always has been.
Most AI test generators solve the oracle problem by reading the code and inferring what it returns. That works when the code is correct. When the code has a bug, the model often writes an assertion that expects the buggy output. The test passes. Coverage rises. The bug is now protected by a green check.
Michael Konstantinou, Renzo Degiovanni and Mike Papadakis tested this directly across 24 Java repositories. Their finding was that LLM-based test generation approaches mainly capture the actual program behaviour, making bug detection difficult. They also found accuracy below 50% in some cases, and concluded that every LLM oracle suggestion needs human inspection.
There is a fair counterpoint in the same paper. LLM-generated assertions reached up to 2.96% higher mutation scores than EvoSuite's assertions. So the model is not worse than the older automated baseline. It is just not the independent check on correctness that a QA team assumes a test to be.
I would draw a hard line here. A test generated from the implementation is a change detector, not a correctness check. That is valuable for refactoring safety. It is not a substitute for a tester who knows what the system is supposed to do, and the problem gets worse when the same agent writes both the code and its tests. We covered the review side of that loop in the analysis of how AI code volume is overloading code review.
Published benchmarks for AI testing tools, side by side
The table below collects every figure in this post that measures fault detection or test quality, rather than adoption. Treat the rows as directional. They use different languages, models, mutation tools and code bases, so they cannot be ranked against each other.
| Study | Setting | Coverage result | Fault detection result |
|---|---|---|---|
| Meta TestGen-LLM, FSE 2024 | Kotlin, Instagram and Facebook | 25% of test classes gained coverage | Not measured |
| Meta ACH, FSE Companion 2025 | Kotlin, 10,795 classes, 7 platforms | 277 of 571 tests added no line coverage | 15% of mutants killed, against 2.4% for TestGen-LLM |
| MutGen, arXiv 2025 | Java, HumanEval-Java and LeetCode-Java | Above 92% line coverage for all approaches | 89.5% and 89.1% mutation score, mutation-guided |
| Ouédraogo et al., arXiv 2025 | Java, 690 classes, 216,300 tests | Best LLM median line coverage below EvoSuite | 0% mutation score on trivial passing tests |
| TestPilot, 2023 | JavaScript, 25 npm packages | 70.2% median statement coverage | Not measured; 61.4% of tests had non-trivial assertions |
| Konstantinou et al., arXiv 2024 to 2026 | Java, 24 repositories | Not the focus | Up to 2.96% above EvoSuite oracles; weak on existing bugs |
Two of the six studies never measured fault detection at all. That gap is the finding. Several are arXiv preprints rather than peer-reviewed journal versions, and the table says so where it applies.
One pattern holds across every row. When a paper reports both numbers, coverage is high and roughly flat across methods, while fault detection varies widely. If you only look at the coverage column, every tool looks about the same. They are not.
How to measure AI testing tools by bugs caught
You do not need a research lab to run this. You need one module, one mutation tool and about a week. The method below is the one I would run before signing any AI QA tools contract, and it doubles as an internal check on tests your developers already generate with coding assistants.
Step 1: baseline the mutation score on one module
Pick a module that matters, has an existing human-written test suite, and is small enough to mutate in under an hour. Run a mutation tool against it and record the score. PIT is the common choice for Java, Stryker covers JavaScript, TypeScript and C#, and mutmut covers Python.
Record three numbers: line coverage, mutation score and total mutants. Also record the runtime. Mutation testing reruns the suite once per mutant, so cost is the first thing that will surprise your team.
Step 2: add the AI tests and measure the delta per test
Run the AI tool on the same module and keep only tests that compile and pass reliably. Then rerun the mutation analysis. The metric that matters is new mutants killed by the AI tests that no human test already killed.
Divide that by the number of AI tests added. A tool that adds 200 tests and kills 4 new mutants has given you 196 tests of maintenance cost and very little protection. Also count AI tests that kill zero mutants. Those are the coverage-only tests, and they are candidates for deletion.
Step 3: read the surviving mutants by hand
The surviving mutants are the most useful output of the whole exercise. Each one is a specific change to your code that no test noticed. Some will be equivalent mutants, meaning the change does not alter behaviour, and those can be ignored. The rest are gaps.
Google's experience is the warning here. When it first surfaced mutants in code review, developers rated about 85% of them unproductive. Six years of filtering based on developer feedback raised the productive share from 15% to 89%. Your first run will be noisy, and that noise is a tuning problem rather than a verdict.
| Weak metric | Why it misleads | Metric that works |
|---|---|---|
| Line coverage gained | Counts lines run, not results checked | New mutants killed that no existing test kills |
| Number of tests generated | Rewards volume and adds maintenance cost | Mutants killed per test added |
| Test pass rate | A test asserting buggy output also passes | Share of tests that kill at least one mutant |
| Developer acceptance rate | Reviewers accept tests that look plausible | Accepted tests that kill a mutant, plus flake rate over 5 runs |
The last row needs one honest concession. Acceptance rate is not useless. Meta's 73% acceptance across both tools shows engineers will merge machine-written tests, which matters for adoption. It just answers whether people will keep the tests, not whether the tests protect anything.
Where the mutation testing argument is weakest
Every argument in this post leans on mutation score as the better yardstick. It has real weaknesses, and a QA lead should know them before using it to reject a tool.
Mutants are a proxy too
The 2014 Just study supports mutants as a substitute for real faults, but correlation is not identity. Mutation operators generate small syntactic changes. Many production bugs are missing features, wrong requirements or integration failures between services. No unit-level mutant represents those, and no unit test, human or AI, will catch them.
A high mutation score on a module also says nothing about whether the module does what the customer needed. It only says the tests notice when the code changes.
The benchmarks are small and the field moves fast
The strongest academic numbers come from HumanEval-Java and LeetCode-Java, which are self-contained algorithmic functions. Real services have state, dependencies and side effects that are much harder to test. Results on toy benchmarks routinely shrink on real repositories, which is the gap we measured for coding agents in the comparison of benchmark scores against real repository work.
The models are also changing faster than the papers. Most studies here used Llama 3, GPT-3.5 or GPT-4o-mini. A 2026 frontier model may write better oracles. I would still expect the coverage and fault-detection gap to persist, because it comes from where the assertion is derived, not from model quality. That is a judgement, not a measured result.
Finally, mutation testing is expensive. Google's whole system exists because computing a full mutation score across its codebase was not practical. For a large monorepo, you will need to mutate changed code only, as Google does, rather than the whole repository.
What to ask an AI QA tools vendor before you sign
Most AI testing tools report coverage gained, tests generated and time saved. None of those tells you whether you are better protected. Ask for numbers that do, and judge the answer by whether it comes with a sample.
Ask for the mutation score delta on a named reference customer's code, with the mutation tool, the module size and the time window stated. Ask what share of generated tests kill zero mutants. Ask how the tool decides what the correct output is, and whether it ever writes a test that fails against current code.
That last question sorts the field quickly. A tool that only emits passing tests is a regression test generator. That can be worth buying. It should be priced and governed as one, and not sold to your leadership as a bug-finding tool. The broader pattern of tools scoring well on the metric they choose is the subject of the breakdown of the three scoreboards AI coding tools are judged on.
If a vendor refuses the pilot in Step 2, treat the refusal as data. A vendor confident in fault detection has every reason to let you measure it.
The same logic applies to the AI tests your developers already write with coding assistants. If the test suite and the code come from the same prompt session, the assertions inherit the same misunderstanding. The longer-run cost of code nobody fully reviewed is laid out in the estimate of the cleanup bill for vibe-coded software. And the evaluation discipline here mirrors the one in the guide to building an LLM eval suite: decide what failure looks like, then measure whether you catch it.
Frequently asked questions
Do AI generated unit tests actually catch bugs?
Sometimes, and coverage will not tell you which. In a 2025 study, Llama-3.3 wrote tests for one HumanEval-Java subject with 100% line and branch coverage that killed only 4% of injected mutants. Mutation-guided generation raised average scores to about 89% on the same benchmark. The method matters more than the model, so measure the tool on your own code before trusting either result.
What is mutation testing in software testing?
Mutation testing makes small deliberate changes to your code, such as flipping a greater-than to a less-than, and reruns your tests. Each changed version is a mutant. If any test fails, the mutant is killed. The mutation score is the share of mutants killed. It measures whether tests check behaviour, while coverage only measures whether code ran.
Is mutation score better than code coverage?
For judging test quality, yes. A 2014 ICSE study of 31,000 generated test suites across five Java systems found only a low to moderate correlation between coverage and fault detection once suite size was controlled. Coverage still has a job: it shows which code no test touches. Use coverage to find gaps, and mutation score to judge the tests that fill them.
Which mutation testing tools work with AI testing tools?
PIT, also called PITest, is the standard choice for Java and other JVM languages, and most academic studies of AI-written tests use it. Stryker covers JavaScript, TypeScript and C#, and mutmut covers Python. All three run on whatever tests exist, so you can score an AI tool's output without the vendor's cooperation. Start with one module, because full runs are slow.
How do I evaluate AI QA tools before buying?
Pick one module with an existing test suite and record its mutation score. Run the AI tool on that module, then rerun the mutation analysis. Count the additional mutants the new tests kill, and how many new tests kill none. Ask the vendor for the same numbers from reference customers. A vendor who only reports coverage has not measured fault detection.
Why do AI generated tests pass on buggy code?
Most AI test generators write assertions by observing what the code currently does. If the code is wrong, the test records the wrong answer as correct. A University of Luxembourg study found LLM test generation mainly captures actual program behaviour, which makes bug detection difficult. Meta's TestGen-LLM discards any test that fails on its first run for a related reason: it cannot tell a bug from a bad assertion.
Where to start this week
Pick the one module your team would least like to see break in production. Run a mutation tool on it on Monday and write down three numbers: coverage, mutation score and runtime. That baseline is the thing nobody in your vendor conversations will have.
Then take the last 20 AI-written tests merged into that module and check how many kill at least one mutant. If the answer is fewer than half, you have found your real coverage number, and the next tool decision gets much easier.
Related on trust in AI-written code
Developers already doubt much of what coding assistants produce. The survey evidence on that gap is in the analysis of why developer trust lags AI usage.
References
- Wang, Xu, Briand and Liu, Mutation-Guided Unit Test Generation with a Large Language Model, arXiv 2506.02954, first posted June 2025, revised 2026. Used for the 100% coverage and 4% mutation score case, MutGen scores and baselines. Preprint.
- Foster, Gulati, Harman et al., Mutation-Guided LLM-based Test Generation at Meta, FSE Companion 2025, arXiv January 2025. Used for ACH scale, the 277 of 571 tests, acceptance rates and the TestGen-LLM comparison.
- Alshahwan et al., Automated Unit Test Improvement using Large Language Models at Meta, FSE 2024. Used for TestGen-LLM build, pass and coverage rates and the 73% acceptance figure.
- Petrovic, Ivankovic, Fraser and Just, Practical Mutation Testing at Scale: A view from Google, IEEE TSE, 2021. Used for productive mutant rates and scale.
- Inozemtseva and Holmes, Coverage Is Not Strongly Correlated with Test Suite Effectiveness, ICSE 2014. Used for the coverage and effectiveness correlation.
- Just, Jalali, Inozemtseva, Ernst, Holmes and Fraser, Are Mutants a Valid Substitute for Real Faults in Software Testing?, FSE 2014. Used for the mutant and real fault correlation.
- Konstantinou, Degiovanni and Papadakis, Can we find bugs using LLM-generated oracles?, arXiv 2410.21136, 2024, revised 2026. Used for the oracle finding. Preprint.
- Ouédraogo et al., Large-scale, Independent and Comprehensive study of the power of LLMs for test case generation, arXiv, 2025; and Schäfer, Nadi, Eghbali and Tip, An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation, arXiv 2302.06527, 2023. Used for compilation, coverage and assertion figures.
The weakest part of this source base: four of the eight sources are preprints, the strongest fault-detection numbers come from small algorithmic benchmarks, and no vendor has published a mutation score on customer code that could be checked. Figures are current as of 10 October 2026.
Related reading