An AI model wrote 69 tests for a Python module. Every one passed. Together they caught zero of the eleven bugs I had deliberately planted in that module.
A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten.
That contrast is the whole project.
What I built
A mutation-testing harness. It changes source code in small ways, flips a comparison, alters a constant, deletes a raise, then runs the existing test suite and records which changes the suite fails to notice. A change nothing catches is a fault your tests cannot detect.
Coverage tells you a line ran. This tells you whether anything would have failed if that line had been wrong. Those are very different numbers. A toy module with one happy-path test showed 47% line coverage and caught 2 of 21 mutations.
Then I pointed an agent at the mutations the tests missed. For each one it writes a single test, and that test is kept only if it both passes on clean code and fails on that specific mutation. Ground truth is a subprocess exit code. No model judges any outcome.
What I found
Three approaches, same model, same token ceiling, twelve widely used Python libraries including cachetools, toolz, tenacity and boltons. 455 mutations generated, 53 of which survive the existing tests and sit on lines those tests actually execute.
| approach | caught |
|---|---|
| one prompt, "write more tests" (556 tests) | 9 of 53 |
| one test per call, no targeting (53 tests) | 2 of 53 |
| targeted at the specific fault, with a pass/fail gate | 44 of 53 |
I hand-audited the nine it missed. Seven are provably unkillable, six of those being type annotations inside TYPE_CHECKING blocks that never execute at runtime. So 44 of 46 that could be caught at all.
Three findings I think matter more than that number.
Most undetected faults are unreached code, not weak assertions. Only 53 of 133 surviving mutations sit on a line the tests execute at all. I suspected my test commands were scoped too narrowly, so I widened every one of them by 6 to 40 times. The count went from 54 to 53. Down. In mature, human-written code, tests do not mostly execute code without checking it. They mostly do not execute it.
The gate never once rejected a broken test. Across 74 attempts, every rejected draft was a valid, passing test that simply failed to detect the fault. Not one was broken. The gate turns out to have exactly one job in practice, and it is not the job I designed it for.
The tests it writes catch the fault they were shown and almost nothing else. Zero cross-function transfer across 44 kept tests. 36 of the 44 catch exactly one mutation. That is the uncomfortable result, and it belongs next to the 44 rather than underneath it.
Why it actually mattered
Building the agent took a weekend. The rest of the time went into discovering that my measuring instrument kept lying to me.
Eleven bugs in the harness itself. Editable installs that made mutations invisible. Parallel execution corrupting one target. A classifier running on the wrong unit. A bytecode cache serving stale results. An outcome bucket that had never once filled, because it matched on a string my version of pytest does not emit. I had published that empty bucket as a finding.
Two things they had in common.
Every one made the results look better, or made an absence look like evidence. That is selection rather than conspiracy. Debugging is triggered by surprise, and a pleasing result is not surprising. So the filter that removes measurement bugs gets applied unevenly, hard against results you dislike and softly against results you like. The ones that flatter you survive to publication.
Not one was found by reading the code. Every single one was caught by running a check whose outcome I had predicted in advance and getting the wrong answer.
Then I published, and three readers found three more. All of them by pointing at my checks, never at my numbers. Nobody has disputed a single result. Every correction landed on the instrument.
What I would tell you
If you build evaluations for your own work, the harness is the part worth publishing. A result is a claim people can take or leave. An instrument is something they can attack, and the attacks are what tell you whether it works.
I got more out of three comment threads than out of forty hours of my own review.
Repo, with all eleven bugs documented and every result reproducible from a clean clone: killcheck repo
The longer write-ups, if you want them:
- The eleven bugs, and why they all pointed one way
- Six checks to run before trusting your own harness
- What happened when readers audited the checks
Thanks to Vinh Nguyen (@vinhnguyenthanhdn), Ahmet Özel (@ahmetozel) and Zain Dana Harper (@zaindanaharper). Three findings, three checks improved, zero numbers moved.
Top comments (5)
This is the trap I see teams fall into after they adopt an agent: the test count goes up, the score looks great, and the actual signal goes down. Tests that pass because the agent learned the test shape, not the problem shape, are worse than no tests. I started maintaining a held-out set that the agent never sees during training runs. How do you decide which tests are the ones worth keeping?
The trap is real, but the mechanism is different here, and the difference changes the answer.
There's no training loop; each call is stateless: one fault, one prompt, one test, no memory between them. So the agent can't learn a test shape across examples. What you're describing arrives another way: the model writes a test shaped like a test rather than like the problem, independently, every time.
I measured that. Built a deliberately vacuous suite, imports, calls, asserts nothing meaningful, expected near zero, got 7 of 51. Every kill was assert x is not None catching a function that started returning None. A real detector for exactly one fault class and nothing else. That's test-shape, and it passes any naive gate.
What I keep: a test survives only if it passes on clean source and fails on the specific fault it targets. Necessary, and weaker than it sounds, across 74 attempts the gate never once rejected a broken test. Every rejection was a valid, passing test that failed to detect the fault. It filters for sensitivity to one edit, not for whether the test specifies anything.
The better criterion came from an experiment a reader suggested last week, close to your held-out practice. I froze the 44 kept tests, generated 92 fresh faults on the same functions that nothing had seen, and scored the frozen tests alone. 34 of 53 reachable.
The split is the useful part. Tests asserting a concrete value: 28 kills. Existence checks: 4. And of the 44, exactly 22 caught nothing unseen while 22 caught at least one. Half probes, half something broader, same gate, two different artifacts.
So a keep rule I'd actually use: gate first, then assertion class as a quality signal, with existence-only tests flagged for review rather than trusted.
One thing worth passing on. I tried your held-out approach twice and failed both times, held out by mutation operator, pooled denominator of 1; held out by position, rounded to 0-1 per target. Both failed because I was partitioning a fixed set. Generating a fresh population instead of slicing the existing one is what made it work.
Caveats: 92 faults against a floor of 100 I'd pre-registered, so the floor wasn't cleared, and two targets supply most of the volume.
The zero cross-function transfer result left me wondering about fresh mutations in the same functions. Have you tried freezing the 44 retained tests, then generating mutations the agent never saw? I'd be curious whether the assertions generalize within each function, even when they don't transfer across functions.
This is the control I couldn't build, and I think you've found the way around the problem that killed my first two attempts.
I tried two holdout designs before publishing and abandoned both. Held out by operator family: pooled denominator of 1. Held out by position within the module: rounds to zero or one per target. Both failed for the same reason: I was slicing a fixed set of 53 survivors, and 53 doesn't partition into anything you can compute a rate over.
Generating fresh mutations sidesteps that completely. The denominator stops being a slice of 53 and becomes however many new mutants I can generate on the functions the 44 kept tests actually cover. That's a few hundred rather than single digits, which is the first time this question has had enough data behind it to answer.
It also splits something I currently can't separate. Zero cross-function transfer tells me the tests don't generalise across function boundaries. It says nothing about whether they constrain behaviour inside the function they were written for. A test that catches only the exact mutation it was shown is a probe. A test that catches mutations in the same function it never saw is closer to a specification. That's the sensitivity-versus-specification distinction my README currently says I can't resolve.
I'll run it. Design as I see it: freeze the 44, generate mutants at sites the original run excluded, sites already killed by the existing suite, and sites the kept tests cover that weren't in the survivor set, then score frozen-44 against them with no regeneration and no retries. Pre-registering the prediction before I run: I expect within-function transfer to be low but non-zero, and I expect it to be concentrated in the value-class tests rather than the existence ones.
Writing that prediction down first because four of my eleven bugs were caught by predicting an outcome and hitting a contradiction. If it comes back at zero, that's a much sharper negative than what I have now.
Ran it. You were right that the question was answerable, and the answer corrected a claim I'd published.
My post says the generated tests catch the edit they were shown and essentially nothing else. That's true across function boundaries and false within them. Frozen the 44 kept tests, generated 92 fresh mutants on the same functions that the agent never saw, and scored the kept tests alone against them. 34 of 53 reachable.
The design point you found is the thing that made it work, and it's worth naming. I'd abandoned two holdout controls before publishing - one held out mutation operators and ended up with a pooled denominator of 1, the other held out by position and rounded to zero or one per target. Both failed the same way, because I was partitioning a fixed set of 53 survivors. Generating a fresh population instead of slicing the existing one sidesteps it entirely. A failed holdout's fix is a new population, not a cleverer partition.
The class result held exactly as predicted: 28 of the kills came from value-class tests, 4 from existence-class. That finally separates two things my Limitations section said I couldn't. Value-class tests behave like specifications within their function. Existence-class tests behave like probes. Same gate, same loop, different kinds of artifact coming out of it.
Three mutants were killed by a generated test and missed by the library's own suite - boltons and toolz. n=3, so an existence proof rather than a rate, but it's something I couldn't claim at all before.
Two things I'd want flagged if you quote any of this. The pooled fresh population is 92 against a floor of 100 I pre-registered before writing the experiment. The floor wasn't cleared, and I'm not waiving it after the fact. And two targets supply 30 of the 53 reachable, so it's less evenly distributed than the pooled number suggests.
The part I got wrong: I pre-registered "low but non-zero." Non-zero held. Low did not - 0.64 isn't low by any reading, and I'd have had to retrofit the word to make it fit. A pre-registered prediction that was wrong in the direction that flatters my own tool is exactly what pre-registration is for, so it's in the write-up as a miss rather than absorbed into the result.
Also found two more instrument bugs while building it. One of them, if I hadn't caught it on a single-target dry run, would have collapsed the denominator to only the already-killed mutants and reported a perfect 1.00 transfer rate. The cleanest, most quotable number the project could have produced, and entirely an artifact.
Full write-up is in the README under within-function transfer, with the design and the prediction in METHODOLOGY.md, dated before the experiment existed. Credited you in the changelog. Thank you; this was the control I'd given up on.