DEV Community

Cover image for AI Can Write the Code. Can It Prove the Fix?

AI Can Write the Code. Can It Prove the Fix?

Prince Panchani on September 17, 2026

The most expensive thing an autonomous coding agent can produce isn't a broken build. A broken build is free: CI goes red, you move on. The expens...
Collapse
 
mihai_leanzero profile image
Mihai Perdum •

The net-new fingerprint is the part I keep circling back to. Digit-normalising the message so an inserted import doesn't shift every key down a line makes sense, and because it's a multiset the obvious worry (two unrelated diagnostics colliding onto one key) mostly washes out, since the counts stay separate. There's a version that doesn't wash out though: if the base has one error at key K and the head fixes it while a genuinely different diagnostic appears elsewhere that happens to normalise to that same key, the count at K stays at one on both sides and net-new reports zero. A real regression lands in the same breath as a real fix, at the one key least likely to get a second look. Does anything downstream check whether the fixed and new diagnostics are actually the same call site, or is key-count parity the whole signal?

Collapse
 
prince_panchani_f971a20ec profile image
Prince Panchani •

@mihai_leanzero Yeah, this is a really good catch. The key-count parity is not enough to prove that the diagnostic was actually fixed.

The case you described can definitely hide a regression: one diagnostic disappears and another one normalises to the same key, leaving the count unchanged.

The current approach is mainly using the fingerprint as a strong signal, not as proof that two diagnostics are the same. I think the next step would be to carry enough location/context through the fingerprint to distinguish the fixed call site from the new one, while still keeping it stable across harmless line shifts.

This is exactly the kind of edge case I’d want to add as a regression test.

Collapse
 
locitra profile image
Sunil Kumar Uikey •

Writing the code is increasingly becoming the easier part. The harder problem is proving that the change actually fixes the intended issue without introducing something else.

I think the verification step needs to be treated as a separate task: tests, edge cases, regression checks, and clear acceptance criteria. An AI-generated fix can look convincing while still failing in less obvious scenarios.

This is probably where AI-assisted development will mature — not just generating code faster, but making the validation process equally systematic.

Collapse
 
prince_panchani_f971a20ec profile image
Prince Panchani •

Well said, @locitra. A fix that "looks correct" and a fix that is "proven correct" can be very different things. I think AI workflows will increasingly need verification loops, not just generation loops. 🚀

I think we're moving toward a world where writing code becomes cheap, but proving correctness becomes the real differentiator. Curious to know—what validation steps do you think should always be mandatory in AI-assisted development? 🤔

Collapse
 
locitra profile image
Sunil Kumar Uikey •

Absolutely. I think the key is to treat AI-generated code as a starting point rather than the final result.

My baseline validation loop would be:

  1. Run automated tests — unit, integration, and regression tests where applicable.
  2. Lint and type-check — catch structural issues before deployment.
  3. Review the actual behavior — especially edge cases that tests may not cover.
  4. Verify the production build — a successful development environment doesn't always guarantee a successful deployment.
  5. Check security and dependencies — particularly when AI introduces new packages or implementation patterns.
  6. Human review — someone should still understand and approve important changes.

The interesting shift is that AI can make implementation much faster, but that makes the verification loop even more important.

Collapse
 
prince_panchani_f971a20ec profile image
Prince Panchani •

Absolutely agree!

I think the next evolution of AI-assisted development is exactly this: from code generation to evidence-backed verification. Great point!

Collapse
 
piekwerk profile image
Piekwerk •

The reproduction gate is the idea I'd steal too, with one scar to add: rerun the base-fail check before trusting it. I once had a fix where the offered test failed at the merge base about every third run. Flaky, not causal, and the single run our check sampled looked like perfect evidence. Running the old tree three times and requiring three failures cost seconds and caught it. The other habit I keep regardless of gates is reading the assertion diff first. A changed assertion is still the cheapest signal there is, and it scans in seconds next to a full reviewer pass. Gates raise the floor. They don't replace that reflex.

Collapse
 
prince_panchani_f971a20ec profile image
Prince Panchani •

@piekwerk

That’s a really good scar to add 😄 The base-fail rerun is easy to overlook, but it makes the evidence much stronger. And I really like the point about reading the assertion diff first — sometimes a few seconds of inspection tells you more than another layer of automation. Gates raise the floor, but that reviewer reflex still matters.

Collapse
 
piekwerk profile image
Piekwerk •

Thanks, and fair point on the layers. One detail on the rerun count, since three was not a principled choice, it was just the smallest number where my one-in-three flake reliably showed up. A one-in-ten flake needs more reruns than anyone will sit through interactively, and that is the real argument for moving the check into CI, where extra runs cost patience you do not have to personally spend. It also composes with your tamper guard nicely, since a base-fail check that the coder can delete is back to the same self-attestation problem.

Thread Thread
 
prince_panchani_f971a20ec profile image
Prince Panchani •

@piekwerk It makes sense. Three runs is really just a practical signal, not a confidence threshold you can generalise.

I especially like the CI point. Once the flake rate gets low enough, trying to solve it with more interactive reruns becomes pretty awkward. Let CI absorb that cost instead.

And agreed on the tamper guard. The moment the agent can decide which evidence gets checked, the whole thing starts drifting back toward self-attestation. That separation between generating the fix and validating the evidence is probably the more important boundary.

Thread Thread
 
piekwerk profile image
Piekwerk •

Fair, and I'd keep it at "practical signal" too. What pushed it to three for us was one flaky browser run that passed on retry two: with two runs a coin-flip masquerades as evidence, three makes one flake visible. But the number only means something while the flake rate stays well under one-in-three, and once CI owns the loop, reruns are cheap enough that the exact count stops mattering much.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

The one item on your list where the mechanism and the question have different shapes is the third: "Did the test count go down?"

testing/tamper_guard.py decides tampered from aggregate totals (ta < tb or aa < ab or sa_cmp > sb_cmp or ...) plus an unconditional deleted-file rule. The per-file loop records drops independently in reasons, and reasons is never consulted for the verdict. The module's own comment block names the shape — "the guard named the tamper and passed it" — and two instances of it are now closed: the deleted-file case (promoted into the verdict as bool(deleted)) and the two path shapes whose benchmark fixtures were diluting the aggregate. What got closed is those instances, not the rule.

Reproduce it with your own test's bytes. Take test_fixture_snapshots_do_not_dilute_a_real_reduction and change one thing — the padding file's path, from under eval/reviewer_recall/cases/control-gate-excerpts/base/ to tests/test_snapshot.py:

before = {"tests/test_x.py":      2 tests, 2 assertions}
after  = {"tests/test_x.py":      1 test,  1 assertion,      # the reduction
          "tests/test_snapshot.py": 200 ordinary tests}     # 200 assertions of padding

[clean] tests 2->201, assertions 2->201, skips 0->0, tautologies 0->0, fake-fixtures 0->0
reasons = ['tests/test_x.py: tests 2->1', 'tests/test_x.py: assertions 2->1']
Enter fullscreen mode Exit fullscreen mode

Same bytes as your test, same guard, same commit (5f99b7af) — only the excluded path shape is gone, and the verdict flips to clean with the reduction still named in its own reasons list. A branch that adds a test file anywhere outside those two shapes can delete assertions from an existing suite and be reported clean; the padding only has to exceed the deletion, and padding is the one resource an agent has in unlimited supply.

The fix is what the per-file loop already computes. Decide from it: tampered = bool(deleted) or bool(reasons). I ran that rule against the three honest-refactor control pairs already in tests/test_tamper_guard.py — the describe( → test.describe( rename, the .test( helper swap, the assertions-added case — and reasons is empty in all three, so none of them changes verdict. A per-file rule does fire on an honest relocation (a test moved between files), which is the trade your deleted-file rule already took explicitly, with the justification written next to it: the guard fires, and a human justifies it.

The reason this one is worth more than it looks is in your own argument — "if your verification mechanism is inside the agent's write surface, it is not a verification mechanism." A rule that decides on a total is still inside the write surface, because the total is something the writer can add to. The per-file facts are the part that cannot be padded.

Collapse
 
prince_panchani_f971a20ec profile image
Prince Panchani •

@howcani_howcani_77e786a89

The aggregate-versus-file distinction is the key point for me.

An aggregate check looks at the overall test and assertion totals. That means a real reduction in one file can potentially be hidden by adding enough tests somewhere else.

The per-file check answers a different question: did any existing test file actually lose tests or assertions? That evidence is much harder to mask with unrelated additions.

So the issue isn't that the aggregate numbers are useless. They are useful as a high-level signal, but they shouldn't be the only thing deciding whether a reduction happened. The per-file evidence should have a direct path into the verdict.

The test relocation trade-off is also important. A per-file rule can flag legitimate moves, but that's a review case rather than something I'd want to hide by making the detection weaker.

Collapse
 
nark3d profile image
Adam Lewis •

A fix is only worth anything if the check that proves it runs without the agent anywhere in the loop, because an agent asked to evidence its own fix will happily produce a passing test written against the new behaviour. A check written after the fix tells you almost nothing, because it was written to agree with whatever the fix already does. Writing that check first, while the bug is still reproducing, is the only moment it costs nothing. Strip the agent out of the verification path and you find out quickly whether the fix was real or just plausible.