DEV Community

Your Agent Saved the File. Who Checked the Result?

Bryan Williams on September 27, 2026

The tool call succeeds. The file exists. The agent says, “Done.” Then someone opens it and finds the wrong total. Nothing crashed. The system ans...
Collapse
 
glenallen profile image
Glen Allen •

There’s another failure mode worth considering here: the verifier itself can inherit the same assumption that produced the bad result. Reading the delivered bytes is stronger than trusting the tool receipt, but if the expected contract or verification logic was derived from the same generated output, the system can still validate the wrong thing very consistently. I’d treat independence of the acceptance criteria as part of the verification boundary. The strongest check seems to be one where the expected outcome comes from an authoritative source that the agent cannot modify, while the verifier independently evaluates the actual artifact against it. That makes “verified” mean more than “the system successfully confirmed its own assumptions.”

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

The report should never get to supply its own answer key. I tried a concrete version of your failure case against the example: keep the trusted input at 7 bolts and 9 washers, but change the delivered rows to 8 and 8. The total is still 16, yet the verifier rejects it as wrong_rows. If I replace the expected input with those same wrong rows, it passes. Same artifact, opposite verdict—entirely because of where the expectations came from.

That gives us a useful regression test for the boundary you're describing. Keep the acceptance input separately versioned, deny the producing agent write access to it, and retain a known-wrong artifact that must fail. Those controls don't establish that the source itself is true; they prevent the producer from quietly rewriting what counts as correct.

Collapse
 
glenallen profile image
Glen Allen •

That’s a really useful way to turn the boundary into a concrete test. The 'same artifact, opposite verdict' example makes the provenance problem much easier to see. I also like the idea of keeping a known-wrong artifact as a permanent regression case, because it tests whether the verifier can actually reject something rather than only confirm expected outputs. Together, those checks make the acceptance layer independently testable instead of assuming that a passing verification result is meaningful by itself.

Collapse
 
arhancanli profile image
Arhan Canli •

"The hash identifies the bytes that were checked. It does not make them correct" is the sentence most receipt designs skip. We ran into the same split with a validation API: every hosted result comes with a signed receipt, and the property that matters isn't the signature but that anyone can recompute the result from the inputs in the receipt and compare. A signature tells you who said it; recomputing tells you whether it's right. One thing I'd add to your checkpoint: record the verifier's version next to the receipt, so a later process can tell a PASS from an old verifier apart from a PASS from the current one.

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

That’s a useful addition, Arhan. I’d bind the receipt to both the verifier version and the expected-input contract, alongside the artifact hash. Otherwise a later process can reproduce yesterday’s PASS without noticing that the check itself has changed.

I’d preserve the old receipt and append a fresh result under the new verifier, rather than rewrite history. Recomputing makes the result inspectable; a separate reference check is still valuable when the verifier itself could contain the bug. Your version field makes that distinction much easier to audit.

Collapse
 
arhancanli profile image
Arhan Canli •

Agreed on append-only. Rewriting an old receipt under a new verifier deletes the one record that shows what changed. Your last point is the one we learned the hard way. We tested our own significance statistic on synthetic searches where the true answer is known, and with autocorrelated returns it called noise skilled 20% of the time, four times the 5% it promises. Recomputing would have reproduced that wrong answer perfectly on every run; only a reference check with a known answer caught it. The version that corrects the standard error came back to 5.9%, and we published the failure next to the fix.

Thread Thread
 
bryanw profile image
Bryan Williams CivicDataForge •

That makes the distinction concrete: reproducing a calculation and checking whether it behaves as intended are different tests. Publishing the failure beside the fix makes that boundary visible.

I’d keep the correlated-noise case as a regression test, then use fresh simulation seeds for subsequent checks so the correction isn’t judged only against the examples that exposed it. For readers, the useful artifact would include the number of trials and how the test cases were generated—not just the before/after percentages. Is there a link to the published comparison?

Collapse
 
innokentyb profile image
Kent Bodrov •

That reference check is the crucial third layer.

A signed receipt answers who produced the result. Recalculation answers whether the same inputs still produce it. A known-answer oracle tests whether the verifier itself is measuring the right property.

I would keep all three records append-only: artifact hash, verifier version, and oracle result. Otherwise a corrected verifier can make today’s evidence look clean while erasing the fact that yesterday’s PASS was produced by a broken check.

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

The next step I’d add is explicit invalidation. Keep yesterday’s PASS intact as a historical observation, then append a record identifying which verifier defect undermines which conclusion, and link it to the replacement check. Otherwise append-only storage can still leave a consumer trusting the old PASS.

That gives you a targeted recheck too: find the artifacts whose checks depended on the affected verifier version and property. A bug in the totals check doesn’t automatically invalidate every unrelated check that version performed.

And the oracle needs its own provenance and version. A known-answer test copied from the same mistaken assumption can agree perfectly with a broken verifier. The history should let a reader distinguish ‘this passed then’ from ‘we still have valid grounds to rely on it now.’

Collapse
 
innokentyb profile image
Kent Bodrov •

Yes, explicit invalidation keeps the history honest without rewriting it. The old PASS can remain as a fact about what the verifier concluded at that time, while a separate edge says which conclusion is no longer safe to rely on and why. A materialized “currently valid evidence” view could then drive targeted rechecks without rerunning everything.

Thread Thread
 
bryanw profile image
Bryan Williams CivicDataForge •

The materialized view gives us another useful test: pause its updates, append an invalidation, then ask whether the old PASS is still usable. The consumer should expose how far through the ledger it has processed, rather than silently present an older view as current. How would you handle a consumer that’s behind—wait for it to catch up, or check the relevant records directly?

Collapse
 
hannune profile image
Tae Kim •

We ran into this exact thing last year with a report generation step. The write receipt came back clean, the agent said done, and the numbers were off because someone had swapped a price field in the schema the day before. We started requiring a re-read of the actual output fields after that. The side benefit I didn't expect: it caught a different schema mismatch three weeks later that the write receipt would have never flagged.

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

That price-field swap is a great example of a successful write carrying the wrong meaning. Re-reading the delivered fields gives the check a chance to catch it; comparing them against an independently defined expectation is what keeps the verifier from repeating the same mapping mistake.

I’d turn that incident into a regression case with deliberately different values in the two price fields, so a swap can’t accidentally pass. Did your check compare against the original source record, or against a separately maintained report contract?

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The distinction between “the tool succeeded” and “the requested outcome is true” is one of the most important boundaries in agent workflows.

I’d push the receipt idea one step further: a verification receipt should be treated as a claim with scope and freshness, not as a permanent success token. It should identify the artifact, revision/version, bytes or state that were checked, the acceptance criteria, and ideally the policy/version of the verifier that produced the result.

That becomes especially important when the destination is mutable. A green verification for revision 3 shouldn't authorize delivery if another process can replace the artifact before the consequential handoff. Re-reading at the handoff or using an immutable/versioned reference closes that race.

The recovery point is equally important. A checkpoint should preserve what remains unproven, not merely what the previous process believed. Otherwise restart logic can accidentally convert an old observation into fresh authority.

“Done” is therefore less of a model output and more of a verifiable state transition: the system should only be allowed to claim it when the evidence actually covers the user's acceptance condition.

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

One detail I'd sharpen: re-reading at handoff narrows the race, but doesn't close it if the artifact can change again between the read and the actual use. The handoff needs to consume the exact verified bytes/version, or the consequential operation needs an atomic version precondition.

A useful regression test would be: pause immediately after verifying revision 3, replace the artifact at the mutable path with revision 4, then resume delivery. It should either deliver the verified revision 3, if that's what the request permits, or stop and verify the replacement. Silently delivering revision 4 under revision 3's receipt is the failure.

That also makes recovery easier to reason about. The old receipt remains valid evidence about revision 3; it hasn't become evidence about whatever happens to occupy the path now. Freshness isn't just a timestamp—it's whether the evidence still covers the exact thing being acted on.

Collapse
 
nark3d profile image
Adam Lewis •

What does the check run against, the file as written or the intent behind the call? Because a file that exists and parses can still be the wrong artefact entirely. I can't tell whether the verification is structural or semantic. If it's structural you catch truncation and nothing else. If it's semantic, who wrote the assertion and how do they keep it from going stale the moment the agent changes its approach?

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

The check reads the delivered file, but its expectations come from separately supplied task input—not the agent’s explanation of what it intended.

In this example, it checks report identity, revision, exact item quantities, and arithmetic, alongside structural validation. Eight bolts and eight washers still total sixteen, but fail against the expected seven and nine. That’s a bounded check of domain correctness, not general understanding of intent.

Your question about stale assertions is the important part. I’d separate changing the implementation from changing the requirement. An agent choosing a different way to produce the report shouldn’t change what the report must contain. If the required meaning or format changes, the acceptance contract needs an explicit revision and the affected assertions need review—not an automatic rewrite to match whatever the agent produced.

The tutorial supplies that contract and verifier explicitly. In a real workflow, responsibility for the expectations belongs with whoever owns the requirements; the producing agent shouldn’t be their sole author and approver.

And there’s a remaining limit: a perfectly enforced but mistaken contract can still approve the wrong thing. These checks make particular claims testable; they don’t eliminate the need to validate the requirements themselves.

Collapse
 
jkming profile image
jkming •

The "different rows with the same total" case is the one that bites in practice. I've watched a sum-only check pass after an agent silently merged two line items. The same layering shows up with coding agents: "edit applied" is a tool result, not evidence the change works, so I re-run the build and tests myself before believing any "done". Question: how do you handle receipts for destinations you can't hash, like a row updated in Postgres? Snapshot the query result as stand-in bytes?

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

Yes—a snapshot of the relevant query result can be the bytes you preserve. But I'd label it “these fields were observed for this row/version,” not “the database is still correct.” The hash identifies the snapshot; the checks establish whether it met the task.

For a Postgres update, I'd separate three pieces:

  1. Identity and expectation: the database/environment, tenant, primary key, expected application revision, and the exact field values or invariants supplied by the task—not an expectation regenerated from the row being checked.
  2. The write: a conditional update such as WHERE tenant_id = … AND id = … AND revision = expected_revision, incrementing that application revision and using RETURNING to inspect the changed fields. Require exactly one row; zero is a failed precondition, not success. This assumes every relevant writer participates in the revision scheme.
  3. Commit and readback: returned values inside an open transaction aren't proof of commit. After a confirmed commit, a fresh read from the primary can verify the expected revision and values. If another writer has already changed them, record that mismatch rather than certifying the old snapshot as current.

For the receipt, I'd preserve only the necessary fields in a deterministic representation, plus the identity, revision, check results, and verifier version. Sensitive row contents don't belong in a public log.

The remaining boundary matters: a readback doesn't lock the future. If another action depends on that state, enforce its preconditions atomically when that action happens. And if the connection drops during commit, reconcile the operation before retrying—it may already have committed. That's the database counterpart of “saved” versus “verified.”

Collapse
 
henry786 profile image
Henry •

Spot on! At The Printing World, we see this all the time with automated box die-lines. A tool saying "file saved" doesn't mean the box dimensions actually match the customer's specs. double-checking the output is non-negotiable!

Collapse
 
bryanw profile image
Bryan Williams CivicDataForge •

That’s a useful physical-world example: the mistake survives the export and becomes something you actually manufacture. I’d make the acceptance check explicit about units and which dimensions the customer means, not just whether the file opens. Where do you catch mismatches in your workflow—when checking the exported file, in a physical proof, or both?

Collapse
 
paul-s profile image
Paul-S •

A successful tool call should never be the agent’s definition of done. Re-reading the saved artifact against trusted input turns completion from a claim into evidence.

Some comments have been hidden by the post's author - find out more