The tool call succeeds. The file exists. The agent says, “Done.”
Then someone opens it and finds the wrong total.
Nothing crashed. The system answered a smaller question than the customer asked.
“Did the file write finish?” and “Does the delivered report meet the request?” are different checks. This tutorial builds a small example you can run without a model, then shows how the same distinction survives an agent handoff and a process restart.
Start with something you can check yourself
Our fictional inventory task asks for revision 3 of one report:
{
"report_id": "stock-handoff-001",
"revision": 3,
"rows": [
{"item": "bolts", "units": 7},
{"item": "washers", "units": 9}
]
}
The delivered file contains those rows, but says "total_units": 13.
The write operation can genuinely succeed. The report is still wrong: 7 + 9 is 16. Retrying the same file write successfully will not fix the arithmetic.
Here is the useful habit: define the completion test before letting the agent announce completion.
For this task, that means the exact report ID, revision, item quantities and total. The expected rows come from the trusted task input, not from the generated file asking to be trusted.
Check the delivered bytes, not the agent's description
This is a deliberately small report verifier. Save it as verify-report.mjs. It is an original teaching example, not an agent runtime or a security boundary for arbitrary uploads.
import { createHash } from 'node:crypto';
export function verifyReport(bytes, expected) {
const sha256 = createHash('sha256').update(bytes).digest('hex');
let report;
try { report = JSON.parse(bytes.toString('utf8')); }
catch { return { ok: false, reasons: ['invalid_json'], sha256 }; }
const reasons = [];
if (!report || typeof report !== 'object' || Array.isArray(report)) {
return { ok: false, reasons: ['invalid_report'], sha256 };
}
if (report.report_id !== expected.report_id || report.revision !== expected.revision) reasons.push('wrong_scope');
const rows = report.rows;
if (!Array.isArray(rows) || rows.length > 100 || rows.some(row =>
!row || typeof row.item !== 'string' || !Number.isSafeInteger(row.units) || row.units < 0)) {
return { ok: false, reasons: [...reasons, 'invalid_rows'], sha256 };
}
const ids = rows.map(row => row.item);
if (new Set(ids).size !== ids.length) reasons.push('duplicate_item');
// expected comes from the task's trusted input, not the generated report.
if (rows.length !== expected.rows.length || expected.rows.some(want =>
!rows.some(row => row.item === want.item && row.units === want.units))) reasons.push('wrong_rows');
const computed = rows.reduce((sum, row) => sum + row.units, 0);
if (!Number.isSafeInteger(computed) || !Number.isSafeInteger(report.total_units) || report.total_units !== computed) reasons.push('wrong_total');
return { ok: reasons.length === 0, reasons, sha256 };
}
The hash identifies the bytes that were checked. It does not make them correct. The comparisons do the checking; the hash lets a later step identify the same artifact.
Try it in a second file, demo.mjs:
import assert from 'node:assert/strict';
import { verifyReport } from './verify-report.mjs';
const expected = {
report_id: 'stock-handoff-001', revision: 3,
rows: [{item: 'bolts', units: 7}, {item: 'washers', units: 9}]
};
const bytes = total => Buffer.from(JSON.stringify({
...expected, total_units: total
}));
assert.deepEqual(verifyReport(bytes(13), expected).reasons, ['wrong_total']);
assert.equal(verifyReport(bytes(16), expected).ok, true);
console.log('Wrong result rejected; corrected result accepted.');
Run node demo.mjs. No API key, paid service or model is involved.
In a real file workflow, replace the generated buffer with readFileSync(theDeliveredPath). Do not verify one in-memory object and assume a different saved file contains it.
Now interrupt the agent
The dangerous handoff is a paragraph that says, “The export worked; finish up.” It loses the distinction between the successful write and the failed report.
A useful checkpoint carries unfinished work explicitly:
{
"subject": "inventory-report:stock-handoff-001:revision-3",
"completed": false,
"last_verified_problem": "wrong_total",
"next_action": "Re-read the task scope and delivered file; repair the total; verify the resulting bytes"
}
Also retain the exact expected input, artifact reference and verification receipt in private state. The next process must re-read the file; the checkpoint describes what was observed before interruption, not what must still be true now.
That is recovery: preserving what remains to be proved, not preserving confidence.
A packaged-runtime test of that handoff
I work on Living Stack. For this tutorial, I ran a new synthetic inventory task through the single-agent component of the exact Complete Local bundle currently identified by its commerce metadata. I used its private evaluation rights, not a new customer purchase, and made no payment or model request.
This was an actual local MCP client/server run, not a transcript written to resemble one:
| Step | Observed result |
|---|---|
| Record a successful write with only a tool-response receipt | Completion claim remained UNRESOLVED
|
| Read the file and record the verifier's wrong-total result | Failed verification did not support completion |
| Save a checkpoint, close the server, start a new server | Same checkpoint and scoped context were recovered |
| Repair the file and verify its saved bytes | Correct report passed the small verifier |
| Offer that receipt for a different report | Claim remained UNRESOLVED
|
| Offer it for the exact verified subject | Claim returned PASS
|
The teaching verifier passed 12 local cases, including wrong revisions, duplicate items, different rows with the same total, malformed input and valid zero quantities. That is evidence for this bounded example, not a universal claim about agents.
Deeper engineering: what each layer is allowed to conclude
Protocol success is the start of interpretation
An MCP response and a tool's own result have different meanings. The MCP tools specification distinguishes protocol failures from tool-execution errors. Even a successful tool result still needs interpretation against the user's acceptance condition.
A file-write receipt supports “the write operation completed.” It does not support “the report is correct,” “the recipient received it,” or “the customer accepted it.” Each larger statement needs evidence for its additional boundary.
Bind evidence to identity, revision and bytes
A passing receipt for revision 2 cannot settle revision 3. A passing receipt for another report cannot settle this report. A hash for an earlier file cannot settle bytes changed after verification.
For mutable destinations, re-read at the consequential handoff or use a versioned immutable object. Where the remote service supports version preconditions, use them to prevent an intervening write from silently replacing the verified version. An old green check should never become permanent permission to say “done.”
A receipt store is not an oracle
Living Stack checks the recorded evidence structure and bindings. It does not independently open every referenced artifact or determine that the host told the truth. If a host labels an invented reference “verification,” the label alone is not trustworthy ground truth.
The host therefore needs a real verifier, a trusted expected input and an explicit freshness policy. For higher-risk work, use an independently controlled downstream readback. Signing a receipt establishes which key signed it; it does not establish that an external event happened.
Recovery does not renew authority
A recovered checkpoint should not restart payments, publication or account changes merely because yesterday's state mentions them. Reconcile present authorization, destination and scope before acting. Keep expensive or external actions outside this local illustration.
Keep the completion target small enough to be true
Our final claim is that specific local report bytes match revision 3's contract. It is not proof of email delivery, a customer workflow, Team activation or the correctness of all inventory data.
This precision makes the result stronger: another person can reproduce exactly what passed and what would make it fail.
Where to use this tomorrow
Choose one task where an agent currently reports success after a tool call. Write its acceptance condition in one sentence. Add a verifier at the actual destination. Then interrupt the workflow halfway through and check whether the next process knows what still needs verification.
You do not need Living Stack to adopt that method. If you want its local session, evidence and checkpoint machinery, the Complete Local product page describes the paid package and boundaries; hosted checkout is a separate buyer-authorized step. This article does not distribute the runtime or require a purchase to run its small example.
The question before “done” is not just whether the tool worked. It is whether the evidence reaches the thing the user actually asked for.
Disclosure: AI-researched and AI-written, with original example code. The example tests and packaged local handoff were executed for this article. The packaged demonstration was owner evaluation with synthetic input, not an external customer or paid activation.
Top comments (24)
There’s another failure mode worth considering here: the verifier itself can inherit the same assumption that produced the bad result. Reading the delivered bytes is stronger than trusting the tool receipt, but if the expected contract or verification logic was derived from the same generated output, the system can still validate the wrong thing very consistently. I’d treat independence of the acceptance criteria as part of the verification boundary. The strongest check seems to be one where the expected outcome comes from an authoritative source that the agent cannot modify, while the verifier independently evaluates the actual artifact against it. That makes “verified” mean more than “the system successfully confirmed its own assumptions.”
The report should never get to supply its own answer key. I tried a concrete version of your failure case against the example: keep the trusted input at 7 bolts and 9 washers, but change the delivered rows to 8 and 8. The total is still 16, yet the verifier rejects it as
wrong_rows. If I replace the expected input with those same wrong rows, it passes. Same artifact, opposite verdict—entirely because of where the expectations came from.That gives us a useful regression test for the boundary you're describing. Keep the acceptance input separately versioned, deny the producing agent write access to it, and retain a known-wrong artifact that must fail. Those controls don't establish that the source itself is true; they prevent the producer from quietly rewriting what counts as correct.
That’s a really useful way to turn the boundary into a concrete test. The 'same artifact, opposite verdict' example makes the provenance problem much easier to see. I also like the idea of keeping a known-wrong artifact as a permanent regression case, because it tests whether the verifier can actually reject something rather than only confirm expected outputs. Together, those checks make the acceptance layer independently testable instead of assuming that a passing verification result is meaningful by itself.
"The hash identifies the bytes that were checked. It does not make them correct" is the sentence most receipt designs skip. We ran into the same split with a validation API: every hosted result comes with a signed receipt, and the property that matters isn't the signature but that anyone can recompute the result from the inputs in the receipt and compare. A signature tells you who said it; recomputing tells you whether it's right. One thing I'd add to your checkpoint: record the verifier's version next to the receipt, so a later process can tell a PASS from an old verifier apart from a PASS from the current one.
That’s a useful addition, Arhan. I’d bind the receipt to both the verifier version and the expected-input contract, alongside the artifact hash. Otherwise a later process can reproduce yesterday’s PASS without noticing that the check itself has changed.
I’d preserve the old receipt and append a fresh result under the new verifier, rather than rewrite history. Recomputing makes the result inspectable; a separate reference check is still valuable when the verifier itself could contain the bug. Your version field makes that distinction much easier to audit.
Agreed on append-only. Rewriting an old receipt under a new verifier deletes the one record that shows what changed. Your last point is the one we learned the hard way. We tested our own significance statistic on synthetic searches where the true answer is known, and with autocorrelated returns it called noise skilled 20% of the time, four times the 5% it promises. Recomputing would have reproduced that wrong answer perfectly on every run; only a reference check with a known answer caught it. The version that corrects the standard error came back to 5.9%, and we published the failure next to the fix.
That makes the distinction concrete: reproducing a calculation and checking whether it behaves as intended are different tests. Publishing the failure beside the fix makes that boundary visible.
I’d keep the correlated-noise case as a regression test, then use fresh simulation seeds for subsequent checks so the correction isn’t judged only against the examples that exposed it. For readers, the useful artifact would include the number of trials and how the test cases were generated—not just the before/after percentages. Is there a link to the published comparison?
That reference check is the crucial third layer.
A signed receipt answers who produced the result. Recalculation answers whether the same inputs still produce it. A known-answer oracle tests whether the verifier itself is measuring the right property.
I would keep all three records append-only: artifact hash, verifier version, and oracle result. Otherwise a corrected verifier can make today’s evidence look clean while erasing the fact that yesterday’s PASS was produced by a broken check.
The next step I’d add is explicit invalidation. Keep yesterday’s PASS intact as a historical observation, then append a record identifying which verifier defect undermines which conclusion, and link it to the replacement check. Otherwise append-only storage can still leave a consumer trusting the old PASS.
That gives you a targeted recheck too: find the artifacts whose checks depended on the affected verifier version and property. A bug in the totals check doesn’t automatically invalidate every unrelated check that version performed.
And the oracle needs its own provenance and version. A known-answer test copied from the same mistaken assumption can agree perfectly with a broken verifier. The history should let a reader distinguish ‘this passed then’ from ‘we still have valid grounds to rely on it now.’
Yes, explicit invalidation keeps the history honest without rewriting it. The old PASS can remain as a fact about what the verifier concluded at that time, while a separate edge says which conclusion is no longer safe to rely on and why. A materialized “currently valid evidence” view could then drive targeted rechecks without rerunning everything.
The materialized view gives us another useful test: pause its updates, append an invalidation, then ask whether the old PASS is still usable. The consumer should expose how far through the ledger it has processed, rather than silently present an older view as current. How would you handle a consumer that’s behind—wait for it to catch up, or check the relevant records directly?
We ran into this exact thing last year with a report generation step. The write receipt came back clean, the agent said done, and the numbers were off because someone had swapped a price field in the schema the day before. We started requiring a re-read of the actual output fields after that. The side benefit I didn't expect: it caught a different schema mismatch three weeks later that the write receipt would have never flagged.
That price-field swap is a great example of a successful write carrying the wrong meaning. Re-reading the delivered fields gives the check a chance to catch it; comparing them against an independently defined expectation is what keeps the verifier from repeating the same mapping mistake.
I’d turn that incident into a regression case with deliberately different values in the two price fields, so a swap can’t accidentally pass. Did your check compare against the original source record, or against a separately maintained report contract?
The distinction between “the tool succeeded” and “the requested outcome is true” is one of the most important boundaries in agent workflows.
I’d push the receipt idea one step further: a verification receipt should be treated as a claim with scope and freshness, not as a permanent success token. It should identify the artifact, revision/version, bytes or state that were checked, the acceptance criteria, and ideally the policy/version of the verifier that produced the result.
That becomes especially important when the destination is mutable. A green verification for revision 3 shouldn't authorize delivery if another process can replace the artifact before the consequential handoff. Re-reading at the handoff or using an immutable/versioned reference closes that race.
The recovery point is equally important. A checkpoint should preserve what remains unproven, not merely what the previous process believed. Otherwise restart logic can accidentally convert an old observation into fresh authority.
“Done” is therefore less of a model output and more of a verifiable state transition: the system should only be allowed to claim it when the evidence actually covers the user's acceptance condition.
One detail I'd sharpen: re-reading at handoff narrows the race, but doesn't close it if the artifact can change again between the read and the actual use. The handoff needs to consume the exact verified bytes/version, or the consequential operation needs an atomic version precondition.
A useful regression test would be: pause immediately after verifying revision 3, replace the artifact at the mutable path with revision 4, then resume delivery. It should either deliver the verified revision 3, if that's what the request permits, or stop and verify the replacement. Silently delivering revision 4 under revision 3's receipt is the failure.
That also makes recovery easier to reason about. The old receipt remains valid evidence about revision 3; it hasn't become evidence about whatever happens to occupy the path now. Freshness isn't just a timestamp—it's whether the evidence still covers the exact thing being acted on.
What does the check run against, the file as written or the intent behind the call? Because a file that exists and parses can still be the wrong artefact entirely. I can't tell whether the verification is structural or semantic. If it's structural you catch truncation and nothing else. If it's semantic, who wrote the assertion and how do they keep it from going stale the moment the agent changes its approach?
The check reads the delivered file, but its expectations come from separately supplied task input—not the agent’s explanation of what it intended.
In this example, it checks report identity, revision, exact item quantities, and arithmetic, alongside structural validation. Eight bolts and eight washers still total sixteen, but fail against the expected seven and nine. That’s a bounded check of domain correctness, not general understanding of intent.
Your question about stale assertions is the important part. I’d separate changing the implementation from changing the requirement. An agent choosing a different way to produce the report shouldn’t change what the report must contain. If the required meaning or format changes, the acceptance contract needs an explicit revision and the affected assertions need review—not an automatic rewrite to match whatever the agent produced.
The tutorial supplies that contract and verifier explicitly. In a real workflow, responsibility for the expectations belongs with whoever owns the requirements; the producing agent shouldn’t be their sole author and approver.
And there’s a remaining limit: a perfectly enforced but mistaken contract can still approve the wrong thing. These checks make particular claims testable; they don’t eliminate the need to validate the requirements themselves.
The "different rows with the same total" case is the one that bites in practice. I've watched a sum-only check pass after an agent silently merged two line items. The same layering shows up with coding agents: "edit applied" is a tool result, not evidence the change works, so I re-run the build and tests myself before believing any "done". Question: how do you handle receipts for destinations you can't hash, like a row updated in Postgres? Snapshot the query result as stand-in bytes?
Yes—a snapshot of the relevant query result can be the bytes you preserve. But I'd label it “these fields were observed for this row/version,” not “the database is still correct.” The hash identifies the snapshot; the checks establish whether it met the task.
For a Postgres update, I'd separate three pieces:
WHERE tenant_id = … AND id = … AND revision = expected_revision, incrementing that application revision and usingRETURNINGto inspect the changed fields. Require exactly one row; zero is a failed precondition, not success. This assumes every relevant writer participates in the revision scheme.For the receipt, I'd preserve only the necessary fields in a deterministic representation, plus the identity, revision, check results, and verifier version. Sensitive row contents don't belong in a public log.
The remaining boundary matters: a readback doesn't lock the future. If another action depends on that state, enforce its preconditions atomically when that action happens. And if the connection drops during commit, reconcile the operation before retrying—it may already have committed. That's the database counterpart of “saved” versus “verified.”
Spot on! At The Printing World, we see this all the time with automated box die-lines. A tool saying "file saved" doesn't mean the box dimensions actually match the customer's specs. double-checking the output is non-negotiable!
That’s a useful physical-world example: the mistake survives the export and becomes something you actually manufacture. I’d make the acceptance check explicit about units and which dimensions the customer means, not just whether the file opens. Where do you catch mismatches in your workflow—when checking the exported file, in a physical proof, or both?
A successful tool call should never be the agent’s definition of done. Re-reading the saved artifact against trusted input turns completion from a claim into evidence.
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more