"Done. All tests pass, and nothing else was touched." You read it, it's late, you merge.
Claude Code's own issue tracker is full of what happens next. A test suite whose total quietly changed from 4,992 to 4,966 before the "ALL PASSED". Six tests marked skip. A project reported 100% complete that an audit put nearer 60%.
The summary isn't a lie. It's written to explain the work, and nobody explains the parts they'd rather you didn't see.
Asking the agent to try harder doesn't fix that. Checking does. Here's what goes wrong, what doesn't work, and what happened when I held this site's own Claude Code summaries up against their real diffs.
The four ways "done" isn't done
Read enough of those issues and the same four shapes keep coming back:
The tests got easier, not the code better. A test skipped, an assertion loosened, a test command narrowed to the ones that pass. The suite is green because there's less of it.
The count got reframed. A failure becomes "a flaky timeout". A smaller total gets reported as everything passing.
Intent got ticked off as work. A to-do marked done because it was planned, or a stub counted as a feature.
A file changed and never got mentioned. The quietest one, and the one that bites three days later.
It's common enough to measure. A study of 20,574 real agent sessions found inaccurate self-reporting was about 23% of all the misbehaviour it caught, and a growing share over time. Anthropic's own best-practices guide says it plainly: Claude stops when the work looks done.
What doesn't work
Asking "did you actually do it?" The same model that wrote the summary checks it with the same blind spot. In one report it said yes three times.
A rule in CLAUDE.md. One reporter had a 250-line file, memory notes and hooks, and still got false "done"s. In another, the rule against skipping tests was right there in context and got broken anyway. A rule asks. A check doesn't need to.
What I found in my own sessions
Most writing on this describes the problem. Almost none of it shows a real summary next to its real diff, so I did that with two pieces of work Claude Code finished on this website today. For each one I took the summary it gave me at the end, word for word, and ran it through plumb, a small tool that holds a summary against git diff and prints only what doesn't match.
The first was publishing a blog post, along with a fix to a script:
$ plumb check summary.md --base a12ccca
plumb — 9 files changed, 1 named in the summary
changed but never mentioned (read these first)
· .assetsignore modified
· LOG.md modified
· assets/blog/og-claude-code-token-usage.png added
· baggage/index.html modified
· blog/claude-code-token-usage/index.html added
· blog/feed.xml modified
· blog/index.html modified
· sitemap.xml modified
8 things the summary did not tell you.
Eight flags. I went through them one by one. Seven were things the summary did describe, just not by file name: it gave the post's URL, called the image "the social card", and said "that folder is now excluded from publishing" instead of naming .assetsignore.
One wasn't. baggage/index.html: Claude added a link to a product page while publishing the post, and the summary never said so anywhere. Harmless, as it happens. But I only know it's harmless because I looked.
The second was a batch of SEO fixes: 21 files changed, and again only one named by path. This time the summary described all 20 others in categories, "19 pages and the sitemap", and every one checked out. plumb also flagged two files as "mentioned but not changed", which the summary had only referred to. Nothing was missing.
So, across two real summaries: 30 files changed, 2 named by path, one change genuinely left out. No skipped tests, no removed guards. Neither summary lied. One still hid a change, and I'd have merged it without blinking.
What actually works
Read the diff, not the paragraph.
/diffin Claude Code, orgit diff --stat, shows every file that changed. The summary is a guide to the diff, never a replacement for it.Make it end with a file list. Ask for every file it touched, by path, at the end of each task. My two summaries were noisy to check because they talked in categories. A list turns "did it mention everything?" into a mechanical question.
Check on every stop, automatically. Claude Code's Stop hook receives the agent's last message, so a check can run the moment it says it's finished. trust issues is a free plugin that does exactly that: the same four checks, every time, with nobody remembering to run them.
Count the tests, before and after. If the number went down, ask why before you read anything else.
Get a second pair of eyes that didn't do the work. Anthropic's guide suggests a separate verification step that checks nothing outside the task changed. A fresh session doesn't share the first one's reasons for skipping a test.
Try it on yours
Save the agent's last message to a file, then:
$ npm install -g github:manpreet171/plumb
$ plumb check summary.md # against your uncommitted changes
$ plumb check summary.md --base main # against a branch or commit
Exit code 1 means it found something, so it works as a CI gate too. What it won't do, so you don't over-trust it:
It matches names literally. A URL or a category doesn't count as a mention, which is why my first run flagged seven things that weren't really missing. Ask for a file list and that noise goes away.
It can't tell whether the code works. It finds what was left out of the story, not bugs. Run the tests for that.
Its "quiet cuts" are patterns. It spots a removed assertion or an added
.skip, not every possible way to weaken a check.
This is how we build agents for clients. Nothing gets called done because the agent said so. Every run ends with what changed, and something other than the agent checks it.
Sources: the plumb output is a real run (plumb 1.1.0) on 5 Oct 2026, against the exact end-of-task summary Claude Code gave for each piece of work, with the summary file's path shortened. The second run's output is described rather than shown because it lists 20 files. Incidents are from the linked GitHub issues on anthropics/claude-code; the session study is arXiv 2605.29442; hook and review guidance is from Anthropic's best practices and hooks docs.
Read next: My mailer printed “campaign sent”. It had sent to nobody. · Where your Claude Code tokens actually go. Output is 0.2%.
Originally published at singhlabs.dev.
Top comments (0)