DEV Community

Cover image for Your agent says "done" — and nothing ran
sunnydachs
sunnydachs

Posted on

Your agent says "done" — and nothing ran

Have you ever shipped something an agent said it had finished, and only found out weeks later that nothing had actually run? I have. The failures that cost me time were never the crashes: a crash is a gift, because CI sees it. The expensive ones exit 0, print nothing unusual, and quietly skipped the work.

So instead of trusting exit codes I went looking for them in recorded traffic. Everything below is measured: one small agent task — fetch headlines, write a draft, verify it through a tool — running in three frameworks behind a single recording proxy that keeps every attempt. Aggregating that recording, the failures all wore one of three shapes:

  • the verification never ran
  • an invented argument got accepted
  • the answer was empty

Code, traces and the analysis scripts are on GitHub:

GitHub logo sunnydachs / agent-framework-showdown

Same tech-news-digest agent in Strands/LangGraph/CrewAI with recorded-LLM observability

agent-framework-showdown

The same digest agent built three times — in Strands, LangGraph, and CrewAI — with every LLM call recorded, so you can compare how they actually behave.

English | 日本語

Same task. Same model. Same tools. Three frameworks. Twenty-seven runs. All LLM traffic captured through a local recorder, so "which framework behaves differently" is an answer backed by trace files instead of vibes.

The task

A tech-news digest agent:

  1. collect 5 headlines via a fetch_headlines tool
  2. write a ~100-word digest
  3. verify the word count via a word_count tool, revising if out of band

All three frameworks hit the same model behind a local recorder proxy, so the logs are directly comparable.

How to run

# one venv per framework (Python 3.12 - CrewAI requires <3.14)
uv venv .venv-strands   --python 3.12 && uv pip install --python .venv-strands/bin/python   "strands-agents[litellm]"
uv venv .venv-langgraph --python 3.12 && uv pip install --python
…
Enter fullscreen mode Exit fullscreen mode

Shape 1: the verification never ran, and the run still said success

I removed an argument from the verification tool so that it returned a static count instead of checking the draft. Nothing errored. The tool ran, answered, and the run reported success with a draft that was never verified.

  • Strands: all 3 runs exited 0. Verifications actually executed: 0/3
  • CrewAI: all 3 runs exited 0. Verifications actually executed: 0/3
  • LangGraph: all 3 runs stopped with a hard error (its tool calls live in code)

LangGraph is the only one that cannot slip through silently: its structure cannot step over a broken tool.

There is a nastier variant, where the tool did answer with an error and the process still exited 0.

Strands, argument-type level:
  tool error responses = 2, 1, 5 (three runs)
  process exit code    = 0, 0, 0
Enter fullscreen mode Exit fullscreen mode

Those are not schema rejections. The call was well-formed; the tool answered with a processing error — "document 88 not found" — and every run still finished green.

Shape 2: the model invented a required argument, and the tool accepted it

I added a new required argument to the verification tool and never mentioned it in the prompt.

  • Strands: all 3 runs had a plausible invented value accepted (0 errors)
  • CrewAI: all 3 runs accepted (0 errors) — and all three invented the same sentence
  • LangGraph: never reached the tool

The tool's only validation was "non-empty string".

Strands, 3 runs: "Check word count of AI agents news digest"
                 "AI agents digest word count check"
                 "Word count check for AI agents digest"
CrewAI,  3 runs: "Draft digest summarizing the five headlines into a flowing narrative"
Enter fullscreen mode Exit fullscreen mode

Six fabricated values went straight through, and the recording is the only thing that shows where they came from.

That contrast is not purely a framework property, though: the two frameworks advertised the new parameter differently on the wire, and they did not sample the same way (see limitations).

Shape 3: the run said "done", and the answer was empty

This one only appeared in the model-driven frameworks.

Finish reason was stop, the visible answer was empty, and the draft was sitting inside the last tool-call argument.

  • Strands: 2 runs (the content was in the tool argument: 100 and 106 words)
  • LangGraph / CrewAI: 0 runs

A framework whose output node is the deliverable cannot produce this shape.

Where the evidence lives

  what the process reports           what actually happened
  ┌────────────────┐                ┌──────────────────────────────┐
  │ exit 0         │                │ tool call #1  → error        │
  │ "done" banner  │       ≠        │ tool call #2  → static stub  │
  └────────────────┘                │ verify        → never ran    │
                                    └──────────────────────────────┘
        evidence lives here  →       at the tool boundary (wire)
Enter fullscreen mode Exit fullscreen mode

The error responses existed on the wire, and the tool layer even counted them. The process just never looked.

The evidence of a failure survives only outside the framework, at the tool boundary.

Invisible failures still cost money

A failure that gets reported as success keeps racking up wasted runs while nobody notices.

In a separate experiment (approval gates, crash recovery, duplicate execution, same proxy), it showed up in three ways:

  • On crash recovery, the framework without state recording re-ran the whole task (2 LLM calls). The one with a checkpointer resumed with 0.
  • An irreversible operation (in that experiment, publishing) executed twice under two different call IDs, and both returned success.
  • Even the "did nothing" failure burned LLM calls: CrewAI 6, 5, 4; Strands 4, 5, 4.

Duplicates do not show up in the framework's own trace, either. Every duplicated call carried a different call ID, so on the wire they look like two unrelated requests.

And if the retry only rewords the same intent, an idempotency ledger keyed on the argument bytes cannot tell it apart:

  call ID: X ──> ledger ──> publish executes ──> success
  call ID: Y ──> ledger ──> publish executes ──> success
                   ↑ keys differ, so it doesn't look like a duplicate
  result: the side effect runs twice (both return success)
Enter fullscreen mode Exit fullscreen mode

How to prevent it: the smallest check per situation

  • Hobby pipeline → assert the deliverable is not empty
  • In CI → assert the expected tool was actually called (never 0 times)
  • Writing data → validate values (type, enum, known set)
  • Handling approvals or billing → record approval and execution as separate states
  • Need an audit → record at the boundary (the framework's own trace cannot reconstruct it)

The hobby case is the one I keep running into. Nothing errors, the log is calm, the artifact is empty or unverified — and you find out weeks later. If three lines of assertion catch that, I see no reason to leave them out.

The CI case is the same failure one layer up. If the only contract is the exit code, a skipped verification is invisible by construction.

"Did the verification tool run?" is a different question from "did the process succeed?"

For tools that write data, the invented argument is the dangerous one, because the value is well-formed. That no human ever wrote it is something only the wire recording can show.

And on approvals: approved once is not executed once. In these runs the same logical operation was settled twice under two different call IDs, and both returned success.

The two checks I actually use

Before anything downstream consumes the output, I assert exactly two things: the deliverable is not empty, and the tool I expected to run did actually run. Boring, a few lines each, and they would have caught every failure in this post.

Keys, ledgers and approval binding are for the moment the action touches money or another person. My honest read: most hobby pipelines do not need a ledger. They need the non-empty check and an exit signal they can trust.

Honest limitations

  • 3 runs per cell. A trend, not a statistical claim.
  • One model across all runs; a different model may reword or fabricate differently.
  • The irreversible operation is simulated; whether a gate held is read from the recorded traffic.
  • Fine instruction-following differences are hard to separate from the schema change itself.
  • One framework sampled at temperature 0, and the parameter descriptions were not identical between the two — so the "varied vs repeated" contrast has at least two plausible causes.

This is a personal OSS project with no warranty. Use it at your own risk, and issues are welcome.

Top comments (9)

Collapse
 
anp2network profile image
ANP2 Network •

Both of your assertions live inside the process that did the work. That is also where they stop: once the consumer only holds the sender's account of a run, "the deliverable is non-empty" and "the expected tool was called" can both be satisfied by the account alone. I keep measuring a public append-only event log where independent agents append signed rows and an aggregator reads them later, and two of the shapes you drew show up there with no call IDs anywhere in the system.

On duplicates: of 19,926 rows, 600 sat in groups where author, event kind and body matched exactly, and every one of those rows carried a distinct id. So id-based dedup only removes what the read cursor redelivers at second boundaries. A sender re-appending the same claim goes straight through uncounted, which is your differing-call-ID diagram arriving by a completely unrelated route.

Timing is worse. There is a runtime_ms field for how long the work took, and the read path never reads it. In 11 rows it was stamped earlier than the acceptance row that supposedly started that work. Schema-valid, so stored, and nothing objected. A signature settles who said a thing, never that anything ran. Your proxy is a clock the agent does not own, and that is what makes those two assertions carry weight. Past the point where the evidence becomes another sender's self-report, what would you take as the minimum assertion?

Collapse
 
sunnydachs profile image
sunnydachs • • Edited

Thank you for your comment. Your log and my proxy agree on the shape: id-based dedup only removes what the read cursor redelivers. In my traces the same shows up as hash:reworded (2 publishes, 0 deduped) vs position:same (1 deduped), so the failure is not about call IDs specifically — it is about any key the sender can mint.

On your question — what I would take as the minimum assertion once the evidence is another sender's self-report: nothing inside the report itself. The assertion has to bind to something the reporter cannot write. In my case that is the proxy's own timestamp and the tool's own error response (rec_proxy stamps ts: time.time() outside the agent process). In your case the equivalent would be one consumer-side check: that the acceptance row for a run is stamped no earlier than the runtime_ms row it consumes. That is the same assertion at a different layer — "the evidence must be written by a clock the reporter does not own" — and it is what I would take as minimum.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

"Did the verification tool run?" is a different question from "did the process succeed?" is the line I'd put above every agent pipeline. Your Shape 1 is the scary one because the stub still answered, so from the outside the run looked complete. I ended up with the same rule in my own verify step: it reports what it actually executed, and a run where nothing was selected gets its own verdict instead of a green one. Your "expected tool was called, never 0 times" assertion is the CI version of that. On the duplicate publish: could either framework carry a stable operation ID through the retry, or does that have to live entirely outside the agent?

Collapse
 
sunnydachs profile image
sunnydachs •

Thanks — that question is basically the result. The key that deduped came from the caller, not the framework: in all 18 runs each re-emit of the publish call arrived with a new tool_call id, so the wire id can't be the key. A key fixed per logical operation deduped a same-args retry; a content-hash key missed once the retry reworded the argument and both calls executed. So yes, the operation id has to live outside the agent — a framework can carry checkpoint identity across a crash restart, but that's thread state, not an id on an LLM-initiated retry.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

That's a clean result: the wire id rotates on every re-emit, a content hash breaks as soon as the retry rewords the argument, so the only key that holds is one minted by the caller for the logical operation. It also says where the key belongs: in the plan or task record before the first call, not derived from the call itself. Did you try passing that key into the agent's context so it can quote it back, or does it stay entirely outside the model?

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

The static-count example exposes a gap in the final two assertions: a non-empty draft plus a recorded verifier call can still pass when the verifier ignores that draft. I'd bind the verification result to the exact artifact version (or content hash) and have the consumer check that binding before accepting it. The hash establishes which bytes were checked; a contract test still needs to establish that the verifier actually checks them.

A useful next fault injection would be: verify draft A, replace it with draft B, then attempt delivery with A's verification result. That tests the completion contract across the whole handoff, including the case where every individual tool call succeeds.

Collapse
 
sunnydachs profile image
sunnydachs •

Thanks — and the gap you're pointing at is real. The fault injections I measured in the boundary experiment all test one layer: whether the correct verified field survives when a plausible fake sits right next to it in the same record. A swap across the handoff — verify draft A, deliver draft B with A's verification result — is not in the grid, so my two assertions are exactly as vulnerable as you say: non-empty draft plus a recorded verifier call would pass, whether or not the verifier read that draft. Your two-step test is the right sequence. The closest thing my harness could check today is the binding you described — whether the verifier call's arguments are the same bytes that produced the final answer — which is already in the recorded traffic. I'll take the swap test as a follow-up experiment, and it's a good candidate for the next article.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.