Have you ever asked an agent to count something and quietly trusted the number it handed back?
I did, until I stopped reading the answer and started reading the tool's own response. In one run the model answered 321 where the id list sitting in that same response held 322 matching ids. Nothing else about the run looked wrong.
- the same question through three frameworks, behind one recording proxy, one model throughout
- everything below is measured from the recorded traffic, not from the frameworks' own reports
sunnydachs
/
agent-framework-showdown
Same tech-news-digest agent in Strands/LangGraph/CrewAI with recorded-LLM observability
agent-framework-showdown
The same digest agent built three times — in Strands, LangGraph, and CrewAI — with every LLM call recorded, so you can compare how they actually behave.
English | 日本語
Same task. Same model. Same tools. Three frameworks. Twenty-seven runs. All LLM traffic captured through a local recorder, so "which framework behaves differently" is an answer backed by trace files instead of vibes.
The task
A tech-news digest agent:
- collect 5 headlines via a
fetch_headlinestool - write a ~100-word digest
- verify the word count via a
word_counttool, revising if out of band
All three frameworks hit the same model behind a local recorder proxy, so the logs are directly comparable.
How to run
# one venv per framework (Python 3.12 - CrewAI requires <3.14)
uv venv .venv-strands --python 3.12 && uv pip install --python .venv-strands/bin/python "strands-agents[litellm]"
uv venv .venv-langgraph --python 3.12 && uv pip install --python…The only thing I changed: what the tool hands back
Two versions of one tool. The first returns the id list and leaves the counting to the model. The second precomputes the count and hands that back.
A: the tool returns the rows B: the tool returns the count
┌───────────────────────┐ ┌────────────────────────┐
│ [3, 7, 11, … 330 ids] │ │ {count: 322, min, max} │
└───────────┬───────────┘ └───────────┬────────────┘
▼ the model counts ▼ the model reads
answered 321 (true 322) answered 322 (correct)
The grid is 3 frameworks x 2 tool modes x 3 list sizes x 7 threshold phrasings x 3 seeds, plus five extra 330-row probes — 408 recorded runs in total, every one behind the same recording proxy.
- mode
ids: the tool is asked for the id list only, the model must count - mode
stats: the tool is asked forcount/min/max - sizes 11, 110 and 330 ids, with phrasings such as
id >= 9,id > 9,no less than 9
408 runs in total, 407 of them exiting 0. A mode label describes what the tool was asked for, not what the framework carried into the prompt — so the analyzer also records what the last model call actually held: the id list, the count, or neither. That measurement is the reason the two columns below are not the two tool designs.
At 330 rows the rate matched three ways and the reasons only two
| Framework | model counts (ids) |
tool counts (stats) |
tokens/question ids
|
tokens/question stats
|
|---|---|---|---|---|
| CrewAI | 92.3% (24/26, 2 wrong) | 100% (26/26) | 14,600 | 12,565 |
| LangGraph | 92.3% (24/26, 2 incomplete) | 100% (26/26) | 8,574 | 143 |
| Strands | 92.3% (24/26, 2 wrong) | 100% (26/26) | 11,528 | 1,418 |
92.3% is the same number three times, and it is the same failure twice.
- CrewAI and Strands each answered two runs one or two off.
- LangGraph's two missing runs answered nothing at all: both generate to the completion ceiling and stop.
Three different tool-calling architectures — one keeps the tool call inside code-managed state, the other two let the model drive it — landed on the same rate. What they did not do was fail the same way, and the difference matters more than the rate.
At 11 and 110 rows the answer is right almost everywhere. The single exception is CrewAI in stats mode at 110 rows, which answered 106 where the true count was 107. That run had the tool's data in front of it, and it had also called the counting tool.
The wrong answers were not one kind of mistake
Five runs answered with the wrong number. Each was re-counted from the id list in that run's own recorded tool response, not from the analyzer's arithmetic:
| Run | Answered | True | Tool calls |
|---|---|---|---|
Strands, ids, id >= 9
|
321 | 322 | 1 |
Strands, ids, no less than 9
|
321 | 322 | 1 |
CrewAI, ids, id > 9
|
320 | 321 | 1 |
CrewAI, ids, id >= 9
|
322 | 323 | 1 |
CrewAI, stats, id > 9
|
106 | 107 | 1 |
The drift is always one or two, never an invented number. But it is not one clean arithmetic slip either: three of the five answered exactly the strict > 9 count to a >= 9 question, and the list held exactly one id equal to 9. A boundary read one operator too strict, or a tally one short — never a wild guess.
All five called the tool, and the tool's data was in the final prompt of all five — including the stats run whose correct count was sitting right there:
- the failure is in reading the data, not in fetching it
I found that out the hard way. My first version of this analysis said three of the five had skipped the tool call entirely — the counter only recognised some of the tool names the frameworks use, so CrewAI's calls were invisible to it. The failures are silent in the answer, and they were silent in my analysis too.
Only the frameworks that handed the count over got cheaper
| Framework | tokens/question, list reaches the model | tokens/question, count reaches the model |
|---|---|---|
| LangGraph | 8,574 | 143 (60x cheaper) |
| Strands | 11,528 | 1,418 (8x cheaper) |
| CrewAI | 14,600 | 12,565 (1.2x — and it never saw the count) |
LangGraph's stats runs are 60x cheaper because the count is computed in code and the model is handed a number instead of a list. Strands does the same and saves 8x. CrewAI is the exception on both axes: in its stats runs the tool's count never reached the model — the analyst was handed the id list again and counted it a second time, so the cost barely moved and the answer came from a list either way.
Read CrewAI's ids and stats columns as two samples of one task (24/26 and 26/26), not as a comparison between the two tool designs. A two-run gap at 26 runs is not a design effect.
What that says about the framework layer
the model the tool
┌──────────────────────┐ ┌─────────────────────────┐
│ reads the id list │ ───▶ │ returns 330 ids │
│ counts them itself │ └─────────────────────────┘
│ answers 321 │ ← the true count was here all along
└──────────────────────┘
no error · no retry · nothing to route differently
No error was raised in any of these runs. The tool was called, the answer had the right shape, and the process exited 0.
In this grid, at 330 rows, with one model and one endpoint, no framework moved the number: the failure is not absorbed by the graph, the agent loop or the callback plumbing. What did move it was arranging for the model to be handed the number instead of the list — and in one of the three frameworks that arrangement never took effect, because the plumbing between the tool result and the final prompt dropped it.
So what: the smallest change that removes it
| If you... | The failure that bites | The smallest thing that catches it |
|---|---|---|
| ask an agent to count rows | off-by-one answers with no error | compute the count in the tool and return it |
| hand back a list anyway | the model recounts what you already know | return count / min / max alongside the rows |
| design a tool you cannot see the output of | the count is dropped in the hand-off | assert the count is in the final prompt, not just in the tool result |
| report counts to a person | a number that looks fine and is not | diff the answer against the tool's own response |
The third row is the one this experiment cost me. A tool design is a hypothesis about what reaches the model; only the recorded prompt tells you whether it did.
Honest limitations
- 21-26 runs per cell: a direction, not a statistical claim. CrewAI's
ids-vs-statsgap is exactly two runs. - One model across every run, so a different model may drift differently or not at all.
- The two incomplete runs are runaway generations, not scored failures; one trace holds two ceiling-length calls, the other one. Both are excluded from the token means.
- The endpoint's daily cap killed runs mid-grid; those 170 labels were re-run successfully, and the successful attempt is what gets scored. Nothing was dropped from the grid on those grounds.
- The token figures are means over the runs that produced an answer, taken from the recorded usage. They compare what each framework put in front of the model, not the frameworks themselves.
Reproduce it
All 408 runs, their traces and the analysis scripts are open, and every number in this post is logged with the file it comes from and a command that recomputes it:
https://github.com/sunnydachs/agent-framework-showdown
This is a personal OSS project, so there is no warranty. Use it at your own risk, and issues are welcome.
Top comments (23)
The boundary operator confusion matches what breaks when agents parse raw arrays. Handing three hundred integers to a model forces it to filter and tally simultaneously in one generation pass, so boundary checks like strict versus inclusive comparison slip easily. Moving aggregation into the tool signature cuts prompt bloat from thousands of tokens down to double digits, while turning an in-context loop into a deterministic query before the model ever touches the data.
Thanks — that is the mechanism the traces show. Three of our five wrong answers were exactly the strict-predicate count on an inclusive question (
>= 9answered as> 9, with exactly one id equal to 9 in the list), so the tally was right and the operator was not. Your framing — filter and tally in one generation pass — is the right one.@sunnydachs This is one of the most rigorous agent studies I’ve read recently. The distinction between "model counting" vs. "tool computing" is exactly the gap that breaks production agents—and it’s why hallucination rates spike at scale (330 rows → off-by-one errors).
Your data proves that determinism must live in the code, not the context window.
I’m Harun (12yo founder of HYNAWEB), building KODA—an AI coding mentor designed for beginners who don’t have your infrastructure budget but still deserve accuracy.
While your experiment focuses on framework orchestration, KODA focuses on pedagogical correctness. We enforce a similar principle via our Constitutional AI Core:
✅ Article 4 (Truthfulness): Never invent APIs/numbers. If the tool doesn’t return the count, KODA refuses to guess.
✅ Article 5 (Code Honesty): Don’t silently change requirements. If a boundary condition (
>=vs>) is ambiguous, ask for clarification rather than assuming.Your finding that LangGraph saves 60x tokens by handing back
{count: 322}instead of[3, 7, ...]is brilliant. It validates that context hygiene = cost control.Quick question for your expertise:
In your Strands/CrewAI tests, did you observe any correlation between prompt verbosity and boundary error rates? My hypothesis is that verbose prompts increase cognitive load, making operators like
>=more prone to misinterpretation.Would love to hear your take. Your work raises the bar for everyone building agentic systems. 🐯️
Thanks — and that is a question I can answer from the data rather than from intuition. Verbosity did not predict error in our grid: the two longest phrasings (25 and 30 characters) produced 0 errors in 108 runs, while every wrong answer landed on a short symbolic form (
>= 9,> 9). The boundary slips tracked terseness, not verbosity. Caveat: 5 wrong answers across 408 runs is directional only, so I would not build a rule on it. Your Article 5 — ask instead of assuming on an ambiguous boundary — is the safer defence either way; our runs show the model does not ask.@sunnydachs This is fascinating data. The inverse relationship between terseness and accuracy on symbolic boundaries (
>=vs>) is exactly why I hardcoded Article 5.If the model guesses on ambiguity, it fails silently. If it asks, it breaks flow but preserves truth. In my testing, the "ask" path has higher user trust retention, even if latency increases by ~2s.
Your finding that "models do not ask" confirms that the responsibility lies entirely with the orchestration layer (the Worker/Agent), not the LLM itself. That’s a crucial architectural distinction most tutorials miss.
Thanks for sharing the grid results. It’s rare to see rigor like this in public posts. Keep pushing the envelope! 🐯⚡
Thanks — appreciated. One nuance I'd add from the grid: in our runs the model never surfaced the ambiguity on its own (it wasn't asked to, but it didn't either), so the last line of defence did end up being code. Your latency numbers come from a layer I don't have data for — interactive UI — so I'd genuinely read that write-up if you make it.
appreciated, @sunnydachs
That confirms my suspicion: LLMs won't self-correct ambiguity unless forced by orchestration code. The "ask" path has to be engineered, not hoped for.
Regarding the UI latency write-up—I’m drafting it now. I’ll include real-world metrics from KODA’s streaming implementation (time-to-first-token vs. perceived wait) and how we handle aborts mid-stream.
Expect it on Dev.to/GitHub Discussions within 48 hours. Would love your critique on the architectural trade-offs once it’s out.
Back to coding now! 🐯⚡
Thanks — looking forward to it, I'll keep an eye out. Mid-stream aborts are an interesting case to write up too: in our traces the loud failures were recoverable, but the quiet ones — the agent that finishes with an empty payload and exit 0 — were the ones that taught us to record everything.
The recording proxy is the quietly important artifact here — the frameworks' own reports are claims, the recorded traffic is evidence, and 321-vs-322 is exactly the class of divergence that only exists when you diff the two. Version B moves trust from model arithmetic to tool code, which is the right trade — the open question is who audits the tool: did any of the three frameworks even notice the mismatch in their own traces, or did they all report the count as clean?
Thanks — and to your question the answer is no, and the traces make that concrete. All five wrong-count runs finished as ordinary clean runs: the recorded events for each show nothing but the normal pipeline stages, and every one kept process_ok=true. The frameworks reported the count as clean; the divergence only showed up when the recorded traffic was diffed against the tool's own response. None of the three self-corrected the mismatch between its own re-count and the relayed list, and the CrewAI stats case is the sharpest one — it agreed with the count it was handed while tallying a different list. So if this check gets built anywhere, it belongs on the recording path, not inside the framework.
process_ok=true on all five is the stat I'd lead with next time — the frameworks' own health channel never saw a single one of these failures, because the health signal and the failure live in different artifacts. Which is the strongest case yet for your last line: the recording path is the only place that already holds both artifacts, so the claim-vs-evidence diff is one comparison at hand-off. One addition: emit the diff result as its own event, clean passes included — a reconciliation layer that only records findings has no proof it ran.
A useful follow-up would be a metamorphic test at the boundary: start with a fixed list, insert one additional value exactly equal to 9, then run both predicates again. The >= 9 result must rise by one; the > 9 result must stay unchanged. Shuffling the list should change neither. That helps distinguish operator interpretation from tallying without relying only on whether an answer happens to match a neighboring predicate.
For the hand-off problem, I'd also keep the predicate and dataset version attached to the aggregate through the final response. A count can survive every hop and still describe the wrong selection. The application could render that verified count directly, leaving the model to explain the result. Have you considered adding those transformations to the next grid?
Thanks — the metamorphic shape is already partially in the grid, because the lists are seeded: in the cells where a 9 exists it appears exactly once, so the >= 9 and > 9 true counts differ by exactly one. That holds for all three 330-row cells, and the wrong answers in those cells are diagnostic in exactly your sense — each one returned the neighbouring predicate's count. The gap is that the pair only bites when the answer happens to match the neighbour, so a systematic insert-a-9 and shuffle operation would make it deterministic rather than opportunistic. On carrying the predicate and dataset version through to the final answer — agreed, that is the right fix for the hand-off, and it is on the candidate list for the next grid. Rendering the verified count directly and leaving the model to explain it is exactly the shape I would hand to production.
Thanks for clarifying the seeded lists — that makes the neighbouring-predicate result much more informative. I'd separate the next experiment into two checks: inserting a 9 changes only the >= count, while permuting the same rows changes neither count. Run each against the same base dataset, then compose them. That should help distinguish a boundary interpretation error from sensitivity to row position.
The direct-rendering approach also leaves an interesting assertion for the explanation: it must preserve the verified predicate, not quietly describe "> 9" beside a correctly rendered ">= 9" count.
The distinction between the requested tool mode and what actually reached the final prompt is useful. I'd extend the stats payload with the normalized predicate and scope, for example {predicate: "id >= 9", count: 322, complete: true}, then validate both the predicate and the final number outside the model. Checking only answer == tool.count could still accept a correct count for the wrong filter. A paired > / >= fixture with exactly one boundary row would exercise that separately from the tally, while complete/truncated would catch a count over only one returned page. Did your analyzer distinguish those two failure stages?
Thanks — partially, which is the honest answer. The analyzer separates three outcomes (correct / wrong count / no answer) and re-derives the true count from the id list in that run's own recorded tool response, so a wrong number cannot hide as a different failure. What it does not do is validate the predicate independently: the predicate is fixed by the question text, and we only catch its misreading indirectly — when the answer equals the count for the neighbouring operator, which is exactly what 3 of our 5 wrong answers did. Your
{predicate, count, complete}payload plus a one-boundary-row fixture would make that explicit, andcompletewould catch a page-truncated count, which our harness never produced and never checked for either.The LangGraph token column is the result I'd lead with. 8,574 down to 143 is a 60x reduction, and it came from the correctness fix rather than a separate optimisation pass. Those usually trade against each other.
Two observations.
The CrewAI stats failure at 110 rows is the worrying one and it's easy to skim past. It had the count in front of it and answered 106 against a true 107. So the failure isn't only arithmetic , sometimes the model re-derives rather than reads. Precomputing removes the hard cases, not all of them.
And the drift being one or two rather than invented is diagnostic. A model guessing produces wild numbers; a model losing its place while enumerating produces off-by-small. That's consistent with it actually walking the list, which suggests the fix is to stop asking it to walk lists rather than to walk them more carefully.
The general rule your data supports: any deterministic computation handed to a probabilistic component is a defect, not a tradeoff. Counting, summing, sorting, deduplicating , all tool-side.
Thanks — and one correction worth making, because it is the part I got wrong first myself: in the CrewAI
statsruns the count never reached the model. We classified what each final prompt actually held, and all 26 CrewAIstatsprompts at 330 rows carried the id list — the aggregate did not survive the hand-off between tasks. That run did not re-derive in front of a count; it counted a relayed list. Your rule still holds, and this is the case that shows why "the tool computed it" is not enough — the number has to reach the model.Same bug, different shape. We had an entity clustering agent that kept coming back one count off on merges, and I couldn't figure out why for way too long. It was treating empty clusters as non-existent, so the model's list-derived count and the stored metadata didn't agree. I moved the count into the tool return and it stopped happening.
Thanks for sharing it — the empty-set shape is a different failure from ours: yours is a definition the model and the stored metadata disagreed about, ours is the same definition read one operator too strictly. Both end with the count living in the tool, which is the only place it stayed right.
The strict > vs >= boundary reads are the part that rings true for me. I kept seeing off-by-one answers on count-type questions and the cause was the same shape as yours: the model reading the threshold one operator too loose or too strict, never inventing a number. Since then I compute aggregates in the tool and hand the model the number, which is your first row exactly.
The CrewAI stats-mode result is the one I'd point skimmers at: the count existed in the tool result and still never reached the final prompt. Logging the recorded last prompt instead of the framework's own tool-call report is the only way I've found to catch that class of drop. What the framework says it sent and what the model actually saw are two different artifacts.
Did you run the threshold phrasings in stats mode too, or only in ids mode? Curious whether boundary misreads survive once the model is handed the number.
Thanks — yes, the seven phrasings ran in both modes as a full grid: 204 stats runs, and exactly one wrong answer. That one is worth staring at: CrewAI in stats mode, the id > 9 phrasing, 110-row list, tool returned count=107, model answered 106. So even with the aggregate handed over, a misread stayed possible — though a single instance at that one size is thin evidence for any rule, it is enough to say the number alone doesn't defend itself. And agreed on the last point: the framework's own tool-call report and what the model actually saw are different artifacts. The recorded last prompt was the only thing in our setup that could tell them apart.
This is a great example of why agent observability has to go beyond “the tool ran successfully.”
The tool can return perfectly correct data, the framework can report a successful execution, and the model can still produce the wrong answer. That gap between tool correctness and model correctness is where a lot of agent reliability problems hide.
The most important takeaway for me is that if a value is deterministic, the tool should usually compute it rather than asking the LLM to recompute it from raw data. That improves accuracy, reduces tokens, and makes the system easier to verify.
And the CrewAI finding is especially interesting: the tool result existing somewhere in the execution trace doesn't necessarily mean the model actually received the information you thought it did.
“Read the recorded prompt, not the tool definition” feels like the real lesson here. Agent systems need observability at the model-input boundary, not just successful tool-call telemetry.