The agentic RAG pipeline that was faster and cheaper — and no more accurate than no agent at all
I spent a few weeks building four retrieval architectures over the same graph database, pointing them at the same 150 questions, and measuring what each one cost.
The headline numbers:
| accuracy | LLM tokens/q | mean latency | tool calls/q | |
|---|---|---|---|---|
| P1 · RAG | 30% | 1,284 | 27.0 s | 1.00 |
| P2 · GraphRAG | 89% | 1,694 | 32.2 s | 2.72 |
| P3 · Agentic | 100% | 552 | 10.2 s | 1.95 |
| P4 · Deterministic control | 100% | 0 | 0.1 s | 1.57 |
One model — Qwen3.8-27B, with its reasoning trace enabled — served every pipeline, including the agentic orchestrator's planner and verifier.
Only LLM tokens were counted; GSQL and Python were treated as zero-cost.
The finding I did not expect
I built P4 as a sanity check, not a competitor: strip out the model entirely, parse the question with regexes, and run one deterministic query. I expected it to score maybe 70% and lose badly.
It scored 100%.
That is the most useful result in the whole benchmark, and it changes how you should read every other row.
The agentic pipeline is not better than the deterministic control. It is more expensive — at 552 tokens per question — for identical accuracy.
So why keep it?
Because the control is a boundary marker, not a competitor. It tells you exactly how many questions never needed a model.
And because the benchmark questions turned out to be templated, that number was all of them — which is precisely the caveat I'll come back to at the end.
Why GraphRAG jumps 59 points
The 150 questions come from five templates. Four of them are structured database queries wearing question marks:
| template | P1 · RAG | P2 · GraphRAG | P3 · Agentic |
|---|---|---|---|
| lookup | 63% | 100% | 100% |
| temporal | 27% | 55% | 100% |
| multi_hop | 32% | 96% | 100% |
| aggregation | 5% | 100% | 100% |
| superlative | 20% | 100% | 100% |
Plain RAG scores 63% when the answer is in one document and 5% when the question asks how many events meet a condition.
That is not a retrieval-quality problem.
A vector retriever cannot count. It returns chunks, the model reads them, and the model is bad at counting over a set — so no amount of chunk quality fixes it.
GraphRAG fixed it by pushing aggregation into the database.
The GSQL does count with a predicate and returns a number. The model's job shrinks from "compute this" to "read this", which is the job models are actually good at.
This is the real dividing line between RAG and GraphRAG, and it has nothing to do with graph databases being fashionable.
It's arithmetic.
Why the agent got cheaper
The agentic pipeline costs 67% fewer tokens than GraphRAG while being 11 points more accurate.
Agents are normally assumed to be expensive — more calls, more context, more chances to ramble.
Mine did less work.
Three design decisions did that.
1. The planner chooses the tool; deterministic parsing fills the slots
The orchestrator emits STEP: count_above and nothing else.
sport, year, season, and threshold come from a regex parser.
An earlier version let the planner's parameters override the parsed ones. Because questions say "biathlon" while the graph says "Biathlon", every aggregation silently answered zero.
Slot filling is mechanical.
The plan is the part that varies with the question, and that's the part worth spending tokens on.
2. Stopping is an explicit decision with a recorded reason
When a tool returns a definite value and the deterministic verifier agrees, the loop accepts it and stops rather than paying for another round trip.
72 of 100 public questions ended that way.
3. A repeated query ends the investigation
A planner re-issuing the same query is looping, not reasoning.
Detecting the repeat and telling it to move on is what keeps the average at ~2 tool calls instead of burning the step budget on one empty result.
The reason P3 is also the fastest pipeline is the same mechanism.
GraphRAG's 1,694 tokens are spent almost entirely on the model doing sums that the database could have done, and every token is also latency.
TigerGraph specifics worth knowing
If you build on TigerGraph, some of this will save you days.
vectorSearch() does not work in interpreted queries
It fails with:
GSQL-2500 Unsupported Statement | TOPK_VEC_SEARCH_FUNC
It has to live in an installed query, which means one CREATE QUERY + INSTALL QUERY.
The USE GRAPH prefix also has to be repeated in every call because it does not persist between them.
Interpreted queries refuse parameters
INTERPRET QUERY (sp=STRING) is rejected whether you pass a dict or a query string.
So values get inlined as escaped literals.
If you build this, make sure every inlined value originates from your own parser and not from user input.
This deployment runs GSQL syntax v2
It is stricter than you'd expect:
- A bare
SELECT e FROM Event:eis rejected — everySELECTneeds a vertex-set variable. - You cannot traverse from a filtered vertex set.
Games:g <-IN_GAMES- Event:edoesn't parse. Run it forwards and filter the target in the sameSELECT. -
FOREACHdoesn't mix with v2 accumulators, andMaxAccumcan't be assigned to a global. - There is no
coalesceand noIFF. - A null in an accumulator is silently dropped from that column while its neighbours still get the row — so your columns shift and you return confidently wrong answers. I now write sentinel values instead of skipping attributes.
Adding a vector attribute needs a SCHEMA_CHANGE JOB
A bare ALTER ... ADD VECTOR ATTRIBUTE is not valid GSQL.
And if any type in your schema is declared global, the job must be global too.
The full list is 31 footguns in the repo, each one of which produced a silently wrong answer before it was found.
The bugs my tests caught that my summary table didn't
This is the part I'd most want other people to take away.
My results summary reported 100/100 and errors: 0 — both true — while 26 of 150 agentic answers were a whole sentence where a bare number was due.
For example:
"'Men's foil' has nations=29"
instead of:
"29"
The cause: my answer extractor's markers were all spaced — " is ", " = ", ": ".
The tool emitted an unspaced key=value tail.
All three markers missed, the function fell through to return line.strip(), and the clause became the answer.
I found it by grepping the artefacts for the shape of an answer, not by reading the summary.
The summary was accurate about what it measured and completely silent about what it didn't.
The scorer had a similar blind spot
My scorer returned True for:
is_correct("c", ["Chen Ding"])
because containment had no length floor, so any single character landing inside the gold scored as correct.
A needle now has to be 4+ characters or a standalone token, which keeps "29" matching inside "29 nations" while rejecting "4" inside "24".
Both defects were scored correct the whole time.
Containment hid the first one, and the second had nothing to fire on because the pipelines emit corpus-exact strings.
I re-scored the entire benchmark with the hardened scorer and nothing moved.
That is the point.
They were latent, not active.
Others in the same family
- Diacritics were not folded when scoring, though the rules state they are normalized.
- Abbreviated month names (
"Aug 7 (prelim), Aug 10 (final)") yielded no date tokens at all, silently costing that event its multi-hop match. - A report generator referenced an element ID the markup never defined, so one section rendered blank while everything else looked fine.
The general lesson:
Aggregate metrics are blind to shape defects.
A summary can be entirely correct about accuracy and tell you nothing about whether your answers are in the right form to be graded.
I ended up with 74 tests using Python's standard-library unittest.
They need neither the graph nor the model — which matters more than it sounds, because the model endpoint was down a lot during this build.
What I did not prove
Two things, stated plainly because the results invite the wrong conclusion.
The agent's recovery path is untested
changed_strategy is 0/100.
Every one of 211 tool calls returned data — no question ever made the first move fail, so the re-planning machinery never fired.
The agent demonstrably runs multi-step plans and gets them right, but I never observed it recover from a bad step.
An adversarial probe set — misspelled sport, invented venue, impossible Games edition — would exercise that, and I didn't build one.
Worth noting that such a probe set probably wouldn't show the agent winning anyway.
Both the agentic and deterministic pipelines share the same regex parser and the same entity linker, so corrupting an entity breaks both equally.
My mental model of "the router fails, the agent recovers" was wrong about this architecture.
Latency is not a system property
All three pipelines call a 27B reasoning model over a tunnel, so those 27–32 second figures are dominated by network round-trips to a model runtime, not by retrieval.
The graph itself answers in well under a second.
The zero-token control finishes a full 150-question pass in 0.1 s average.
If you see 32 s for a graph database, ask what fraction of it was the model.
Where this goes next
If I were picking this up again, the interesting work isn't squeezing accuracy — that's saturated, and the deterministic control already proves it.
It's building the adversarial probe set, because the only untested claim in the whole system is whether the agent can recover when its first move fails.
Everything else is measured.
Repo
Full source, traces, and the hidden-50 outputs:
https://github.com/Kultzuki/TigerGraph
Live results page (self-contained, no dependencies):
https://kultzuki.github.io/TigerGraph/olympic-graphrag/results/report.html
Setup is one dependency:
bash
pip install -r requirements.txt
Top comments (0)