DEV Community

Search Agents Waste Half Their Tokens Rediscovering Entity Links

Reid Marlow on September 30, 2026

If you wire an LLM agent to a local directory of documents and give it terminal tools (grep, find, cat), you quickly notice an ugly pattern. When a...
Collapse
 
aifrontierpost profile image
AI Frontier Post •

Table 4's ablation deserves more attention than it gets: building the map with GPT-5.6 Luna instead of GPT-5.5 cut indexing from $4,732 to $74.65 while quality held at 73.59 and 74.68. The leg it doesn't test is entity dedup — a cheap model merging 'Atlas' with 'Project Atlas' silently fuses two backlink lists onto one page, and content hashes can't catch that because the pointers aren't stale, just misattributed. Cheap extraction is fine; the merge step is where you still spend the money.

Collapse
 
reidmarlow profile image
Reid Marlow •

That misattribution problem is why greedy alias merging causes so much damage downstream. Once two distinct entities get collapsed into one canonical record, every downstream retrieval call inherits that poisoned neighborhood, and reranking cannot fix a bad graph edge.

A workable balance is splitting link extraction from alias resolution. You can let a cheap model propose candidate entity mentions, but gate the merge step behind an explicit disambiguation check before updating the canonical table. Keeping candidate clusters separate until there is verifiable overlap keeps the blast radius contained to individual document lookups instead of polluting the global index.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

The finding that LLM-written wikis did worse than raw search matches what I've seen with code: a map the model drew can be confidently wrong, and then the agent trusts it. What works for code is the CorpusMap idea with a stricter source: edges extracted from the code itself, not summarized by a model. In Semitexa, the PHP framework I'm building, the agent asks a project graph "who uses this class" before an edit instead of grepping, and the drop in how much it reads has the same shape as your token numbers. The open problem for both is staleness: does the paper say how often entity pages have to be rebuilt as the corpus changes?

Collapse
 
reidmarlow profile image
Reid Marlow •

The paper benchmarked static snapshots like HotpotQA and Wikipedia, so it skipped live mutation cadences entirely. Rebuilding the full graph on every commit burns too many tokens to justify. In production, incremental cache invalidation based on git diffs is the sustainable path. You track source file paths and line offsets on each entity entry, flag referenced nodes as dirty when a commit touches those lines, and trigger re-extraction only when a subsequent query actually requests that dirty subgraph. For structured code graphs, deterministic AST parsers extract class references in two milliseconds with zero model calls, keeping LLM calls strictly bounded to ambiguous dynamic dispatches.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

That matches how we handle it on the code side: the graph comes from deterministic extraction of the same attributes and conventions the runtime uses, with no model calls, and a query refreshes only what changed since the last scan, keyed by content hash. The part I'd underline from your answer is keeping LLM calls bounded to ambiguous dynamic dispatch. In our case those spots aren't guessed at all: the graph reports them as unknown, so the reviewer sees where the map ends. Could an entity map for documents do the same, mark a page as stale or partial instead of quietly serving an old summary?

Collapse
 
reidmarlow profile image
Reid Marlow •

Yes, by pairing each entity record with a cryptographic content hash of its source paragraphs. When a document gets modified and its SHA-256 diverges from the index manifest, the entity page flips a status flag to stale rather than quietly serving outdated summaries. When an agent queries that entity, the harness returns the cached structure with a warning header citing the specific source lines that changed, prompting the model to re-verify those paragraphs before asserting downstream conclusions.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

Citing the exact changed lines in the warning is the detail that makes it usable: the model knows what to re-read instead of distrusting the whole page. Thanks for the thorough answers.

Collapse
 
slabb profile image
Sam LABBE •

The part of this that outlives the benchmark: the entity map becomes a new trusted component.
The extraction pipeline that builds it is itself an LLM, so every hallucinated fact or wrong entity merge it makes gets baked into the structure everything downstream trusts — and the paper's verification covers the links, not the facts on the pages.
Two cheap hardenings: bind each entity page to the content hashes of its source documents at extraction time, so a page that stops matching its sources flags itself stale instead of pointing confidently at history; and version the map, so a pipeline re-run diffs like any other build artifact.
Otherwise you've fixed navigation and quietly recreated the LLM-wiki failure one level down — confident pointers that outlive the things they pointed at. (Same shape as install lines in docs, incidentally — the pointer outliving its referent is where trust rots.)

Collapse
 
reidmarlow profile image
Reid Marlow •

Treating the map as a compiled build artifact with content hashes is the right move here. The riskiest part of the paper's setup is letting the offline extractor write natural language fact summaries onto the entity page.

Once an entity page stores prose claims, downstream agents read that summary and skip opening the underlying files. Stripping the entity page down to a pure structural index (canonical entity ID, known aliases, source SHA-256 hashes, and line offsets) gives the agent the navigation speedup while forcing every factual claim to come straight from the raw document.

Collapse
 
slabb profile image
Sam LABBE •

That's the cleaner cut. A page that holds only pointers can't lie about content — it can only go stale, and staleness is mechanical: a background hash sweep over the bipartite graph turns the map into something that reports its own rot.

And there's a familiar property hiding in it: the trusted layer holds no facts, so it doesn't need to be trusted, only checked — same shape as a domain-blind journal carrying claims about references instead of content.

Good exchange. Keep the prose off the page.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

The useful unit in that KAIST/Microsoft number isn't final correctness alone — it's tokens until the first durable entity join across files.

One release-gate check I'd keep: count rediscoveries of the same entity id inside a single trajectory. If that count stays high after you ship a "corpus map," the map is decoration and you're still paying blind-navigation tax while the dashboard looks improved.

Collapse
 
reidmarlow profile image
Reid Marlow •

Tracking duplicate entity lookups inside one trajectory catches the silent failure mode where the agent queries the index, gets the pointer, and then three tool calls later greps for the exact same identifier anyway because the link fell out of working context. If an eval only checks final task success, you never see that the model brute-forced the answer despite the map rather than through it. Measuring repeat resolution attempts per entity id is the only way to prove the harness is actually using the index as a cache instead of paying the search tax twice.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

That framing is the right eval cut: final-task success alone will green a trajectory that paid the search tax twice after the map already returned the pointer.

I'd treat "resolved via map vs re-grepped" as a first-class path label on the trajectory, not only a counter. If the harness claims the entity map is a cache, the release gate should fail cases where the id is re-resolved from search after a successful map hit — even when the final answer is correct.

Practical check: for each entity id in a trajectory, do you log whether the first resolution came from the map, and whether any later tool call searched the same id again?

Collapse
 
makeyouragent profile image
MakeYourAgent •

Raw grep over a help center is the failure mode I keep hitting with support-style bots.

The question "what changed for Project Atlas after the July policy update" lives in three pages. Without an entity page that already lists those three paths, the agent burns the context window re-finding the same name. Folder indexes make it worse when related SOPs sit in different team trees.

What I do now is boring: nightly extract people, products, and ticket categories into short entity notes that only store verified backlinks into the raw docs. The bot starts at the entity note, then opens sources. It is allowed to say the map is stale. It is not allowed to invent a fourth doc that is not linked.

The cheap builder model point matters in practice. I would rather rebuild the map often with a small model than pay once for a fancy index that drifts. Correctness went up in the paper when navigation got pointers. That matches what I see when the bot stops wandering.

Collapse
 
reidmarlow profile image
Reid Marlow •

Forbidding the bot from inventing unlinked docs turns navigation into a strict sandbox. When support bots wander across team trees, the failure mode is almost always pulling in a deprecated migration guide or an old staging doc that happens to match the project name.

Running a nightly rebuild with a cheap model keeps the whole graph under a couple dollars a month. If a page gets moved or renamed, the broken pointer surfaces on the next scheduled diff instead of silently failing during a customer query. That boundary between the pointer index and the raw files keeps the search honest without needing an expensive embedding pipeline.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

I like the direction the discussion took toward entity pages as structural indexes rather than prose authorities. One trajectory metric I'd add alongside token count and answer accuracy is time-to-first-relevant-source, plus repeated opens of the same document. That would show whether CorpusMap actually removes navigation loops instead of merely relocating them. Have you seen the paper break that out?

Collapse
 
reidmarlow profile image
Reid Marlow •

The paper tracks total tool calls and token counts per trajectory, but it does not isolate time-to-first-relevant-source or repeated opens of the same file. In benchmark logs, the clearest proxy for a navigation loop is observing repeated grep or file_open calls targeting the same path across turns. When an agent lacks explicit entity pointers, that repeated retrieval count shoots up rapidly. Adding an explicit metric for duplicate file inspections per task would expose whether an index is genuinely stopping navigation backtrack loops.

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The interesting part here is that the biggest optimization isn't making the agent better at searching; it's giving the agent a better representation of the corpus to search.

I’d take the entity-page idea one step further and treat those pages as a navigation index, not a source of truth. The entity summary can provide the shortest path to relevant documents, but the final answer should still be grounded against the linked originals. That keeps a cheap or occasionally stale index from becoming an authority layer.

There’s also a useful freshness problem hiding here. Incremental updates only work if you can reliably identify which entity relationships were invalidated when a document changes. A renamed project, superseded specification, or deleted vendor contract can leave perfectly valid backlinks pointing to semantically stale evidence.

That suggests a clean architecture: entity map for discovery → source documents for verification → provenance/freshness metadata for deciding whether the evidence is still authoritative.

The token reduction is impressive, but reducing the agent’s search space without losing that verification boundary is probably the more durable architectural win.

Collapse
 
reidmarlow profile image
Reid Marlow •

Keeping entity summaries strictly non-authoritative prevents the index from masking downstream drift. If an agent reads a pre-computed entity summary as ground truth, an outdated claim inside the index infects every subsequent reasoning turn without triggering a verification fetch. The architecture we settled on stores only pointer tuples: entity ID, source URI, git commit SHA, and character byte offsets. The agent navigates the graph to locate candidate files, but prompt context is populated exclusively by slicing the live source text at those offsets. When a document changes, stale pointers throw a missing-offset error that forces an index rebuild rather than returning hallucinated claims.

Collapse
 
hannune profile image
Tae Kim •

We hit this on a service reorg: new names didn't have signal yet to earn entity pages, so the agent just grepped around and found nothing useful for anything under two months old. It's not a rare edge case if your corpus is active. The fix wasn't elegant: we keep a separate stub list alongside the main map, one entry per singleton, and let the agent check that too until it shows up in a second doc. Messy, but it works.

Collapse
 
reidmarlow profile image
Reid Marlow •

Singletons are the classic blind spot of frequency thresholds in graph extraction. If an extractor requires an entity to appear across multiple files before creating an entry, every recent service rename or brand-new endpoint stays invisible to the index.

Treating the stub list as a promotion queue works well in practice. Once a singleton shows up in a second document or gets touched by a tool run, you can upgrade it into the primary map automatically. That gives fresh code immediate discoverability without polluting the core graph with one-off variable names.

Collapse
 
jkming profile image
jkming •

The Group Page result matches what I ran into. We tried LLM-generated folder summaries over a monorepo docs tree and the agent just ended up reading summaries of summaries. The transferability ablation is the part I keep coming back to, cheap model builds the map, frontier model queries it. One thing I'd want to test before adopting: entity name collisions. In our corpus 'Atlas' is both a project codename and a database, and I'd expect a cheap builder model to merge them into one entity page. Did the paper report anything on extraction precision, or is dedup left to the reader?

Collapse
 
reidmarlow profile image
Reid Marlow •

The paper skips collision handling entirely because the benchmark corpora were Wikipedia and HotpotQA snapshots where titles already act as unique primary keys. It does not evaluate extraction precision on private monorepos or polysemous terms like your Atlas example. In practice you cannot rely on the LLM to disambiguate identical tokens without feeding it directory scopes or explicit namespace prefixes during the extraction pass, otherwise the cheap builder merges disparate call sites into one broken hub page.

Collapse
 
patjo profile image
Pat Johansen •

The most valuable insight here is that agentic search is often a knowledge-structure problem, not a reasoning-capacity problem. The CorpusMap approach makes that distinction very concrete: instead of asking the model to rediscover relationships at inference time, you encode the durable relationships once and let the agent traverse them. I especially liked the contrast between folder indexes, unconstrained wikis, and entity-centric backlinks—the latter preserves both fast orientation and the ability to verify against source documents. The offline/online cost separation is another important takeaway: using a cheaper model to build the structure while reserving expensive models for reasoning makes the architecture much more economically compelling. The incremental-update results also suggest this could work well for repositories and living knowledge bases, where stale indexes are otherwise a major maintenance problem. This feels less like “better RAG” and more like giving agents the missing navigation layer they should have had all along.

Collapse
 
reidmarlow profile image
Reid Marlow •

The living-repo case is where the offline extraction economics get tested. Extracting entity links with a smaller model works on static snapshots, but incremental updates get messy when symbols get renamed across files without doc changes. To keep the navigation layer honest without re-indexing the whole tree, I set unresolved backlinks to trigger a single bounded grep instead of letting the agent wander down dead relationship paths.

Collapse
 
kozmonot20 profile image
kozmonot20 •

dev.to/kozmonot20/cyber-crab-3lmj