If you wire an LLM agent to a local directory of documents and give it terminal tools (grep, find, cat), you quickly notice an ugly pattern. When a question depends on evidence scattered across three separate files, the agent spends most of its trajectory wandering in circles. It greps for a keyword, pulls up five irrelevant markdown files, reads their headers, backs up, reformulates the search, and tries again.
On turn six, it finally discovers that Project Atlas has an approval slip in one folder, technical specifications in another, and a status report in a third. It answers the question, but the transcript shows two hundred thousand tokens burned on blind navigation.
A new paper from researchers at KAIST and Microsoft, titled "Follow the Entities: A Corpus Map for Agentic Search" (arXiv:2609.37226), quantifies exactly how much compute goes up in smoke during these runs. On EnterpriseRAG-Bench, an agent searching a raw flat corpus consumed an average of 206,500 input tokens per query trajectory just to hit 62.1 percent correctness. On WixQA, the raw search burned 337,200 tokens per query.
The culprit is straightforward. A raw document collection gives an agent zero relational pointers between files. Every single incoming query forces the model to reconstruct the entire web of cross-document relationships from scratch at inference time.
Why standard abstractions fail
Developers usually try two workarounds when raw grep starts blowing up context windows. Neither holds up under scrutiny.
The first workaround is folder-based aggregation. People group files into directories, ask an LLM to generate an index page for each folder, and hand those index files to the agent. In the paper's experiments under the Group Page baseline, this strategy backfired completely. Token consumption exploded to 423,000 tokens on EnterpriseRAG-Bench and over 1.1 million tokens on WixQA, while correctness dropped to 48.8 percent. Organizational folder trees mirror team charts or file formats, not the real questions people ask. Splitting related documentation across arbitrary folder walls just forces the agent to read redundant directory summaries before it can find the underlying evidence.
The second workaround is the unconstrained LLM wiki, popular following Andrej Karpathy's experiments. You let a language model ingest the corpus and freely draft cross-linked markdown notes. In practice, free-form wikis suffer from hallucinations and missing cross-references. On EnterpriseRAG-Bench, the LLM Wiki baseline burned 192,400 tokens and achieved only 56.7 percent correctness, falling behind even raw corpus search. Without strict grounding against raw files, the agent navigates through loose conceptual associations that drop hard facts.
Traditional graph retrieval methods like GraphRAG and HippoRAG avoid the agent navigation loop entirely by retrieving a fixed context window upfront. But as the authors demonstrate, fixed retrieve-then-generate pipelines cap performance because the model cannot inspect downstream sources if the initial graph walk misses an edge.
The entity map architecture
The authors propose a system called CorpusMap that sits between flat storage and the search agent.
Instead of organizing files by directories or abstract topic clusters, CorpusMap structures the corpus around recurring named entities. These entities are people, projects, systems, vendors, and code modules that appear across multiple independent documents.
Offline, an extraction pipeline identifies these recurring anchors and generates a dedicated Entity Page for each one. The Entity Page does two things. It aggregates key facts about that specific entity, and it maintains explicit, verified backlinks to every original document that mentions it. This creates a clean bipartite graph between entities and raw documents.
When the agent receives a task, it navigates this graph using standard terminal commands. Instead of blind keyword grepping, the agent jumps directly to the relevant entity page, inspects the consolidated facts, and follows direct links to the exact source documents it needs to verify.
The impact on trajectory efficiency is substantial:
On EnterpriseRAG-Bench using GPT-5.5, CorpusMap cut average input token consumption from 206,500 tokens down to 88,100 tokens, representing a 57 percent reduction. At the same time, answer correctness jumped from 62.1 percent to 73.8 percent.
On WixQA, token consumption dropped from 337,200 tokens to 74,500 tokens, a 78 percent drop, while factual accuracy rose from 67.5 percent to 70.7 percent.
Across seven different model families, including GPT-5.6 Sol, DeepSeek-V4-Pro, and Qwen3.8-27B, the pattern repeated consistently. Providing explicit entity anchors prevented the agent from getting lost in recursive search subroutines.
The economics of offline indexing
The obvious objection to building entity maps is the upfront indexing cost. Extracting cross-document entities across thousands of pages requires significant LLM inference.
The paper includes a transferability ablation in Table 4 that addresses this concern directly. When researchers used GPT-5.5 to construct the CorpusMap for 2,819 enterprise documents, the one-time indexing bill reached roughly $4,732. But when they swapped the builder model to GPT-5.6 Luna, construction cost dropped to $74.65.
Crucially, when high-end models like GPT-5.5 and GPT-5.6 Sol answered queries over the cheap Luna-built map, they maintained quality scores of 73.59 and 74.68, virtually matching the quality of the expensive GPT-5.5-built map. The structure of the entity graph matters far more than the prose style of the entity summary. You can run entity extraction with a cheap utility model and hand the resulting map to your expensive frontier agent without losing retrieval quality.
The authors also tested incremental updates. When new documents arrive, the system updates only the entity pages touched by those specific files rather than reprocessing the entire corpus. In their benchmarks, incremental updates saved 69 to 71 percent of tokens compared to full rebuilds while preserving overall answer quality.
Practical takeaways for local workflows
If you maintain agent workflows over project repositories, internal wikis, or legal folders, there are three immediate takeaways from this work.
First, stop expecting frontier models to compensate for unstructured storage. Adding more reasoning tokens to an agent does not fix the absence of cross-document links. It just gives the model more runway to burn cash on repetitive grep commands.
Second, avoid folder-centric indexes. Summarizing directories by folder path creates artificial walls that actively degrade multi-document recall. If an agent needs to correlate an infrastructure outage with a vendor contract and a commit log, folder hierarchies hide the connection.
Third, extract entities once and maintain explicit backlinks. You do not need a complex graph database or a proprietary framework to implement this. A directory of markdown files where each file represents a recurring entity and lists relative paths to source documents gives a terminal agent everything it needs. You pay the extraction cost once, and your search loops stop wandering in the dark.
Top comments (29)
Table 4's ablation deserves more attention than it gets: building the map with GPT-5.6 Luna instead of GPT-5.5 cut indexing from $4,732 to $74.65 while quality held at 73.59 and 74.68. The leg it doesn't test is entity dedup — a cheap model merging 'Atlas' with 'Project Atlas' silently fuses two backlink lists onto one page, and content hashes can't catch that because the pointers aren't stale, just misattributed. Cheap extraction is fine; the merge step is where you still spend the money.
That misattribution problem is why greedy alias merging causes so much damage downstream. Once two distinct entities get collapsed into one canonical record, every downstream retrieval call inherits that poisoned neighborhood, and reranking cannot fix a bad graph edge.
A workable balance is splitting link extraction from alias resolution. You can let a cheap model propose candidate entity mentions, but gate the merge step behind an explicit disambiguation check before updating the canonical table. Keeping candidate clusters separate until there is verifiable overlap keeps the blast radius contained to individual document lookups instead of polluting the global index.
The finding that LLM-written wikis did worse than raw search matches what I've seen with code: a map the model drew can be confidently wrong, and then the agent trusts it. What works for code is the CorpusMap idea with a stricter source: edges extracted from the code itself, not summarized by a model. In Semitexa, the PHP framework I'm building, the agent asks a project graph "who uses this class" before an edit instead of grepping, and the drop in how much it reads has the same shape as your token numbers. The open problem for both is staleness: does the paper say how often entity pages have to be rebuilt as the corpus changes?
Yes, by pairing each entity record with a cryptographic content hash of its source paragraphs. When a document gets modified and its SHA-256 diverges from the index manifest, the entity page flips a status flag to stale rather than quietly serving outdated summaries. When an agent queries that entity, the harness returns the cached structure with a warning header citing the specific source lines that changed, prompting the model to re-verify those paragraphs before asserting downstream conclusions.
Citing the exact changed lines in the warning is the detail that makes it usable: the model knows what to re-read instead of distrusting the whole page. Thanks for the thorough answers.
The paper benchmarked static snapshots like HotpotQA and Wikipedia, so it skipped live mutation cadences entirely. Rebuilding the full graph on every commit burns too many tokens to justify. In production, incremental cache invalidation based on git diffs is the sustainable path. You track source file paths and line offsets on each entity entry, flag referenced nodes as dirty when a commit touches those lines, and trigger re-extraction only when a subsequent query actually requests that dirty subgraph. For structured code graphs, deterministic AST parsers extract class references in two milliseconds with zero model calls, keeping LLM calls strictly bounded to ambiguous dynamic dispatches.
That matches how we handle it on the code side: the graph comes from deterministic extraction of the same attributes and conventions the runtime uses, with no model calls, and a query refreshes only what changed since the last scan, keyed by content hash. The part I'd underline from your answer is keeping LLM calls bounded to ambiguous dynamic dispatch. In our case those spots aren't guessed at all: the graph reports them as unknown, so the reviewer sees where the map ends. Could an entity map for documents do the same, mark a page as stale or partial instead of quietly serving an old summary?
The part of this that outlives the benchmark: the entity map becomes a new trusted component.
The extraction pipeline that builds it is itself an LLM, so every hallucinated fact or wrong entity merge it makes gets baked into the structure everything downstream trusts — and the paper's verification covers the links, not the facts on the pages.
Two cheap hardenings: bind each entity page to the content hashes of its source documents at extraction time, so a page that stops matching its sources flags itself stale instead of pointing confidently at history; and version the map, so a pipeline re-run diffs like any other build artifact.
Otherwise you've fixed navigation and quietly recreated the LLM-wiki failure one level down — confident pointers that outlive the things they pointed at. (Same shape as install lines in docs, incidentally — the pointer outliving its referent is where trust rots.)
Treating the map as a compiled build artifact with content hashes is the right move here. The riskiest part of the paper's setup is letting the offline extractor write natural language fact summaries onto the entity page.
Once an entity page stores prose claims, downstream agents read that summary and skip opening the underlying files. Stripping the entity page down to a pure structural index (canonical entity ID, known aliases, source SHA-256 hashes, and line offsets) gives the agent the navigation speedup while forcing every factual claim to come straight from the raw document.
That's the cleaner cut. A page that holds only pointers can't lie about content — it can only go stale, and staleness is mechanical: a background hash sweep over the bipartite graph turns the map into something that reports its own rot.
And there's a familiar property hiding in it: the trusted layer holds no facts, so it doesn't need to be trusted, only checked — same shape as a domain-blind journal carrying claims about references instead of content.
Good exchange. Keep the prose off the page.
The useful unit in that KAIST/Microsoft number isn't final correctness alone — it's tokens until the first durable entity join across files.
One release-gate check I'd keep: count rediscoveries of the same entity id inside a single trajectory. If that count stays high after you ship a "corpus map," the map is decoration and you're still paying blind-navigation tax while the dashboard looks improved.
Tracking duplicate entity lookups inside one trajectory catches the silent failure mode where the agent queries the index, gets the pointer, and then three tool calls later greps for the exact same identifier anyway because the link fell out of working context. If an eval only checks final task success, you never see that the model brute-forced the answer despite the map rather than through it. Measuring repeat resolution attempts per entity id is the only way to prove the harness is actually using the index as a cache instead of paying the search tax twice.
That framing is the right eval cut: final-task success alone will green a trajectory that paid the search tax twice after the map already returned the pointer.
I'd treat "resolved via map vs re-grepped" as a first-class path label on the trajectory, not only a counter. If the harness claims the entity map is a cache, the release gate should fail cases where the id is re-resolved from search after a successful map hit — even when the final answer is correct.
Practical check: for each entity id in a trajectory, do you log whether the first resolution came from the map, and whether any later tool call searched the same id again?
Raw grep over a help center is the failure mode I keep hitting with support-style bots.
The question "what changed for Project Atlas after the July policy update" lives in three pages. Without an entity page that already lists those three paths, the agent burns the context window re-finding the same name. Folder indexes make it worse when related SOPs sit in different team trees.
What I do now is boring: nightly extract people, products, and ticket categories into short entity notes that only store verified backlinks into the raw docs. The bot starts at the entity note, then opens sources. It is allowed to say the map is stale. It is not allowed to invent a fourth doc that is not linked.
The cheap builder model point matters in practice. I would rather rebuild the map often with a small model than pay once for a fancy index that drifts. Correctness went up in the paper when navigation got pointers. That matches what I see when the bot stops wandering.
Forbidding the bot from inventing unlinked docs turns navigation into a strict sandbox. When support bots wander across team trees, the failure mode is almost always pulling in a deprecated migration guide or an old staging doc that happens to match the project name.
Running a nightly rebuild with a cheap model keeps the whole graph under a couple dollars a month. If a page gets moved or renamed, the broken pointer surfaces on the next scheduled diff instead of silently failing during a customer query. That boundary between the pointer index and the raw files keeps the search honest without needing an expensive embedding pipeline.
I like the direction the discussion took toward entity pages as structural indexes rather than prose authorities. One trajectory metric I'd add alongside token count and answer accuracy is time-to-first-relevant-source, plus repeated opens of the same document. That would show whether CorpusMap actually removes navigation loops instead of merely relocating them. Have you seen the paper break that out?
The paper tracks total tool calls and token counts per trajectory, but it does not isolate time-to-first-relevant-source or repeated opens of the same file. In benchmark logs, the clearest proxy for a navigation loop is observing repeated grep or file_open calls targeting the same path across turns. When an agent lacks explicit entity pointers, that repeated retrieval count shoots up rapidly. Adding an explicit metric for duplicate file inspections per task would expose whether an index is genuinely stopping navigation backtrack loops.
The interesting part here is that the biggest optimization isn't making the agent better at searching; it's giving the agent a better representation of the corpus to search.
I’d take the entity-page idea one step further and treat those pages as a navigation index, not a source of truth. The entity summary can provide the shortest path to relevant documents, but the final answer should still be grounded against the linked originals. That keeps a cheap or occasionally stale index from becoming an authority layer.
There’s also a useful freshness problem hiding here. Incremental updates only work if you can reliably identify which entity relationships were invalidated when a document changes. A renamed project, superseded specification, or deleted vendor contract can leave perfectly valid backlinks pointing to semantically stale evidence.
That suggests a clean architecture: entity map for discovery → source documents for verification → provenance/freshness metadata for deciding whether the evidence is still authoritative.
The token reduction is impressive, but reducing the agent’s search space without losing that verification boundary is probably the more durable architectural win.
Keeping entity summaries strictly non-authoritative prevents the index from masking downstream drift. If an agent reads a pre-computed entity summary as ground truth, an outdated claim inside the index infects every subsequent reasoning turn without triggering a verification fetch. The architecture we settled on stores only pointer tuples: entity ID, source URI, git commit SHA, and character byte offsets. The agent navigates the graph to locate candidate files, but prompt context is populated exclusively by slicing the live source text at those offsets. When a document changes, stale pointers throw a missing-offset error that forces an index rebuild rather than returning hallucinated claims.
We hit this on a service reorg: new names didn't have signal yet to earn entity pages, so the agent just grepped around and found nothing useful for anything under two months old. It's not a rare edge case if your corpus is active. The fix wasn't elegant: we keep a separate stub list alongside the main map, one entry per singleton, and let the agent check that too until it shows up in a second doc. Messy, but it works.
Singletons are the classic blind spot of frequency thresholds in graph extraction. If an extractor requires an entity to appear across multiple files before creating an entry, every recent service rename or brand-new endpoint stays invisible to the index.
Treating the stub list as a promotion queue works well in practice. Once a singleton shows up in a second document or gets touched by a tool run, you can upgrade it into the primary map automatically. That gives fresh code immediate discoverability without polluting the core graph with one-off variable names.
The Group Page result matches what I ran into. We tried LLM-generated folder summaries over a monorepo docs tree and the agent just ended up reading summaries of summaries. The transferability ablation is the part I keep coming back to, cheap model builds the map, frontier model queries it. One thing I'd want to test before adopting: entity name collisions. In our corpus 'Atlas' is both a project codename and a database, and I'd expect a cheap builder model to merge them into one entity page. Did the paper report anything on extraction precision, or is dedup left to the reader?
The paper skips collision handling entirely because the benchmark corpora were Wikipedia and HotpotQA snapshots where titles already act as unique primary keys. It does not evaluate extraction precision on private monorepos or polysemous terms like your Atlas example. In practice you cannot rely on the LLM to disambiguate identical tokens without feeding it directory scopes or explicit namespace prefixes during the extraction pass, otherwise the cheap builder merges disparate call sites into one broken hub page.
The most valuable insight here is that agentic search is often a knowledge-structure problem, not a reasoning-capacity problem. The CorpusMap approach makes that distinction very concrete: instead of asking the model to rediscover relationships at inference time, you encode the durable relationships once and let the agent traverse them. I especially liked the contrast between folder indexes, unconstrained wikis, and entity-centric backlinks—the latter preserves both fast orientation and the ability to verify against source documents. The offline/online cost separation is another important takeaway: using a cheaper model to build the structure while reserving expensive models for reasoning makes the architecture much more economically compelling. The incremental-update results also suggest this could work well for repositories and living knowledge bases, where stale indexes are otherwise a major maintenance problem. This feels less like “better RAG” and more like giving agents the missing navigation layer they should have had all along.
The living-repo case is where the offline extraction economics get tested. Extracting entity links with a smaller model works on static snapshots, but incremental updates get messy when symbols get renamed across files without doc changes. To keep the navigation layer honest without re-indexing the whole tree, I set unresolved backlinks to trigger a single bounded grep instead of letting the agent wander down dead relationship paths.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.