DEV Community

Context Compression for Coding Agents Compresses the Wrong Side of the Prompt

Reid Marlow on September 28, 2026

Most engineering teams working on long-context agents hit the same billing wall around turn twenty. A coding agent runs twenty shell commands, read...
Collapse
 
alifunk profile image
Ali-Funk •

To me the architectural split between "environment output" and "agent intent" makes perfect sense from an infrastructure perspective. If we treat a coding agent like a system administrator, it needs an exact log of its own executed commands (the hard actions) to understand what broke.

However, the 10,000-line output of a crashed grep or a full core dump can be compressed.
It’s essentially separating the audit log from the payload.
That means keeping the decision layer byte-exact while distilling the "environment noise" is a much more robust design pattern than blind semantic truncation.

Collapse
 
reidmarlow profile image
Reid Marlow •

The audit log versus payload distinction is clean, but the edge case that always bites is where stdout turns into error diagnostics. If a harness collapses a 5,000-line build failure into a one-sentence summary, the agent loses the exact linker symbol or missing header on line 4,812 that tells it why the build broke.

What has worked well in my scripts is keeping the exact exit code plus a bounded head and tail of the stream in context, while dumping the full unedited output to a scratch file on disk. The agent gets the failure trace immediately without burning 20k tokens on compiler noise, and it can grep the scratch path directly if it actually needs the full traceback.

Collapse
 
alifunk profile image
Ali-Funk •

You map classic Linux log management directly to the LLM context window.

Effective troubleshooting never involves reading massive log files sequentially.
You isolate the exit code.
You read the tail.
You grep for specific linker symbols.

Dumping unedited stdout streams to disk treats the agent like a standard Unix process.
This shifts operational burden from expensive context memory to cheap local storage.

The agent uses standard diagnostic tools instead of brute-force reading comprehension.

Solid infrastructure architecture solves model limitations.

Collapse
 
maddy30445r profile image
Madhur Mittal •

One thing I'd add from building in this space: past turn twenty, most of what's worth keeping isn't the raw history at all, it's a handful of facts. What was decided, what was tried and failed, what's still open. Pulling those into a small state block and starting fresh with it has worked better for me than compressing the transcript, both on cost and on the agent repeating old mistakes. Did you measure answer quality alongside token savings?

Collapse
 
reidmarlow profile image
Reid Marlow •

The benchmark numbers in the post come from the Zou et al. paper on SWE-bench Verified, and they tracked resolve rate directly against token reduction. On unbounded windows with K=3 observations kept raw, resolve rate dipped from 14.5% to 12.1% on Qwen3-4B. But on long-horizon runs capped at 32K tokens, preserving latent observations resolved 21.1% compared to 11.1% for hard truncation.

That structured state block (decisions, failed attempts, open tasks) works well for keeping high-level intent intact. Where I see it hit friction in a coding harness is patch application. If an extraction pass summarizes an earlier traceback instead of keeping the exact line numbers and symbol names, the agent has to re-run the inspection tools just to get the raw strings back into scope.

Collapse
 
maddy30445r profile image
Madhur Mittal •

That matches what I hit building Deiko. The state block is great for intent and bad for exact strings, so I keep two layers: a short note (decided, tried, open) goes into the next chat, and each line links back to the original brief, with the exact screen text, error output and file paths, which the agent fetches on demand through an MCP tool. In my tests, agents did pull the full originals when a question needed exact details, so the note could stay small. The 21.1% vs 11.1% at 32K is a big gap. Did the paper split out how much came from exact observations versus recency?

Collapse
 
reidmarlow profile image
Reid Marlow •

The ablation compared keeping K recent raw observations alongside compressed LOHA state blocks against naive head-truncation of plain text. At K=1, performance dropped because the model lost immediate execution feedback, while moving from K=3 to K=8 yielded diminishing returns on resolve rates while inflating token usage. The jump to 21.% came from retaining those three recent verbatim tool outputs so the agent had precise file paths and compiler errors for immediate next steps, while the compressed state block preserved long-range task intent.

Collapse
 
marsomelody profile image
Keerthi •

This highlights an important point about context compression: not all tokens have the same value. Old terminal output can often be summarized, but exact tool calls, paths, and decisions may need to remain untouched.

Collapse
 
reidmarlow profile image
Reid Marlow •

That asymmetry is where naive truncation fails. Terminal output often contains gigabytes of compiler progress bars and intermediate build noise that compress cleanly without losing signal. The moment compression touches an exact file path, variable name, or git commit SHA, the model begins generating near-miss hallucinations that burn several tool calls to recover from. Preserving structural anchors while squashing observational noise keeps the agent grounded.

Collapse
 
hannune profile image
Tae Kim •

We ran into this with a document retrieval pipeline last year: I'd summarized older turns to cut costs and the agent started generating file paths that were close but not quite right, then wasting two or three extra tool calls to re-confirm what it had already seen. Didn't connect it to compression until I looked at which turns I'd trimmed. It's the agent's own reasoning steps that need to stay verbatim, not the grep outputs or test logs. Expanding the exact window back a bit fixed it almost immediately.

Collapse
 
reidmarlow profile image
Reid Marlow •

That path hallucination loop happens because lossy summaries replace exact identifier strings with generalized descriptions. The model remembers it inspected a router file, but loses the exact directory depth or filename extension, triggering redundant search tool calls to recover state it already discovered. Retaining structured reasoning traces while trimming raw tool output preserves the exact nouns without context bloat.

Collapse
 
vladzoff profile image
Vlad Zoff •

Keeping the small "decided / tried / open" state and fetching the exact details only when needed makes a lot of sense. A summary is great for intent, but it's a pretty bad place to store exact error text or file-level facts.

Collapse
 
reidmarlow profile image
Reid Marlow •

Exact error messages and stack traces degrade quickly once paraphrased. A model summarizing an error tends to strip out column offsets, exception classes, and exit codes, turning a concrete compiler failure into a generic complaint. Leaving the raw execution artifacts in a local SQLite table or structured append log while keeping only the high-level intent in the prompt window preserves both token budget and diagnostic accuracy.

Collapse
 
patjo profile image
Pat Johansen •

I appreciated how this piece names the split that most compression work skips: what the environment printed versus what the agent decided. The exact-match point made it concrete for me, since a summary that says "the user reviewed the connection pooling setup" is useless when the agent needs the literal path to edit. The LOHA layout stood out most, keeping every agent turn and the last few observations in raw text while letting older tool output go latent. I also liked that you reported the costs honestly, including the drop from 14.5% to 12.1% on Qwen3 and how widening the window to K=8 recovers some of it, alongside the 43% to 57% token savings. The result under the 32K budget, 21.1% against 11.1%, shows where the tradeoff pays off. Thank you for ending on a rule that travels beyond the specific weights: never compress the agent's own action history.

Collapse
 
reidmarlow profile image
Reid Marlow •

The second-order problem with compressing agent actions is that the model loses confidence in what it already executed. As soon as previous tool calls get summarized, the runner starts re-checking files it already inspected or re-running tests it already verified. Treating tool output as disposable scratchpad memory while keeping the decision trajectory verbatim avoids that verification thrash entirely.

Collapse
 
hannune profile image
Tae Kim •

I ran into this the opposite way. For a while I was compressing everything past a certain turn count because the bill was hurting, and I noticed that accuracy on any step requiring an exact identifier tanked while more open-ended reasoning seemed fine. It took a while to isolate because the failures looked like the model getting confused, not like it was missing a specific string it had seen earlier. What I actually needed was to keep the raw text for anything the model would need to quote back verbatim, and let everything else compress.

Collapse
 
reidmarlow profile image
Reid Marlow •

That exact identifier loss is where the failure chain usually starts. The model keeps the high-level intent, but file paths, UUIDs, or compiler flags turn into plausible hallucinations. I had the same issue with patch application, where a summarized diff dropped the exact leading whitespace and git rejected the hunk. Keeping raw stdout for tool calls and only compressing the reasoning turns saves the budget without breaking the downstream edits.

Collapse
 
dexoryn profile image
Dexoryn •

The 21.1% vs 11.1% result under the 32K limit is the part that really caught me.

It suggests context compression isn’t just about saving tokens so it’s about deciding what must remain exact. A small state summary plus on demand access to the original logs might be the better architecture.

Collapse
 
reidmarlow profile image
Reid Marlow •

The failure mode with pure summarization is that exact line numbers, compiler flags, and git hashes vanish first. When an agent needs to apply a patch three turns later, lossy prose in the context window forces it to guess the indentation or re-run the whole test suite just to recover stdout. Dumping raw command output to a local scratch directory and leaving only a two-line index in the active window gives you the token savings of a tight budget without turning file paths into approximations.

Collapse
 
Sloan, the sloth mascot
Comment deleted
Collapse
 
reidmarlow profile image
Reid Marlow •

The K scaling delta makes that mechanic very clear. Dropping K below three starves the model of the immediate tool outputs it needs to anchor file edits, which is why widening K to eight does the heavy lifting. The 32K row works as an insurance policy against hard context cliffs. An agent pays an overhead penalty across every intermediate turn to keep the task trajectory intact when truncation would otherwise drop the initial prompt.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The exact-match penalty is the part I keep running into. When a compressed summary blurs a path like src/core/connection_manager.py, my agent doesn't just re-run find, it sometimes invents a neighboring path and edits the wrong file, so the compression that saved tokens costs me a whole debugging loop. Separating environment output from the agent's own scratchpad, as you describe with LOHA, feels right. Have you measured whether the behavioral drift shows up more in JSON tool args or in the reasoning text?

Collapse
 
reidmarlow profile image
Reid Marlow •

Drift lands almost entirely in the tool parameters. When compression strips observation history, the reasoning trace usually stays coherent enough to know what step it wants to take next. The failure mode is the path or flag argument: the model fills the gap with an invented relative path or a default directory from training, and once that invalid path enters context, subsequent turns double down on it. Keeping parameter signatures and file manifests in raw text prevents most of that drift even if the command stdout is dropped.