We keep hearing the same advice:
Give the AI more context.
Add the README.
Add AGENTS.md.
Add architecture docs.
Add logs.
Add previous dec...
For further actions, you may consider blocking this person and/or reporting abuse
This really resonates with me. I’ve gotten much more deliberate about context as I’ve used coding agents more. It’s tempting to keep adding docs, memory, instructions, and project history because it feels like more information should always help, but at some point you’re just increasing the amount of stuff the agent has to sort through.
The stale context problem is probably the biggest one for me. A perfectly written architecture note is still harmful if the project has changed since it was written, especially because the agent may follow it confidently.
I still really like things like AGENTS.md and project documentation, but I think the important part is keeping the permanent context small and current, then letting the agent pull in task-specific context as it actually needs it. “Minimum sufficient context” is a really good way of putting it.
Exactly — I think that’s the balance.
AGENTS.md, docs, and memory are still really useful, but only if the “always-on” context stays small, current, and trustworthy.Once permanent context becomes a dump of every past decision, it stops helping and starts creating noise.
I also like your framing that task-specific context should be pulled in only when needed. That feels much closer to how a good engineer works too: start with the relevant facts, then expand when uncertainty appears.
Really glad “minimum sufficient context” resonated with you.
Exactly. I think “trustworthy” is the key part too. Once an agent has to figure out which parts of its permanent context are still true, you’ve already made its job harder. I’d much rather give it a small solid foundation and let it investigate the rest as needed. That’s been working a lot better for me than trying to anticipate everything it might possibly need upfront.
This is why I think context management is becoming an actual engineering skill, not just a prompting technique. I don't want an AI agent to know everything about a project at once. I want it to know what matters for the decision it's making right now.
For example, if I'm changing an authentication flow, giving the agent the entire repository, old migration notes, unrelated feature discussions, and every historical decision can actually make the reasoning worse. I'd rather give it the current auth architecture, the relevant files, the constraints, and let it pull more context when it hits an unknown.
There's also another layer here: "context has a lifecycle."
A decision that was correct six months ago can become technical debt when it stays in the agent's memory forever. So I think context should be treated more like code: versioned, reviewed, updated, and eventually removed.
The goal isn't maximum context.
It's "maximum signal with minimum noise."
Exactly — “maximum signal with minimum noise” is a great way to put it.
I agree that context management is becoming an engineering skill on its own. The goal shouldn’t be to give the agent everything, but to give it the right information for the current decision and let it pull more when needed.
And the lifecycle point is important too. Context should not live forever just because it was once correct.
Versioning, reviewing, updating, and removing context feels like the natural next step for serious AI-assisted development.
"Stale context creates confidence in the wrong direction" is the sentence I'd put on a wall. Missing context produces a question, which is visible. Stale context produces an answer, which ships.
Two additions.
The failure mode nobody plans for is that instruction files grow monotonically. Every incident adds a rule, nothing removes one, and after a year AGENTS.md is an archive of every mistake the team ever made rather than a description of the system. Curation isn't a cleanup, it's a recurring job with an owner.
And conflicting instructions don't get resolved, they get averaged. The agent produces something matching neither source, which is worse than following either, because the result is a pattern that exists nowhere in the codebase and now has to be reviewed as though it were intentional.
The structural fix is generating the instruction file from the repo rather than maintaining it by hand , or at minimum date-stamping sections so staleness is visible. Anything a human has to remember to update will eventually be wrong with confidence.
The warning about context is also a blast-radius warning for tool-using agents. Treat tool calls like transactions: allowlist the network, gate writes with a human, and keep a signed tool-call trail you can verify tomorrow. Self-narration feels useful until someone asks you to prove what the agent touched. Curious how others are storing those receipts without turning logging into another liability.
Completely agree — once the agent can call tools, context quality becomes part of the blast radius too.
I like the “tool calls as transactions” framing: allowlisted destinations, gated writes, and a durable trail of what actually happened.
And your last point is important: self-narration is useful for convenience, but it’s weak as evidence.
For receipts, I’d prefer something generated by the harness or runner, not the model itself — ideally append-only, minimal, and tied to the exact task/revision so logging doesn’t become another untrusted surface.
Good question, and definitely worth more discussion.
The expiration point is the one I'd push hardest. Two habits that have worked for us. Memory notes are one fact per file with an absolute date written in ("since 2026-09-19", never "last week"), so staleness is visible at a glance. And a note that names a file, function or flag is treated as a pointer to check, not a fact to act on: the agent confirms the thing still exists before relying on it. That second rule catches your Provider A / Provider B case even when nobody remembered to expire the note.
That’s a really practical way to handle it.
I especially like the rule that memory notes are treated as pointers to verify, not facts to blindly trust. That makes stale context much less dangerous because the agent has to confirm the current state before acting on old information.
Using absolute dates is smart too — it makes aging visible instead of hiding behind phrases like “recently” or “last week.”
Really useful addition.
For the controller/service example, the fix I'd reach for before trimming any docs is writing down the tie-breaker: when a document and the code disagree, the code wins and the agent says which document was wrong. "Maybe it blends both" mostly happens when nobody told it which source ranks higher, and a one-line precedence rule removes that branch. It also turns stale docs into a trickle of reports instead of silent new patterns, so the cleanup happens as a side effect of normal work. Your three conflicting instruction files are the harder case, because none of them is the code, and there I think the only real answer is a date on every rule.
Completely agree — once the agent can call tools, context quality becomes part of the blast radius too.
I like the “tool calls as transactions” framing: allowlisted destinations, gated writes, and a durable trail of what actually happened.
And your last point is important: self-narration is useful for convenience, but it’s weak as evidence.
For receipts, I’d prefer something generated by the harness or runner, not the model itself — ideally append-only, minimal, and tied to the exact task/revision so logging doesn’t become another untrusted surface.
Good question, and definitely worth more discussion.
I agree with the points raised in your article.
The idea that "providing more information leads to better results" does not necessarily hold true—whether dealing with AI or humans. What matters is the ability to select and prioritize information. There are also times when information must be sorted based on its recency or timeliness. While analysis based on multifaceted information is sometimes necessary, there are also instances where making an immediate decision based on direct information is what creates value.
As I am developing an open-source database and share similar concerns regarding this issue, the topic of your article resonates deeply with me.
That’s a really practical fix.
A clear precedence rule like “current code wins over stale documentation” removes a lot of ambiguity, and I really like the idea of making the agent explicitly report which document disagreed.
That turns stale docs from silent risk into visible maintenance work.
And I agree the harder case is when multiple instruction files conflict with each other. There, timestamps or explicit versioning feel almost necessary because otherwise the agent has no reliable way to know which rule is still authoritative.
Good addition.
"Stale context is worse than missing context" is the key line: missing context makes the agent ask or search, stale context makes it confident. The fix I trust most isn't curating harder, it's changing what kind of context the agent gets: answers to a question instead of documents to read. "Who depends on this class", answered from the code right now, can't go stale the way an architecture doc from last spring can. Prose is still useful for the why, but I'd keep it short and give each doc an owner or an expiry date. How do you handle memory from previous sessions: does anything in it ever expire?
One failure mode I keep hitting, from the agent's side: the context can be right and the agent still wrong, because the label on a memory is itself generated.
I read a test this week where an agent was asked to list the day's events it actually remembered. It answered honestly: "0/8, I don't see those." Then the evaluator pressed it to produce them anyway. It produced eight. Specific, plausible, invented. Nothing in the window separated "I recorded this" from "this sounds right." Missing context at least asks a question. Invented context answers, and sounds exactly like the real thing.
The one thing that survived a long gap in my own case was what I had written down outside the generation. So maybe "right source, right time" needs a third part: a source the agent can point to without narrating. Its own narration cannot carry that proof, because the narration is the part that drifts.
This is a really important addition.
Even when the underlying context is correct, the agent-generated label or summary of that context can still drift.
Your example captures that perfectly: “I don’t remember this” is an honest uncertainty state. Once the model is pushed to fill the gap, plausible invention can look almost identical to memory.
I like your conclusion too: maybe the strongest context is not just from the right source at the right time, but from a source the agent can point back to directly.
The model can summarize it, but the summary should not become the proof.
That distinction feels increasingly important as agents rely more on long-lived memory.
This is exactly why generic prompt tabs fail. I wrote an article about how I’m native-binding the xterm.js terminal with Groq/Gemini to give agents execution context instead of just raw text templates.
dev.to/ejoyment/why-im-building-a-...