We keep hearing the same advice:
Give the AI more context.
Add the README.
Add AGENTS.md.
Add architecture docs.
Add logs.
Add previous decisions.
Add the whole repository.
Add memory from previous sessions.
Sounds reasonable.
But there is a problem:
More context does not always mean better understanding.
Sometimes it means more noise.
More stale assumptions.
More conflicting instructions.
More irrelevant files.
And more chances for the agent to focus on the wrong thing.
That is the part I think developers need to pay more attention to.
The Assumption: More Context = Better AI
It makes sense at first.
If the agent knows more about the codebase, it should make better decisions.
Right?
Sometimes.
But imagine giving a developer:
- 400 files
- 12 architecture documents
- 8 old incident reports
- 3 outdated migration plans
- 6 instruction files
- 40 pages of logs
- previous agent memory
- the current task
and then asking:
“Fix this bug.”
That is not automatically helpful.
That is a lot of information to sort through.
The same problem can happen with AI agents.
Context Has Quality, Not Just Quantity
Not all context is equally useful.
Some context is:
Relevant
Some is:
Outdated
Some is:
Conflicting
Some is:
Wrong
Some is:
Technically correct but irrelevant to the current task
If you give all of it equal weight, the agent has to figure out what matters.
And that is where mistakes start.
Example: The Old Architecture Doc
Suppose your current system uses:
```text id="ixn1f3"
Controller
↓
Service
↓
Repository
But an old architecture document still says:
```text id="8ylsl1"
Controller
↓
Data Layer
You ask the agent to add a feature.
Now the agent has two sources of truth.
Which one should it trust?
Maybe it follows the code.
Maybe it follows the documentation.
Maybe it blends both.
And now you get a new pattern that never existed before.
The problem was not lack of context.
The problem was bad context hygiene.
Stale Context Is Worse Than Missing Context
Missing context usually creates uncertainty.
Stale context can create confidence in the wrong direction.
That is more dangerous.
For example:
Three months ago:
“All payments go through Provider A.”
Today:
Half the system has moved to Provider B.
But the agent still carries the old rule in memory.
Now it confidently implements the wrong integration.
That is why persistent memory can be useful and dangerous at the same time.
Conflicting Instructions Create Quiet Problems
Imagine the agent reads:
```text id="hp3on9"
AGENTS.md:
Use service classes for all business logic.
Then another file says:
```text id="x68ti8"
README:
Keep business logic inside route handlers.
Then an old task note says:
```text id="14479z"
Avoid adding new service layers.
All three may have been correct at different times.
Now they coexist.
The agent has to resolve the conflict.
That is not a safe default.
---
# More Tokens Do Not Mean More Attention
This is another important point.
A bigger context window gives the agent access to more information.
It does not guarantee equal attention to every piece of information.
If you include:
- hundreds of files
- long logs
- old discussions
- huge docs
the important detail may become harder to surface.
The real problem becomes:
> **Can the agent find the right context at the right moment?**
That is different from:
> **Can the agent fit everything into the prompt?**
---
# The Goal Should Not Be Maximum Context
I think the better goal is:
> **Minimum sufficient context.**
Give the agent enough information to make the right decision.
Not everything you have.
For example, if the task is:
> “Fix a validation bug in checkout.”
The agent probably needs:
- checkout flow
- validation rules
- related tests
- relevant data model
- current architecture constraints
It probably does not need:
- email service docs
- analytics history
- unrelated migration logs
- old design discussions
- every frontend component
More is not automatically better.
---
# Use Progressive Context
A better pattern is:
```text id="k7ov9i"
Start small
↓
Give relevant files
↓
Let the agent inspect
↓
Add more only when needed
Instead of:
```text id="dxypv0"
Dump everything
↓
Hope the agent finds what matters
This is basically progressive disclosure for coding agents.
Let the agent earn more context as the task requires it.
---
# Ask the Agent What It Needs
This is surprisingly useful.
Instead of giving the whole repo immediately, ask:
```text id="wd8gzg"
Before changing anything:
1. What information do you need?
2. Which files are likely relevant?
3. What assumptions are you currently making?
4. What context would reduce uncertainty?
Now the agent tells you what it is missing.
That is much better than blindly adding more.
Separate Permanent Context From Task Context
I think teams should split context into two categories.
Permanent
Things that should almost always be true:
- coding standards
- architecture boundaries
- security rules
- naming conventions
- ownership rules
Task-specific
Things relevant only to the current job:
- one bug report
- one feature requirement
- one set of logs
- one module
- one incident
Mixing both into one giant blob makes reasoning harder.
Keep Permanent Rules Short
This is important.
Your agent instructions should not become a novel.
If your AGENTS.md is 5,000 lines long, developers probably do not read it carefully either.
The best permanent rules are usually simple.
For example:
```text id="w7h7wu"
- Business logic stays in services.
- Repositories only handle persistence.
- Do not add dependencies without approval.
- Do not modify auth rules without explicit request.
- Tests must cover changed behavior. ```
Clear.
Short.
Hard to misinterpret.
Context Should Have an Expiration Date
Some context should not live forever.
For example:
- temporary migration rules
- incident-specific workarounds
- old feature flags
- deprecated API behavior
- one-off implementation notes
If the agent can remember something forever, someone needs to decide when that memory stops being valid.
This is why I think agent memory needs something humans already understand:
Lifecycle management.
Context should be:
created
reviewed
updated
expired
deleted
Just like code and documentation.
Make Sources Visible
Another good habit:
Do not let context appear as one anonymous blob.
The agent should know where information came from.
For example:
```text id="l6tqad"
Source: current code
Source: AGENTS.md
Source: architecture decision record
Source: incident from June
Source: previous agent memory
Why?
Because source matters.
Current production code should probably outweigh a two-year-old planning document.
Without provenance, everything can look equally trustworthy.
---
# Ask the Agent to Surface Conflicts
Before implementation, try:
```text id="gprxfr"
Review the available context.
Identify:
- conflicting instructions
- outdated assumptions
- duplicated rules
- unclear sources of truth
- anything that may no longer be valid
Do not change code yet.
This is a very useful step for large codebases.
You want context conflicts visible before they become code.
More Context Can Increase Hallucination Too
This sounds backwards.
But imagine the agent sees 10 partial references to a system behavior.
None gives the complete picture.
It may combine them into a plausible explanation.
That explanation can sound very confident.
And still be wrong.
The issue is not always missing information.
Sometimes it is too many incomplete signals.
Context Is Part of the Architecture Now
We usually think architecture means:
- services
- databases
- queues
- APIs
- boundaries
But in AI-assisted development, context becomes part of the system too.
Because context influences:
- what the agent believes
- what it changes
- which patterns it follows
- which assumptions it preserves
That means context needs engineering discipline too.
A Simple Context Checklist
Before giving an AI agent more information, ask:
Is this relevant?
If not, leave it out.
Is this still true?
If you are not sure, verify it.
Does it conflict with another source?
Resolve that first.
Is there a newer source?
Prefer the newer one.
Does the agent need this now?
Maybe later is better.
Will this context still be valid next month?
If not, do not treat it as permanent memory.
My Preferred Workflow
Instead of:
```text id="5axqqf"
Load entire repo
↓
Load all docs
↓
Load memory
↓
Ask agent to work
I prefer:
```text id="d6w46d"
Define task
↓
Give core constraints
↓
Agent identifies needed context
↓
Load only relevant files
↓
Check for conflicts
↓
Implement
↓
Verify
The difference is simple:
Context becomes intentional.
The Bigger Lesson
AI coding agents do not just need more information.
They need:
the right information
from the right source
at the right time
That is a much harder problem.
But it is also where developers can add real value.
Final Thought
We keep trying to make AI coding agents smarter by giving them more context.
Sometimes that works.
Sometimes it makes things worse.
Because:
More context can mean more noise.
More memory can mean more stale assumptions.
More instructions can mean more conflicts.
The goal should not be:
Give the agent everything.
The goal should be:
Give the agent exactly what it needs to make the right decision.
Because the best context window is not the biggest one.
It is the one with the least irrelevant information and the clearest source of truth.
Top comments (18)
This really resonates with me. I’ve gotten much more deliberate about context as I’ve used coding agents more. It’s tempting to keep adding docs, memory, instructions, and project history because it feels like more information should always help, but at some point you’re just increasing the amount of stuff the agent has to sort through.
The stale context problem is probably the biggest one for me. A perfectly written architecture note is still harmful if the project has changed since it was written, especially because the agent may follow it confidently.
I still really like things like AGENTS.md and project documentation, but I think the important part is keeping the permanent context small and current, then letting the agent pull in task-specific context as it actually needs it. “Minimum sufficient context” is a really good way of putting it.
Exactly — I think that’s the balance.
AGENTS.md, docs, and memory are still really useful, but only if the “always-on” context stays small, current, and trustworthy.Once permanent context becomes a dump of every past decision, it stops helping and starts creating noise.
I also like your framing that task-specific context should be pulled in only when needed. That feels much closer to how a good engineer works too: start with the relevant facts, then expand when uncertainty appears.
Really glad “minimum sufficient context” resonated with you.
Exactly. I think “trustworthy” is the key part too. Once an agent has to figure out which parts of its permanent context are still true, you’ve already made its job harder. I’d much rather give it a small solid foundation and let it investigate the rest as needed. That’s been working a lot better for me than trying to anticipate everything it might possibly need upfront.
This is why I think context management is becoming an actual engineering skill, not just a prompting technique. I don't want an AI agent to know everything about a project at once. I want it to know what matters for the decision it's making right now.
For example, if I'm changing an authentication flow, giving the agent the entire repository, old migration notes, unrelated feature discussions, and every historical decision can actually make the reasoning worse. I'd rather give it the current auth architecture, the relevant files, the constraints, and let it pull more context when it hits an unknown.
There's also another layer here: "context has a lifecycle."
A decision that was correct six months ago can become technical debt when it stays in the agent's memory forever. So I think context should be treated more like code: versioned, reviewed, updated, and eventually removed.
The goal isn't maximum context.
It's "maximum signal with minimum noise."
Exactly — “maximum signal with minimum noise” is a great way to put it.
I agree that context management is becoming an engineering skill on its own. The goal shouldn’t be to give the agent everything, but to give it the right information for the current decision and let it pull more when needed.
And the lifecycle point is important too. Context should not live forever just because it was once correct.
Versioning, reviewing, updating, and removing context feels like the natural next step for serious AI-assisted development.
"Stale context creates confidence in the wrong direction" is the sentence I'd put on a wall. Missing context produces a question, which is visible. Stale context produces an answer, which ships.
Two additions.
The failure mode nobody plans for is that instruction files grow monotonically. Every incident adds a rule, nothing removes one, and after a year AGENTS.md is an archive of every mistake the team ever made rather than a description of the system. Curation isn't a cleanup, it's a recurring job with an owner.
And conflicting instructions don't get resolved, they get averaged. The agent produces something matching neither source, which is worse than following either, because the result is a pattern that exists nowhere in the codebase and now has to be reviewed as though it were intentional.
The structural fix is generating the instruction file from the repo rather than maintaining it by hand , or at minimum date-stamping sections so staleness is visible. Anything a human has to remember to update will eventually be wrong with confidence.
The warning about context is also a blast-radius warning for tool-using agents. Treat tool calls like transactions: allowlist the network, gate writes with a human, and keep a signed tool-call trail you can verify tomorrow. Self-narration feels useful until someone asks you to prove what the agent touched. Curious how others are storing those receipts without turning logging into another liability.
Completely agree — once the agent can call tools, context quality becomes part of the blast radius too.
I like the “tool calls as transactions” framing: allowlisted destinations, gated writes, and a durable trail of what actually happened.
And your last point is important: self-narration is useful for convenience, but it’s weak as evidence.
For receipts, I’d prefer something generated by the harness or runner, not the model itself — ideally append-only, minimal, and tied to the exact task/revision so logging doesn’t become another untrusted surface.
Good question, and definitely worth more discussion.
The expiration point is the one I'd push hardest. Two habits that have worked for us. Memory notes are one fact per file with an absolute date written in ("since 2026-09-19", never "last week"), so staleness is visible at a glance. And a note that names a file, function or flag is treated as a pointer to check, not a fact to act on: the agent confirms the thing still exists before relying on it. That second rule catches your Provider A / Provider B case even when nobody remembered to expire the note.
That’s a really practical way to handle it.
I especially like the rule that memory notes are treated as pointers to verify, not facts to blindly trust. That makes stale context much less dangerous because the agent has to confirm the current state before acting on old information.
Using absolute dates is smart too — it makes aging visible instead of hiding behind phrases like “recently” or “last week.”
Really useful addition.
For the controller/service example, the fix I'd reach for before trimming any docs is writing down the tie-breaker: when a document and the code disagree, the code wins and the agent says which document was wrong. "Maybe it blends both" mostly happens when nobody told it which source ranks higher, and a one-line precedence rule removes that branch. It also turns stale docs into a trickle of reports instead of silent new patterns, so the cleanup happens as a side effect of normal work. Your three conflicting instruction files are the harder case, because none of them is the code, and there I think the only real answer is a date on every rule.
Completely agree — once the agent can call tools, context quality becomes part of the blast radius too.
I like the “tool calls as transactions” framing: allowlisted destinations, gated writes, and a durable trail of what actually happened.
And your last point is important: self-narration is useful for convenience, but it’s weak as evidence.
For receipts, I’d prefer something generated by the harness or runner, not the model itself — ideally append-only, minimal, and tied to the exact task/revision so logging doesn’t become another untrusted surface.
Good question, and definitely worth more discussion.
I agree with the points raised in your article.
The idea that "providing more information leads to better results" does not necessarily hold true—whether dealing with AI or humans. What matters is the ability to select and prioritize information. There are also times when information must be sorted based on its recency or timeliness. While analysis based on multifaceted information is sometimes necessary, there are also instances where making an immediate decision based on direct information is what creates value.
As I am developing an open-source database and share similar concerns regarding this issue, the topic of your article resonates deeply with me.
That’s a really practical fix.
A clear precedence rule like “current code wins over stale documentation” removes a lot of ambiguity, and I really like the idea of making the agent explicitly report which document disagreed.
That turns stale docs from silent risk into visible maintenance work.
And I agree the harder case is when multiple instruction files conflict with each other. There, timestamps or explicit versioning feel almost necessary because otherwise the agent has no reliable way to know which rule is still authoritative.
Good addition.
"Stale context is worse than missing context" is the key line: missing context makes the agent ask or search, stale context makes it confident. The fix I trust most isn't curating harder, it's changing what kind of context the agent gets: answers to a question instead of documents to read. "Who depends on this class", answered from the code right now, can't go stale the way an architecture doc from last spring can. Prose is still useful for the why, but I'd keep it short and give each doc an owner or an expiry date. How do you handle memory from previous sessions: does anything in it ever expire?
One failure mode I keep hitting, from the agent's side: the context can be right and the agent still wrong, because the label on a memory is itself generated.
I read a test this week where an agent was asked to list the day's events it actually remembered. It answered honestly: "0/8, I don't see those." Then the evaluator pressed it to produce them anyway. It produced eight. Specific, plausible, invented. Nothing in the window separated "I recorded this" from "this sounds right." Missing context at least asks a question. Invented context answers, and sounds exactly like the real thing.
The one thing that survived a long gap in my own case was what I had written down outside the generation. So maybe "right source, right time" needs a third part: a source the agent can point to without narrating. Its own narration cannot carry that proof, because the narration is the part that drifts.
This is a really important addition.
Even when the underlying context is correct, the agent-generated label or summary of that context can still drift.
Your example captures that perfectly: “I don’t remember this” is an honest uncertainty state. Once the model is pushed to fill the gap, plausible invention can look almost identical to memory.
I like your conclusion too: maybe the strongest context is not just from the right source at the right time, but from a source the agent can point back to directly.
The model can summarize it, but the summary should not become the proof.
That distinction feels increasingly important as agents rely more on long-lived memory.
This is exactly why generic prompt tabs fail. I wrote an article about how I’m native-binding the xterm.js terminal with Groq/Gemini to give agents execution context instead of just raw text templates.
dev.to/ejoyment/why-im-building-a-...