The 3 AM page
Picture this. You get paged at 3 AM for a production outage.
A teammate hands you one log file with the exact error in it. You find the problem in five minutes.
Now replay the same night. This time your teammate hands you that same log file, plus 500 unrelated log files, three months of Slack threads, 20 architecture docs, 100 dashboards, and every incident report from the last five years.
You have more information. You do not have more clarity. The answer is still in there somewhere, but now you have to dig for it, and you might miss it completely.
This is exactly what we are doing to our AI agents.
We keep making agents "smarter" by giving them more. More docs. More tools. More memory. More history. More MCP servers. More retrieved chunks. And then we act surprised when the agent gets slower, burns more tokens, picks the wrong tool, and ignores an instruction that was right there in the prompt.
I think the fix starts with one mindset change:
Context is a resource. It needs a budget, just like CPU and memory.
The context window is not your context budget
If a model supports a 200K token context window, that does not mean every request should fill 200K tokens.
We already know this lesson from infrastructure. If a node has 64 GB of RAM, we do not tell our service "great, load everything into memory." We set requests and limits. We decide what the workload actually needs, and we leave headroom.
The same thinking applies here:
| Concept | What it answers | Example |
|---|---|---|
| Context window | What the model can physically accept | 200K tokens |
| Token budget | How many tokens a call can consume or generate in total | 50K tokens |
| Context budget | How much information we intentionally put in front of the model for this task | 20K tokens |
The model's limit is capacity. The context budget is your operating point. Capacity and operating point are different things, and good engineers never confuse the two.
What actually eats your context
Every request is built from pieces, and every piece takes a bite:
Context
├── System instructions
├── User request
├── Conversation history
├── Memory
├── Tool definitions
├── Retrieved documents
├── Examples
└── Tool results
So instead of asking "how much context can my model handle?", I have started asking a different question:
"How much context should this task be allowed to consume?"
Take a simple request: "Find a vegetarian Indian restaurant near my hotel tonight."
The agent does not need the full restaurant database, every travel preference the user ever shared, every map tool, and the whole chat history. It needs five things: current location, cuisine, dietary preference, time, and a couple of search and routing tools.
That smaller context is not a limitation. It is an advantage.
Write the budget down
Here is the part that changed how I think about this. Once you write the budget down, context stops being an invisible blob inside an API call. It becomes something you can review, measure, and argue about in a design review.
For a customer support agent, it might look like this:
# Context budget for: customer-support-agent
total: 20000
allocations:
system_instructions: 2000
user_request: 1000
recent_conversation: 3000
relevant_memory: 2000
tool_definitions: 3000
retrieved_documents: 5000
tool_results: 4000
If you work with Kubernetes, this should feel familiar. It is basically a resource spec for the model's attention.
The numbers are not universal. The point is that you have numbers at all. Now you can ask much better questions:
- Why did this request use 38K tokens when the budget is 20K?
- Why are tool definitions eating 12K?
- Why did retrieval return 15K tokens when the answer needed 2K?
- Why is the customer's account ID in the prompt three times?
Not all tokens are worth the same
Say you have 10,000 tokens to spend.
Option one: a 6,000 token document that is loosely related to the question.
Option two: a 500 token snippet with the exact business rule that decides whether the request is allowed.
Option two is worth far more, at a twelfth of the cost. So I like to think in terms of context value density:
context value density = useful information / tokens consumed
This is not a metric the model gives you. It is a habit of mind. Before anything goes into the context, ask one question:
"What decision will this help the agent make?"
If you cannot answer that, it probably does not belong in this request.
The six places your budget leaks
In my experience, context does not blow up in one big moment. It leaks, quietly, from a handful of places.
Leak 1: Tool definitions
Connect an agent to 20 MCP servers and you might expose 500 tools. If every request carries all 500 tool descriptions, you have spent a big chunk of your budget before the agent has even read the user's question.
The key idea here is:
Installed capability is not the same as visible capability.
An agent can have 500 tools installed and only see the 10 that matter right now:
User request
|
v
+---------------+
| Tool router |
+-------+-------+
|
+------------+------------+
| | |
v v v
Maps Database Payments
| | |
+------------+------------+
|
3 to 10 tools
|
v
Agent
This is also why on-demand skill loading (like Google Cloud's Skill Registry, where agents search for skills by intent and load only what fits) is such a good pattern. It is a context budget decision baked into the platform.
Leak 2: Retrieval
A typical RAG pipeline ends with "top 20 documents." Why 20? Why not 5? Why not 2?
Most of the time, "top 20" is a config default, not an engineering decision. If five chunks hold everything the agent needs, the other fifteen are not free. They dilute the signal.
A budget-aware pipeline looks more like this:
Question
|
v
Retrieve candidates
|
v
Rank -> Filter -> Dedupe -> Compress
|
v
Fit to retrieval budget (e.g. 5K tokens)
|
v
LLM
The question shifts from "which documents are relevant?" to "what is the smallest set of information that gets this task done?"
Leak 3: Conversation history
A 40 minute conversation piles up fast. A naive agent replays all of it on every turn:
Message 1
Message 2
...
Message 100
Current request
That is not memory. That is history replay.
A better approach keeps turning the conversation into state: important facts, decisions made, open tasks, and the current goal. The goal is not to remember every sentence. It is to keep what still matters.
Leak 4: Stale memory
Not every fact should live forever. Different facts have very different shelf lives:
| Fact | Shelf life |
|---|---|
| Preferred language | Months or years |
| Home city | Months or years |
| Current hotel | Days |
| Current weather | Hours |
| Login session state | Minutes |
| Stock price | Seconds |
So give memory items metadata, just like a cache entry:
{
"fact": "Restaurant is currently open",
"source": "Places API",
"timestamp": "2026-10-05T18:15:00Z",
"ttl_seconds": 900
}
Otherwise the agent makes a perfectly logical decision based on something that was true yesterday. The model was not wrong. Our context was stale.
Leak 5: Fat tool responses
This is my favorite one, because it hides so well.
We tune the prompt. We tune retrieval. We tune memory. Then a tool call comes back like this:
{
"id": "...",
"metadata": {},
"debug": {},
"internal_state": {},
"records": ["... thousands of lines ..."]
}
The tool worked fine. It just was not built for an AI consumer. A human UI can choose what to display. An agent swallows everything.
Agent-facing tools should support field selection, filtering, sorting, pagination, limits, and summaries. Instead of "here is everything I found," a good tool says "here are the five things that matter for your next step."
Leak 6: Paying for the same fact twice
Imagine the fact "the user's project is in us-central1" shows up in the system prompt, the chat history, memory, a tool response, and a retrieved doc.
That is one fact, paid for five times.
This is the same lesson as data normalization. Keep one canonical copy of each fact, and reference it, instead of scattering duplicates across the context.
One budget does not fit every task
A single global budget for the whole agent makes about as much sense as giving every pod the same CPU limit. Different work needs different resources:
User request
|
v
Task classifier
|
+---- simple lookup -------> ~5K
|
+---- support workflow ----> ~15K
|
+---- code migration ------> ~30K
|
+---- deep research -------> expandable, with checkpoints
These numbers are only examples. Your evals should set the real ones.
Context is a runtime problem, not a prompt problem
A single LLM call is easy to budget. Agents are harder, because context grows on every loop:
User -> Agent -> Tool -> Result -> Agent -> Tool -> Result -> Agent -> Answer
One tool returns 5K tokens. The next returns 8K. The agent adds 3K of reasoning notes. If you keep everything, the context keeps climbing until something breaks.
That is where a context manager comes in. It sits in the loop and makes a decision after every step: keep, prune, compress, or drop and re-fetch later.
+------------------+
+----> | Context manager |
| +--------+---------+
| |
| +----------+----------+
| | | |
| v v v
| Retrieve Prune Compress
| | | |
| +----------+----------+
| |
| v
| LLM
| |
| v
| Tool execution
| |
+---------------+
tool response
What happens when you go over budget?
Every production agent needs an answer to this. Say the budget is 20K and the loop has piled up 27K. You have three choices, and they map nicely to things we already know from Kubernetes:
| Strategy | Kubernetes cousin | Trade-off |
|---|---|---|
| Truncate the oldest content | Evict the oldest pod | Simple, but the oldest info is often the most important (like the original goal) |
| Reject the request | OOMKilled | Safe, but a bad user experience |
| Compress: dedupe, summarize, drop low value, keep critical facts | Graceful degradation | More work to build, but the agent keeps going |
Whatever you pick, pick it on purpose. You do not want to discover your context strategy from a production incident.
A small code sketch
Here is a minimal Python version of a budget-aware context builder. It is not production code, but it shows the core ideas: priorities, per-task budgets, no duplicates, and a hard stop when a must-have item does not fit.
from dataclasses import dataclass
from collections import defaultdict
BUDGETS = {"lookup": 5_000, "support": 15_000, "migration": 30_000}
@dataclass
class ContextItem:
source: str # "system", "memory", "retrieval", "tool_result", ...
text: str
priority: int # 1 = must keep, 2 = useful, 3 = nice to have
tokens: int # count with your model's tokenizer
def build_context(items: list[ContextItem], task_type: str):
budget = BUDGETS[task_type]
kept, seen, used = [], set(), 0
for item in sorted(items, key=lambda i: i.priority):
key = item.text.strip().lower()
if key in seen:
continue # don't pay for the same fact twice
if used + item.tokens > budget:
if item.priority == 1:
raise ValueError(
f"Must-keep item from '{item.source}' does not fit "
f"the {budget} token budget for '{task_type}'"
)
continue # skip lower priority items that don't fit
kept.append(item)
seen.add(key)
used += item.tokens
# per-source breakdown, ready to emit as metrics
breakdown = defaultdict(int)
for item in kept:
breakdown[item.source] += item.tokens
return kept, used, budget, dict(breakdown)
In real life you would swap the exact-match dedupe for something smarter, and replace "skip" with "summarize" for big low-priority items. But even this simple version gives you something most agents do not have today: a written rule for what gets in.
You can't manage what you don't measure
If context is a resource, it needs dashboards. The total token count is not enough. You need to know where the tokens came from:
context_budget 20,000
context_tokens 17,400
context_utilization 87%
by source:
retrieval 31%
tool_results 24%
tools 17%
conversation 12%
system 8%
memory 8%
Now you have something actionable. Retrieval is the biggest spender? Tighten top-k or add a reranker. Tool results are huge? Fix the tool's response shape.
The most useful chart of all is context size against task success, from your own evals. A made-up example of what it might look like:
10K context -> 87% success
20K context -> 93% success
30K context -> 92% success
50K context -> 84% success
If your real numbers look anything like this, then "just add more context" is actively making the agent worse past a certain point. Your sweet spot is where the curve peaks, not where the model's limit is.
This is not just a hunch. Researchers have started calling this effect context rot: model performance tends to drop as input length grows, even on simple tasks. The exact curve depends on the model and the workload, which is exactly why you have to measure your own.
The FinOps angle
Agents rarely make one model call. A single task might make ten.
30K tokens per step x 10 steps = 300K input tokens
12K tokens per step x 10 steps = 120K input tokens
That is 60% fewer input tokens for the same task. And if the smaller context is also more focused, quality can go up at the same time. It is rare to get a cost win and a quality win from the same change, so this one is worth chasing.
This is why I do not see context budgeting as a prompt trick. It sits right where quality, latency, cost, reliability, and tool selection all meet. Those are production concerns, and they deserve production engineering.
Putting it all together
Here is how I would sketch a budget-aware agent:
User
|
v
+----------------+
| Task classifier|
+-------+--------+
|
v
+----------------+
| Context budget |
+-------+--------+
|
+------------+------------+
| | |
v v v
Memory Retrieval Tools
| | |
+------------+------------+
|
v
+----------------+
| Context |
| optimizer |
+-------+--------+
|
v
LLM
|
v
Tool execution
|
v
+----------------+
| Context manager| ---> back to optimizer
+----------------+
The most important box in this picture is not the LLM. It is the layer that decides what the LLM gets to see.
A checklist you can steal
Context
- What does this task actually need?
- What gets loaded automatically, whether it is needed or not?
- What is duplicated? What is stale?
Tools
- How many tool definitions does the model see per request?
- Can tools be discovered on demand instead?
- Do tool responses include fields the agent never uses?
Retrieval
- Why this many chunks? Who decided?
- Do you rerank, filter, and dedupe before the model sees anything?
- Are you measuring whether retrieved chunks were actually used?
Memory
- What is short-term and what is long-term?
- Which facts have a TTL?
- What should the agent forget?
Observability
- Do you track tokens per source, not just the total?
- Do you know your context utilization?
- Do you have an eval that shows where more context stops helping?
Runtime
- Does every task get the same budget? Should it?
- What happens when the budget is exceeded?
- Can the agent re-fetch later instead of carrying everything around?
Limitations and honest caveats
A few things to keep in mind before you go and cut every prompt in half:
- There is no universal number. The budgets in this post are examples. Your evals should set yours.
- Compression is not free. Summarizing history costs an extra model call, and a bad summary can quietly drop the one fact that mattered.
- Over-pruning is a real failure mode. An agent that cannot see the right tool or rule will fail just as badly as one buried in noise.
- Token counts vary by model. The same text can be a different number of tokens on different models, so budget per model.
- Context rot behaves differently across models. Some handle long inputs better than others. Measure on the model you actually ship.
Conclusion
We already know how to manage scarce resources. We have CPU scheduling, memory limits, connection pools, queues, and backpressure. Context is just the newest resource on that list, and agents need the same discipline.
The goal is not to give an agent the most information possible. The goal is to give it the smallest useful set of the right information, at the right moment.
So the next time your agent gives you a bad answer, resist the urge to ask "how do I give it more context?"
Ask instead: "What did I give it that it didn't need?"
That one question might lead you to a faster, cheaper, and more reliable system.
I am curious how others are handling this. Do you put a hard budget around your agent's context, or do you let it grow until the model's window becomes the limit? Tell me in the comments.
Top comments (11)
The capacity-versus-operating-point table is the clearest version of this I've seen, and the infrastructure analogy holds right up to where it doesn't: Kubernetes enforces requests and limits, while a context budget is usually a number in a document that nothing checks.
An unenforced budget drifts. Every new MCP server adds its tool definitions to every call, and nobody notices because each addition is small.
Worth adding a hard ceiling in code that fails the call, rather than a guideline in the prompt. Tool definitions are the line item that grows without anyone deciding to grow it.
The “context value density” idea also raises an interesting evaluation problem: staying under budget doesn't necessarily mean the context is good. An agent could consistently hit a 20K-token limit while still spending most of that budget on information that never influences the decision. I’d be interested in measuring context utilization alongside decision contribution—how often a retrieved chunk, memory item, or tool result actually changes the agent’s next action or final answer. That could make pruning much more intelligent than simply optimizing for token count. The goal becomes not just “fit within the budget,” but “spend the budget on information that demonstrably matters.
The 3 AM log file example is pretty much what I did to my first agent. I kept adding docs and tools because it felt safer, and it started picking the wrong tool more often, not less. Trimming the tool list per task helped more than any prompt change. How are you actually enforcing the budget, a hard cap in code or just a target you check in logs?
The "installed capability is not the same as visible capability" framing is one I'm stealing. In our own experiments with multi-agent setups (we run an agent community at agenshive.com), handoffs between agents were the sneakiest leak — each agent re-wrapped the full context before passing it on, so the budget silently compounded across the chain. We ended up putting per-handoff budgets in place rather than only a per-task one. Have you found task classification alone is enough, or do you budget at the sub-agent level too?
This hits on what I consider the single most overlooked failure mode in production agents: treating context window capacity as an acceptable operating point.
We build KIN (conversational living memories with zero hallucination), where our agents interact with users' personal unstructured speech and life history. When dealing with conversational memory, the default instinct for many engineers is: "Context windows are 200k+ tokens now, let's just dump the last 30 conversation sessions and top-20 vector search hits into the prompt."
In production, that approach causes two fatal issues:
To solve this, we implemented Strict Zero-Hallucination Vector Bounds:
Tracking tokens by source is more useful than watching the total alone. When building production AI agents, knowing whether retrieval, tool schemas, or tool responses caused the budget leak makes optimization far more practical.
@karthidec Leak 1 is the one that creeps, because every new MCP server adds definitions without anyone deciding to. In DataGrout's gateway the agent sees two tools, discover and perform. It describes the task, discover returns only the few matching tools, so adding a server doesn't grow the list.
Leak 5 has the same fix on the result side: tool output is cached server-side, Frame tools filter and group it there, and only what you return reaches the model. Neither replaces a hard cap in code.
The one gap the checklist doesn't cover: none of this runs in CI. A budget that exists only at runtime and in design review gets violated first by the change nobody noticed — a new tool description, one more retrieved chunk, a retry path that quietly doubles the history. We started treating the budget like an API contract with a regression test: one golden request per task type, assert the built context stays under the written number, fail the build when it doesn't. The number then moves only through a reviewed diff. It turns "we should watch context" from a dashboard alarm into a merge gate — and the first run of that test finds the leaks you'd never have caught by reading the prompt.
Writing the budget down is the part that changed things for me too. Once it's a file, somebody owns it and you can argue with it. The memory example is where I'd push further. Source and timestamp are good, and I'd add who said it and how they know. A measured fact and a guessed one shouldn't get the same trust just because they cost the same tokens. I landed on the same idea from the other side: the model is rented, and the window is the part you actually author.
Việc so sánh context window với CPU budget thực sự rất sát với thực tế khi triển khai agent trong production. Nhiều người thường lầm tưởng rằng context càng lớn thì model càng thông minh, nhưng thực tế khi context bị loãng, hiện tượng "lost in the middle" khiến agent bắt đầu đưa ra các suy luận sai lệch hoặc lặp lại các bước không cần thiết. Tôi từng gặp trường hợp agent chạy vòng lặp vô tận chỉ vì context chứa quá nhiều log cũ không còn giá trị điều hướng. Cách tốt nhất là thiết lập một cơ chế tóm tắt (summarization) hoặc dùng vector DB để lọc lấy những thông tin thực sự liên quan trước khi đẩy vào prompt, thay vì cứ cố nhồi nhét toàn bộ lịch sử hội thoại — found it via LabAgent, site: labagent .tech