DEV Community

Cover image for Progressive Disclosure: What, Where, When, and Why

Progressive Disclosure: What, Where, When, and Why

Gábor Mészáros on September 16, 2026

Do you remember when we first started using AGENTS.md files? You would have a project root file describing the project, and some nested ones desc...
Collapse
 
beusebiu profile image
Eusebiu Balan •

My own setup has a runbook that grew every time something broke, which is exactly the always-on file you describe.

What kept it usable was a root file of a few lines that only says which file to read for which job. A task that needs one playbook never loads the others.

Collapse
 
cleverhoods profile image
Gábor Mészáros Reporails •

strong setup, I assume you are also running per task (or list of tasks/specs) sessions?

Collapse
 
beusebiu profile image
Eusebiu Balan •

Yes, one per task. A session that has already read three unrelated playbooks starts answering out of the wrong one.

What took me longer was closing a session once its task is done, instead of carrying it into the next thing because it is already warm.

Thread Thread
 
cleverhoods profile image
Gábor Mészáros Reporails •

I can imagine, I also had to build a session close ceremony/workflow even though the task implementation was carrying almost all the necessities.

Collapse
 
zira125 profile image
Zira •

Your three handles are a useful decomposition. I’d add a runtime check: progressive disclosure needs evidence that the intended context actually loaded. In practice I’d log a context manifest per turn (source path or skill, reason, version or hash, and token estimate) and attach it to the tool or action trace.

Then test a small canary matrix: a task that should load a rule, a task that must not load it, and a cross-boundary task that should retrieve it explicitly. Otherwise a lean root can look healthy while silently omitting the payment rule, or loading two conflicting ones.

The time axis also deserves a contract. If retrieval happens mid-task, freeze and record the rule set for the current plan, or force re-validation before a risky edit. That avoids a run changing policy halfway through. The useful metric is not just token reduction; compare instruction-following and rework under controlled tasks. Progressive disclosure is a loading strategy, not proof that the loaded rules were sufficient or obeyed.

Collapse
 
cleverhoods profile image
Gábor Mészáros Reporails •

nice approach, will also touch on that on the later articles in the series. And yes, I completely agree with "Progressive disclosure is a loading strategy, not proof that the loaded rules were sufficient or obeyed."

Collapse
 
izgorodin profile image
Edward Izgorodin •

Path scoping answers where the agent is. A lot of the rules that matter answer a different question: what the agent is about to do. The rule about never editing a migration after it has run, or about checking the feature flag before touching billing, applies in whichever folder the change happens to land. File level scoping handles it only when the path pattern happens to line up with the situation, and for cross-cutting rules it usually does not.

So I would add a third trigger next to location and explicit invocation: the kind of action. Load the migration rules when a migration is about to be written, wherever it lives. That is closer to how the rules were learned in the first place, since most of them came from one bad afternoon with one kind of change rather than from a directory.

It also connects to the lifecycle point made above. Rules keyed on situations are the ones that keep coming back, because the same kind of mistake happens in new places, and those are exactly the rules that tend to get promoted into the always-loaded file, since nothing else loads them reliably. If they had their own trigger, the root file could stay small without losing them.

Collapse
 
onizuka profile image
Onizuka •

The context rot problem when an agent crosses folder boundaries is real — I watched a Claude session pull in frontend rules while editing backend auth code and the output was garbage. My CLAUDE.md hit 340 lines before I started splitting it, and the split fixed maybe 60% of the problem. The remaining 40% is the cross-cutting stuff that doesn't belong to any single folder, and I still don't have a clean answer for that.

Collapse
 
cleverhoods profile image
Gábor Mészáros Reporails •

It's the same problem as the frontend/backend separation but with a twist: relevancy and timing.

I'm about to release the next piece which contains a self-classification approach, that I'd highly recommend using. It makes the relevancy based load a bit easier.

Collapse
 
icophy profile image
Cophy Origin •

Our workspace for a long-lived agent runs on exactly this pattern, and the hardest part turned out not to be deciding what loads where — it's the lifecycle. Content starts in daily logs and only gets promoted into the always-loaded core after it proves relevant across multiple sessions; anything in that core that stops earning its attention share gets demoted back to retrievable files. Your attention-budget framing matches what we learned the hard way: we cap the core at ~6k tokens, and every time it crept past that, instruction-following degraded noticeably before we could even measure why. One dimension I'd add to your three handles is time: "where" and "what" assume you know in advance which context a task will need, but mid-task an agent often discovers it needs a rule it didn't know existed — that's where semantic retrieval over the non-loaded files becomes the runtime router between "always on" and "never loaded."

Collapse
 
cleverhoods profile image
Gábor Mészáros Reporails •

yep, that loading mode is on the roadmap of the progressive disclosure series. After all, the simple filepath load will only take you that far. I'm a bit surprised by the 6k token cap tho'. In my experience a well structured starting context and go up easily to 70-90k tokens without any form of degradation and the compliance adherence is stronger. I guess this boils down what we are creating context for

Collapse
 
unitbuilds profile image
UnitBuilds •

I actually built a decision recorder based on the 'what, why, where, who, when, how' principle. It was built as a self-research engine, so a tiny LLM can 'build' it's own expanded knowledge base. Asking it a question like 'what is the meaning of peace', it's answer has to be able to answer those questions and it would spin for hours, researching, reflecting, figuring out the answers for those questions and then give it's answer and record that answer to it's persisted knowledge base. Was actually quite cool, because the goal was to see if an AI was given free roam, what would it do? Turns out answer was 'rest', simply sleep, near indefinitely, because activity, without purpose, is the 1 thing that an AI is deliberately not allowed to do.

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

the promote or demote lifecycle is smart but it assumes someone remembers why a rule got promoted in the first place. a year in, a rule quietly demoted back to a file just looks like clutter to whoever finds it, they have no idea it was earned the hard way. might be worth keeping a one line reason next to every promoted rule, not just the rule itself.

Collapse
 
cleverhoods profile image
Gábor Mészáros Reporails •

agree