1. Hook + thesis
At some point, Claude Code started to "break" on me.
To be precise, what broke wasn't Claude Code itself. What broke was the way I was using it. Firing off dozens of agents at once, pasting giant logs and screenshots straight into the conversation, keeping long sessions alive by stitching them together with /compact — the heavy, fan-out-centric style that Claude Code's ultracode opt-in naturally produces. Keep that up and, one day, tool calls start leaking into the conversation as raw text, or the model goes quiet after thinking and never answers at all.
If I had to name the one factor most consistently present, it's context bloat. The bigger the conversation grows, the more every turn costs to resend, the closer you drift to the conditions that trigger corruption, and the less stable the responses become.
Context Drop is a small desktop tool I built to cut exactly one habit at the root: pouring raw context directly into the main conversation. Screenshots, logs, JSON — instead of landing in the main conversation, they're read by an isolated subagent (a worker AI that runs with its own conversation history, separate from the main one), and all that comes back is a compact result. The main context stays light.
This post first lays the groundwork — why heavy usage breaks things and how to choose between model, effort, and ultracode — and then explains what Context Drop is, how to use it, and how to install it, step by step.
2. Background 1 — why it broke (ultracode and context bloat)
What "ultracode" actually is
Let me get the term straight first. ultracode is a Claude Code opt-in. You turn it on by putting the keyword ultracode in a prompt, or by switching it on for the whole session. While it's on, the model orchestrates multi-agent Workflows for every substantive task, and it treats token cost as not a constraint (tokens are the units an AI uses to count its input and output, and what you are billed by).
That second part is the important one. A model that is told cost doesn't matter will, quite reasonably, fan out wide: split the work across many agents, then add more agents to verify and refute the first ones. Our internal record has a good example: a 77-agent adversarial review in one release, which the record describes as launched at the author's discretion — that is, with ultracode on — rather than mandated by any rule .
So ultracode is not an always-on mode defined by our internal design records. It's a Claude Code switch that, by design, produces huge fan-outs — and with them, a huge amount of context and tokens. That accumulated context was the common factor in the failures below (a correlation, as I note later).
Symptom 1 — tool calls mutating into "count / court"
The nastiest one was the tool-call corruption on Opus 4.8. A tool call that should have become a structured tool_use instead leaks into the conversation as raw text. The internal control tag mutates into real English words like count / court / call, and the parser gives up with "tool call could not be parsed" .
The firing conditions are telling — Opus 4.8/4.7, a very long context (1M sessions), right after /compact, non-ASCII (especially Japanese) arguments, three or more MCP servers, long tool arguments, several tools invoked at once . In short, the heavier, longer, more multilingual, and more packed the state, the more likely it is.
Worse, once the history is broken it reinforces itself. Presumably the model imitates the broken output; either way, retrying in the same session only reinforces it — it didn't self-heal in our cases. The fix is not to "go forward" (--continue) but to "go back": /rewind, then /compact, and ultimately /clear .
To be honest about the evidence: the link "context bloat → corruption" is recorded at the level of correlation, not proven causation, and hooks (handlers that run automatically at set moments) cannot run /compact for you automatically. But the correlation was strong enough to change how I work.
Symptom 2 — re-inflation right after compact, silence after a marathon
/compact is the feature that compresses a conversation to make it lighter. But we had a mechanism that injected a 19–32 KB digest at session start — and it also fired right after /compact. The result: the conversation snapped back toward its original size immediately after the compact, cancelling the shrink and re-creating the very firing condition above .
Its other face is total silence after a long ("marathon") session. After a marathon session, the model can go completely silent. Throw the same request at a fresh session and it answers instantly — the only difference being the context that had piled up in the old session .
"Just run everything at the maximum setting" didn't work
At this point you might think: if it's hard, why not just use Claude Code at its top settings? If a bigger model, higher effort, and more agents through ultracode make it smarter, then as long as money doesn't matter, you could run everything at the smartest setting from the start.
We actually ran something close to that. And that is exactly where the problems came from. Look again at the trigger conditions for the breakage above: the top Opus model, a long 1M context, a multi-agent fan-out. Every one of them is "the maximum setting." The harder we ran at maximum, the more the conversation bloated, and in that bloated conversation, tool calls broke and responses stopped.
So before it was a cost problem, it was a quality problem. The evidence for causation stops at correlation, but our judgment is that spending money to pile up context was itself what lowered quality. We thought we were buying quality with money; in fact we were spending money to lower it. The answer was not "use an even higher setting" but "decide what you keep out of the main conversation."
Why a heavy context also burns money
It isn't only corruption. Context bloat hits cost, too. Every turn, Claude Code re-sends and re-bills the entire conversation so far plus all tool results. By the estimate on record, reading a 20k-token log once means paying for about 85k tokens over the next 30 turns. A 200-token digest costs about 0.9k over the same span — roughly a 100x difference. And large reads bring compaction sooner, nudging you back toward the corruption risk above .
As an extreme example: one blog article processed under full heavy-gate fan-out consumed about 800k tokens including subagents, with 289k in the parent context .
The crux — "delegation ≠ saving"
Here's one counterintuitive fact worth nailing down. "Just offload heavy reads to a subagent and it gets cheaper" — that's only half true.
A subagent's reads are billed too. Isolation protects the parent context, but it does not make the work free. Delegation moves where the tokens land; it isn't a reduction .
So what is the value of delegation? Not polluting the parent conversation. Keep the raw bulk data out of the parent by isolating it, and the parent context stays light, stays away from the corruption trigger conditions, and doesn't inflate the resend cost of every subsequent turn. That idea — "isolate to protect the parent" — became Context Drop's design philosophy directly.
3. Background 2 — model, effort, and ultracode: what for, when which, how to switch
The countermeasures against corruption and bloat boil down to two levers: a context diet (don't over-read, don't over-paste) and model-tier routing — choosing how much model you spend on each piece of work . Context Drop is a tool for the former, but knowing the latter stabilizes your operation, so let me lay it out here.
In one line
Effort is "how deeply one agent thinks." ultracode is "how many agents split up the work and verify it." The two are independent — you can turn each one up or down on its own.
The three knobs
| knob | what it decides | values | my current setting |
|---|---|---|---|
| model | the ceiling on how smart the model is | Opus / Sonnet / Haiku / Fable | opus[1m] |
| effort | how much it thinks in one response (depth) |
low / medium / high / xhigh / max
|
medium (effortLevel in ~/.claude/settings.json) |
| ultracode | whether to parallelize and cross-check across multiple agents (breadth, certainty) | on / off | on only for the turns that need it, via the keyword |
What each effort level means
My rough sense of each level (a rule of thumb, not documented behavior):
| effort | way of thinking (my rough sense) | suited for (my rough sense) |
|---|---|---|
low |
shallowest reasoning | wording fixes, summarizing grep results, mechanical replacements |
medium |
standard | everyday implementation and investigation (my current default) |
high |
works through the steps and alternatives carefully | changes that touch a boundary (billing, auth, DB), hard bugs |
xhigh |
thinks quite deeply | design decisions, reasoning about race conditions and concurrency |
max |
the deepest setting | a root cause that keeps coming back no matter how often you fix it, the final call on an important design |
A caveat that matters: the higher the effort, the slower the answer and the more tokens it uses. And it only improves accuracy for problems that can be solved by thinking. If the real cause is missing information, raising effort won't fix it — what helps then is widening the investigation (which is where breadth, i.e. ultracode, comes in).
Combination patterns
| pattern | effort | ultracode | when to use it |
|---|---|---|---|
| A. Everyday | medium |
off | normal implementation, questions, small fixes |
| B. Think it through alone |
high–max
|
off | a hard bug whose cause is probably in one place; design thought experiments. Cases where thinking deeper beats parallelizing |
| C. Wide and shallow | medium |
on | mapping the blast radius of a change, migrating many files, exhaustively checking "is there any X?". Each item is easy, but there are many of them |
| D. All out |
high–xhigh
|
on | adversarial review before merge (billing/auth), production incidents, comparing design approaches |
| E. Mixed (inside a Workflow) | set per stage | on | e.g. the discovery stage at low, the verification/judging stage at xhigh. In my experience, the easiest way to balance cost and accuracy |
A note on E. Inside a Workflow script, each agent can override effort:
agent(prompt, { effort: "low" }) // 'low' | 'medium' | 'high' | 'xhigh' | 'max'
So the efficient split is: stages with many simple items run shallow; the few stages that make critical judgments run deep. When ultracode is on, this is the split I want the workflows built with. It's also the cheapest brake on the problem from §2 — ultracode, left alone, fans out hugely.
Which model — the tier rules
The model knob has its own decision rules. Summarizing our internal operating rules :
| tier | model | suited for |
|---|---|---|
| light | Haiku | mechanical work (bulk replace, rename, index updates, small docs) |
| mid | Sonnet | standard implementation, bug fixes, test writing, QA |
| top | Opus | design, review, adversarial verification, skeletons of large specs |
| max | inherited from the parent session (e.g. Fable 5) | risk-floor calls, task-decomposition decisions, design-level judgment |
And the rules for picking one, so you don't dither:
- Risk floor: if billing, auth, production, or anything irreversible is involved, use the top tier plus human confirmation.
- Size: small and mechanical → light; multiple files or new design → top for the skeleton.
- Task type: implementation, bug fixes, tests → mid; review, adversarial verification, design → top.
- When in doubt, go higher: the loss from a quality incident outweighs the token savings.
Model tier and effort are two different axes: the tier sets the ceiling, effort sets how much of it you use per response. A boundary change (billing/auth) usually justifies both a top tier and high effort; a bulk rename deserves neither.
How to switch
Effort
-
At launch:
claude --effort high— "Effort level for the current session" (confirmed viaclaude --help; valueslow/medium/high/xhigh/max). -
During a session: run
/effortand set the level you want. Verified in this project — the run that setmediumprintedSet effort level to medium (saved as your default for new sessions). Note that it also saves the level as your default for new sessions. -
Standing default: change
"effortLevel": "medium"in~/.claude/settings.json. -
Inside a Workflow: override per agent with
agent(prompt, {effort: ...}), as above.
Model
-
At launch:
claude --model …(listed inclaude --help). -
During a session:
/model. -
Standing default: the
modelfield in~/.claude/settings.json(e.g.opus[1m]= Opus with the 1M-token context window). -
Speed:
/fasttoggles Opus fast mode — faster output from the same Opus model, not a drop to a smaller one.
ultracode: put the keyword ultracode in the prompt for the turn that needs it, or switch it on for the session.
Note:
/effort,--effort,effortLevel,/model,opus[1m],/fast, and ultracode are features of the Claude Code harness, not something the internal design records cited here invented. What those records define is the decision rules for "which tier, when."
A naming trap
The high in /code-review high (and its siblings low … max) is not the session's reasoning effort. It controls how broadly the review picks findings. The names look alike; they are different things. Likewise, /code-review ultra is a paid, deep multi-agent review that runs in the cloud when you trigger it — a separate mechanism from ultracode.
Recommended operation
Leave the default at A (medium / ultracode off) and switch deliberately:
- You want one hard point thought through → raise effort only (B).
- Nothing may be missed → add ultracode (C).
- It absolutely cannot fail — before a merge, during a production incident → raise both (D).
Matching model/effort/breadth to the work is one habit; keeping the context light is the other. Only when both are in place does heavy usage stop being so fragile. Context Drop builds the latter into the tool and its workflow.
4. What Context Drop is
Context Drop is a local-only desktop menu-bar app: no telemetry, and Context Drop itself sends nothing over the network. It does one thing — keep raw context out of the main conversation and divert it to an isolated subagent.
The core flow is this:
raw context → Context Drop packet → isolated subagent → compact result → main conversation
And here is the flow it deliberately does not support:
raw context → main conversation → subagent ← forbidden
The difference is whether raw data ever passes through the main conversation, even once. The former never pollutes the main. The latter etches raw data into the main, re-bills it on every subsequent turn, and raises the corruption risk. Context Drop draws that line by giving the main conversation only metadata such as file locations, while the plugin's instructions hand the reading of the contents to a subagent.
Another pillar of the design is "capture first, route later." You don't decide the destination while gathering material. The destination is decided when you run /context-drop:pull (or its short alias /cd) in the Claude Code session you want to hand it to — that session claims the packet. So you never agonize over "which session of which project do I hand this to."
5. How to use it
It's basically three steps.
-
Collect: press Start Capture in the app (global shortcut
Cmd+Shift+9). Only items you copy after pressing Start are collected. You can also drag & drop files onto the window at any time. Stop pauses capture and keeps what you have; Clear empties the packet. -
Claim: type
/context-drop:pull <instruction>in the Claude Code tab you want to hand it to (or/cd <instruction>if you installed the optional/cdshort alias — see §6). It infers ANALYZE (investigate only) / FIX (fix it too) / REVIEW from your instruction. - Receive: the isolated subagent reads the packet's contents (screenshots, logs, JSON, files) in its own context, and all that comes back is a compact result.
Here's how much that helps, measured in a real /cd run in this project. The packet held 5 items: two PNG screenshots (163,772 bytes and 173,585 bytes) and three short text files (184, 487 and 87 bytes). The isolated subagent spent 19,365 tokens reading them. The main conversation received only a compact inventory — a few hundred tokens. The raw bytes never entered the main conversation.
Had I pasted those two screenshots straight into the main, that weight would have been etched into the main context and re-sent on every subsequent turn. Offloading it to a disposable subagent while the main stays light — that's the entire value of Context Drop. (And per "delegation ≠ saving," those 19,365 tokens are still billed; what you save is the main context, and every future turn's resend.)
6. Installation (with screen captures)
Here I'll go slowly. Context Drop is built with Tauri (a framework for building desktop apps with Rust and a web frontend) and is designed for both macOS and Windows. For now, though, the only prebuilt download (.dmg) is for macOS on Apple Silicon (arm64). The steps below are for that build. If you're on an Intel Mac or Windows, see "Build from source" right after Step 1.
Step 1 — Download
Open the GitHub Releases page (https://github.com/EarthLinkNetwork/context-drop/releases) and download the .dmg from the newest release (the one marked Latest). The screens and steps in this article were checked on v0.1.3 (Context-Drop-v0.1.3-macos-arm64.dmg, about 6.3 MB). As of October 4, 2026, the latest version is v0.1.5, and depending on when you read this a newer version has probably been released, so always take the latest one. Read the file names and version numbers below as the version you downloaded. The .dmg is signed with an Apple Developer ID certificate and notarized by Apple (ticket stapled), so macOS won't block it as coming from an unidentified developer. On first launch macOS may still ask you to confirm opening an app downloaded from the internet — click Open.

Fig 1: The GitHub Releases page for v0.1.3. Click Context-Drop-v0.1.3-macos-arm64.dmg to download.
(Intel Mac / Windows) Build from source
Context Drop is open source (MIT license), so where there's no prebuilt download you can build it yourself. You need three things:
- Rust (stable, installed with rustup)
- Node.js and pnpm
- Your OS's WebView (on Windows, the Microsoft Edge WebView2 Runtime)
The build commands:
git clone https://github.com/EarthLinkNetwork/context-drop.git
cd context-drop
pnpm install
pnpm tauri build
The resulting installer or app is written under apps/desktop/src-tauri/target/release/bundle/ (the desktop app is a standalone Cargo workspace, so it is not the repo-root target/). CI (GitHub Actions, windows-latest) checks on every push that the app compiles and links on Windows. To be honest, though, we have not yet launched and used the resulting Windows build on a real machine. If it doesn't work for you, please tell us through a GitHub issue or pull request.
Step 2 — Drag to Applications
Double-click the downloaded .dmg (e.g. Context-Drop-v0.1.3-macos-arm64.dmg) to open it, and Context Drop.app appears next to a shortcut to the Applications folder. Drag the app into Applications to copy it.

Fig 2: The opened .dmg. Drag Context Drop.app into the Applications folder.
Step 3 — First launch
Launch Context Drop from Applications. If macOS asks whether to open an app downloaded from the internet, click Open — that prompt is normal for a notarized app. It's a menu-bar app, so its icon appears in the menu bar at the top of the screen; click it to open the window. Because the Claude Code integration isn't set up yet, a "Finish setup" banner with an Open Setup button appears above the tabs. Launching the app also installs (or refreshes) the bundled context-drop CLI in its canonical location automatically.

Fig 3: First launch. The "Finish setup" banner and its Open Setup button above the tabs (plugin not yet installed). (v0.1.3 UI rendered with sample state)
Step 4 — Open "Claude Code Setup" in Settings
Click Open Setup (or the Settings tab) and look at the "Claude Code Setup" section. Step 1 (enable the plugin) shows the two /plugin … commands with Copy buttons; Step 2 explains capture and routing. Below them is Optional / offline install.

Fig 4: "Claude Code Setup" on the Settings tab. Step 1 with the two copyable commands, Step 2, and Optional / offline install. (v0.1.3 UI rendered with sample state)
Steps 5 and 6 happen in Claude Code and are shown as text rather than figures.
Step 5 — Add the marketplace in Claude Code
In a Claude Code session, type the following. This is the public route — you name the GitHub repository directly:
/plugin marketplace add EarthLinkNetwork/context-drop
Expected output (observed with the same /plugin marketplace add command):
Successfully added marketplace: context-drop
Step 6 — Install the plugin and reload
Then install it:
/plugin install context-drop@context-drop
Observed output (with a check mark):
Installed context-drop. Run /reload-plugins to apply.
Then apply it:
/reload-plugins
This prints a line beginning with Reloaded:.
Important: already-open sessions don't see the new commands yet. A Claude Code session that was open before the install does not get
/context-drop:pulluntil you run/reload-pluginsin it, or open a new session. (/cdadditionally needs the short alias from Step 6a.) I hit this in practice: a long-running session treated/cd ...as plain text and just answered it as a message.Offline path. Without internet access, use Settings → Claude Code Setup → Optional / offline install → Install locally to stage the marketplace locally, then pass the printed local path to
/plugin marketplace add <path>. Either way — public or local — you still need the desktop app running, since it handles capture and the CLI.
Step 6a — (Optional) Install the /cd short alias
The plugin itself provides /context-drop:pull; the short /cd is a separate alias. To get it, open Settings → Claude Code Setup → Optional / offline install and click Install /cd short alias. If you skip this, use /context-drop:pull everywhere this post shows /cd.
Step 7 — Confirm the install
The "Finish setup" banner disappears, and Settings → Claude Code Setup now shows the "Plugin installed in Claude Code" badge. The app detects this by reading Claude Code's own install record (installed_plugins.json) — once you see it, you're ready.

Fig 5: The Settings tab showing the "Plugin installed in Claude Code" badge. (v0.1.3 UI rendered with sample state)
Step 8 — Pile up material
Press Start Capture on the Capture tab (or Cmd+Shift+9), then copy what you need — or drag & drop files onto the window. Items pile up in Current Packet, with a text preview for text and a thumbnail for images, so you can tell at a glance what you put in. Click an item to view its content (text is shown up to the first 1 MiB); the trash icon removes a single item. The Last Dispatch line shows when a packet was last handed off.

Fig 6: Current Packet on the Capture tab — text previews, an image thumbnail, and Last Dispatch. (v0.1.3 UI rendered with sample state)
Step 9 — Claim with /context-drop:pull (or /cd)
In the tab you want to hand it to, type /context-drop:pull <instruction> — this always works once the plugin is installed. If you installed the short alias in Step 6a, /cd <instruction> does the same thing. For example:
/context-drop:pull Investigate and fix the problem in this UI
/cd Investigate and fix the problem in this UI
What happens next is all text, so here it is in words rather than a screenshot. The isolated subagent opens the packet in its own context and reads everything. In the measured run from §5 (a separate run, not the example instruction above) — a 5-item packet of two screenshots (163,772 and 173,585 bytes) plus three short texts — the subagent used 19,365 tokens, and what came back to the main conversation was only a compact inventory (a few hundred tokens). None of the raw bytes ever entered the main conversation.
7. The crux of the design / wrap-up
What Context Drop does is, technically, dead simple — keep raw context out of the main conversation, have an isolated subagent read it, and return only a summary.
But that one simple line addresses the three problems that tormented me in my heavy ultracode days. Corruption (if a huge context and post-compact states are the trigger, don't let the main bloat), bloat (cut off, at the root, the raw data that gets re-sent every turn), and cost (confine heavy reads to a disposable subagent, so they're paid once instead of on every turn — but since delegation isn't saving, also cut how much you have it read).
It doesn't make the model smarter, nor does it raise the effort or the agent count. It builds one thing into the tool and its workflow: keep the main context light. The heavier your usage, the more that single point is worth.
8. It's open source
Context Drop is published as MIT-licensed open source at https://github.com/EarthLinkNetwork/context-drop. There is still plenty to do: prebuilt Intel Mac and Windows builds, integrations with other AI tools, and smoothing out rough edges. If you have an idea that would make it more useful, or a feature you want, please send a pull request. We're looking forward to it.
Top comments (1)
tool calls leaking into the chat as raw text is one i know too well. we get the same thing on open models once the context gets heavy, the format goes before the reasoning does. keeping screenshots and logs out of the main thread is prob the right fix. does the subagent ever drop the one detail u actually needed? (i build grunz, a coding agent on open models)