You hit the usage limit before lunch. So you do what everyone online says: make the AI talk less. Ban the "Great question!", cut the summaries, install one of the skills that trims its replies.
I wanted to know where the tokens actually go, so I counted. Not estimated. Counted, from the transcripts Claude Code already keeps on your machine: every session still on mine, 173 of them, about 60,000 turns.
Everything Claude wrote back to me was 0.2% of the tokens. The rest was my own session, fed back in.
That number is real, but it isn't the whole story either, because most of those re-sent tokens are cheap ones. This post is both halves: where the tokens go, what they really cost, and the few things that actually move the number.
The count
Here's the summary across every project, straight from the terminal:
$ baggage --all
baggage — 173 sessions, all projects, 60,454 turns
before you typed anything 75k tokens
system prompt, every tool schema, every skill, CLAUDE.md.
paid again on all 60,454 turns = 4.5B tokens, 17% of the bill,
whether you called any of it or not.
conversation you can see 27.8M tokens
what the API billed 26.0B tokens
a 934x gap. Some of it is the fixed cost above; the rest
is everything you picked up being re-sent on every later turn.
output tokens 56.6M 0.2%
re-sent context 25.3B 97.5% <- the bill
Two lines matter. Output (every reply, every line of code Claude wrote, its thinking included) is 56.6 million tokens. Re-sent context is 25.3 billion. That's roughly 450 to 1.
I'm not the first to see this shape. A GitHub issue on Claude Code found the same thing in 30 days of someone else's transcripts, and a dev.to post measured 0.5% output across 32 sessions. Different people, different work, same picture. What none of them show is which things in a session cost the most, and what any of it means for the bill. That's the rest of this post.
Why it works like that
The model has no memory between turns. Every time you press enter, Claude Code sends the whole conversation again: the system prompt, the tool definitions, every file it read, every command output, every reply so far, and then your new message.
Think of a suitcase you have to carry up every flight of stairs. Pack a brick on the second floor of a 200-floor building and you carry it up 198 more flights. A 4,000-token build log that lands at turn 12 of a 200-turn session isn't 4,000 tokens. It's 4,000 sent 188 more times.
So the cost of anything in a session isn't its size. It's its size times the number of turns it stays.
Caching makes it cheaper, not smaller
This is the part most guides get half right. Claude Code caches the conversation, and Anthropic prices a cache read at a tenth of normal input (a twentieth on Opus 5.5, a fortieth on Fable 5.1). Output costs five times input. So 97.5% of the tokens is not 97.5% of the money.
I priced my own split at those published rates, with the one-hour cache a subscription uses. Output comes to 7 to 14% of the cost, depending on the model. Re-reading and re-caching the session is the other 86 to 93%.
Less dramatic than 0.2%. Still the opposite of what "make it talk less" assumes.
On a Pro or Max plan you never see dollars, you see a limit. Anthropic's cost docs say re-read history still draws on your usage, at the cached rate, so a one-line question late in a long session draws usage for the whole session. How heavily a cached token counts against the limit isn't published. I've seen people online insist it's full weight, and others insist it's free. Nobody I found has shown either, so I'm not going to guess.
The floor you pay before you type
That first block of the output is the one nobody shows you. Before you've typed a word, a session already carries the system prompt, the built-in tools, the descriptions of every skill you've installed, your CLAUDE.md files and anything else that loads at start. A typical session on my machine starts at about 75,000 tokens.
It rides along on every turn. Across all my sessions, that floor alone is 17% of everything billed, whether I used any of it or not. It's also the one line you can fix in thirty seconds: every paragraph of CLAUDE.md you don't need is paid again on every turn of every session, and so is the description of every skill you never call.
The heaviest things I carry
Here's one project, this website, with the list of what's costing the most rent, trimmed for length:
$ baggage
baggage — 13 sessions in singhlabs, 9,307 turns
before you typed anything 84k tokens
system prompt, every tool schema, every skill, CLAUDE.md.
paid again on all 9,307 turns = 779.0M tokens, 18% of the bill,
whether you called any of it or not.
output tokens 7.1M 0.2%
re-sent context 4.2B 98.0% <- the bill
HEAVIEST THINGS YOU ARE STILL CARRYING
(rent = its size x the turns it stayed in context)
304.4M 19.4% 3769x assistant reply
282.4M 18.0% 1376x your message
244.2M 15.6% 1415x claude-in-chrome · browser_batch
54.6M 3.5% 145x WebSearch
25.6M 1.6% 275x claude-in-chrome · javascript_tool
21.5M 1.4% 17x claude-in-chrome · get_page_text
14.9M 0.9% 159x Agent
Two things surprised me.
Claude's own replies are the biggest line: 19.4%. Not because they were expensive to write. Writing all of them was part of the 0.2%. They're expensive because each one stays in the session and is sent again on every turn after it. So the "talk less" skills aren't wrong. They're right for the wrong reason: a short reply saves you far more in re-sends than it ever cost to write.
Browser automation is close behind: 15.6%. And that's the text alone: page contents, element lists, logs of each click. baggage doesn't count images, so every screenshot rides along on top of that figure, uncounted. A session that clicks through a website carries a stack of pictures of it.
What actually moves the number
Ranked by what the counts above say, not by what's easiest to write about:
Start fresh between unrelated tasks.
/cleardrops everything you've been carrying, and Anthropic's docs say it costs nothing. The brick stays on floor two.Lower the floor. Uninstall skills you don't use, keep
CLAUDE.mdto what Claude can't work out on its own, and run/contextonce to see what loads before you type.Keep big output out of the main session. Send a long log to a file and search it, instead of printing it into the chat where it's carried to the end. Hand noisy exploration and browser work to a subagent, so only its answer comes back.
Don't break the cache mid-session. Per the caching docs, switching model, changing effort, turning on fast mode or changing MCP servers can throw it away, and then the whole session is written to cache again at the higher rate. Decide those at the start.
Shorter replies, for the right reason. Ask for the answer without the essay. It helps, because the essay gets carried.
On a paid plan, /usage now shows which skills, subagents and MCP servers used your allowance. Worth a look before changing anything.
Count your own
The tool that printed everything above is called baggage. It's free, it reads the transcripts already on your machine, and nothing leaves it. The name on npm belongs to someone else, so install it from the repo:
$ npm install -g github:manpreet171/baggage
$ baggage # this project
$ baggage --all # every project
What it doesn't tell you, so you don't over-read it:
Tokens, not dollars. The totals are the exact counts the API reported. Pricing them is your model and your plan.
Item sizes are estimates. The list ranks things with a rough four-characters-a-token ruler, the same ruler for everything. The totals above it are exact.
Only what's still on your machine. Claude Code keeps about 30 days of history by default.
One heavy user. These are my numbers. Yours will differ, which is the point of running it.
This is how we work. Measure what's really happening before changing anything, then change the thing the numbers point at. It's the same rule we use when an AI system we build for a business starts costing more than it should.
Sources: both terminal blocks are real baggage runs on my own machine on 5 Oct 2026, the second trimmed to whole lines for length. Prices and cache multipliers are from Anthropic's pricing page; the cost split weights my exact token counts by them. Cache and usage behaviour is from Claude Code's costs and prompt caching docs.
Read next: The best CLAUDE.md rules are hiding in your chat history · I researched loop engineering to build a product. I built nothing.
Originally published at singhlabs.dev.
Top comments (1)
Counting from transcripts is the right call, and the 934x gap is the part most people refuse to look at because it indicts their own habits.
One thing that measurement can't see when it's taken in aggregate: the ratio is a property of the caller, not of the session. A one-shot subagent that gets a task and returns has almost no re-sent context. A long-lived parent that runs forty tool calls carries the whole prefix forty times. Pool them and you get a number that describes neither population, and it's a very persuasive number because it looks like a property of the model.
We found this the expensive way. Our own "85% cache hit rate" was an average over subagents and long-lived parents with completely opposite prefix shapes. One group was caching aggressively, the other was thrashing, and the aggregate moved for reasons that had nothing to do with either. Once we tracked hit rate per caller class the regression was visible immediately, and it wasn't where the aggregate had pointed. Same prompt, different caller, wildly different bill.
Your fixed-prefix arithmetic has a consequence worth stating plainly, because nobody prices it that way. You have 75k tokens of system prompt, tool schemas, skills and
CLAUDE.mdre-paid on all 60,454 turns. So the marginal cost of one more tool is not the size of its schema. It's the schema times the number of turns that remain. Adding a tool to a session you then run for two hours is a per-turn tax for two hours. Adding the same tool to a 4-turn session is nearly free. Same decision, wildly different bill, and nothing in a token-per-tool count will show you that.The lever I'd actually pull, given your number: the fixed prefix is fixed only until it isn't. Trimming the prompt is the obvious move and mostly rearranging. What actually moved our number was changing who calls what, so that the long-lived parents got shorter prefixes and the throwaway work stopped paying to be remembered.
One limit on my own numbers, since your post is careful about this: ours come from a self-hosted gateway on a phone, so the absolute figures are not comparable to a managed API. The caller-mix shape should be, though, because it's a property of how the agent is structured rather than of what it's running on.