If your LLM API spend climbed after you added retrieval, a longer system prompt, or tool definitions, the cause is almost always input tokens being reprocessed at full price on every request — not the model you picked. Check the cache fields in the response usage object first: if cached reads are zero across repeated requests that share a prefix, you have a bug, not a pricing problem. Fixing prefix stability and breakpoint placement is free; downgrading the model costs you quality.
Why is the bill rising when output length didn't change?
The first thing to do is stop reasoning about the bill and start logging per-request token accounting. On the Anthropic API, every response carries a usage object that splits input tokens three ways:
import anthropic
client = anthropic.Anthropic()
resp = client.messages.create(
model="claude-opus-5",
max_tokens=2048,
system=[
{
"type": "text",
"text": STABLE_PREAMBLE, # instructions, schema, few-shot examples
"cache_control": {"type": "ephemeral"},
}
],
messages=[{"role": "user", "content": question}],
)
u = resp.usage
print(
f"uncached={u.input_tokens} "
f"written={u.cache_creation_input_tokens} "
f"read={u.cache_read_input_tokens} "
f"output={u.output_tokens}"
)
The field that trips people up is input_tokens: it is the uncached remainder, not the prompt size. Total prompt size is input_tokens + cache_creation_input_tokens + cache_read_input_tokens. I have watched a team conclude their prompt was small because input_tokens read 4K, while the other two fields accounted for ten times that. OpenAI's API exposes the analogous number as usage.prompt_tokens_details.cached_tokens; the arithmetic lesson is the same.
Run the same request twice in a row and compare. A healthy second request reads the shared prefix and writes only the delta. If the second request shows a cache write roughly equal to the full prompt and a read of zero, your prefix is not stable.
Takeaway: before you argue about model pricing, prove whether your repeated prefix is being read from cache or reprocessed from scratch.
What silently breaks prompt caching?
Prompt caching is a prefix match over the exact rendered bytes. The rendering order is tools → system → messages, and any byte that changes early invalidates everything after it. That makes the failure mode sneaky: the request still succeeds, the output is still correct, and only the bill notices.
The classic offender is a system prompt that looks like this:
# Every request produces a different prefix — nothing after this ever caches
system = f"""You are a support assistant.
Current time: {datetime.now().isoformat()}
User: {user.email} (plan: {user.plan})
...500 lines of instructions...
"""
Three invalidators in four lines. The fix is to freeze the prefix and push volatile values past the last breakpoint — into a later message, where a change at turn five invalidates nothing before turn five:
system = [{"type": "text", "text": STATIC_INSTRUCTIONS,
"cache_control": {"type": "ephemeral"}}]
messages = [
{"role": "user", "content": [
{"type": "text", "text": f"Current time: {now}. Plan: {user.plan}."},
{"type": "text", "text": question},
]},
]
Other invalidators worth grepping for in any code that builds a prompt: json.dumps() without sort_keys=True, iteration over a set, tool lists assembled per user (tools render at position zero, so a per-user tool set means no cross-user reuse), and conditional system sections where every feature-flag combination becomes its own distinct prefix. Switching models invalidates too — caches are model-scoped.
Takeaway: anything that varies per request belongs after the last cache breakpoint, and "per request" includes the clock.
Where should the cache breakpoint actually go?
Marking the end of the prompt feels right and is usually wrong. If the last block is the user's unique question or freshly retrieved rows, the breakpoint lands after bytes that will never repeat — so every request pays the write premium and no request ever reads it back. The signature in your logs is a cache write on every request while reads never cover the shared prefix.
| Prompt shape | Where the breakpoint goes | Failure if you get it wrong |
|---|---|---|
| Big fixed preamble, varying question | End of the shared preamble | Write premium on every request, reads near zero |
| Growing multi-turn conversation | Last block of the newest turn | Whole history reprocessed each turn |
| Prefix differs from byte one per request | Nowhere — don't cache | Pure surcharge, no reads |
| Sections with different change rates | One breakpoint per stability boundary | Daily-changing context invalidates never-changing tools |
Two more constraints that bite in practice, as of late 2026: there is a minimum cacheable prefix (model-dependent, and not monotonic across generations — newer models cache shorter prefixes than some older ones), and below it caching silently no-ops with no error. And parallel requests over an identical prefix all miss, because an entry only becomes readable once the first response starts streaming. For fan-out, send one request, wait for its first token, then fire the rest.
The economics are a ratio, not a mystery: a cache write costs more than a plain input token and a cache read costs a fraction of one, so a prefix read back even twice is already ahead. Longer-lived cache entries cost more to write, which means they only pay off across gaps that the default short-lived entry can't bridge. If you want the version of this that requires no breakpoint bookkeeping at all, the Anthropic API's automatic caching places the breakpoint for you and is the right default for ordinary multi-turn chat — it places exactly one breakpoint, so prompts with several stability boundaries still need explicit markers.
Takeaway: put the breakpoint at the end of what repeats, not at the end of the prompt.
When is the batch endpoint worth it?
For any workload where a human isn't waiting on the response, the batch endpoints from the major providers run asynchronously at roughly half the standard per-token rate (as of late 2026 — confirm the current discount in your provider's pricing docs). Nightly summarization, backfills, evaluation runs, bulk classification, and embedding generation all qualify.
Three things to plan for. Results arrive in arbitrary order, so key them by your own custom_id and never by position. Completion is a window, not a latency target — build the job so a delayed batch degrades a dashboard rather than blocking a user request. And batch does not compose with everything: pairing it with caching tricks or streaming-dependent code paths usually means reshaping the request.
Takeaway: if no human is blocked on the answer, half-price asynchronous processing is the highest-leverage change you can make without touching prompt quality.
Should you route to a smaller model?
Eventually, maybe — but it belongs last in the order, after the free wins, because it is the only lever that trades quality. Two things make cascades less attractive than the napkin math suggests. First, caches are model-scoped, so splitting traffic across two models forfeits cache reuse between them; a cascade that saves 40% on paper can lose most of it to cold prefixes. Second, the honest unit is cost per completed task, not cost per request — a cheaper model that needs a retry, a longer chain, or human cleanup is not cheaper.
Measure the simpler alternative first: the same capable model at a lower effort or reasoning setting, on the same traffic. On current-generation models, a reduced-effort run often lands where the previous generation's maximum effort did, and you keep one cache namespace. Tune that per route — classification and extraction routes usually hold quality at the low end, while coding and long-horizon agent loops do not.
For attributing spend to routes before you cut anything, an LLM observability layer such as Langfuse gives you per-trace token and cost breakdowns without writing your own aggregation, at the cost of running one more service (or paying for the hosted tier) and adding a logging dependency to your request path.
Takeaway: prove the cheap model wins on cost per completed task — including retries and lost cache reuse — before you ship the cascade.
FAQ
Why is cache_read_input_tokens always 0?
Something in the prefix changes between requests. Log two consecutive full request bodies, strip the cache_control markers, and diff them — the first difference inside the overlapping region is your invalidator. Timestamps, UUIDs, unsorted JSON, and per-user tool lists cause most of these.
Does prompt caching reduce latency or just cost?
Both, in practice: cached prefix tokens are not reprocessed, so time to first token drops on long prompts. Cost is the more reliable win; latency improvement depends on how much of the prompt was cacheable.
Is the LLM batch API worth it for a small app?
Yes, for any job a user isn't waiting on — roughly half price for the same model and prompt, as of late 2026. It is the rare optimization with no quality tradeoff, only a latency one.
Bottom line
Start with measurement: log input_tokens, cache writes, and cache reads per request, and treat a zero read rate on a repeated prefix as a bug with a ticket. Fix prefix stability and breakpoint placement before anything else — they cost nothing in quality and often account for most of the surprise. Move every job with no human in the loop to the batch endpoint. Only then consider lower effort settings, and only after that a smaller model, judged on cost per completed task rather than cost per call.
Top comments (1)
The tool definition ordering point bites hard with Pydantic schemas, but another trap in multi-turn agent loops is tool output metadata. A lot of agent harnesses dump tool results containing temporary scratch paths, process IDs, or elapsed microsecond timers directly into the conversation history. That turns the entire preceding conversation into a dirty prefix for sibling branches and retry turns. Stripping runtime runner paths and timers down to raw stdout and exit codes keeps consecutive turns hitting cache.