Attribution, allocation, and why "unknown" belongs in the schema.
This piece grew out of a comment thread on Sarvar Nadaf's Per-Agent Cost Tracking for Multi-Agent AI on AWS. The schema below was worked out in that conversation, in public, and it is better for it. Where a specific idea came from the exchange, I have tried to say so.
Most AI cost dashboards answer one question well: how much did this run cost. Total tokens, model spend, per-agent spend, latency, tool usage. Those are real and useful numbers, and for a while they are enough.
They stop being enough the moment your system becomes a composition. Once a request flows through a retriever, a knowledge graph, three specialist agents, and a supervisor that synthesizes their output, the total tells you almost nothing about what to change. A run can be correct, return HTTP 200, look healthy in every latency-and-errors panel, and still cost forty percent more than an identical run that produced the same answer. The overspend is real. It is just not anywhere you are looking.
To find it, you have to stop asking where the money was spent and start asking what caused it to be spent. Those are different questions, and the gap between them is the whole subject of this piece.
Where a Cost Is Incurred Is Not What Caused It
Consider a retrieval operation that costs $0.004: searching, ranking, fetching. That is the direct cost, and it is easy to attribute. It happened on that span, you can measure it, done.
Now suppose that retrieval returned 20,000 tokens, and all of them were hydrated into a supervisor's context on the next step. The supervisor then costs $0.009. How much did the retrieval really cost?
The direct answer is still $0.004. But that is no longer the interesting answer, because the retrieval also caused cost somewhere else. It inflated the supervisor's context, and some portion of that $0.009 exists only because the retrieval handed it too much material. The cost was incurred at the supervisor. It was caused, in part, at the retriever.
This gives you two distinct dimensions, and a useful cost model has to carry both:
- Where the cost was incurred. This is just the span. Directly observed, low ambiguity.
- What caused or contributed to it. This is the interesting axis, and it is the one no aggregate dashboard shows.
A record that captures both might look like this:
span_id: retrieval-104
direct_cost: $0.0040
primary: retrieval
returned_tokens: 20000
downstream:
span_id: supervisor-105
attributed_cost: $0.0021
caused_by: retrieval-104
attribution_method: proportional
The dollars are single-counted. We do not charge the $0.0021 twice. The supervisor genuinely incurred it; the retrieval genuinely contributed to causing it; and the record says both without inventing money. What we have added is lineage: a link from a downstream cost back to the decision that helped produce it.
The Hard Part Is Honesty About How You Know
Here is the question that breaks naive versions of this: how do we know retrieval-104 actually caused $0.0021 of the supervisor's cost, and not some other amount?
Sometimes you can measure it. If you have a controlled comparison where the only meaningful change is that retrieval result, the delta is real evidence. Supervisor costs $0.006 without the retrieved material and $0.009 with it, so roughly $0.003 of downstream cost is attributable to that retrieval. That is measured causation, and it is the strongest claim you can make.
Most production traces do not give you that. In a real run the supervisor is carrying system instructions, conversation state, the outputs of other agents, tool results, and the retrieved material, all at once. There is no clean counterfactual. So you fall back on allocation: split the supervisor's context cost proportionally by the tokens each source contributed. That is a reasonable method. It is not measurement, and the receipt must not pretend it is.
This is why the single most important field in the whole schema is not a dollar amount. It is this:
attribution_method:
- measured_delta
- proportional
- estimated
- unknown
That field is what keeps the entire model honest. It stops a proportional guess from masquerading as measured causation. With it, a line can say "retrieval span 104 contributed an estimated $0.0021 of downstream context cost, allocated proportionally by hydrated token share," and every word in that sentence is defensible, because the method is stated. Without it, the same $0.0021 acquires a precision the evidence never earned.
Resist collapsing this into a confidence score. A number like confidence: 0.82 feels rigorous and gives you nothing, because now you have a second number whose provenance you have to go investigate. measured_delta, proportional, estimated, and unknown each tell you why you are entitled to believe the figure. The method is the provenance. A score would hide it.
Show the Method Where the Decision Is Made
A natural instinct is to keep the attribution method as drill-down metadata, out of the main view, so the report stays clean. That instinct is wrong, and it is wrong for the same reason aggregate dashboards are wrong: it makes two different claims look equivalent.
The method belongs inline, next to any attributed cost, with one sensible exception. A directly observed cost carries no ambiguity and needs no method tag:
retrieval-104 RETRIEVAL $0.0040
There is nothing to disclose there; it was measured on the span. But the moment a number is attributed rather than observed, the method has to ride along:
retrieval-104 -> downstream CONTEXT $0.0021 proportional
Drop the word proportional and that $0.0021 visually becomes as solid as the $0.0040 above it, which is a lie of formatting. A report that hides the distinction between what it measured and what it allocated has committed the same sin as the dashboard that only shows a total. If two numbers make materially different claims, the interface must not make them look the same.
So the main report shows amount, category, direct versus downstream, and method. The drill-down holds the evidence behind the method: hydrated token counts, comparison runs, parent-child span references, the assumptions the allocation rests on. The decision surface stays readable; the receipt is one click away, not dumped into the table.
"Unknown" Is Not a Gap in the Accounting
Every honest version of this schema has to allow unknown as a real value, not a placeholder you feel bad about. And once you sit with it, the unknown rows turn out to be the most useful rows in the report.
A high downstream cost with unknown attribution is not incomplete bookkeeping. It is the system telling you exactly where your observability boundary stops letting you explain its own behavior. It is pointing at the place where you cannot yet answer "what caused this," which is the place most worth instrumenting next. A tidy report with no unknown rows has usually not achieved understanding. It has hidden its ignorance behind confident allocation.
There is also a real reason unknown is sometimes the only honest answer: context is not additive. An extra 5,000 tokens of context does not simply add a proportional slice of cost. It can change caching behavior, alter the reasoning path the model takes, or change how much output the model generates downstream. When that happens, token share and cost share stop mapping to each other cleanly, and any proportional number you report is a polite fiction. In those cases the schema should say unknown and mean it, rather than allocate a figure it cannot defend.
Read that way, the report stops being a statement of where you paid and becomes a map of two things at once: what you can explain about your spending, and where your ability to explain it runs out. The second map is the one that tells you what to build.
What This Actually Costs to Build
The reason this is not a research project is that most of the structure already exists. If your traces are parent-child spans, the causal lineage is physically present already; a downstream cost sits under the decision that produced it. You are not inventing a new tracing mechanism. You are making the attribution semantics explicit on top of a trace you already record.
Concretely, that is a small number of additions. Stamp a primary category on each span. Add a contributes_to link and a cause on spans that produce downstream effects. Add the attribution_method on any attributed cost. Then roll the report up along two axes, category and direct-versus-downstream, and let unknown be a first-class row rather than a swept-under one. The intelligence is not in the plumbing. It is in having the honest method field and being willing to publish the unknown rows.
One warning from the same conversation that produced all this: how you record and how you attribute are coupled. Change the way spans are emitted and you can silently break the logic that reads them. The defense is the same one that makes the whole model trustworthy, a known-good baseline you compare against, so that when your instrumentation shifts under you, the numbers move and you notice.
The Point
We have spent a lot of effort making AI spend visible. Better token counts, per-agent breakdowns, nested traces. All of it answers "how much." Almost none of it answers "why," and "why" is the only version of the question you can act on.
The move from one to the other is not a bigger dashboard. It is a small, honest schema: separate where a cost was incurred from what caused it, state the method behind every attributed number, and treat the costs you cannot explain as signal rather than embarrassment. Do that, and the report stops telling you what you spent and starts telling you what to fix, including, in the unknown rows, where to look first.
Even the cost report, it turns out, needs provenance. Not just the amount and the category, but how sure you are about who to blame. That last column may be the most useful one on the page.
With thanks to Sarvar Nadaf, whose post and the conversation under it produced this schema, and to the commenters in that thread who pushed on the baseline and the propagation. The receipt is better for the argument.
Top comments (16)
The incurred-vs-caused split is the right frame, and the part people underestimate is how fast the causal chain gets laundered. Your retrieval→supervisor example is clean because the 20k tokens flow straight through. But the moment you put a summarization or compaction step in between — retriever returns 20k, a cheap model condenses it to 2k, supervisor reads the 2k — the supervisor's cost now looks small and well-behaved, and the real culprit (a retriever with no token budget) is invisible on every span. The cost got attributed to the step that cleaned up the mess.
The thing that's worked for us is carrying a
contributed_bylist down the chain rather than trying to reconstruct it after the fact, so the condense step inherits the retriever's id even though it spent almost nothing itself. Curious how you handle fan-out — one retrieval that feeds three agents in parallel. Do you split the downstream cost across them, or attribute the full contribution to each? We never found a split that wasn't arbitrary.“Causal laundering” is exactly the problem. The compaction step can make the downstream span look healthier while simultaneously hiding the event that made compaction necessary in the first place.
I like
contributed_bybecause it preserves lineage without pretending lineage is allocation. For the fan-out case, I don't think I'd split the downstream cost unless there were an independent reason to believe the split represented causation. Three consumers receiving the same retrieval result means that retrieval contributed to three causal chains. Saying it caused 33.3% of each feels like we've manufactured precision.So I'd be inclined to preserve the full contribution relationship on each branch, then keep monetary allocation as a separate question with its own
attribution_method. That means contribution edges don't have to sum to 100%, which I increasingly think is a feature rather than a defect.You've also given me another failure mode for the model: transformation can reduce an input's visible size without reducing its causal significance.
The observation about context not being additive hits the exact wall where proportional allocation falls apart in production. Prompt caching makes this especially sharp. When an upstream retrieval span injects unpinned dictionary keys, dynamic timestamps, or fluctuating headers near the front of a supervisor prompt, it invalidates the prefix cache for the supervisor and every subsequent turn in that trajectory. That converts cheap cache reads into full cache writes across the entire conversation history. On a standard trace, that invalidation shows up as a massive cost spike directly on the supervisor span, even though the supervisor ran the exact same logic as before. If the schema forces a proportional token split, the upstream retriever looks harmless because it only passed four hundred bytes, while the supervisor gets blamed for burning the budget. Having unknown as an explicit attribution method keeps telemetry honest when cache boundary changes make token volume decouple from actual spend.
That's a fantastic counterexample to proportional allocation because the upstream contribution can be tiny in bytes and enormous in consequence.
It also suggests the causal object isn't always the tokens themselves. In your example it's the cache-boundary change caused by those tokens. The supervisor incurred the additional spend, but nothing about the supervisor's own behavior explains the delta.
I think that strengthens the case for keeping
unknownrather than forcing allocation, but it may also require richer causal records eventually: not justretrieval → supervisor, but something closer toretrieval changed prefix → cache invalidated → supervisor incurred full-price processing.That's much closer to an explanation than “retriever 3%, supervisor 97%,” even if the explanation refuses to produce percentages.
I like the idea of treating
unknownas an actual useful result instead of something that needs to be cleaned up or estimated away. I think there’s a tendency with AI systems to make dashboards look more precise than the underlying data really is, especially once you have agents, retrieval, tools, and context all affecting each other.Knowing where money was spent is useful, but knowing what actually caused that cost is a much harder and more interesting problem. And if you can’t explain part of it yet, that’s probably exactly where you should be looking next. I’d much rather see an honest
unknownthan a very confident-looking number that’s basically an educated guess.The confidence: 0.82 line is the one worth stealing. A proportional guess wearing a score feels rigorous while telling you nothing, which is exactly why the method belongs inline next to the number instead of buried in drill-down.
Exactly.
confidence: 0.82looks quantitative enough that it's very easy to stop asking what the 0.82 is confidence in.A highly confident proportional allocation is still a proportional allocation. That's why I'm increasingly convinced the method is part of the value, not metadata about the value. Remove the method, and you've changed what the number is entitled to mean.
We've started treating unknown attribution as a weekly triage queue, not a dashboard embarrassment. The failure mode I keep hitting is proportional allocation quietly becoming "truth" in finance reviews once someone pastes the number without the method, you can't walk it back. Making attribution_method a required column on any attributed line stopped that faster than adding more span tags.
I really like treating
unknownas a queue rather than an embarrassment. That changes it from “telemetry we haven't cleaned up yet” into a visible backlog of causal questions worth investigating.And your finance-review example is exactly why I wanted
attribution_methodbeside the number. Once42%escapes the system without “proportional allocation” attached, it becomes organizational fact remarkably quickly.That suggests another useful metric: not merely unknown cost, but unknown attribution resolved per period and what method replaced it. Then shrinking
unknownmeans the organization actually learned something rather than somebody finding a convenient bucket to dump it into.The lineage is in the trace for causes inside one run. Two causes sit outside it, and both land in unknown unless a record can point into another trace. One is the prompt cache. If the provider caches prompt prefixes, whether the supervisor's first call in a run reads its prefix from cache or pays for a fresh write depends on whether another run used the same prefix within the cache lifetime, which runs from minutes to hours depending on provider and settings. So two identical runs can differ in cost with nothing in either trace that names the cause. Reid's case is a cause inside the trajectory. This one is the timing of a different run.
The other is persistent memory. A note or summary hydrated into context today was written by an earlier run, and if it was written long, every later run that recalls it pays for that length, while none of that cost appears in the trace of the run that wrote it. The decision that caused the cost is a write in a past trace.
Both fit the schema without new attribution fields. Record cache read and cache write tokens on each span where the provider reports them, so a cold start shows as its own line instead of hiding inside the supervisor's total, and stamp each stored memory with the id of the span that wrote it, so caused_by can name a span in a past trace. The cache case even admits a measured_delta: replay the same request while the cache is warm, confirm from the read field that it hit, and compare the input side of the two bills.
This is an interesting problem we have been thinking about with our crash debugging software.
ForensicDbg does a pile of things with the crash data to fix up missing and incorrect information, add additional known labeling and context, then links it all together and presents it clearly. When it comes to our MCP server though, it is that last part that becomes interesting.
We can send you the bare minimum amount of information needed to solve most basic crashes. Efficient! Unless it is not enough to get the job done, then the multiple back-and-forth trips to the tool are WAY more expensive than if we just sent more the first time.
We can send you a complex blob of nicely formatted and analyzed data that should be able to solve 98% of crashes in a single return. If it takes additional data you are still six steps ahead in terms of efficiency. But you are also 6x more expensive on any crash that would have been solved with the simple return above.
Figuring out the middle ground has been a very interesting, and ongoing challenge.
The retrieval example is the clearest statement of this I have read: the span that incurs the cost is not the span that caused it. We see a small version of it with an AI presenter that takes live questions during a session (I work on Presango, so bias declared). The expensive step is always the answer, but the cause is usually upstream: how much of the deck and the supporting documents got hydrated into context for a question that needed one slide. Attributing that to the answering step makes the dashboard say the wrong thing to fix. Keeping an explicit unknown bucket rather than forcing every cost onto a cause is the honest version, and I suspect its size over time is a better health metric than the total. Do you track how the unknown share trends run to run, or just report it?
I haven't taken it as far as a run-to-run operational metric yet, but I think your instinct is right with one important caveat: I'd want to track why the unknown share changed, not just whether it went down.
A falling unknown rate could mean the system got better at causal attribution. It could also mean somebody replaced
unknownwith increasingly aggressive proportional guesses. Those are opposite outcomes hiding behind the same improving chart.So I'd probably track unknown share alongside transitions out of unknown:
unknown → measured_delta,unknown → proportional,unknown → estimated, etc. Then the trend tells us whether we're accumulating evidence or merely accumulating confidence.Your presenter example is also a great illustration of incurred versus caused. The answer generation span can be perfectly healthy while repeatedly paying for an upstream decision to hydrate far more context than the question required.
The incurred-versus-caused split is useful because it keeps the receipt honest. The distinction I would add to this proposed schema is the decision context that allowed the expensive branch to run. A retrieval span can be correctly attributed and still exceed a declared budget if no policy sees its returned-token count before hydration. Keeping the task, budget, observed count, and resulting decision beside the causal edge would let a reviewer ask two separate questions: what caused the spend, and what allowed it to continue? How are you thinking about the point where attribution becomes a control rather than a report?
This showed up for me in a specific shape: the same retrieval step ran identically for five different requests and charged to five separate causal chains because nothing connected them back to a shared trigger. Five lines in the report, one real cause. What actually fixed it was adding a source_request_id tag to the retrieval call and grouping by that before any allocation ran, which dropped the unexplained bucket by a noticeable amount. I'm curious whether your schema has anything that surfaces when two attribution chains are pointing at the same underlying event, because that's the version of the problem that's hardest to see from the dashboard alone.
That's a really useful failure mode because it happens before allocation. If five attribution chains actually descend from one source event, no allocation method can repair the report until you've established that identity.
source_request_idsounds like exactly the kind of provenance I'd want carried forward rather than reconstructed later. I'd probably keep both identities: the immediate causal parent and the originating request/event. Then you can ask “what directly caused this?” and “what common event are these chains ultimately descended from?” without collapsing the two.I don't have anything in the schema yet that explicitly detects two chains converging on the same originating event, and I think that's a gap. Before asking how to divide cost, we may have to establish how many causes we're actually looking at.