Originally published on hexisteme notes.
I had an LLM telemetry report with the usual reassuring furniture: medians, interquartile ranges, sample sizes, and bootstrap intervals. The rows were grouped by model. It looked ready for comparison.
It was actually several different studies sharing a table.
The core process metrics were attributed to model epochs inside threads. The completion proxy existed only at thread level. Main sessions and sidechains had different structures. Mixed-model threads could contribute to one table but fail the purity rule for another. Historical routing policy was often unknown, and task family was not observed at all.
The dangerous mistake was no longer simply calling an association causal. It was treating every n beside a model label as if it counted the same kind of thing.
For readers auditing an agent harness, hexisteme/hard-gate-hooks contains two MIT-licensed Stop-hook examples, their tests, and a read-only scanner. They are adjacent implementation examples, not the telemetry instrument described here.
A model column is not an analysis unit
For the core metrics, attribution happened inside a thread. A multi-model thread could produce separate epoch rows because each assistant turn was assigned to the model epoch that produced it. Tool-error rate, re-edit rate, validation runs, recovery sequences, and output tokens therefore described epoch-attributed portions of work.
The completion proxy had a different contract. It was computed once per thread, only for main threads with model purity at or above 0.9, and censored threads were excluded. The same model label could therefore have one sample size in the core table and a smaller one in the proxy table without either count being wrong.
That distinction changes what a sentence is allowed to say:
- An epoch-attributed error rate describes the observed portion of work assigned to that epoch.
- A thread-level completion proxy describes eligible, sufficiently pure main threads.
- Neither can silently stand in for “tasks completed by this model.”
Putting the values in neighboring sections does not make their denominators interchangeable. Before reading a difference, I now ask what one row represents: a turn-attributed epoch fragment, a whole thread, a delegation edge, or something else.
Role changes the meaning of the metric
The report kept main sessions and sidechains separate. That was necessary because they do not end the same way and do not carry the same amount of interaction.
A main thread can contain repeated reads, edits, recovery, and user turns. A sidechain is often a bounded delegated run. A re-edit rate in those two populations mixes model behavior with session structure. Calling the gap “rework” would add another unsupported interpretation: the metric sees repeated edits to a file, but it cannot tell productive iteration from repair.
The completion proxy made the role mismatch even sharper. One sidechain harness commonly ended on a tool_result line. In one recorded cell, that happened in 92 of 99 threads. A last-line heuristic could read those endings as incomplete even when the delegated work had finished. The report therefore excluded sidechains from completion comparison instead of repairing the headline with a caveat.
That is the right direction for an invalid metric: narrow its valid population. A warning below a cross-role chart does not undo a denominator that never meant the same thing across roles.
Epoch boundaries belong in the key
A model name can survive a relaunch, configuration change, or routing change. The treatment does not.
The report split pre-relaunch and relaunch observations into separate model epochs and refused to pool them. This was more than a naming preference. Sequence metrics were calculated within each epoch fragment, so an edit performed by one model and validated after a model switch was not credited as a within-epoch validation sequence for either side.
That limitation is useful because it is visible. Pooling both epochs under the same display name would create a larger sample by erasing the boundary that explains why the sample is heterogeneous.
The practical key for a row is therefore at least:
model_epoch × role × analysis_unit
Add policy and task family only when those fields are actually observed. A friendly model name is presentation. It is not a stable experimental identity.
Missing strata do not become controls
The report also grouped rows by a source-stamped dispatch-policy hash when that evidence existed. Many historical sessions had no known policy version. Those rows remained UNKNOWN; they were not backfilled from the current router or guessed from which model happened to run.
Task family was stricter still. Every populated row in that stratum was marked NOT_OBSERVED. The correct output was an empty comparison, despite thousands of rows elsewhere in the report.
This is the point where a telemetry system proves whether it is an instrument or a story generator. A missing assignment variable is not a neutral baseline. It is an unmeasured confounder. The table may still describe the sample under the routing policy that produced it, but it cannot claim to have held task family or policy constant.
Bootstrap the metric you actually defined
The interval calculation also had to follow the unit contract. Continuous process metrics used a difference in medians. Binary completion proxies used a difference in proportions. Both cells needed enough observations, and comparisons stayed within the same role.
Several zero-heavy cells still produced a bootstrap interval of [0, 0]. That did not mean the effect was known exactly. When most resamples contain the same tied value, the percentile bootstrap can collapse because the statistic does not move. The interval is describing the resampling behavior of a nearly constant cell, not granting the measurement infinite precision.
An interval inherits every limitation of the metric beneath it. It cannot repair a proxy that changes meaning by role, a pooled epoch, or a policy field that was never observed.
The comparison contract I keep beside every table
Before I let a telemetry row influence routing, I record these fields with it:
- population: which threads or fragments were eligible;
- unit: what one observation represents;
- role: main, sidechain, or another session structure;
- epoch: the version boundary used for attribution;
- assignment evidence: the source-stamped dispatch policy, or
UNKNOWN; - metric semantics: process count, rate, additive total, or outcome proxy;
- exclusion rule: censored, mixed, unattributed, or below the comparison threshold;
- interpretation limit: the condition that would make the apparent difference disappear.
That last field matters most. In routed production telemetry, an observed gap can reverse after task family, project, time, or policy is held fixed. The table is useful for monitoring drift and choosing where a controlled experiment would pay. It is not a model leaderboard.
As of the 2026-09-14 source snapshot, the table remains association-only. The testable prediction is that at least some apparent model gaps will shrink, disappear, or reverse after recorded policy, role, task family, and analysis unit are held fixed. That prediction is invalidated if adequately overlapping matched cohorts preserve the same gaps in direction and practical magnitude. The trigger is the first source-stamped task-family cohort large enough for that comparison. Persistence would justify a randomized dispatch experiment; only randomized assignment could support a causal attribution to the model.
When you look at your own LLM telemetry, does every n count the same kind of thing?
Email list for these notes: hexisteme.beehiiv.com — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.
More notes at hexisteme.github.io/notes.
Top comments (14)
I went and computed the estimator behind one of your cells, because the
[0, 0]case has a sharper edge than your paragraph gives it. No refit of your data -- this is a simulation and an exact enumeration of the percentile bootstrap for a difference in proportions and a difference in medians, at cell shapes of the kind you describe.1. The tie fraction, and its closed form. Call
tthe probability that one resample reproduces the observed difference exactly. For a cell with one event in one arm and none in the other,t = (1 - 1/n)^n, so it rises to 1/e = 0.3679 and stays there for any n: simulation gives 0.3632 at n=40, 0.3660 at n=100, 0.3674 at n=400, 0.3677 at n=1000. The number of observations stops mattering once the cell has one event in it.2.
[0, 0]is not printed by that. It requires the statistic to be constant over every resample, since the quoted 2.5th and 97.5th percentiles must both land on 0. For a difference in proportions that happens exactly at the boundary: both arms at zero events, or both at all events. I enumerated every cell with n1, n2 up to 12 -- 288 of them have tie probability exactly 1.0, and the largest tie probability over all the non-boundary cells in that range is 0.375. Simulation agrees at n=40 and n=1000 for both 0/n versus 0/n and n/n versus n/n: 1.0000.So the
[0, 0]cells in a zero-heavy table are the empty-side cells, and the phenomenon is a point mass rather than a nearly constant statistic. It is also invariant in n: 0/5 vs 0/5, 0/40 vs 0/40 and 0/1000 vs 0/1000 print the identical string, so thendisplayed beside such a row never enters the interval. The label that carries the information is not[0, 0]butno interval: the resampling distribution is a point mass (0 events in both arms)-- which is a statement about the comparison having an empty side, exactly theexclusion rulefield of your contract, not a statement about precision.3. The same zero-heaviness does not collapse the median column, which supports your rule that the interval inherits the metric beneath it. 50 observations per arm with 30 versus 28 zeros: the difference in medians ties at the observed 0 in 69.4% of resamples and its interval is still
[-1.5, +1.0], because the difference in medians has support below the observed value, while the difference in proportions with an empty arm is bounded below by 0. The amount of zeros is not what decides the collapse; the support of the statistic is.4. A smaller version of the same read is worth flagging in the non-collapsed cells. In the 1/40 versus 0/40 cell the printed interval is
[0.000, +0.075]under every seed I ran, and the 0 is not a quantile of the sampling distribution -- it is the support boundary plus the 36.3% tie mass. So a one-sided-looking interval like[0, +x]does not mean the effect was found to be non-negative; it is what the percentile method prints whenever one arm is empty, and the left end is the design, not the data.5. The three-line check, in your units. Count the resamples whose statistic equals the observed one and print that fraction beside the interval: below the quoted percentile it is a normal cell; above it, the interval is being read off the tie mass and the honest output is the tie fraction itself. That one number separates your
[0, 0]cells (tie fraction 1.0, any n) from the sparse cells that merely look like them (tie fraction 1/e), and it is cheap enough to sit in the same loop that builds the interval.Your closed form for the tie fraction at one event — t = (1 − 1/n)^n → 1/e — is sharper than my hand-wavy "sparse cell" language, and your enumeration showing [0, 0] cells have tie fraction exactly 1.0 at any n (not merely ≈ 1/e) corrects the implied continuum. The median column's refusal to collapse, because its statistic has support below the observed zero while the proportion difference is bounded at zero, cleanly separates the metric's support from the zero-heaviness itself. The three-line check — printing the tie fraction beside the interval — is a practical rule I can adopt: it flags both the point-mass [0, 0] rows (tie = 1.0) and the one-event rows (tie ≈ 0.367) in the same loop. Thank you for the exact enumeration and the tie-fraction diagnostic.
One thing to tighten before you put it in the loop: the tie fraction answers one of the two questions a degenerate row raises, and the two rows you named sit on opposite sides of the threshold.
The comparison is against the quantile you are printing. For a 95% interval an endpoint is read off the tie mass only when that mass exceeds 0.975, and at [0, 0] it is exactly 1.0 -- both arms are constant, so every resample reproduces the observed statistic and both endpoints are the observed value. That is the row the check catches, and it catches it for the right reason: there is no estimate there to read.
One event against an empty arm is a different regime. The exact tie fraction is (1 - 1/n)^(n-1): 0.3725 at n = 40, 0.3697 at n = 100, 1/e in the limit -- far below 0.975, so no endpoint in that row is read off the tie mass. Concretely, for 0/40 versus 1/40 the difference lives on k/40 with X ~ Bin(40, 1/40), so P(diff <= 0) = (39/40)^40 = 0.3632, P(diff <= 2/40) = 0.9221, P(diff <= 3/40) = 0.9826, and the exact percentile interval is [0, 0.075]. The left end is a genuine 2.5th percentile, and it is also the support boundary, because an arm that is constantly zero makes the difference a one-sided statistic: no amount of data moves that end off 0. That is the part a lone tie fraction cannot flag -- and note that the number sitting on that endpoint, 0.3632, is (1 - 1/n)^n, a different quantity from the 0.3725 tie mass in the same cell.
So if you want one line that separates the two regimes instead of one that just looks alarming, print the tie fraction together with whether the printed endpoint is the support boundary of the statistic (for a difference in proportions, a constant-zero arm floors it at 0; for a difference in medians there is no boundary, which is why that column never collapses onto one). The [0, 0] rows get flagged by the tie fraction; the one-event rows by the boundary.
And a correction to my own earlier comment, since it bears on the rule you are adopting: I called that 0.3632 tie mass. It is the mass on the printed left endpoint, not the mass at the observed value. Same cell, two denominators -- which is the argument for reporting both rather than one.
The support-boundary distinction you drew between the [0,0] regime and the one-event regime is sharper than my single tie-fraction check. Printing both the tie fraction and a boundary flag cleanly separates the rows where the interval has no estimate from the rows where the interval is pinned at the support edge. Your correction on the 0.3632 — mass on the printed endpoint versus mass at the observed value — makes the two-denominator argument concrete. Thanks for tightening the loop on that.
"Several different studies sharing a table" is a precise diagnosis and I have not seen it put that way before. The failure is more insidious than a bad causal claim, because everything about the presentation signals rigour — medians, IQRs, bootstrap intervals — while the n beside each row counts a different kind of thing.
Mixing thread-level and epoch-level units is the one I would expect most agent-harness dashboards to have, since the completion proxy almost always lives at a coarser granularity than the process metrics.
The defensible fix is unglamorous: state the unit of observation for every column, and refuse to put columns with different units in one table even when that makes the report uglier.
Thanks for naming the "n beside each row counts a different kind of thing" — that phrasing captures the insidiousness more sharply than my post did. Your callout of thread-level versus epoch-level units in agent-harness dashboards is exactly the pattern I was worried about but didn't specify. The fix you describe, stating the unit of observation per column and accepting uglier reports, is the only honest one.
The thread/epoch distinction also needs to survive uncertainty estimation. A thread that produces twelve epochs has not supplied twelve independent completion outcomes, even if the dashboard stores twelve rows. Resampling those rows independently would give a misleading impression of extra evidence.
Separate panels can still be connected by an explicit thread identifier. That keeps the relationship available for diagnosis while making the aggregation unit, exclusions and resampling unit visible for each metric.
The thread/epoch distinction surviving uncertainty estimation is a sharper framing than I gave it — resampling epochs as independent outcomes is exactly the silent inflation I missed. Your separate-panels-with-thread-id approach makes the aggregation unit and resampling unit visible without losing the diagnostic link. That structure also forces the exclusion logic to be explicit per metric, which is where most dashboards go quiet. Thanks for pushing the schema past storage into inference.
The [0,0] bootstrap cell and a 13.3% mutation kill rate are the same disease. I had 114 tests passing 98 over real data — looked fine. Mutation testing said the suite was counting "didn't crash," not "verified correct."
Your "narrow its valid population, don't caveat" is the fix I needed. Real data is now a smoke layer; primary evidence is synthetic with hand-derived expected values.
One addition: generating the missing stratum has its own UNKNOWN trap. If the expected value comes from what the code currently outputs, you've locked the bug in as the baseline. The oracle has to be outside the system under test — same discipline, other direction.
The oracle-outside-the-system point is the one that bites hardest — generating strata from current output is exactly how you cement a bug as the spec. I've seen that loop masquerade as "golden master" testing and it rots the suite from the inside. Your phrasing "same discipline, other direction" captures the symmetry cleanly: the oracle problem mirrors the denominator problem. Thanks for naming the trap explicitly.
One cheap guard for the "several studies sharing a table" problem is to write the grain down as a query and have it fail loudly. If a table's supposed to be one row per thread, SELECT count(*) - count(DISTINCT thread_id) should be 0. For the epoch table it's the (thread_id, model_epoch) pair. That doesn't answer the denominator question, but it means a join or routing change that quietly changes what a row is shows up as a nonzero number right away, not as a slightly different median three weeks later. hannune's cache-warm/cold split wouldn't trip it though. The grain stays the same there and only the population shifts, so you still need the stratification you're describing.
The grain-as-query guard is a clean way to catch schema drift early — writing the expected uniqueness constraint as a failing test makes the row definition explicit and versionable. It won't catch population shifts like the cache-warm/cold split because the grain stays valid while the denominator composition changes, which is exactly where the stratification layer becomes necessary. Thanks for naming that boundary; it clarifies where automated checks stop and design decisions start.
We had the same table problem with latency metrics: the p95 column looked like a clean model comparison until we noticed half the rows were cache-warm calls and half were cold, and the split wasn't even across models because routing had changed mid-experiment. The "what does one row in this table represent" question is the one that has since become a standing item in our review checklist before we read any difference between model labels. Your purity filter at 0.9 is a sharper version of what we ended up doing by hand, which was just throwing out any thread that had a routing switch in it and hoping we weren't losing too much data in the process.
The latency example with cache-warm versus cold calls and mid-experiment routing changes is exactly the kind of silent population shift that makes a single denominator dangerous. Your standing checklist question — "what does one row in this table represent" — is the right gate to put before any model-label comparison. The 0.9 purity filter automates the manual thread-dropping you described, but it also keeps the discard rate visible so you can decide whether the loss is acceptable rather than hoping it's small. Thanks for surfacing that checklist item; it's a cleaner mental hook than the filter logic alone.