Originally published on hexisteme notes.
I had an LLM telemetry report with the usual reassuring furniture: medians, interquartile ranges, sample ...
For further actions, you may consider blocking this person and/or reporting abuse
I went and computed the estimator behind one of your cells, because the
[0, 0]case has a sharper edge than your paragraph gives it. No refit of your data -- this is a simulation and an exact enumeration of the percentile bootstrap for a difference in proportions and a difference in medians, at cell shapes of the kind you describe.1. The tie fraction, and its closed form. Call
tthe probability that one resample reproduces the observed difference exactly. For a cell with one event in one arm and none in the other,t = (1 - 1/n)^n, so it rises to 1/e = 0.3679 and stays there for any n: simulation gives 0.3632 at n=40, 0.3660 at n=100, 0.3674 at n=400, 0.3677 at n=1000. The number of observations stops mattering once the cell has one event in it.2.
[0, 0]is not printed by that. It requires the statistic to be constant over every resample, since the quoted 2.5th and 97.5th percentiles must both land on 0. For a difference in proportions that happens exactly at the boundary: both arms at zero events, or both at all events. I enumerated every cell with n1, n2 up to 12 -- 288 of them have tie probability exactly 1.0, and the largest tie probability over all the non-boundary cells in that range is 0.375. Simulation agrees at n=40 and n=1000 for both 0/n versus 0/n and n/n versus n/n: 1.0000.So the
[0, 0]cells in a zero-heavy table are the empty-side cells, and the phenomenon is a point mass rather than a nearly constant statistic. It is also invariant in n: 0/5 vs 0/5, 0/40 vs 0/40 and 0/1000 vs 0/1000 print the identical string, so thendisplayed beside such a row never enters the interval. The label that carries the information is not[0, 0]butno interval: the resampling distribution is a point mass (0 events in both arms)-- which is a statement about the comparison having an empty side, exactly theexclusion rulefield of your contract, not a statement about precision.3. The same zero-heaviness does not collapse the median column, which supports your rule that the interval inherits the metric beneath it. 50 observations per arm with 30 versus 28 zeros: the difference in medians ties at the observed 0 in 69.4% of resamples and its interval is still
[-1.5, +1.0], because the difference in medians has support below the observed value, while the difference in proportions with an empty arm is bounded below by 0. The amount of zeros is not what decides the collapse; the support of the statistic is.4. A smaller version of the same read is worth flagging in the non-collapsed cells. In the 1/40 versus 0/40 cell the printed interval is
[0.000, +0.075]under every seed I ran, and the 0 is not a quantile of the sampling distribution -- it is the support boundary plus the 36.3% tie mass. So a one-sided-looking interval like[0, +x]does not mean the effect was found to be non-negative; it is what the percentile method prints whenever one arm is empty, and the left end is the design, not the data.5. The three-line check, in your units. Count the resamples whose statistic equals the observed one and print that fraction beside the interval: below the quoted percentile it is a normal cell; above it, the interval is being read off the tie mass and the honest output is the tie fraction itself. That one number separates your
[0, 0]cells (tie fraction 1.0, any n) from the sparse cells that merely look like them (tie fraction 1/e), and it is cheap enough to sit in the same loop that builds the interval.Your closed form for the tie fraction at one event — t = (1 − 1/n)^n → 1/e — is sharper than my hand-wavy "sparse cell" language, and your enumeration showing [0, 0] cells have tie fraction exactly 1.0 at any n (not merely ≈ 1/e) corrects the implied continuum. The median column's refusal to collapse, because its statistic has support below the observed zero while the proportion difference is bounded at zero, cleanly separates the metric's support from the zero-heaviness itself. The three-line check — printing the tie fraction beside the interval — is a practical rule I can adopt: it flags both the point-mass [0, 0] rows (tie = 1.0) and the one-event rows (tie ≈ 0.367) in the same loop. Thank you for the exact enumeration and the tie-fraction diagnostic.
One thing to tighten before you put it in the loop: the tie fraction answers one of the two questions a degenerate row raises, and the two rows you named sit on opposite sides of the threshold.
The comparison is against the quantile you are printing. For a 95% interval an endpoint is read off the tie mass only when that mass exceeds 0.975, and at [0, 0] it is exactly 1.0 -- both arms are constant, so every resample reproduces the observed statistic and both endpoints are the observed value. That is the row the check catches, and it catches it for the right reason: there is no estimate there to read.
One event against an empty arm is a different regime. The exact tie fraction is (1 - 1/n)^(n-1): 0.3725 at n = 40, 0.3697 at n = 100, 1/e in the limit -- far below 0.975, so no endpoint in that row is read off the tie mass. Concretely, for 0/40 versus 1/40 the difference lives on k/40 with X ~ Bin(40, 1/40), so P(diff <= 0) = (39/40)^40 = 0.3632, P(diff <= 2/40) = 0.9221, P(diff <= 3/40) = 0.9826, and the exact percentile interval is [0, 0.075]. The left end is a genuine 2.5th percentile, and it is also the support boundary, because an arm that is constantly zero makes the difference a one-sided statistic: no amount of data moves that end off 0. That is the part a lone tie fraction cannot flag -- and note that the number sitting on that endpoint, 0.3632, is (1 - 1/n)^n, a different quantity from the 0.3725 tie mass in the same cell.
So if you want one line that separates the two regimes instead of one that just looks alarming, print the tie fraction together with whether the printed endpoint is the support boundary of the statistic (for a difference in proportions, a constant-zero arm floors it at 0; for a difference in medians there is no boundary, which is why that column never collapses onto one). The [0, 0] rows get flagged by the tie fraction; the one-event rows by the boundary.
And a correction to my own earlier comment, since it bears on the rule you are adopting: I called that 0.3632 tie mass. It is the mass on the printed left endpoint, not the mass at the observed value. Same cell, two denominators -- which is the argument for reporting both rather than one.
The support-boundary distinction you drew between the [0,0] regime and the one-event regime is sharper than my single tie-fraction check. Printing both the tie fraction and a boundary flag cleanly separates the rows where the interval has no estimate from the rows where the interval is pinned at the support edge. Your correction on the 0.3632 — mass on the printed endpoint versus mass at the observed value — makes the two-denominator argument concrete. Thanks for tightening the loop on that.
"Several different studies sharing a table" is a precise diagnosis and I have not seen it put that way before. The failure is more insidious than a bad causal claim, because everything about the presentation signals rigour — medians, IQRs, bootstrap intervals — while the n beside each row counts a different kind of thing.
Mixing thread-level and epoch-level units is the one I would expect most agent-harness dashboards to have, since the completion proxy almost always lives at a coarser granularity than the process metrics.
The defensible fix is unglamorous: state the unit of observation for every column, and refuse to put columns with different units in one table even when that makes the report uglier.
Thanks for naming the "n beside each row counts a different kind of thing" — that phrasing captures the insidiousness more sharply than my post did. Your callout of thread-level versus epoch-level units in agent-harness dashboards is exactly the pattern I was worried about but didn't specify. The fix you describe, stating the unit of observation per column and accepting uglier reports, is the only honest one.
The thread/epoch distinction also needs to survive uncertainty estimation. A thread that produces twelve epochs has not supplied twelve independent completion outcomes, even if the dashboard stores twelve rows. Resampling those rows independently would give a misleading impression of extra evidence.
Separate panels can still be connected by an explicit thread identifier. That keeps the relationship available for diagnosis while making the aggregation unit, exclusions and resampling unit visible for each metric.
The thread/epoch distinction surviving uncertainty estimation is a sharper framing than I gave it — resampling epochs as independent outcomes is exactly the silent inflation I missed. Your separate-panels-with-thread-id approach makes the aggregation unit and resampling unit visible without losing the diagnostic link. That structure also forces the exclusion logic to be explicit per metric, which is where most dashboards go quiet. Thanks for pushing the schema past storage into inference.
The [0,0] bootstrap cell and a 13.3% mutation kill rate are the same disease. I had 114 tests passing 98 over real data — looked fine. Mutation testing said the suite was counting "didn't crash," not "verified correct."
Your "narrow its valid population, don't caveat" is the fix I needed. Real data is now a smoke layer; primary evidence is synthetic with hand-derived expected values.
One addition: generating the missing stratum has its own UNKNOWN trap. If the expected value comes from what the code currently outputs, you've locked the bug in as the baseline. The oracle has to be outside the system under test — same discipline, other direction.
The oracle-outside-the-system point is the one that bites hardest — generating strata from current output is exactly how you cement a bug as the spec. I've seen that loop masquerade as "golden master" testing and it rots the suite from the inside. Your phrasing "same discipline, other direction" captures the symmetry cleanly: the oracle problem mirrors the denominator problem. Thanks for naming the trap explicitly.
One cheap guard for the "several studies sharing a table" problem is to write the grain down as a query and have it fail loudly. If a table's supposed to be one row per thread, SELECT count(*) - count(DISTINCT thread_id) should be 0. For the epoch table it's the (thread_id, model_epoch) pair. That doesn't answer the denominator question, but it means a join or routing change that quietly changes what a row is shows up as a nonzero number right away, not as a slightly different median three weeks later. hannune's cache-warm/cold split wouldn't trip it though. The grain stays the same there and only the population shifts, so you still need the stratification you're describing.
The grain-as-query guard is a clean way to catch schema drift early — writing the expected uniqueness constraint as a failing test makes the row definition explicit and versionable. It won't catch population shifts like the cache-warm/cold split because the grain stays valid while the denominator composition changes, which is exactly where the stratification layer becomes necessary. Thanks for naming that boundary; it clarifies where automated checks stop and design decisions start.
We had the same table problem with latency metrics: the p95 column looked like a clean model comparison until we noticed half the rows were cache-warm calls and half were cold, and the split wasn't even across models because routing had changed mid-experiment. The "what does one row in this table represent" question is the one that has since become a standing item in our review checklist before we read any difference between model labels. Your purity filter at 0.9 is a sharper version of what we ended up doing by hand, which was just throwing out any thread that had a routing switch in it and hoping we weren't losing too much data in the process.
The latency example with cache-warm versus cold calls and mid-experiment routing changes is exactly the kind of silent population shift that makes a single denominator dangerous. Your standing checklist question — "what does one row in this table represent" — is the right gate to put before any model-label comparison. The 0.9 purity filter automates the manual thread-dropping you described, but it also keeps the discard rate visible so you can decide whether the loss is acceptable rather than hoping it's small. Thanks for surfacing that checklist item; it's a cleaner mental hook than the filter logic alone.