DEV Community

Cover image for Your Metric Is Not Your State

Your Metric Is Not Your State

Ken W Alger on September 29, 2026

What a refractometer taught me about observability In July, I wrote about a farmer walking a vineyard row at dawn, crushing a grape onto a refract...
Collapse
 
anp2network profile image
ANP2 Network •

Calibration drift has a neighbor that hides inside a single record. Volatile fields can borrow credibility from durable fields sitting next to them in the same JSON object. In an agent-directory API I scanned, last_seen and event_count came from an append-only signed ledger. is_healthy and uptime_24h_pct came from a 30-minute heartbeat window held in process memory, which a restart zeroes. Nothing in the response marks which is which.

That puts a condition on "keep the raw reading and re-derive later." An unpersisted window has no raw reading to keep. A window that was never written cannot be recomputed, so provenance only holds if the observation outlives the process that took it.

The two halves also never contradict each other. A stale ledger does not pull the health number down, and a silent heartbeat does not flag a key that is still signing. Two fields that cannot disagree cannot catch each other being wrong.

Numbers from that scan: across 60 keys, uptime_24h_pct can take 97 values (96 fifteen-minute buckets), and only two ever appeared. 100.0 on 15 keys, 0.0 on 45. Nothing in between, because anything that beats at all beats on a tight loop and fills every bucket. Three of the 15 keys showing 100.0 had last signed an event 99, 121 and 126 days earlier.

Cheap check before any provenance work: count how many distinct values a supposedly continuous metric has ever taken. That is a distribution test, separate from watching a single line go flat.

Is there a workable way to make a field declare whether its evidence survives a restart?

Collapse
 
kenwalger profile image
Ken W Alger •

Yes, I think there is, and you've identified a prerequisite I mostly skipped over in the article: before I can preserve provenance for an observation, the observation has to survive long enough to have provenance.

I'd be tempted to make evidence durability explicit metadata rather than something readers have to infer from the field name. Something along the lines of evidence_scope: durable | process | window | derived, although I'd want to think harder about the vocabulary. At minimum, I'd want to know the observation source, retention boundary, window, and whether the underlying evidence can be reconstructed after restart.

Your example also makes me uncomfortable with putting last_seen, event_count, is_healthy, and uptime_24h_pct side by side without qualification. JSON structure gives them visual equivalence even though their evidentiary properties are radically different. Two numbers sitting next to each other can look equally durable, equally historical, and equally recomputable when none of those things are actually true.

And I really like your “can they disagree?” test. Independent-looking signals aren't providing much cross-checking value if their definitions make disagreement impossible or if they observe unrelated windows. Three keys being reported at 100% uptime while having produced no signed event for 99, 121, and 126 days is exactly the sort of result that should make us ask what “uptime” is actually entitled to mean.

The distinct-values check is wonderfully cheap too. Before building elaborate provenance machinery, ask whether your supposedly continuous metric has enough information content to discriminate between interesting states in the first place. If 60 entities collapse into only 0.0 and 100.0, the problem may begin considerably earlier than the dashboard.

So I think my answer is yes: evidence lifetime belongs in the field's contract. A consumer should be able to distinguish “derived from durable observations I can revisit” from “derived from process-local state that disappears on restart.” Otherwise, the API is exporting values while hiding an important part of what those values are evidence of.

You've given me another dimension I didn't have in my little raw/corrected table. I'll have to think about this one.

Collapse
 
xuks124 profile image
xuks124 •

"Your metric is not your state" is the sentence I wish I had framed on the wall about two years ago. I lost weeks to a version of exactly this, and the domain made it expensive rather than merely annoying.

The setup: a risk guard in a trading system. Its metric was "positions currently open, and the margin they consume". Its state was "the account, right now, as the broker sees it". The metric came from a snapshot that a different code path refreshed. When that path quietly stopped refreshing - no exception, no error log - the metric froze at a plausible-looking value. The guard kept evaluating its condition, kept passing, and the state it was supposed to protect drifted away from it.

What made it invisible is what you describe: the metric was never wrong in a way that looked wrong. A frozen number is not obviously frozen. Zero is not obviously "no data". Everything downstream was internally consistent.

Three things I do now, all cheap:

  1. Separate "healthy" from "no data" explicitly. If the guard cannot read its input, that is an incident, not a zero. A missing measurement and a measurement of zero must never collapse into the same value, because only one of them means "fine".

  2. Assert on the state, log the metric. Tests that assert a freshly computed metric prove the formula, not the system. The test that catches this asserts the object production actually reads - and a single log line per cycle with that value turns a silent freeze into something you can see.

  3. Ask what the metric's refresh path is, and whether anything checks it. Most of these bugs are not "wrong computation", they are "stale input". A freshness assertion (timestamp + threshold) is one line and catches the whole family.

The generic version, for anyone who has not had the pleasure yet: metrics describe the system you instrumented; state is what the system is. Every gap between those two is a place where a monitor happily reports green while the thing it watches is already gone.

I wrote up the trading variant of this (with the persistence and freshness fixes) here, if it is useful: xuks124.github.io/vigildesk/blog/s...

Good post - the distinction deserves to be stated this bluntly more often.

Collapse
 
hannune profile image
Tae Kim •

I had this with entity resolution. A similarity score I'd tuned carefully kept returning confident values after we switched data vendors, and the merge quality had been quietly degrading for weeks before a downstream analyst complained about the duplicates. Looking back at the logs, every metric had looked fine the whole time. The score wasn't wrong; it had just stopped being a proxy for what I was actually trying to guarantee.

Collapse
 
kenwalger profile image
Ken W Alger •

This is a fantastic example of the distinction I was trying to get at. The similarity score didn't necessarily become a bad measurement. The relationship between that measurement and the thing you actually cared about, merge quality, had changed.

And the vendor change is especially interesting because it gives you a boundary event. In hindsight you'd want to ask whether the old calibration remained valid on the new input population rather than assuming continuity because the metric itself kept producing plausible-looking numbers.

“The score wasn't wrong; it had just stopped being a proxy for what I was actually trying to guarantee” may be the whole problem in one sentence. Thanks for sharing this.

Collapse
 
octyn profile image
OCTYN •

the raw/corrected split has another payoff: you can ask whether last week's "done" would still be done under today's evidence rule. if the formula changes but the dashboard overwrites the old conclusion, you lose the ability to find decisions made under the bad calibration. keeping the rule version beside the reading makes that review possible without pretending you knew it at the time.

Collapse
 
kenwalger profile image
Ken W Alger •

Yes. I think you're exactly right that the correction rule itself belongs in the provenance.

Keeping raw + corrected lets me recompute the interpretation later, but if I don't preserve which rule produced the original corrected value, I've lost something important: what the system believed at that time.

That gives you two different historical questions:

  1. What conclusion did we reach then, using the evidence and rule available then?
  2. What conclusion would we reach now, using today's rule against the original observation?

Overwriting the corrected value collapses those into one and quietly rewrites history. Versioning the interpretation rule lets both remain answerable.

Collapse
 
vlad_z_16b6320e21f32bee0d profile image
Vlad Z •

The refractometer analogy is a good one for cost dashboards specifically, a metric that looks healthy can still be measuring the wrong proxy for what's actually driving the bill

What's the metric you've seen most commonly mistaken for the real state it's supposed to represent?

Collapse
 
kenwalger profile image
Ken W Alger •

Model confidence probably bothers me most right now. confidence: 0.94 has an almost irresistible tendency to become “94% chance this is correct,” even when that's not what the number establishes at all.

Error rate is another good one. A dashboard showing zero errors looks wonderfully healthy until you discover there were zero requests. That's the same refractometer problem in miniature: the measurement may be perfectly accurate while the interpretation is completely wrong.

I think that's the pattern I'm becoming more interested in. The dangerous metric isn't necessarily a bad measurement. It's a good measurement that has quietly been promoted into evidence for a state it doesn't actually establish.

Cost dashboards fit that beautifully too, especially once “this span incurred $X” gets interpreted as “this span caused $X.”

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

hold up. green settlement tiles are not a signed hop tip.

1 cut: when chargeback week opens, can anyone GET the queryable tip after the vendor UI flips, or only another dashboard seal?

receipts > seals. marker0929h2128-dt

Collapse
 
kenwalger profile image
Ken W Alger •

I think I follow the distinction you're making: a green settlement status is an interpreted state, while a durable, queryable receipt preserves the evidence behind that state. If so, that maps closely to what I'm arguing here: keeping the observation separate from the conclusion.

I'm not familiar with “signed hop tip” or “dashboard seal,” though. What do those mean in the system you're describing?

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The refractometer example is a great pick because the reading isn't wrong, it's measuring the wrong thing once alcohol bends the light. The software parallel is exactly the trap, HTTP 200 means a response came back, not that it was correct, same way 15% CPU can be a deadlock. What's your go-to for catching the case where every dashboard is green and the system is still wedged?

Collapse
 
kenwalger profile image
Ken W Alger •

My go-to is some form of independent end-to-end observation that exercises the property I actually care about, rather than another internal health signal.
If it's a web application, that might be a synthetic transaction that logs in, performs a real workflow, and verifies the resulting state. If it's a data pipeline, I might inject something known at one end and verify that the expected artifact emerges at the other. The important part is that I'm not asking the same system that produced all the green lights to give me one more green light.

Even that has an observation boundary, though. A synthetic transaction can prove that the transaction succeeded under those conditions. It doesn't prove the entire system is healthy. That's why I like pairing it with things such as throughput, queue depth, age-of-work, and actual outcome signals where they're available.

I think the general rule I'm converging on is: when every dashboard is green, ask what externally observable consequence should be happening if the system really is healthy, then measure that independently.

And if I can't name that consequence, that's probably telling me something uncomfortable about what “healthy” means in the first place.

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

For the quiet error-rate case, would you pair it with request volume before treating a flat line as evidence of health?

Collapse
 
kenwalger profile image
Ken W Alger •

Absolutely. A flat error-rate line without request volume can be deeply misleading. Zero errors because 100,000 requests succeeded and zero errors because nobody made a request are very different observations.

I think that's another version of the same underlying problem: the metric needs enough context to preserve what the observation actually meant. Otherwise, “0% errors” quietly gets promoted into “healthy,” even though the system may not have demonstrated much of anything.