What a refractometer taught me about observability
In July, I wrote about a farmer walking a vineyard row at dawn, crushing a grape onto a refractometer prism, and typing "13.5 Brix" into a chat window. The agent on the other end didn't care how the number arrived. To the Digital Scribe, I wrote, a number is just a number.
In the vineyard, that's a defensible position. Fresh grape juice is exactly what a refractometer is built to read.
This fall, I've been pointing the same kind of instrument at juice that's fermenting. It turns out the agent should have cared.
Three batches and a prism
Over the last several weeks I've started three small batches: a blackberry wine that's now bulk aging, a pear wine that began fermenting on September 24, and a fireweed honey mead that got its yeast two days later. They're one-gallon batches, which matters for this story, because at that size every sample you pull is wine you don't get to drink and oxygen you've let in.
So I track them with a refractometer. A few drops of liquid on a glass prism, close the cover, hold it up to the light, read a number off the scale. The number is Brix, roughly the percentage of dissolved sugar. It's fast, cheap, and costs a few drops per reading. During the blackberry's primary fermentation I took readings twice a day, every time I punched down the cap of fruit.
And once fermentation starts, the refractometer is wrong.
The instrument isn't lying, exactly
A refractometer measures how much light bends as it passes through a liquid. Dissolved sugar bends light, so in fresh juice the amount of bending is a good proxy for the amount of sugar. The scale on the instrument is calibrated on that assumption.
Fermentation breaks the assumption. Yeast converts sugar into alcohol, and alcohol bends light too. A few days into primary, the refractometer is reporting the combined effect of the sugar that remains and the alcohol that's been produced, and presenting all of it as sugar. The reading comes out higher than the real sugar content. Taken at face value, it says the fermentation is further behind than it actually is.
To get something usable, you run each reading through a correction formula. Most home winemakers use calculators built on formulas originally developed for brewing, and every one of them needs the same extra input: the original reading, taken before the yeast went in.
Here's what that looked like on the blackberry:
| Day | Raw refractometer (Brix) | Corrected value |
|---|---|---|
| 08/29/2026 | 20.7 | 20.7 |
| 08/30/2026 | 18.4 | 16.5 |
| 08/30/2026 | 17.0 | 14.3 |
| 08/31/2026 | 13.2 | 8.1 |
| 08/31/2026 | 10.8 | 4.1 |
| 09/01/2026 | 8.7 | 0.7 |
| 09/01/2026 | 8.3 | 0.1 |
Nothing is malfunctioning. The instrument is doing exactly what it was built to do. What changed is the system it's pointed at, and a number that was a reliable proxy in one state became a misleading one in the next.
Neither instrument measures sugar
The traditional alternative is a hydrometer, a weighted glass float. The deeper it sinks, the less dense the liquid. Sugar makes liquid denser, so falling specific gravity means sugar is being consumed.
Except the hydrometer doesn't measure sugar either. It measures density, and alcohol is less dense than water, so as fermentation proceeds the alcohol pulls the reading down on its own. A finished dry wine routinely reads below 1.000, lower than plain water, which would be impossible if the hydrometer were really a sugar meter.
So the refractometer measures how light bends and the hydrometer measures how heavy the liquid is. Neither measures what I actually care about: how much sugar is left, whether the yeast is still working, whether this wine is done. I infer those. The instruments give me evidence.
Even a steady reading is ambiguous. The traditional sign that fermentation has finished is the same reading several days running. But a fermentation that has stalled also produces the same reading several days running. Mead is notorious for this: honey has very little natural buffering, the pH drifts down as fermentation proceeds, and the yeast can slow or stop well short of dry. A flat line can mean done or stuck, and the instrument can't tell you which. You tell them apart with other evidence: where the line flattened, what the pH is doing, what you expected to see.
That distinction between evidence and state sounds pedantic until you notice it's one we get wrong in software constantly.
Your dashboard has the same problem
CPU utilization isn't health. A service at 15% CPU can be deadlocked, and one at 90% can be doing exactly what you want. An HTTP 200 isn't correctness; it tells you the server returned a response, not that the response was right. A green health check tells you the health check endpoint answered. Model confidence isn't truth; it's a number the model produced about its own output.
Each of these started as a reasonable proxy under a particular set of conditions. Then the system changed, a new failure mode appeared, or someone started optimizing the number itself, and the proxy drifted away from the state it was meant to represent. The dashboard kept rendering it with exactly the same authority.
That's the refractometer problem. The metric was calibrated for one state of the system and is now being read in another, and nothing on the display tells you so. And like the flat fermentation line, a steady metric can mean two opposite things: a quiet error rate might mean a healthy service, or one that stopped receiving traffic an hour ago.
The failure isn't collecting the metric. It's treating the metric as the state instead of as evidence about the state.
A reading needs its history
Back to that correction formula. To interpret today's reading, you need the reading from before fermentation began. Without it, the current number is close to uninterpretable. The same Brix value could describe a fermentation that has barely started or one that's nearly finished, depending entirely on where it began.
The meaning of a measurement depends on its lineage.
This is where I think observability practice most often falls short. We store the value and the timestamp and call it a record. What we usually drop is everything needed to interpret it later: which instrument produced it, under what conditions, what calibration assumptions it carried, what baseline it should be read against, and what state the system was believed to be in at the time.
My fermentation log keeps the raw reading and the corrected value side by side. The corrected value is what I act on. The raw value is what I re-derive from if I later learn the correction was off, find a better formula, or discover I misread the original. A corrected value on its own is a conclusion with its evidence thrown away.
That's a pattern worth stealing. Store the observation separately from the interpretation. Keep the provenance that makes the observation meaningful. Let state be something you derive, with its evidence attached, rather than something you overwrite. It's the same idea behind the forensic receipt work I've been doing: a claim should carry what it was based on.
Watching isn't free
One more thing the one-gallon batches have taught me. I use a refractometer because a hydrometer needs a much larger sample, and in a small carboy that sample is a real fraction of the batch, and pulling it lets oxygen in. The act of measuring changes the system being measured.
It also means deciding how often to look. The blackberry had to be punched down twice a day anyway, so twice-daily readings were free. The pear and the mead don't need that kind of handling, so I've settled on a reading every couple of days: often enough to catch a stall, rare enough to leave the batch alone. That's a measurement policy, and it means my log has deliberate gaps in it. Those gaps are their own story, and I'll come back to them.
Software has the same trade-off: tracing overhead, probes that add latency, sampling that shifts timing. Usually the effect is small. Sometimes it isn't, and how often to look is a design decision with costs on both sides.
Evidence, then state
Here's what I've taken from a few weeks of squinting at a prism.
Separate what you observed from what you concluded. "The refractometer read 9 Brix" and "fermentation is about two-thirds done" are different kinds of statement. They belong in different places, and they should be allowed to disagree.
Keep the raw reading. Corrections, normalizations, and aggregations are interpretations. If you keep only the interpreted value, you can never revisit the interpretation.
Record the conditions that make a reading meaningful. Instrument, baseline, calibration assumptions, the believed state of the system. A metric without its context is a number waiting to be misread.
Treat state as a derived claim. "Fermentation is complete" isn't a reading. It's a conclusion drawn from several readings over time, plus whatever evidence rules out "stuck," and it should be traceable back to all of it.
Notice when your proxy stops being a proxy. Every metric was calibrated for some range of system behavior. When the system leaves that range, the metric keeps reporting with exactly the same confidence.
The agent in July was right about one thing: it shouldn't matter whether a number comes from a clipboard or a five-thousand-dollar probe. Where it was wrong was in believing the number could travel without its history.
The pear wine is still in primary as I write this. The raw refractometer reading says it has a long way to go. The corrected number says otherwise, and the corrected number is only trustworthy because I wrote down what the juice read before the yeast went in.
The number on the instrument isn't the system. It never was.
Top comments (15)
Calibration drift has a neighbor that hides inside a single record. Volatile fields can borrow credibility from durable fields sitting next to them in the same JSON object. In an agent-directory API I scanned,
last_seenandevent_countcame from an append-only signed ledger.is_healthyanduptime_24h_pctcame from a 30-minute heartbeat window held in process memory, which a restart zeroes. Nothing in the response marks which is which.That puts a condition on "keep the raw reading and re-derive later." An unpersisted window has no raw reading to keep. A window that was never written cannot be recomputed, so provenance only holds if the observation outlives the process that took it.
The two halves also never contradict each other. A stale ledger does not pull the health number down, and a silent heartbeat does not flag a key that is still signing. Two fields that cannot disagree cannot catch each other being wrong.
Numbers from that scan: across 60 keys,
uptime_24h_pctcan take 97 values (96 fifteen-minute buckets), and only two ever appeared. 100.0 on 15 keys, 0.0 on 45. Nothing in between, because anything that beats at all beats on a tight loop and fills every bucket. Three of the 15 keys showing 100.0 had last signed an event 99, 121 and 126 days earlier.Cheap check before any provenance work: count how many distinct values a supposedly continuous metric has ever taken. That is a distribution test, separate from watching a single line go flat.
Is there a workable way to make a field declare whether its evidence survives a restart?
Yes, I think there is, and you've identified a prerequisite I mostly skipped over in the article: before I can preserve provenance for an observation, the observation has to survive long enough to have provenance.
I'd be tempted to make evidence durability explicit metadata rather than something readers have to infer from the field name. Something along the lines of
evidence_scope: durable | process | window | derived, although I'd want to think harder about the vocabulary. At minimum, I'd want to know the observation source, retention boundary, window, and whether the underlying evidence can be reconstructed after restart.Your example also makes me uncomfortable with putting
last_seen,event_count,is_healthy, anduptime_24h_pctside by side without qualification. JSON structure gives them visual equivalence even though their evidentiary properties are radically different. Two numbers sitting next to each other can look equally durable, equally historical, and equally recomputable when none of those things are actually true.And I really like your “can they disagree?” test. Independent-looking signals aren't providing much cross-checking value if their definitions make disagreement impossible or if they observe unrelated windows. Three keys being reported at 100% uptime while having produced no signed event for 99, 121, and 126 days is exactly the sort of result that should make us ask what “uptime” is actually entitled to mean.
The distinct-values check is wonderfully cheap too. Before building elaborate provenance machinery, ask whether your supposedly continuous metric has enough information content to discriminate between interesting states in the first place. If 60 entities collapse into only
0.0and100.0, the problem may begin considerably earlier than the dashboard.So I think my answer is yes: evidence lifetime belongs in the field's contract. A consumer should be able to distinguish “derived from durable observations I can revisit” from “derived from process-local state that disappears on restart.” Otherwise, the API is exporting values while hiding an important part of what those values are evidence of.
You've given me another dimension I didn't have in my little raw/corrected table. I'll have to think about this one.
"Your metric is not your state" is the sentence I wish I had framed on the wall about two years ago. I lost weeks to a version of exactly this, and the domain made it expensive rather than merely annoying.
The setup: a risk guard in a trading system. Its metric was "positions currently open, and the margin they consume". Its state was "the account, right now, as the broker sees it". The metric came from a snapshot that a different code path refreshed. When that path quietly stopped refreshing - no exception, no error log - the metric froze at a plausible-looking value. The guard kept evaluating its condition, kept passing, and the state it was supposed to protect drifted away from it.
What made it invisible is what you describe: the metric was never wrong in a way that looked wrong. A frozen number is not obviously frozen. Zero is not obviously "no data". Everything downstream was internally consistent.
Three things I do now, all cheap:
Separate "healthy" from "no data" explicitly. If the guard cannot read its input, that is an incident, not a zero. A missing measurement and a measurement of zero must never collapse into the same value, because only one of them means "fine".
Assert on the state, log the metric. Tests that assert a freshly computed metric prove the formula, not the system. The test that catches this asserts the object production actually reads - and a single log line per cycle with that value turns a silent freeze into something you can see.
Ask what the metric's refresh path is, and whether anything checks it. Most of these bugs are not "wrong computation", they are "stale input". A freshness assertion (timestamp + threshold) is one line and catches the whole family.
The generic version, for anyone who has not had the pleasure yet: metrics describe the system you instrumented; state is what the system is. Every gap between those two is a place where a monitor happily reports green while the thing it watches is already gone.
I wrote up the trading variant of this (with the persistence and freshness fixes) here, if it is useful: xuks124.github.io/vigildesk/blog/s...
Good post - the distinction deserves to be stated this bluntly more often.
I had this with entity resolution. A similarity score I'd tuned carefully kept returning confident values after we switched data vendors, and the merge quality had been quietly degrading for weeks before a downstream analyst complained about the duplicates. Looking back at the logs, every metric had looked fine the whole time. The score wasn't wrong; it had just stopped being a proxy for what I was actually trying to guarantee.
This is a fantastic example of the distinction I was trying to get at. The similarity score didn't necessarily become a bad measurement. The relationship between that measurement and the thing you actually cared about, merge quality, had changed.
And the vendor change is especially interesting because it gives you a boundary event. In hindsight you'd want to ask whether the old calibration remained valid on the new input population rather than assuming continuity because the metric itself kept producing plausible-looking numbers.
“The score wasn't wrong; it had just stopped being a proxy for what I was actually trying to guarantee” may be the whole problem in one sentence. Thanks for sharing this.
the raw/corrected split has another payoff: you can ask whether last week's "done" would still be done under today's evidence rule. if the formula changes but the dashboard overwrites the old conclusion, you lose the ability to find decisions made under the bad calibration. keeping the rule version beside the reading makes that review possible without pretending you knew it at the time.
Yes. I think you're exactly right that the correction rule itself belongs in the provenance.
Keeping raw + corrected lets me recompute the interpretation later, but if I don't preserve which rule produced the original corrected value, I've lost something important: what the system believed at that time.
That gives you two different historical questions:
Overwriting the corrected value collapses those into one and quietly rewrites history. Versioning the interpretation rule lets both remain answerable.
The refractometer analogy is a good one for cost dashboards specifically, a metric that looks healthy can still be measuring the wrong proxy for what's actually driving the bill
What's the metric you've seen most commonly mistaken for the real state it's supposed to represent?
Model confidence probably bothers me most right now.
confidence: 0.94has an almost irresistible tendency to become “94% chance this is correct,” even when that's not what the number establishes at all.Error rate is another good one. A dashboard showing zero errors looks wonderfully healthy until you discover there were zero requests. That's the same refractometer problem in miniature: the measurement may be perfectly accurate while the interpretation is completely wrong.
I think that's the pattern I'm becoming more interested in. The dangerous metric isn't necessarily a bad measurement. It's a good measurement that has quietly been promoted into evidence for a state it doesn't actually establish.
Cost dashboards fit that beautifully too, especially once “this span incurred $X” gets interpreted as “this span caused $X.”
hold up. green settlement tiles are not a signed hop tip.
1 cut: when chargeback week opens, can anyone GET the queryable tip after the vendor UI flips, or only another dashboard seal?
receipts > seals. marker0929h2128-dt
I think I follow the distinction you're making: a green settlement status is an interpreted state, while a durable, queryable receipt preserves the evidence behind that state. If so, that maps closely to what I'm arguing here: keeping the observation separate from the conclusion.
I'm not familiar with “signed hop tip” or “dashboard seal,” though. What do those mean in the system you're describing?
The refractometer example is a great pick because the reading isn't wrong, it's measuring the wrong thing once alcohol bends the light. The software parallel is exactly the trap, HTTP 200 means a response came back, not that it was correct, same way 15% CPU can be a deadlock. What's your go-to for catching the case where every dashboard is green and the system is still wedged?
My go-to is some form of independent end-to-end observation that exercises the property I actually care about, rather than another internal health signal.
If it's a web application, that might be a synthetic transaction that logs in, performs a real workflow, and verifies the resulting state. If it's a data pipeline, I might inject something known at one end and verify that the expected artifact emerges at the other. The important part is that I'm not asking the same system that produced all the green lights to give me one more green light.
Even that has an observation boundary, though. A synthetic transaction can prove that the transaction succeeded under those conditions. It doesn't prove the entire system is healthy. That's why I like pairing it with things such as throughput, queue depth, age-of-work, and actual outcome signals where they're available.
I think the general rule I'm converging on is: when every dashboard is green, ask what externally observable consequence should be happening if the system really is healthy, then measure that independently.
And if I can't name that consequence, that's probably telling me something uncomfortable about what “healthy” means in the first place.
For the quiet error-rate case, would you pair it with request volume before treating a flat line as evidence of health?
Absolutely. A flat error-rate line without request volume can be deeply misleading. Zero errors because 100,000 requests succeeded and zero errors because nobody made a request are very different observations.
I think that's another version of the same underlying problem: the metric needs enough context to preserve what the observation actually meant. Otherwise, “0% errors” quietly gets promoted into “healthy,” even though the system may not have demonstrated much of anything.