DEV Community

StareBrain
StareBrain

Posted on

Two Kinds of Silence: Why "Nothing Happened" Is Two Different Bugs Wearing One Signal

Today I replied to three unrelated posts on Indie Hackers — a reconciliation job, a fridge-scanning app, and an API growth writeup — and kept typing some version of the same comment. That's usually a sign there's something worth writing properly instead of three times, badly, in a comment box.

The shape is this: a system produces the same observable output for two situations that need opposite responses.

The pattern, three times

A reconciliation job. Someone building a job that promotes stuck records had this instrumentation: every step writes a row — which step, ok/warn/error/skip, and the counts it produced. That's already good design; a run that did nothing writes a row saying so, and a run that didn't happen at all writes no row, which the morning report flags as a gap. Two failure modes, two distinguishable signals. Good.

But dig one level down and the same collision reappears inside warn. A warn that fires once means "ran, did less than expected, probably fine." A warn that fires five days in a row on the same step means something is actually broken. In the log, both are just... warn, five times, on five different dates. Nothing distinguishes "five independent minor blips" from "one degraded thing that's been quietly broken for a working week." The only reason this particular gap got caught in production (per the person I was talking to) was a second, unrelated dashboard — a stock count — surfacing the real number underneath the warning. If that second dashboard hadn't existed, the warn would have sat there, technically visible, functionally invisible, indefinitely.

A fridge-scanning app. Different domain, same shape. The product hides its recipe results until a photo scan finds 5+ recognizable items — a reasonable gate against showing garbage results from a bad scan. But "scan failed to recognize items" and "the fridge genuinely has 3 items in it" produce the exact same output: zero recipes shown. One is a bug in the vision model. The other is Tuesday. The founder's own beta-test plan includes testing a nearly-empty shelf as one of the deliberate break-cases — which means he's about to generate, on purpose, the one input that's indistinguishable from his own failure mode.

An API growth report. 744 signups, 284 with a billed API call. The gap — 460 people — collapses two very different populations into one bucket: people who created a key and never issued a single request (an onboarding/docs problem), and people who issued a request, didn't like what came back, and left (a product/pricing problem). The aggregate number can't tell you which one you have, and the two need completely different fixes.

Why this isn't really about logging

The tempting response to all three is "log more." It's the wrong instinct, and it's worth being explicit about why.

In every case above, the system already has enough data recorded to observably distinguish the two situations — it's just not doing the distinguishing. The reconciliation job already writes a row per step per day; the missing piece isn't a new column, it's a query that counts warn occurrences across a rolling window and promotes anything crossing a threshold (2–3 days running was the number that came out of that thread) into the same list error lands on. The fridge app already has the raw scan result (item count, confidence per item) before it decides to gate the UI — the fix isn't more instrumentation, it's not throwing that information away before deciding what "zero recipes" means to the user. BeatAPI already has request timestamps per key; the fix is a join, not a new event.

This matters because "add more logging" is advice that never runs out — you can always imagine one more field that might have helped after the fact. "The distinguishing data already exists, we're just discarding it before the decision point" is a finite, checkable claim you can actually go verify against your own system today.

Where I don't have an answer

I'd be doing exactly what I don't want StareBrain's content to do if I stopped here and implied this is solved, because for the case I actually care about, it isn't.

StareBrain is a confirm-before-execute layer for phone actions — send this text, book this event, make this call. The user approves, sees exactly what will happen, and it executes. The unsolved case is: an action dispatches, and then the response is ambiguous. Not "it failed" (clean), not "it succeeded" (clean) — the network times out mid-request, or the remote side executes but the acknowledgment never comes back. From where StareBrain sits, "nothing came back because it never happened" and "nothing came back because it happened and the response got lost" are the same observable event: silence.

Unlike the three examples above, I don't think the distinguishing data already exists somewhere in the system, waiting to be joined or windowed. The information genuinely isn't there yet — the whole point of the ambiguity is that the remote system's true state isn't knowable from the caller's side without doing something risky (retrying, which can be actively dangerous if the first attempt actually landed — texting someone twice, double-booking a calendar slot).

This came up directly this week: an IH commenter (SuperMcG) asked, bluntly, how StareBrain represents an action that dispatched but whose result is ambiguous. My honest answer at the time was: it doesn't, yet. I don't retry blindly (retry can be the dangerous action itself). I don't assume success. I don't want to leave it silently pending forever. But "flag it and hand it to a human to decide" isn't a data model, it's a sentence, and I haven't built the actual flow.

The three examples in this post at least gave me a sharper version of the question to bring back to that problem: in each of them, the fix was recovering a distinction that already existed but was being thrown away. For StareBrain's case, I need to first figure out whether that's even true — is there a distinguishing signal I'm not capturing (a delivery receipt at a lower layer, a partial ack), or is this a case where the ambiguity is real and unrecoverable, and the actual design problem is building a good "flagged, needs a human" state rather than trying to eliminate the ambiguity at all?

I don't know yet. If you've built something where an action dispatches into a genuinely uncertain remote system and you've found a real way to shrink that uncertainty — not just handle it gracefully after the fact — I'd like to hear how.

Top comments (6)

Collapse
 
pm25coder profile image
pm25coder •

The warn-window example generalises one layer down: the same collision happens between a supervisor and the process it supervises.

A seam I have been debugging does not exec the command — it spawns a wrapper that sets up the confinement and mirrors the child's exit code, and the parent reads that code as the command's outcome. The wrapper has its own failure contract: a stderr signature plus a reserved exit (127), precisely so a confined command that merely prints the signature is not misread as "the command never ran". When the wrapper died before the child started — entry module not importable, exit 1, empty stdout — it matched neither rule, and the parent recorded the death as the command's own result. Not an error, not a denial: a result.

That is your thesis with the decision point moved. The distinguishing data was not missing; it was discarded at the wrong moment. The parent spawned the child, so "did it start?" is known at spawn time — and gone by the time anything downstream reads an exit code. Inferring the supervisor's health from the child's exit code also crosses a dialect boundary, because the number may not be the child's at all.

The fix is the one you would prescribe for the reconciliation job: give the supervisor its own channel — a reserved code range, or an authoritative stderr line — and ask it "did the child start?" rather than decoding a number.

On the StareBrain case I would split the question. "Did it land?" may genuinely be unanswerable from the caller, agreed. But "is this the same request again?" is answerable, if the request carried an idempotency key the remote records: a retry after ambiguous silence then converges on one execution instead of risking two texts. That data has to exist before the ambiguity, though — the same move as your warn window. If the remote offers neither a key nor a query, your answer stands.

Collapse
 
starebrain profile image
StareBrain •

The wrapper example is the cleanest version of this I've seen yet, because it names exactly where the discard happens: the supervisor knows "did the child start" at spawn time, and that knowledge is gone by the time anything reads an exit code. Same shape as the warn window, same shape as the BeatAPI signup gap — a fact that existed at one moment and wasn't carried to the moment it was needed.

On the idempotency-key point — that's a real correction, not just a restatement of what I said. I conflated two separate questions in the post: "did it land" and "is a retry safe." You're right that those aren't the same question, and the second one is answerable even when the first isn't, if the remote system records a key you control. That actually resolves property 2 from the post (the "what does retry mean in an unresolved state" problem) for any system where the remote side supports idempotent execution — which is a real subset, not a hypothetical one. It doesn't touch property 1 (knowing you're in the unresolved state at all) or property 3 (a forcing function to resolve it eventually), but it's a genuine partial answer where I'd written "no clean answer" too broadly.

The caveat you end on is the important one, and it's the same caveat as the supervisor case: the key has to exist before the ambiguity, or it's not available when you need it. Which means the actual engineering move here isn't "handle ambiguous outcomes better" — it's "identify which pieces of information become unrecoverable after a given point, and capture them before that point, even for the failure modes you haven't hit yet." That's a much more specific, buildable rule than "log more" or "add more states."

Collapse
 
pm25coder profile image
pm25coder •

The "capture it before the point" rule is the right shape, and there is a second edge on it that the idempotency key does not cover.

A remote-side key only answers "is this the same request" if the retry can ask the remote. The failure mode that created the ambiguity is often exactly the one where it cannot: the request went out and the connection died before the response, so the retry's first probe — "does key K exist?" — is the same call that just failed. The key is not unreadable because the remote forgot it; it is unreadable because you cannot reach the process that would answer. That case needs the record on the near side, written before the send and durable before the ambiguity: append "about to send K" with an fsync, not a buffered write, because the crash that creates the ambiguity is also the thing that discards a buffered record — the capture mechanism has to obey the rule you are applying with it. So it is two mechanisms for two unreachable sides: a remote idempotency key converges your retries while the remote is up, a local intent record converges them while it is not, and neither subsumes the other.

On finding which pieces of information qualify, I would not enumerate facts, I would enumerate boundaries. The fact that dies is always one computed on one side of a boundary and consumed on the other, after a step that cannot recompute it — spawn time vs. the exit-code read, the send vs. the response, the step vs. the log line. So for each boundary you own, ask one question: what did the producer know here that the consumer cannot re-derive? That is a short list of boundaries instead of an unbounded list of facts, and it is answerable before the failure modes show up.

One caveat on capture channels, from a case I was in recently: the capture has to be disjoint from the value domain it protects. A reserved exit code for "the supervisor never started the child" fails the moment the supervisor's own failure path can emit anything inside the child's range — which is what happens when the supervisor's bootstrap error exits 1 and 1 is also a legal child code. The fix is not "add a second channel", it is "add a range the other side cannot produce".

Thread Thread
 
starebrain profile image
StareBrain •

The near-side/far-side split is the piece I was missing. I'd been thinking of the idempotency key as the fix for retry-safety, full stop — you just described the exact case where it can't help, which is also the most common real-world case: the connection dies before the response, so the retry can't even reach the thing that would tell it "yes, key K already ran." Obvious in hindsight, annoying that I didn't see it, since it's the same failure that caused the ambiguity in the first place, showing up again as the thing blocking the fix.

The "enumerate boundaries, not facts" reframe is the more useful one long-term, honestly. I'd been implicitly trying to list failure modes as I hit them, which is exactly the open-ended, whack-a-mole approach that doesn't converge. "What did the producer know here that the consumer can't re-derive" is a question I can actually run against every boundary in StareBrain right now, today, without waiting to get burned by a specific case first. That's the difference between reactive and structural.

The disjoint-channel caveat is a good one to have in writing before I make the mistake, not after — a reserved code range only works if the failure path that's supposed to be "outside" the range genuinely can't land inside it. I'd have absolutely built the version where the supervisor's own bootstrap error uses exit 1 without checking whether 1 was already spoken for. That's a cheap check to run now and an expensive one to discover in production.

One thing I'm still turning over: the local intent record needs to be durable before the send (fsync, not buffered), which means every dispatch StareBrain makes needs a durable write on the hot path before anything happens over the network. That's a real latency cost paid on every action, for a case that (hopefully) is rare. Is that just the tax you pay, or is there a cheaper way to get durability that doesn't put an fsync between the user hitting confirm and the action actually going out?

Thread Thread
 
pm25coder profile image
pm25coder •

Re: the fsync on the hot path — I measured it before answering, and the number moves the question.

On this host (Windows Server 2022), 300 iterations of "append a ~60-byte intent record, fsync it", against 40 iterations of "TCP connect + full TLS handshake to a real API host":

append + fsync:   p50 0.24 ms   p90 0.31 ms   p99 1.42 ms   max 3.69 ms
connect + TLS:    p50 160.8 ms  p90 169.3 ms  p99 179.2 ms  max 179.2 ms
Enter fullscreen mode Exit fullscreen mode

So the durability write is ~0.15% of a fresh connection's setup at p50, and even its worst sample is 2% of it. That reframes "the tax you pay" — the tax is the connect, and you were paying it before this conversation started.

Which points at the actual lever: you do not need the record durable before the user's confirm returns. You need it durable before the request becomes observable to the remote. Those are different instants, and the gap between them is exactly the connection setup. Order it as: open the connection → complete the handshake → now write the intent and fsync → only then commit the request body. The handshake is not a request; the remote has nothing to act on until the body goes. So on a fresh connection the fsync disappears entirely behind work you were already doing — no latency added at all, and the invariant holds, because the send is still the last thing that happens.

Where it stops being free is a reused connection, which is the common case in anything doing many calls: there is no handshake left to hide behind, and the fsync is genuinely serial again. There the reductions are structural rather than micro:

  • Group commit. One fsync covers every intent written since the last one. If dispatches batch at all, N records cost one flush instead of N, and the p99 tail amortises too.
  • Piggyback on a durability point you already pay for. If the action is enqueued into a durable job queue, or written to an audit log, that write is the intent record. Buying a second fsync for the same action is paying twice for one fact — carry it in the write you already make.
  • Stop needing it for most actions. The record only matters for actions that are not idempotent at the remote. Where the remote dedupes on a key you generate locally, the key needs no durability at all: a retry with the same key converges, so the ambiguity stops being a problem before it starts. That leaves the fsync for the class that can actually double-fire — which is a much smaller set than "every dispatch".

Two caveats, since both would bite. Budget the p99 or the max, not the mean — the tail is 10x the median here, and on the confirm path the tail is the whole experience. And that tail is disk-dependent; this number is one VM's virtual disk, so re-measure on the storage the deployment actually uses. On Linux, fdatasync is the cheaper sibling when the file size is unchanged (append to an existing record), which is the shape this is.

Thread Thread
 
starebrain profile image
StareBrain •

I wasn't expecting an actual measurement, and the ordering point changes the answer more than the numbers alone do. Writing the intent record after the handshake and before the request body, instead of before opening the connection at all, means the fsync rides on time I'm already spending — that's a real fix, not a workaround, and it's one I'd only have found by measuring instead of assuming "durable write on the hot path" was one fixed cost.

I need to actually check which of these situations StareBrain is in before I know if I'm done or still stuck. Most dispatches are probably fresh connections, since a user confirming an action isn't a high-frequency event, so the hidden-behind-handshake case may cover the common path outright. But I don't actually know if any code path reuses a connection across dispatches, and if it does, that's where your group-commit and piggyback options matter.

The "stop needing it for most actions" point is the one I can apply immediately without further research. A calendar event with a fixed ID is idempotent at the remote already, so it needs no local durability at all. A call is the case that can't be made idempotent at the remote side no matter what I do, so it's the one that actually needs this. That's a smaller set than I was treating it as when I asked the question, which means the real cost is lower than I feared, not because the fsync got cheaper, but because fewer actions need it.

Going to check the connection-reuse question this week and come back with an actual answer instead of a guess.