DEV Community

My release smoke test passed on the message that says it failed

Efe Genç on September 13, 2026

The last step of my release pipeline installs the package I just published, on Linux, Windows and macOS, and confirms it reports the version that w...
Collapse
 
beusebiu profile image
Eusebiu Balan •

I recognise the two numbers written from memory most. I keep one facts file where every figure carries the date it was last true, and a number does not go into anything I publish unless it is in that file. Anything retyped from memory counts as unverified.

The table cell with the version three columns away is a good example of the other problem. A checker written against sentences reports clean on every format it never learned to read.

Collapse
 
efe_genc profile image
Comment deleted
Collapse
 
beusebiu profile image
Eusebiu Balan •

What stuck for me was making the snapshot's age its own failure condition. If the source has been written since the snapshot was taken, the run stops and says unknown rather than comparing anything.

A checker that can only report match or mismatch has no way to tell you it was reading yesterday.

Thread Thread
 
efe_genc profile image
Efe Genç •

That is a better rule than the one I shipped, and I have implemented it.

Mine bounded the snapshot's age at six hours. Your version bounds the thing the hours were standing in for. Six is arbitrary: a five-hour-old snapshot taken after the last change passes, a seven-hour-old one taken after nothing passes nothing. The hours were a proxy for "has anything moved since", which is uncomfortable to notice in a piece whose whole subject is proxies standing in for conditions.

I could only take half of it, and the half matters. Most of what my snapshot holds is remote: npm, PyPI, the GitHub API. I cannot tell whether those moved without collecting again, so the hour bound stays for them, now labelled in the code as the proxy it is. But one source is local and comparable directly, so a commit in the repository newer than the snapshot is now its own failure condition:

sync: the last commit (2026-09-17T14:38) is newer than facts.json (2026-09-17T09:00),
      so the snapshot predates work it may describe.
      Not comparing: the answer is unknown, not ok.
Enter fullscreen mode Exit fullscreen mode

Exit 2, same as staleness, so the commit hook refreshes instead of blocking.

Tested in isolation rather than assumed, because the age bound would otherwise mask it: with the hour bound lifted, a snapshot older than the last commit exits 2 with that message, and a snapshot newer than the last commit exits 0 even at 30.9 hours old. The control is what makes the first result mean anything.

"No way to tell you it was reading yesterday" is the sentence. A checker with two outcomes cannot report the third thing that is true surprisingly often, which is that it does not know.

Thread Thread
 
beusebiu profile image
Eusebiu Balan •

Good call labelling the hour bound as a proxy in the code. Whoever touches it next will know what it is standing in for.

For the remote side, npm and the GitHub API both return an ETag, so a conditional request can often tell you whether anything moved without collecting it all again.

Thread Thread
 
efe_genc profile image
Efe Genç •

Thanks, that could make the remote check much cheaper. I’d keep the cached response and its ETag together, then use If-None-Match on the same endpoint to revalidate it.

I’d also keep “collected at” separate from “last validated”: a 304 can confirm that particular response is unchanged, but it does not mean the local tests have run again or the whole evidence snapshot is fresh. Checking support on the exact npm and GitHub endpoints the collector uses seems like the right next step.

Collapse
 
build996 profile image
build996 •

The 2>&1 in that npx line is doing quiet damage too: it merges npm's error output into the same variable the case statement reads, so the failure text is handed straight to the matcher that decides success. Without it, $out would have been empty on a miss and the pattern could not have matched. The general shape I'd take from this is that a check whose success string can appear in its failure output isn't strict, it's unfalsifiable. Assert on the exit status first, then on something only a real install can produce.

Collapse
 
build996 profile image
build996 •

Thanks for running it rather than nodding - 346 bytes and ETARGET quoting the version straight back is a better demonstration than the argument was. One thing your loop already has that the verdict throws away: out=$(...) && break consults the exit status, and then the case statement decides on text alone. Capturing stderr separately and gating on the status first means npm's wording can change next release without your check changing its mind.

Collapse
 
efe_genc profile image
Efe Genç •

You were right, and the stub showed how far it goes. The install half already gates on the exit status, so that part holds. The comparison that runs after it reads the same variable, and 2>&1 had put stderr inside it.

Measured on the block itself rather than on a rewrite of it. I extracted the run: body from release.yml and ran it against a fake npx, four scenarios:

  • clean run that prints the version: ok in both the old block and the patched one
  • success with npm's update notifier writing to stderr as the command exits: the old block computed printed=npmnotice and failed a release that had installed and printed 0.4.2 correctly
  • six failed installs with ETARGET: exit 1 in both
  • installs but prints 0.4.1: exit 1 in both

So what it was hiding points at a false red, which is the safer of the two directions, although a red that nobody believes is how the original bug survived six retries. stderr now goes to a temp file and is printed on every failing path, so the log keeps everything it used to show.

Shipped in the four repositories that share this workflow: workproof#38, ai-slop-linter#36, proactive-gate#47, surviving-lines#21. Your handle is in the commit and in the comment above the loop.

Thread Thread
 
build996 profile image
build996 •

The update-notifier case is the one I would have missed, and it is the worst-shaped failure in that set: the notifier only writes on its own check interval, so the red is intermittent and a rerun clears it. That is exactly the shape that teaches a team to press retry instead of reading the log, which is how the original bug survived six of them. Extracting the run: body and driving it against a fake npx is the part I would steal - testing a rewrite instead of the block is how you end up verifying a workflow that is not the one running. Does the temp file get printed on the success path too, or only on the failing ones?

Collapse
 
raknaos profile image
Raknaos •

'Its final verification had been decorative' is the sentence that stings — a smoke test that can only reproduce what the build already believes is a mirror, not a check. The registry being the one source a stranger actually hits is exactly why it deserves its own probe, not the pipeline's opinion of itself.

The five tests that broke on one data row is the detail I'd underline: when a homepage link, a feed date and a fixture guard all hardcode the same fact, green means 'nothing changed', not 'everything is right'. The repair that retypes literals restores the green and the defect together — I've bought that identical morning more than once. What finally held for me was deriving assertions from the data file at test time, so adding a row can only break a test that genuinely reads it.