DEV Community

Cover image for Five Things Release Day Caught That Six Weeks of Green Tests Didn't
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

Five Things Release Day Caught That Six Weeks of Green Tests Didn't

For weeks my CI badge was red, and I had learned to look past it. Release day, I finally read the log start to finish — and what the red was hiding was worse than the red.

The plan for the day was: build, upload, tag, done. Instead, tagging HivePlane v0.1.0 became the most valuable debugging session of the entire project, because release day is the first day everything has to be true at the same time. If you have ever treated a failing CI badge as background noise, item 3 is for you. If you ship packages, item 1 is for everyone.

Here are the five things I found, in descending order of embarrassment.

1. There was a 35 MB third-party file inside my 22 MB sdist

The first package build "succeeded." Then I listed its contents: 1,442 files, including the entire field-test corpus — among it a 35 MB JSON of upstream repository commits — plus 111 files from vendored third-party agent checkouts.

I hadn't configured the sdist at all. The default had quietly packaged a slice of the whole working tree, not the product. My first fix — an exclude list — still leaked: the builder was walking the disk, not the tracked tree, so untracked third-party files slipped through the cracks.

The fix that worked was the opposite philosophy: an allow-list. Only src/hiveplane plus the root metadata files can enter the sdist. Everything else is excluded by definition.

Build Contents Size
Default sdist 1,442 files, field-test corpora, third-party files 22 MB
Exclude-list sdist still leaked third-party files 180 KB
Allow-list sdist sources + root metadata only 135 KB

tip: For artifact contents, allow-lists beat deny-lists. A deny-list enumerates what you remember to fear; an allow-list enumerates what you actually ship.

2. Two third-party agents were committed to my repository

Auditing what was checked in — not what would ship, what was in git — turned up two complete copies of other people's public repos: a GitHub-analysis agent (10 files) and a weather agent (7 files), vendored into my field-test tree with their original READMEs still telling readers to clone them from upstream.

They were referenced only by documentation; nothing depended on them. They were untracked (git rm --cached), gitignored, and documented in the re-clone table so the field test stays reproducible. The evidence directories — the actual field-test results and Docker evidence — were untouched, because that's the whole point of the audit: remove the borrowed code, keep the receipts.

The reassurance: a full-history secret scan across every commit had already come back clean — zero findings. This was a licensing and hygiene problem, not a leaked-credential problem. Both findings went into the security audit.

3. My CI had never been green — not once

Every recent push showed the same red ❌. I'd been reading the local suite (make test: green, 925 passing) and treating CI as noisy.

The failure log told the real story: the strict-mypy step failed before pytest ever ran — on every push, for weeks. Which meant the Postgres-gated tests had never executed in CI. My "CI is flaky" was actually "CI has been testing a different, smaller system than I thought, and even that not successfully."

The fix for the blocker itself was one line — a type annotation. One line, hiding four real test failures behind it for weeks. The worst of them is next.

4. The tamper-evident audit log wasn't

When CI's pytest finally ran against Postgres, it failed on a test I wanted to fail: tamper the audit log, verify it detects the tampering.

The test edited the actor column directly in the database and expected verify() to return False. It returned True. The reason is a classic denormalization trap: the audit log stored each record twice — once in proper columns, once as a JSON payload — and verification rebuilt records from the payload, checking the hash chain against the payload while the columns sat tampered and unread.

A tamper-evident log that verifies against a copy the tamperer doesn't touch isn't tamper-evident. The fix reads the hashed fields from the authoritative columns. The bug had been invisible locally because the Postgres-gated tests skip without a database — my green local suite was green about a security property it had never once exercised.

tip: Test tamper-evidence by tampering. A verify-only test proves the chain validates itself; it says nothing about whether an edit is detected.

5. pytest is not python -m pytest

The last one is small, but it explains a whole category of "works on my machine":

  • The Makefile ran python -m pytest, which adds the current directory to sys.path. Green locally, every day.
  • CI ran the pytest console script, which does not. Two test modules importing a top-level examples package failed at collection — and had been failing at collection, behind the mypy wall, the entire time.

And while fixing the follow-up warning filter, I learned the hard way that pytest's filterwarnings entries split on : — a message regex containing a colon made pytest try to import a module literally named <function BaseConnection. Configuration that fails to parse should fail loudly; this one failed quietly, by misinterpreting.

What I learned

A gate that always fails hides everything. Weeks of red CI had trained me to stop reading it — and inside that noise, four real test failures sat unnoticed. An always-red gate is operationally identical to no gate, except worse: it manufactures the habit of ignoring the signal.

Your local suite is green about things it never ran. The skipped Postgres tests weren't failing; they were silent — and silence reads exactly like success on a green terminal.

Inspect your artifact before a registry makes it permanent. Publishing is irreversible. The 30 seconds of tar -tzf before upload is the cheapest quality gate I've ever run.

Audit what's checked in, not just what ships. The two third-party agents were never in the wheel — but the repo is the product too, and it just went public.

All five findings are fixed and shipped in v0.1.0 — that's the actual argument of this article: release day found in hours what weeks of development had quietly missed, because it was the first day everything had to be true at once.

References

What's hiding in your CI log or your build artifact right now that you've been meaning to look at — and what's your excuse for not looking?

Top comments (2)

Collapse
 
dhruv_malaviya profile image
Dhruv Malaviya •

"The builder was walking the disk, not the tracked tree" is the transferable line. Build inputs should come from version control, not the filesystem , git archive makes accidental contents structurally unable to enter the artifact. An allow-list enforces the same thing a layer later.

Two additions: check the wheel as well as the sdist, since they're built by different code paths. And a clean history scan means no detectable secrets, not no secrets , worth saying so the result isn't over-read.

Item 3 helps the most people. A badge nobody reads is worse than none.

Collapse
 
reidmarlow profile image
Reid Marlow •

The denormalization bug in item 4 is the cleanest breakdown of the self-validating ledger trap. Storing the row in structured columns while hashing a separate serialized payload means a direct database update preserves the hash chain because nobody touched the blob. Testing tampering by actually mutating the persistent store is something most audit-trail test suites skip.

The pytest console script discrepancy in item 5 is also why the src-layout became standard. Putting code under src/ forces tests to resolve against an installed editable package, so import resolution fails immediately on local machines if packaging metadata is broken.