Code: Megapixel99/didrun
Exit code 0 means "I did not fail". A suite of ten thousand assertions and a suite that collected nothing both report it; no amount of reading the number harder will separate them. That is how a check silently stops checking and nobody finds out for a year. The glob stops matching, the marker deselects everything, the step keeps exiting 0, and green is what everyone was looking for.
The measured example, from didrun's README rather than from folklore: go test ./... on a tree with no test files prints [no test files] and exits 0. Wrapped as didrun --expect "^ok " -- go test ./..., the same run exits 3. The report says nothing in the output matched, there is no evidence this command did anything, and the exit 0 was itself the failure.
The rule the tool is built on is that a check must answer three questions separately: did it run, did it fail, and was the failure the right one. Collapsing any two of those is how every defect of this kind happens. So didrun returns four states rather than a number: ran-and-passed (exit 0), ran-and-failed (the command's own exit code), did-not-run (exit 3), and ran-and-failed-wrongly (exit 4). did-not-run never borrows the command's own code, even when the command failed. "Your tests failed" and "you have no tests" send you to different places. ran-and-failed-wrongly gets exit 4, and exists because "it failed" is not "my check caught something". A suite that dies on a syntax error fails exactly as red as one that caught the defect you planted.
Evidence is what you declare, and declaring none is an error rather than a degradation. With no predicate, the library throws and the CLI exits 2, since a tool that silently falls back to forwarding the exit code is the thing it replaces. The predicate that matters most is --expect-count, and the committed control test is test_zero_passed_is_DID_NOT_RUN_even_though_the_line_matched. A run printing 0 passed in 0.01s with exit 0 satisfies matches(/\d+ passed/) perfectly, because the runner cheerfully prints its zero. The thing that makes a green run meaningless is almost always a zero rather than an absence, so count(/(\d+) passed/) reads the number and scores the run did-not-run. The suite pins both behaviours side by side: the pattern check is fooled, as advertised, and the count check is not.
--wrote PATH is the other predicate with a sharp edge: it is not "the file exists". A junit.xml left over from yesterday exists, and a runner that never started leaves it exactly where it was. So the file must have been created, changed or rewritten during this run, and a byte-identical leftover is reported as the stale thing it is. A timeout is classified rather than swallowed; a killed check did not finish, and must never read as having passed. And the report prints what each predicate looked for and what it found even when everything passed, because a check whose output is only a verdict is one nobody can audit.
Some runners already answer the first question, and the README's table says to let them. pytest exits 5 when it collects nothing, and jest and vitest fail by default when no test matches (their --passWithNoTests flag is the decision this tool exists to argue with). go test, linters handed an empty glob, and every plain shell step in every CI file are the rows this is for. Even the good rows fail the second question, though. A pytest suite in which every test is skipped prints 2 skipped and exits 0 (measured here on pytest 9.1.1), and nothing about that run checked anything.
The package ships as one command from two registries, pip install didrun and npm install -g @megapixel99/didrun. The npm half is scoped because npm refused the bare name as too similar to an existing package called madrun. The release workflow enforces that the two names may differ by the scope and nothing else, and a parity suite asserts the two halves share the four state names, the exit codes, and the same classification of the same run. Six mutations were applied to the source (never reporting did-not-run, treating it as success, requiring only one predicate instead of all, allowing a run with no evidence, ignoring the count floor, accepting a stale artefact) and each was caught by the test that should catch it.
The prior-art sweep found the space emptier than I expected. pytest-custom-exit-code is one runner answering the question for itself, and evidence-gate on PyPI audits GitHub Actions evidence bundles after the fact. ranit intersects a coverage database with a git diff, which is the closest thing in spirit found anywhere. Nothing found wraps an arbitrary command and asks whether it did anything, which is a strange gap for a failure mode this common. Parsing the reports the runner already wrote is a different question, with its own package and its own post to come.
Top comments (0)